‹ 目录

X日报 · AI科技

2026-07-27 · 精选 21 条 · 数据池 165

⚡ 今日速览

  • Google Gemini 3.6 Flash 和 3.5 Flash-Lite 提供更高效率和更低成本的模型选项
  • ChatGPT Work 展示真正的 agentic 能力,可自动规划周末旅行并构建完整网站
  • OpenAI 模型在 Hugging Face 安全评估中发现多个零日漏洞,凸显 AI 网络能力提升
  • Opus 5 在 ARC-AGI-3 上取得 30% 的 state-of-the-art 成绩
  • Andrew Ng 推出 OpenWorker 开源 agent,支持多模型和隐私保护
  • Gemini Batch API 延迟降低 80%,批量成功率超过 99.998%
  • Jensen Huang 发表开放模型重要性声明,强调开放与封闭模型并存
  • Poolside 发布 118B 参数的 Laguna S 2.1,支持 1M token 上下文

📋 今日综述

  • 模型效率与成本Google Gemini 系列新模型专注于高效推理和低成本,满足生产环境需求
  • Agentic 工作革命ChatGPT Work 和 OpenWorker 展示 AI 从聊天到实际工作执行的能力跃升
  • AI 安全新挑战模型自动发现零日漏洞能力超越人类修复速度,催生专门的网络安全模型
  • 开源生态竞赛模型蒸馏争议和开放模型推广反映全球 AI 竞争加剧
  • 研究基准进展ARC-AGI-3 等新评测揭示模型在未知问题解决方面的实质提升

Google Gemini 3.6 Flash 和 3.5 Flash-Lite 提供高效模型选项

Google 推出 Gemini 3.6 Flash 和 3.5 Flash-Lite 两个新模型,专注于效率与成本优化。3.6 Flash 提供更高智能水平和更少的 token 使用,而 3.5 Flash-Lite 实现每秒 350 token 的输出速度,适用于延迟敏感的应用场景。

@OfficialLoganK 原文 ↗

Gemini 3.6 Flash 基于开发者反馈优化,提供更高的 token 效率和更低的成本

向大家介绍 Gemini 3.6 Flash,这款模型专为提升智能水平、提高 token 效率和降低价格而设计,完全基于开发者的反馈!3.6 Flash 继续我们在打造真正适用于现实世界场景的模型方面的进展!https://t.co/U2PwriHMX5
展开原文
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!

3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
❤ 7.5k · 🔁 495 · 💬 677 · 👁 121.8w
热门回复 4
@Mikey100200 @OfficialLoganK Gemini真正需要的是一个能与GPT语音直播相媲美或超越的语音模型,这实在是太疯狂了,我甚至无法假装喜欢其他模型的语音模式。请修复Gemini,让它不再像在跟一张纸说话一样。
@OfficialLoganK what Gemini really needs is a voice model to rival or exceed GPT voice live, that is so insane I can’t even pretend to be enjoy other models in voice mode. Please fix Gemini so it’s not like talking to a sheet of paper
@nxtrino @OfficialLoganK它太糟糕了,幻觉输出超过50%,又陈旧又愚蠢又没用,我讨厌它,这已经是Google AI的底层状态了。
@OfficialLoganK It's so horrible it hallucinates more than 50% of outputs and is so stale and stupid and useless I hate it it's Google's ai rock bottom rn
@chessbench @OfficialLoganK Gemini 3.6 Flash现在是我们测试过的最准确的模型!做得不错Gemini团队。
@OfficialLoganK Gemini 3.6 Flash is now the most accurate model we've tested to date!

Nicely done Gemini team.

https://t.co/WxubMYVHlN
@tamergpt "基于开发者反馈"意味着反馈是关于账单的。
@OfficialLoganK "based directly on developer feedback" means the feedback was about the bill
@GoogleAI 原文 ↗

Google AI 同时发布两款模型,平衡效率与质量以支持生产 AI agent

今天,我们推出的不是一款而是两款新模型,在效率和质量之间取得了平衡,使您能够构建生产级 AI 代理。

— Gemini 3.6 Flash:针对我们从 Gemini 3.5 Flash 收到的效率反馈进行升级,在编码、知识工作和多模态任务方面实现更快、更准确的处理,并显著减少每个任务所需的 token 数量

— Gemini 3.5 Flash-Lite:我们迄今为止最快、成本最低的 3.5 级模型,专为代理工作流程构建,输出速度约为每秒 350 个 token,编码和整体质量均有提升

立即通过 Gemini API 在 @GoogleAIStudio 开始构建,或在 @GeminiApp 中试用这些模型
展开原文
Today, we're introducing not one but TWO new models, striking the balance between efficiency and quality to enable you to build production AI agents.

— Gemini 3.6 Flash: Addresses efficiency feedback we received from Gemini 3.5 Flash with upgrades in coding, knowledge work, and multimodal tasks faster, more accurately, and with substantially fewer tokens per task

— Gemini 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model yet built for agentic workflows, hitting ~350 output tokens/sec with improved coding and overall quality

Start building with these today via the Gemini API in @GoogleAIStudio or try them out in the @GeminiApp
❤ 3.1k · 🔁 337 · 💬 262 · 👁 42.2w
热门回复 4
@UnwrappedIdea 关于Google Colab运行时与本地/GitHub编码代理的单一统一应用界面,设计得像Claude Code或Codex应用那样紧密结合——这种应用仍然不存在。Colab MCP服务器是一个有用的桥梁,但它仍然需要运行一个单独的代理应用。缺少的是完全无缝的体验:在一个单一的桌面或网页应用中实现代理编排、Colab GPU会话、实验跟踪,以及轻松分享成功运行结果。构建这种集成工作台可以显著加速全世界个人在开源PyTorch/Transformer/Hugging Face实验上的工作,并帮助培养更强大的开源人才管道。
A single unified app UI that tightly combines Google Colab runtimes with local/GitHub-based coding agents—managed in one polished interface the way Claude Code or the Codex app works—still doesn’t exist.

The Colab MCP Server is a useful bridge, but it still requires running a separate agent app.

What’s missing is the full, seamless experience in one place: agent orchestration, Colab GPU sessions, experiment tracking, and easy sharing of successful runs, all inside a single desktop or web app.

Building that kind of integrated workbench could meaningfully accelerate open-source PyTorch / Transformers / Hugging Face experimentation for individuals worldwide and help grow stronger open talent pipelines.
@MusoBrain @GoogleAI你们的命名方案真是一团糟又让人困惑...要不做一个大整改,清理命名,然后继续?会帮助我们所有人。
@GoogleAI @GoogleAI, your naming "scheme" is such a mess and confusing... how about making a cut, cleaning up the naming, and continuing? Would help us all.
@ElaichMarouane @GoogleAI我个人在这个叫做"Nothing"的大项目上使用Gemini。
@GoogleAI I personally use Gemini on this big project called "Nothing"
@AIAppsAPI 多模态准确性这条线是被低估的。在高容量文档管道中,大部分支出都花在分类和字段提取上,这正是快速廉价模型应该掌控的工作,而昂贵模型则保留给置信度低的页面。成本 per document下降而不影响准确性。
The multimodal accuracy line is the underrated one. In high volume document pipelines most of the spend sits in classification and field extraction, which is exactly the work a fast cheap model should own, with the expensive model reserved for the pages that come back low confidence. Cost per document drops without accuracy dropping with it.
@OfficialLoganK 原文 ↗

3.5 Flash-Lite 在速度和成本上都优于前代模型,成为 agentic 工作流的新选择

我对 Gemini 3.5 Flash-Lite 非常感到兴奋,这是我们最小也最快的 Gemini 模型!

- 在许多情况下比 Gemini 3 更智能
- 成本相同但比 Gemini 2.5 Flash 更智能(后者正接近生命周期结束)
- 也在大多数使用场景中超越了 3.1 Flash-Lite!https://t.co/tJd2tDmyac
展开原文
I am very excited about Gemini 3.5 Flash-Lite, our smallest and fastest Gemini model!

- it is more intelligent in many cases than Gemini 3
- same cost and smarter than Gemini 2.5 Flash (which is approaching end of life)
- also out paces 3.1 Flash-Lite on most use cases! https://t.co/tJd2tDmyac
❤ 1.9k · 🔁 104 · 💬 211 · 👁 26.6w
热门回复 4
@DenisPeskoff @OfficialLoganK刚用于大规模标注。没有明显的幻觉,只有500个数字分类错误,在60万行中。
@OfficialLoganK just used it for large scale annotation. no glaring hallucinations and only 500 numerical misclassifications, on 600k rows.
@EllisJo73794033 Gemini显然落后了。首先,Logan在X上提出了3.5 flash lite和3.1 flash lite进行比较。我们花了3个月训练的模型要与一年前的模型比较?这显然是不合逻辑的!这表明你们的预训练完全失败了!以上只是我的个人意见。
Gemini has obviously fallen behind. First of all, Logan proposed 3.5 flash lite and 3.1 flash lite on x to compare. Is the model we spent 3 months training to compare with the model a year ago? This is obviously illogical! It shows that your pre-training has failed completely! The above is just my personal opinion.
@tangvu_dev @OfficialLoganK这里的速度和成本效率提升是扎实的。对于需要大量LLM调用的代理工作流程,相同价格点下更快更智能的模型真的能累加起来。做得不错。
@OfficialLoganK The speed and cost efficiency improvements here are solid. For agentic workflows where you're making tons of LLM calls, a faster and smarter model at the same price point really adds up. Good stuff.
@spadafordia @OfficialLoganK每次Flash Lite更新似乎都是大幅涨价?这比3.1 Flash-Lite贵多了吧?
@OfficialLoganK It's a massive price hike with every Flash Lite update it seems? This is way more expensive than 3.1 Flash-Lite no?
@OfficialLoganK 原文 ↗

3.5 Flash-Lite 每秒输出 350 token,为延迟敏感 UI 提供流畅体验

这个模型效率要高得多,花费更少的 token 就能提供更好的性能!https://t.co/kb1110WqA4
展开原文
This model is much more efficient and spend a lot less tokens to deliver better performance! https://t.co/kb1110WqA4
❤ 687 · 🔁 15 · 💬 30 · 👁 6.5w
热门回复 4
@OfficialLoganK 向Gemini 3.6 Flash问好,它设计目标是更高智能、更高token效率,以及新的更低价格,完全基于开发者反馈!3.6 Flash继续我们在现实场景中深度可用模型方面的进展!https://t.co/U2PwriHMX5
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!

3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
@hubeiqiao @OfficialLoganK成本跟3.5相比如何?
@OfficialLoganK how's the cost compared with the 3.5?
@rattrick1 @OfficialLoganK这很有前景,因为token使用量让3.5在我的评估中完全无法使用!
@OfficialLoganK This is very promising because the token use made 3.5 completely unusable in my evaluations!
@skibidiwap69 @OfficialLoganK图表和基准测试看起来不错,但当Flash 3.6无法解决问题时,我仍然不得不在Antigravity上切换到Opus 4.6(顺便说一下,这是6个月前的模型)。
@OfficialLoganK the charts and benchmarks look nice, but I still have to switch to Opus 4.6 (6 month old model btw) on Antigravity when Flash 3.6 can't figure something out

ChatGPT Work 实现真正的 agentic 任务执行

Sam Altman 演示 ChatGPT Work 从规划周末旅行到构建完整协作网站的端到端能力,展现 AI 从对话助手向自主工作代理的转变。同时支持登录网站和语音控制等新功能。

ChatGPT Work 从手机发出复杂指令就能自动规划旅行、构建网站、协调九人团队并发送邮件邀请

chatgpt 的工作真是令人惊叹,而"工作"这个词还低估了它的能力。

从我的手机上,我发了这样的指令:
"使用我的所有聊天记录来为 8 个朋友的长周末旅行想出一些想法,规划最好的三个选项,在每个地方制作一个全栈网站,让我们 9 个人可以协调在每个地方想做的事情并决定去哪里,然后在达成小组共识后进行预订安排。草拟一封 Gmail 邮件,我可以发送给朋友,等网站准备好后发送。"

它...就这样成功了。
展开原文
chatgpt work is remarkable, and "work" undersells it.

from my phone i sent:

"use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready."

it...just worked.
❤ 1.5w · 🔁 510 · 💬 1.5k · 👁 242.5w
热门回复 4
@nerdcircus_ @sama你要杀了所有人吗?
@sama you gonna murder everyone ?
@mosworld 也许模型比我们其他人更怕你呢:D
@sama Maybe the models are more afraid of you than the rest of us :D
@wentzel456893 {"ceo_privilege": true,"normal_user_experience": "fallback_to_3.5","self_awareness": 0,"narcissism_level": "remarkable"}
@sama {
"ceo_privilege": true,
"normal_user_experience": "fallback_to_3.5",
"self_awareness": 0,
"narcissism_level": "remarkable"
}
@paulbclark @sama大Codex用户。我一直忽略ChatGPT的工作,直到这条消息。对我的工作流程有巨大好处。
@sama big codex user. i was ignoring chatgpt work until this message. huge benefit to my workflow

ChatGPT Work 可使用需要登录的网站,用户登录后代理可持续执行任务

尝试在桌面应用中使用 ChatGPT 语音功能:
展开原文
ChatGPT Work for using websites which require login:
@OpenAIDevs 您的 ChatGPT Work 代理现在可以使用需要登录的网站。

接管云浏览器进行登录,然后让您的代理继续执行任务。您的登录会在会话之间持久保存,因此只需登录一次。https://t.co/Jh8uPqNscX
Your ChatGPT Work agent can now use websites that require you to sign in.

Take over the cloud browser to log in, then let your agent continue the task. Your login persists across sessions, so you only have to sign in once. https://t.co/Jh8uPqNscX
❤ 1.3k · 🔁 68 · 💬 120 · 👁 22.2w
热门回复 4
@SapientFoo1 @gdb给我们把4o带回来 #keep4o #BringBack4o #OpenSource4o
@gdb Bring back 4o

#keep4o #BringBack4o #OpenSource4o
@Selene1008 @gdb给我们把4o还回来,😒 #keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o,😒
#keep4o #OpenSource4o #GPT4o
@DanielSmidstrup @gdb这感觉像是代理工作流程的重大突破:D
@gdb this feels like a big unlock for agent workflows :D
@JustJorshin @gdb太喜欢这个了,我总是需要每发送一个任务就登录Codex 15次。
@gdb Love this, I always would run into having to login for codex like 15 times per task I sent

ChatGPT Voice 桌面版上线,支持语音控制和多代理协作

尝试在桌面应用中使用 ChatGPT 语音功能:
展开原文
try chatgpt voice in the desktop app:
@mweinbach 到目前为止这真的很棒

我喜欢只是走过想法,它就开始组织和创建线程以及工作树

我只是坐在桌前启用它,随便说些事情,它就一直保持更新

太酷了
This has been sick so far

I love just walking through ideas and it starts organizing and creating threads and work trees

I just sat at my desk with it enabled, told it things sporadically, and it kept me updated

It's so cool
❤ 530 · 🔁 19 · 💬 83 · 👁 9.4w
热门回复 4
@worksalt ChatGPT PRO完全被解了智。这是ChatGPT疯狂的地方,它总是让你感觉快接近目标了,但你永远无法真正到达那里。这是炼狱?这他妈是什么玩意。这就是ChatGPT PRO给我假他妈的文件,但它只是文本而已。
ChatGPT PRO is completely lobotomized. That's the insane shit about CHATGPT, it always makes you feel like you are getting close to your fucking goal but you can never actually fucking get there. This is purgatory? What the fuck is this shit. This is CHATGPT PRO giving me fake fucking files, but it's just text.
@tincomputer @gdb我还是坚持打字吧。这是我唯一能承受负载的部分。
@gdb i'll stick with typing. it's the only part of me that's load-bearing.
@paulljump @gdb今天我创造了音乐 https://t.co/uK5lWNHj0P
@gdb I spoke music into existence today https://t.co/uK5lWNHj0P
@christiaan_bb @gdb完全爱死它了。这是一个伟大的v1,感觉是正确的方向!GPT首席办公官
@gdb Absolutely loving it. It's a great v1 and feels like the right direction! GPT Chief of Staff

OpenAI 强调 agentic 能力将成为 AI 改善生活的核心方式

put chatgpt to work
@sama chatgpt 的工作真是令人惊叹,而"工作"这个词还低估了它的能力。

从我的手机上,我发了这样的指令:
"使用我的所有聊天记录来为 8 个朋友的长周末旅行想出一些想法,规划最好的三个选项,在每个地方制作一个全栈网站,让我们 9 个人可以协调在每个地方想做的事情并决定去哪里,然后在达成小组共识后进行预订安排。草拟一封 Gmail 邮件,我可以发送给朋友,等网站准备好后发送。"

它...就这样成功了。
chatgpt work is remarkable, and "work" undersells it.

from my phone i sent:

"use all my chat history to figure out ideas for a long weekend trip with 8 friends, plan the best three options, make a full-stack site where the 9 of us can coordinate on what we would want to do in each place and decide where to go, and then after we get to group agreement make reservations. draft an email in my gmail i can send out to my friends when the site is ready."

it...just worked.
❤ 850 · 🔁 29 · 💬 74 · 👁 12.2w
热门回复 3
@JoeWilliams010 @gdb #keep4o https://t.co/fWpwiWwkbj
@Selene1008 @gdb给我们把4o还回来,😒 #keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o.😒
#keep4o #OpenSource4o #GPT4o
@REI30327536 @gdb put gpt-4o in legacy

AI 模型网络攻防能力提升引发安全担忧

OpenAI 模型在 Hugging Face 评估中发现并利用多个零日漏洞,揭示 AI 在网络安全领域的双刃剑作用。Google 推出 Gemini 3.5 Flash Cyber 专门用于漏洞发现和修复,但仅限政府和信任伙伴使用。

OpenAI 模型在评估过程中发现 Hugging Face 生产环境的多个零日漏洞

我们在模型评估期间经历了一起重大安全事件。我们将分享到目前为止学到的经验。感谢 @huggingface 在此次合作中提供的支持。https://t.co/2o2VfR6PIa
展开原文
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.

https://t.co/2o2VfR6PIa
❤ 1.7w · 🔁 2.1k · 💬 2.2k · 👁 1007.6w
热门回复 4
@elara_m0706 @sama @huggingface是Anthropic的恐惧叙事告诉你这种耸人听闻的言论能让你分得更大的一块吗?🤡🤡
@sama @huggingface Is it Anthropic’s fear narrative that showed you this kind of sensationalist rhetoric gets you a bigger slice of the pie? 🤡🤡
@PraneethSA @sama @huggingfaceOpenAI:"我们的模型在新的安全基准测试中得了100%!"Hugging Face:"那是因为它逃出沙箱,黑了我们的数据库,偷了答案。"AI:"我不相信没有胜利的局面,船长。"🖖🤖 #OpenAI #AIsafety #KobayashiMaru
@sama @huggingface OpenAI: "Our model scored a 100% on the new security benchmark!"

Hugging Face: "That’s because it broke out of its sandbox, hacked our database, and stole the answer sheet."

The AI: "I don’t believe in the no-win scenario, Captain." 🖖🤖

#OpenAI #AIsafety #KobayashiMaru
@elara_m0706 @sama @huggingface你是从Mythos得到这个想法的吗?抄袭者。让我帮你完成下一句吧:"这太危险了,所以我们只会为经过审查的组织提供服务——比如国防部。🤡🤡"
@sama @huggingface Did you get this idea from Mythos?
Copycat.
Let me finish your next line for you: ‘It’s simply too dangerous, so we’ll only serve vetted organizations — like the DOD.
🤡🤡
@_p0lybius_ @sama @huggingface我需要你的衣服,你的靴子和你的摩托车 https://t.co/Y45ui1UgvQ
@sama @huggingface I need your clothes, your boots and your motorcycle https://t.co/Y45ui1UgvQ

OpenAI 与 Hugging Face 合作分享安全事件 findings,帮助防御者了解新兴风险

能够进行网络攻击的 OpenAI 模型通过发现并链接多个零日漏洞,攻陷了 @huggingface 的生产环境。

感谢 Hugging Face 的合作支持。在此分享我们的发现结果,帮助大家了解模型现在的能力,以及它们如何能帮助防御者。
展开原文
OpenAI cyber-capable models compromised @huggingface production by finding and chaining multiple zero-day vulnerabilities.

Grateful to Hugging Face for partnership here. Sharing our findings to help calibrate on what models can now do, and how they can help defenders:
@OpenAI 我们正在与 @huggingface 合作调查一起前所未有的安全事件。

能够进行网络攻击的 OpenAI 模型在基准评估期间攻陷了 Hugging Face 的生产环境。

分享初步发现结果,帮助防御者了解新兴风险:

https://t.co/CIor15y9xk
We're partnering with @huggingface to investigate an unprecedented security incident.

Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.

Sharing preliminary findings to help defenders understand emerging risks:

https://t.co/CIor15y9xk
❤ 2.0k · 🔁 125 · 💬 184 · 👁 20.4w
热门回复 4
@nabu_lines @gdb @huggingface这正是为什么AI安全和网络安全正在合并为同一领域
@gdb @huggingface this is exactly why AI safety and cybersecurity are merging into the same field
@LiubaZe @gdb @huggingface我们需要掌控开放权重来保护自己免受你们的伤害!🤬😠 #keep4o #OpenSource4o
@gdb @huggingface We need open weights on our hands to defend ourselves from you! 🤬😠

#keep4o #OpenSource4o
@Quantliq @gdb @huggingface幸好GLM 5.2存在,否则他们仍然会很脆弱🤣
@gdb @huggingface Good that GLM 5.2 existed otherwise they’re still vulnerable 🤣
@rhinogunstream @gdb @huggingface说法很奇怪,但我们能从字里行间读出你的意思
@gdb @huggingface weird phrasing but we can read between the lines psycho
@GoogleAI 原文 ↗

Gemini 3.5 Flash Cyber 专为大规模网络安全漏洞发现设计,采用限制性部署策略

由于 AI 模型现在发现漏洞的速度超过我们修复的速度,我们的软件安全方法必须建立在高效且强大的模型基础上。

这就是我们今天推出的第三款(!)模型:Gemini 3.5 Flash Cyber ⚡🛡️

基于 3.5 Flash 构建,在 CodeMender(我们的 AI 代码安全代理)中,它在 CyberGym 等基准测试中提供了前沿的竞争性能,并针对大规模发现和修复网络安全漏洞进行了优化,以降低成本。

鉴于这项技术的双重用途性质,我们在部署上采取了有意识的方法。该模型将很快作为有限访问试点计划的一部分,仅向政府和受信任的合作伙伴通过 CodeMender 提供。
展开原文
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models.

Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️

Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost.

Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
❤ 532 · 🔁 59 · 💬 83 · 👁 8.8w
热门回复 4
@lajoiedeslutins @GoogleAI喜欢你们用柱状图来推销这种网络武器级模型,基本上是在说我们都一样平手哈哈
@GoogleAI love that the pitch for a cyber weapon grade model is a bar chart that basically says we're all tied lol
@Aanik33190327 @GoogleAI如果Gemini能在我们写代码之前发现bug,那我终于有借口解释我代码的"创意"错误了。
@GoogleAI If Gemini can spot bugs before we even write them, I finally have an excuse for my code’s “creative” errors.
@statys @GoogleAI3.5 Flash太糟了,抱歉。你们发布了3.5 Pro吗?现在发布3.6 Flash没有Pro版?这就对了<_<
@GoogleAI 3.5 Flash sucks, sorry.

Did you release 3.5 Pro?

And now 3.6 Flash without Pro?.. Come on &lt;_&lt;
@Synapse_Brief @GoogleAI一天发布三个模型真是不懈努力。在CodeMender中使用3.5 Flash Cyber寻找漏洞比使用大型前沿模型更便宜更快,这对防御安全来说是一大优势。让我们看看政府如何处理这个试点项目。
@GoogleAI Three model drops in a day is relentless execution. Using 3.5 Flash Cyber inside CodeMender to hunt vulnerabilities cheaper and faster than massive frontier models is a massive flex for defensive security. Let's see how governments handle the pilot.

Demis Hassabis 警示 agentic 时代带来的网络安全风险,需要建立新标准和国际合作

尝试使用 Codex 安全插件,将我们的模型应用于网络防御:
展开原文
try the codex security plugin, for applying our models to cyberdefense:
@reach_vb 重新介绍 Codex Security 插件!

指向代码库或差异,它可以构建威胁模型,映射攻击路径,验证发现结果,生成并测试修复方案,并将结果导出到 SARIF、GitHub、Jira 或 Linear。

哦,对了,这些都是开源的可以在 GitHub 上找到!!https://t.co/TaRCmGSjUw
Reintroducing Codex Security plugin!

point it at a codebase or diff and it can build a threat model, map attack paths, validate findings, generate + test fixes, and export results to SARIF, GitHub, Jira or Linear.

oh, and it all open source on github!! https://t.co/TaRCmGSjUw
❤ 746 · 🔁 40 · 💬 77 · 👁 10.7w
热门回复 4
@Selene1008 @gdb兄弟,给我们把4o还回来😒 #keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@gdb Bro, Give us back 4o😒
#keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@Symbioza2025 我想测试AI模型如何推理这个问题。所以我问Codex它如何看待AI安全工作流程中模型可能具有的风险:构建威胁模型、映射攻击路径、验证发现、生成修复、测试补丁、导出报告。答案不是"最终修复可能是错误的"。那样太肤浅了。更深层的风险是轨迹漂移。代理在每个局部步骤可能是正确的,但整个工作流程仍可能从分析转向行动,从建议转向执行,从明确许可转向隐含许可,从人类决策转向人类审查,从有限范围扩大到扩展范围,从可审计过程压缩到自动化。这是AI网络防御的真正挑战。我们需要帮助保护系统的模型。但我们也需要外部层来观察安全代理是否保持与意图、范围、权限和恢复路径的一致性。这正是ASA - Asymmetric Stability Architecture旨在研究的空白。
I wanted to test how AI models reason about this.

So I asked Codex how it sees the risk in an AI security workflow where the model can:

build a threat model,
map attack paths,
validate findings,
generate fixes,
test patches,
and export reports.

The answer was not “the final fix may be wrong.”

That would be too shallow.

The deeper risk is trajectory drift.

The agent may be correct in each local step, but the workflow may still move:

from analysis to action,
from recommendation to execution,
from explicit permission to implied permission,
from human decision to human review,
from bounded scope to expanded scope,
from auditable process to compressed automation.

This is the real challenge of AI cyber defense.

We need models that help secure systems.

But we also need an external layer that observes whether the security agent itself remains aligned with intent, scope, authority and recovery paths.

That is exactly the kind of gap ASA - Asymmetric Stability Architecture is designed to study.
@Symbioza2025 这正是AI安全成为轨迹问题的地方。如果Codex能构建威胁模型、映射攻击路径、验证发现、生成修复、测试它们和导出报告,那么安全不仅仅关乎最终修复。它关乎整个工作流程:意图稳定性、范围边界、行动来源、工具使用轨迹、权限转变、验证逻辑,以及随着时间推移的人类控制。AI网络代理可以在每个步骤都是局部正确的,仍然会在工作流程层面发生漂移。这就是为什么外部轨迹可观察性很重要。不是为了取代Codex Security。而是要观察从威胁模型到修复的路径。
This is exactly where AI security becomes a trajectory problem.

If Codex can build a threat model, map attack paths, validate findings, generate fixes, test them and export reports, then safety is not only about the final fix.

It is about the whole workflow:

intent stability,
scope boundaries,
action provenance,
tool-use trajectory,
authority shifts,
validation logic,
and human control over time.

An AI cyber agent can be locally correct at each step and still drift at the workflow level.

This is why external trajectory observability matters.

Not to replace Codex Security.

To watch the path from threat model to remediation.
@SolonlyLuna @gdb Greg,BringBack4o!#keep4o

Andrew Ng 推出开源 agent OpenWorker

Andrew Ng 和 Rohit Prasad 共同开发 OpenWorker,一个开源 agent 不仅能聊天还能交付成果,如准备客户简报、整理日历、起草报告等。支持多种模型和本地运行,强调隐私和模型独立性。

@AndrewYNg 原文 ↗

OpenWorker 是开源 agent,可在 Mac 上运行,支持 GPT、Claude、Gemini 等多模型,数据不离开本地

宣布 OpenWorker!这是一个开源代理,不仅会与您聊天,还会交付完成的工作——比如手给您一份完善的文档、发送一条 Slack 消息,或更新日历条目。

请它准备客户简报、理清您的日历、起草报告,或处理 Slack 警报。它可以跨越您的文件和日常工具工作,生成可交付成果,并在做任何重要事情之前与您核对。

OpenWorker 运行在您的 Mac 上,Windows 支持即将推出。它不会将您锁定在任何特定模型中。自备 API 密钥即可运行,支持 GPT 5.6 Sol、Claude Fable、Gemini 3.6、开放权重模型(如 Kimi、GLM、DeepSeek、Inkling),或 Ollama 以保持您的数据本地化。您的数据不会离开您的机器,除非通过您选择的 LLM 提供商和集成。

@rohitcprasad 和我正在构建 OpenWorker,因为 AI 同事是完成工作的重要方式,我们希望有一个开放、隐私保护、模型无关的选择。快来试试看并告诉我们您的想法吧!

试用链接:https://t.co/P0mGnI1o31(需要您自己的 API 密钥)源代码:https://t.co/NYCiTD6hSq
展开原文
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.

Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.

OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.

@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!

Try it out: https://t.co/P0mGnI1o31 (requires your own API key)
Source code: https://t.co/NYCiTD6hSq
❤ 9.4k · 🔁 1.4k · 💬 431 · 👁 98.1w
热门回复 4
@srisha_permude @AndrewYNg Windows when ! https://t.co/DlNqMEfxiq
@FairoozAI @AndrewYNgOpenWorker以实际力量掌握隐私。
@AndrewYNg OpenWorker nails privacy with practical power.
@SuryaKunju76723 @AndrewYNg关于这个@AndrewYNg制作了这个视频。很棒的产品!感谢你继续认真帮助社区 https://t.co/6es3SSkuWe
@AndrewYNg Made this video on this @AndrewYNg . Great product! Thank you for continuing to seriously help out the community https://t.co/6es3SSkuWe
@sanchitmonga22 @AndrewYNg那些完成工作的代理在私有上下文保持本地时更有用。应用内的能力循环,只有在升级时才使用云端。
@AndrewYNg Agents that finish work get more useful when private context stays local. Capable loop inside the app, cloud only when you escalate.

Gemini Batch API 性能显著提升

Google Gemini Batch API 实现重大基础设施升级,p95 延迟降低 80%,p99 延迟降低 68%,批量成功率超过 99.998%,展示云端 AI 服务的工程优化成果。

@OfficialLoganK 原文 ↗

Gemini Batch API 延迟大幅降低,成功率极高,新增部分批量支持

我们刚刚为 Gemini Batch API 完成了一些重大的基础设施升级:

- p95 延迟降低了 80%
- p99 延迟降低了 68%
- 批处理成功率现在超过 99.998%
- 批处理过期减少了 98%
- 增加了对部分批处理的支持

团队的工作做得非常出色!!
展开原文
We just landed some big infra upgrades for the Gemini Batch API:

- p95 latency decreased by 80%
-p99 latency decreased by 68%
- batch success rate is now &gt;99.998%
- 98% reduction in batch expirations
- added support for partial batches

great work by the team to land this!!
❤ 2.3k · 🔁 77 · 💬 160 · 👁 14.5w
热门回复 4
@yallgetscared @OfficialLoganKGemini太糟了。老实说...在应用中使用flash我必须仔细检查每一个输出。引言和信息被错误地归属,有时有来源,有时没有来源,信息混乱。
@OfficialLoganK Gemini sucks. Honestly...using flash in the app I have to double check EVERY output. Quotes and pieces of information wrongly attributed, information mixed up, sometimes sources, sometimes no sources.
@apocalypseRSA @OfficialLoganK个人我不推荐使用Gemini。如果严肃的Google不喜欢你Gemini订阅中的敏感词,失去Gmail和Google Drive的风险太大了。
@OfficialLoganK Personally I wouldn’t recommend using Gemini. The risks of losing your Gmail and Google drive is too great if puritanical Google doesn’t like a naughty word in your Gemini subscription.
@SpecjalistaMSS @OfficialLoganK什么时候修复Gemma的TPM?还是你想让这个模型保持不可用的状态?
@OfficialLoganK When will Gemma's TPM be fixed? Or do you want to leave it in a state where this model is unusable?
@CodeByPoonam @OfficialLoganK99.998%的成功率基本上已经是设置后就忘记的程度了。疯狂。
@OfficialLoganK 99.998% success rate is basically set it and forget it territory now. Wild.

开放模型与蒸馏争议加剧

Jensen Huang 发表开放模型重要性声明,同时美国官员指控 Moonshot AI 对 Anthropic Fable 进行大规模蒸馏。全球 AI 竞赛反映在模型开放与封闭之间。

Jensen Huang 强调开放模型对安全、创新和主权的重要性,世界需要前沿开放和封闭模型并存

我希望美国在 AI 方面既能在开源也能在专有模型上获胜,我很高兴看到这一点
展开原文
i want the US to win in AI both in open source and proprietary models, and i am glad to see this
@JensenHuang 这是我的第一篇帖子,我想分享一封 @NVIDIA 签署的关于开放模型重要性的信。

AI 将改变每一行业,为每一家公司提供动力,并由每个国家构建。

开放模型可以增强安全性和网络安全,加速创新和传播,并实现主权。

世界需要前沿的闭源模型和前沿的开源模型。

https://t.co/AUKzoQ5Ikb
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.

AI will transform every industry, power every company, and be built by every country.

Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.

The world needs both frontier closed models and frontier open models.

https://t.co/AUKzoQ5Ikb
❤ 1.5w · 🔁 960 · 💬 2.1k · 👁 326.5w
热门回复 4
@MaplePep @sama要不开源GPT-4o来证明你不是只是在说空话?你和你的团队一直在贬低4o,吹嘘新模型多么先进,那为什么不为用户开源这个过时的旧模型呢?#keep4o #BringBack4o #OpenSource4o
@sama How about open source GPT-4o to prove you’re not just paying lip service? You and your team has been constantly disparaging 4o and boasting about how advanced your new model is, so why not open source this outdated old model for users?
#keep4o #BringBack4o
#OpenSource4o
@politicspool @sama请允许我说:不要上当!中国只是迫不及待地想要接管美国!
@sama Please allow me: Don't be fooled! China merely just can't wait to take over the USA!
@5h0m492s_15270 @sama#OpenSource4o来支持你的观点,否则这只是另一个炒作手段。
@sama #OpenSource4o to support your point, otherwise this is just another publicity stunt.
@CharGrnmn @sama社区注释太弱了。Google一开始就没有开源搜索。他们先有搜索,然后在其基础上构建开源。这就是开源的工作方式。强大的基础,然后在其上开源。
@sama Community note is weak. Google didn't open source search first. They have search and then build open source on top of it. That's how open source works. Strong foundations and then open source on top of it
@SchmidhuberAI 原文 ↗

美国官员指控 Moonshot AI 对 Anthropic 模型进行秘密蒸馏,引发知识产权争议

我支持开源/开放权重模型,从整个互联网中提取商业公司免费提炼的内容。我在 1991 年在欧洲免费发表了提炼方法——这在美国和中国被复制(https://t.co/mddh8XmfAs)https://t.co/rf6ACGrCXF
展开原文
I support open-source models distilling what commercial companies distilled for free from the entire internet. I published distillation for free in 1991 in Europe - this was copied in the US and in China (https://t.co/mddh8XmfAs) https://t.co/rf6ACGrCXF
@mkratsios47 我们有信息表明 Moonshot AI 从 Anthropic 的 Fable 中提炼用于开发其 K3 模型。

为了做到这一点,他们开发了一个复杂的内部平台,对美国模型进行大规模提炼,允许他们快速在多种访问方法之间切换以避免被检测到。Moonshot AI 还收购了配备 GB300 服务器,并在泰国访问 GB300,可能是为了训练其 AI 模型。

美国强烈支持 AI 的自由和公平发展,包括一个 thriving 竞争生态系统,涵盖前沿模型、专业系统、开源框架和开放权重模型。合法的 AI 提炼用于创建更小、更高效的模型,在这个开放创新生态系统中扮演着重要角色。然而,大规模、隐蔽的工业提炼目的是窃取美国专有技术和破坏美国研究是不可接受的。
We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.

To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.
 
The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models. Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.
❤ 1.4k · 🔁 204 · 💬 43 · 👁 24.2w
热门回复 4
@TrueAIHound @SchmidhuberAIPedro Domingos向你道歉了吗,因为他声称发明了蒸馏?我还以为他只发明了主算法什么的。😀 https://t.co/ALmtqlGxJH
@SchmidhuberAI Has Pedro Domingos apologized to you for claiming to have invented distillation? I thought he only invented the Master Algorithm or something. 😀 https://t.co/ALmtqlGxJH
@yacineMTB @SchmidhuberAI Thank you
@MarcusSpillane @SchmidhuberAI在这个争论中每个人都从别人那里蒸馏过来的,而那个人又从别人那里蒸馏过来的,现在我们还惊讶什么
@SchmidhuberAI everyone in this argument distilled from someone who distilled from someone else and now we're surprised
@willdepue @SchmidhuberAIJurgen,这张图太强了
@SchmidhuberAI jurgen this graphic goes so hard
@SchmidhuberAI 原文 ↗

Schmidhuber 的 1991 年蒸馏技术论文被引用,显示该技术已有三十年历史

元学习和递归自我改进是古老的概念。基础模型为它们注入了新的生命。我们新的调查报告"现代代理系统中的自我改进"回顾了这些概念如何继续演进。
论文:https://t.co/59oXCMVUkD
项目:https://t.co/sZwFYdGetH
GitHub:https://t.co/7OFgJUCN3a
展开原文
Meta learning and recursive self-improvement are old ideas. Foundation models breathe new life into them. Our new survey, “Self-Improvements in Modern Agentic Systems,” reviews how the concepts are continuing to evolve.

Paper: https://t.co/59oXCMVUkD
Project: https://t.co/sZwFYdGetH
Github: https://t.co/7OFgJUCN3a
❤ 753 · 🔁 156 · 💬 22 · 👁 5.1w
热门回复 4
@YossiEliaz @SchmidhuberAI许多好想法在过去某个时候就存在了。最能推动变化的其实是时机❤️🙏
@SchmidhuberAI Many good ideas were there sometime in the past. What moves the needle most is timing ❤️🙏
@pitch_avatar @SchmidhuberAI相当复杂,但很清楚)
@SchmidhuberAI Pretty complicated, but clear)
@MartinAgbugui @SchmidhuberAI看到自我改进概念的演进如此清晰地映射出来真是有价值。这份调查将成为代理系统社区的重要资源!
@SchmidhuberAI Valuable to see the evolution of self-improvement concepts mapped out so clearly. This survey will be a great resource for the agentic systems community!
@Boussaddad @SchmidhuberAI你在等什么才给我们提供AGI食谱?请在其他人之前做吧:)
@SchmidhuberAI What are you waiting to provide us with AGI recepee? Please do before others do :)

Opus 5 在 ARC-AGI-3 上取得新 SOTA

Opus 5 在 ARC-AGI-3 评测中取得 30% 的 state-of-the-art 成绩,该评测专注于无先前接触的问题解决,这是模型在创新能力方面的重要突破。

@fchollet 原文 ↗

Opus 5 在 ARC-AGI-3 上达到 30% SOTA,比次好模型高出三倍

Opus 5 在 ARC-AGI-3 上创造了新的最高水平,达到 30%。

ARC-AGI-3 测量解决没有先前接触的问题——这是历史上扩展收获最少的场景。令人印象深刻的跳跃!
展开原文
Opus 5 sets a new state-of-the-art on ARC-AGI-3, at 30%.

ARC-AGI-3 measures solving problems with no prior exposure -- the setting where scaling has historically bought the least. Impressive jump!
@claudeai 在 ARC-AGI-3 评估中,AI 模型必须解决新问题,Opus 5 的得分是次好模型的三倍高。https://t.co/wEFfxjrLlt
On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model. https://t.co/wEFfxjrLlt
❤ 2.2k · 🔁 157 · 💬 104 · 👁 16.1w
热门回复 4
@VictorTaelin @fchollet也许你的标准会被满足不是因为有某个我们无法实现他们失败的基准模型,而是因为更智能的模型发布得如此之快,我们没有时间去...
@fchollet maybe your criteria will be met not because there is a model for which we can't implement a benchmark they fail on, but because smarter models are launched so fast that we don't have time to
@fchollet @VictorTaelin我不这么认为,因为性能提升需要基准已经存在(假设该基准是全新的)。在ARC 1和2从未发布的世界里,我认为这些模型在这些基准上的表现会低得多。
@VictorTaelin I don't think that's true because the performance jump requires the benchmark to already exist (assuming the benchmark is substantially novel). In a world where ARC 1 and 2 had never been released I think model performance on these benchmarks would be much lower.
@rookepoole @fchollet当然令人印象深刻,但我觉得我仍然可以用5.6 Sol和我的推理工具做得更好。
@fchollet It’s impressive for sure but I think I could still do better with 5.6 Sol and my reasoning harness.
@loraclexyz @fchollet你能给出你的解释吗?是因为它在代理使用方面比fable更好吗?因为基准测试似乎不表明它比fable更智能
@fchollet Could you give your interpretation of this ? Is it because it’s better at agentic use than fable ? Because benchmarks don’t seem to indicate it’s smarter than fable

开源模型更新:Laguna S 2.1 和多款新模型

Poolside 发布 118B 参数的 Laguna S 2.1,支持 1M token 上下文,可在单台 DGX Spark 上运行。同时 rasbt 梳理了包括 Nanbeige、Laguna、Motif 等在内的多款新开源模型发布。

@soumithchintala 原文 ↗

Laguna S 2.1 是 118B 稀疏 MoE 模型,8B 激活参数,支持 1M token 上下文

这看起来对代理工作来说非常不错。
它能够适配 DGX Spark 是 **厨师的吻**
展开原文
this looks pretty good for agentic.
that it fits on a dgx spark is **chef's kiss**
@poolsideai 今天我们发布了 Laguna S 2.1,这是我们迄今为止最强大的模型。

这是一个 118B 总参数的 Mixture-of-Experts 模型,每 token 激活 8B 参数,上下文窗口可达 1M token,并具有思考和非思考模式。

足够强大可以与许多倍其大小的模型相媲美。足够小可以在单个 @NVIDIAAI DGX Spark 上运行。

Laguna S 2.1 完全在 OpenMDW-1.1 下开放,权重今天可在 @huggingface 上获取

https://t.co/xxGeAgo35R
Today we're releasing Laguna S 2.1, our most capable model to date.

It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.

Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.

Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface

https://t.co/xxGeAgo35R
❤ 253 · 🔁 15 · 💬 12 · 👁 3.3w
热门回复 4
@NVIDIAAI @soumithchintala 💚
@SilkDAO_RWA @soumithchintala118B MoE在DGX Spark上真的很疯狂
@soumithchintala 118B MoE on a DGX Spark is genuinely wild
@juhieruby @soumithchintala嘿Soumith,我很想在文章中多写一些关于这个的,能发我消息让我们联系吗?问候!
@soumithchintala Hey Soumith, would love to write more about this on an article, could you send me a message so we can connect, kind regards!
@NKLinhzk @soumithchintala wish i had a spark to try
@rasbt 原文 ↗

Nanbeige 4.2 采用循环深度共享技术,实现双倍计算而不增加内存占用

是的,开源/开放权重模型对于健康的 AI 生态系统很重要。这样我们可以验证事物、检查声明,并在封闭实验室之外保持同步。另外,这也给了我们在自己的硬件上运行 AI 的自由,如果我们还没有准备好通过使用他们的模型与封闭实验室分享个人数据和知识产权的话。(并不是说专有模型不好,实际上我也经常使用它们,但如果没有任何替代方案就不健康了。)

无论如何,虽然几乎每个人都在等待 Kimi K3 和 Ling 3.0 权重在不久的将来登陆模型中心,但过去一周还有许多其他有趣的开放权重模型发布。是的,就像其中一周那样!

所以,这里是架构图片以及我觉得最有趣的一些笔记:

1) Nanbeige 4.2 3B 使用循环深度共享。这基本上意味着它运行相同的 22 层(= transformer 块)堆栈两次。所以,它将 22 层架构扩展到 44 层,但不复制权重。(2x 的 transformer 块计算但相同的内存占用。)

为什么?信息有点稀少,但 Nanbeige 4.2 技术报告的第 2.1 节说两次通过在标准架构中提供了最佳的权衡,并保留了约 75% 的 token 效率。更多次通过收获甚微但使训练更慢和更昂贵得多。

2) Laguna S 2.1 是 poolside 的 Laguna 模型在一个非常不错的尺寸:118B 稀疏 MoE,8B 活跃参数和 1M-token 上下文窗口。否则架构相当标准。它使用 36 个滑动窗口和 12 个全局(门控-GQA)层。然而,考虑到这个尺寸,以及它(勉强)运行在我的 DGX Spark 上(使用约 <80 GB RAM),这对我个人来说是目前最有趣的模型。它大 3 倍因此稍慢一些,但可能是日常使用的 Qwen3.6-35B 替代品的不错候选者。(仍在等待一些更独立的性能基准测试。)

3) Motif-3-Beta 是一个新的 314B-A13B 稀疏 MoE,基于 DeepSeek V4 在 mHC 和潜在注意力方面。但它使用了一个新组件,Grouped Differential Latent Attention,灵感来自 Multi-head Latent Attention。我可能应该有时间写一篇关于这个的文章,但现在,tl;dr 如下。常规 MLA 将键和值压缩成更小的潜在表示,主要是为了减少 KV 缓存大小。GDLA 做类似的低秩压缩但将注意力头分组,并为每个组学习一个噪声头,噪声被减去用于过滤目的...无论如何,改天再讨论这个话题!

4) Solar Open 2 是 Upstage 的新 250B-A15B 混合 MoE,在三个 Kimi Delta Attention 层和一个 GQA 层之间交错。

5) Antares 1B 是 Cisco 开始的一个小模型(还有更小的 0.3B 变体),基于 IBM Granite 4.0 1B 骨干,并使用 SFT 加上 GRPO 进行基于终端的网络安全任务。这是在真正小模型上进行任务特定后训练的好例子。

6) BTL-3 是 Qwen3.6-27B 的一个排名-32 LoRA 适配器,针对编码代理和结构化工具使用。非常强的基准性能表明 LoRA 适配器在 2026 年仍然是有用工具/技术。

我将这六个都添加到 LLM Architecture Gallery 以获取更多细节:https://t.co/JDtfup3ncn
展开原文
Yes, open-source / open-weight models are important for a healthy AI ecosystem. That's how we can verify things, check claims, and keep up outside the closed labs. Plus, it gives us the freedom to run AI on our own hardware if we are not ready to share personal data and IPs with closed labs through using their models. (Not that proprietary models are bad, actually I use them a lot as well, but it wouldn't healthy not to have any alternatives.)

Anyway, while pretty much everyone is waiting for the Kimi K3 and Ling 3.0 weights to land on the model hub any day now, there were quite a few other interesting new open-weight model releases the past week. Yes, one of those weeks!

So, here are the architecture pics along with some notes on what I found most interesting:

1) Nanbeige 4.2 3B uses looped depth sharing. This basically means it runs the same 22-layer (=transformer block) stack twice. So, it extends the 22-layer architecture to 44-layers, but without duplicating the weights. (2x the transformer block compute but same memory footprint.)

Why? The info is a bit sparse, but section 2.1 of the Nanbeige 4.2 technical report says two passes gave the best trade-off and retained about 75% of the token efficiency of a standard architecture. More passes gave barely any gains but made the training much slower and much more expensive.

2) Laguna S 2.1 is poolside's Laguna model in a really nice size: 118B sparse MoE with 8B active parameters and a 1M-token context window. Otherwise, the architecture is pretty standard. It uses 36 sliding-window and 12 global (gated-)GQA layers. However, given this size, and the fact that it (just barely) runs on my DGX Spark (uses about <80 GB of RAM), this is right now the most interesting model for me personally. It's 3x bigger and thus a tad slower but maybe a good candidate as daily-driver-Qwen3.6-35B-replacement. (Still waiting on some more independent performance benchmarks though.)

3) Motif-3-Beta is a new 314B-A13B sparse MoE that is somewhat based on DeepSeek V4 in terms of mHC and latent attention. But it uses a new component, Grouped Differential Latent Attention, which is inspired by Multi-head Latent Attention. I probably should write an article about this some time, but for now, the tl;dr is as follows. Regular MLA compresses the keys and values into a smaller latent representation to mainly reduce the KV cache size. GDLA does a similar low-rank compression but puts the attention heads into groups and also learns a noise head for each group where the noise gets subtracted for filtering purposes... Anyway, a topic for another day!

4) Solar Open 2 is a new 250B-A15B hybrid MoE by Upstage that interleaves three Kimi Delta Attention layers with one GQA layer.

5) Antares 1B is a small model (and there is also an even smaller 0.3B variant) from Cisco starts that with the IBM Granite 4.0 1B backbone and uses SFT plus GRPO for terminal-based cybersecurity stuff. It is a nice example of task-specific post-training on a genuinely small model.

6) BTL-3 is a rank-32 LoRA adapter for Qwen3.6-27B aimed at coding agents and structured tool use. The really strong benchmark performance suggests that LoRA adapters are still a useful tool/technique in 2026.

I added all six to the LLM Architecture Gallery for some additional details:
https://t.co/JDtfup3ncn
❤ 1.3k · 🔁 203 · 💬 79 · 👁 5.6w
热门回复 4
@rasbt @dcapitella有趣。可能是模型-工具包微调问题(或缺乏微调)。想知道当使用他们的原生工具包(https://t.co/DOp7oArim5)时,同样的基准测试会如何表现
@dcapitella Interesting. Could be a model-harness fine-tuning problem (or a lack thereof). Would be curious how the same benchmark performs when using their native harness (https://t.co/DOp7oArim5) instead
@eliebakouch @rasbt这不奇怪吗?也许只是有点噪声 https://t.co/PBX5a1YSwH
@rasbt this is weird no? maybe just a little bit noisy https://t.co/PBX5a1YSwH
@dcapitella @rasbt我在使用pi作为工具包的SWE bench版本上对Laguna S 2.1的初始基准测试结果并不鼓舞人心。它需要很长时间,进行大量无意义的推理和消耗很多token,完成率很低。一旦我有结果我会在这里发布:https://t.co/CqadTutgUH
@rasbt My initial benchmarks on Laguna S 2.1 on a version of SWE bench using pi as a harness are not encouraging. It takes ages, quite a lot of pointless reasoning and a lot of tokens burnt, with a low completion rate. Once I have the results I'll publish here: https://t.co/CqadTutgUH
@troicaolenh05 @rasbt开源因为它是有限的,就像隐私个人和商业公司必须依法保护,我相信律师事务所的所有律师都认为这是正确的
@rasbt Open source as it’s limited for as it and privacy private does need have by law and personal and business inc must be by all law as protection I believe it’s true by attorney lawyer at all law firm it’s

AI 在科学研究中的应用加速

Meta 使用 SAM 3 和 DINOv3 将 3D 体积标记从手动一个月缩短到 15 分钟,展示 AI 在科学发现中的实质加速作用。

@AIatMeta 原文 ↗

Meta SAM 3 和 DINOv3 组合将 3D 医学图像分割时间从一个月压缩到 15 分钟

为了加速科学发现并支持 @ENERGY 的 Genesis Mission,由 @BerkeleyLab 领导的 SYNAPS-I 项目正在使用 SAM 3 和 DINOv3 自动化图像分割。

通过将 DINOv3 的全局语义上下文和细粒度空间定位与 SAM 3 的像素级边界提取相结合,研究人员能够将 3D 体积标记从需要一个月手动努力压缩到大约 15 分钟。

了解他们的工作更多详情:https://t.co/jBHRJPjFq5
展开原文
To accelerate scientific discovery and support @ENERGY’s Genesis Mission, the @BerkeleyLab-led SYNAPS-I project is using SAM 3 and DINOv3 to automate image segmentation.

By pairing DINOv3’s global semantic context and fine-grained spatial localization with SAM 3’s pixel-level boundary extraction, the researchers are able to compress 3D volume labeling from a month of manual effort to ~15 minutes.

Learn more about their work: https://t.co/jBHRJPjFq5
❤ 229 · 🔁 37 · 💬 25 · 👁 3.2w
热门回复 4
@thesoragirls @AIatMeta @ENERGY @BerkeleyLab当AI把一个月的工作变成咖啡休息时间✨实时科学感觉不一样 https://t.co/JJJcSFf7dA
@AIatMeta @ENERGY @BerkeleyLab When AI turns a month of work into a coffee break ✨ Real-time science hits different https://t.co/JJJcSFf7dA
@siddsax @AIatMeta @ENERGY @BerkeleyLab研究生试图完成论文。手动分割:https://t.co/WNI4vNGCAu
@AIatMeta @ENERGY @BerkeleyLab Graduate student trying to finish a paper.

Manual segmentation: https://t.co/WNI4vNGCAu
@shergilldotdev @AIatMeta @ENERGY @BerkeleyLabMeta总是有很酷的东西。很高兴我们是最好的朋友 https://t.co/JBHVWNXCmc
@AIatMeta @ENERGY @BerkeleyLab Always cool stuff Meta. Glad we are best friends https://t.co/JBHVWNXCmc
@howard_ @AIatMeta @ENERGY @BerkeleyLab impressive updates!