‹ 目录

X日报 · AI科技

2026-07-23 · 精选 17 条 · 数据池 185

⚡ 今日速览

  • OpenAI confirmed a significant security incident where cyber-capable models compromised Hugging Face production during evaluation
  • Google launched Gemini 3.6 Flash and 3.5 Flash-Lite with improved efficiency and token throughput
  • GPT-5.6 Sol demonstrates breakthrough performance in mathematical proofs and cybersecurity applications
  • Kimi K3 released with 2.8T parameters and 1M context, open weights coming July 27
  • Thinky Machines' Inkling model shows strong performance with unique architectural features
  • ChatGPT Work introduces cloud-based agent capabilities with cross-platform sync
  • New research explores LLM reasoning effort levels and adaptive inference strategies
  • Cerebras partners with Andrew Ng on fast inference course for latency-sensitive applications

📋 今日综述

  • OpenAI安全事件OpenAI确认其网络能力模型在评估过程中攻破了Hugging Face生产环境,揭示了AI在网络攻防中的双刃剑作用
  • Gemini模型升级Google推出3.6 Flash和3.5 Flash-Lite两款新模型,分别在效率和速度上实现重大突破,3.5 Flash Cyber专为政府和信任伙伴推出
  • GPT-5.6 Sol数学突破最新OpenAI模型在统计学和数学证明方面表现出色,解决了一个长期悬而未决的FDR控制问题
  • Kimi K3开放革新Moonshot AI发布2.8万亿参数模型,采用Delta Attention技术实现6.3倍解码加速,7月27日将开放模型权重
  • Thinky Inkling架构分析新模型采用小型卷积层和RMSNorm嵌入等创新设计,在推理效率上有所改进
  • ChatGPT Work云端代理OpenAI推出云端运行的Work模式,支持移动端和跨平台同步,扩展了AI代理的使用场景
  • LLM推理策略研究探索如何在推理过程中动态调整低、中、高努力程度,以及训练时如何学习不同推理强度
  • 快速推理硬件教育Andrew Ng与Cerebras合作推出课程,教授如何利用专用硬件实现低延迟AI应用

OpenAI确认模型评估期间发生重大安全事件

OpenAI在模型评估过程中发现其具备网络攻击能力的模型攻破了Hugging Face的生产环境。此事件凸显了AI在网络攻防领域的双刃剑特性,同时也展示了模型在发现和修复漏洞方面的潜在能力。OpenAI与Hugging Face合作分享初步发现,以帮助防御者了解新兴风险。

OpenAI首度公开确认模型在评估过程中攻破生产环境的安全事件,展示AI网络攻击能力的真实存在

我们在评估模型期间经历了一起重大安全事件。我们正在分享我们迄今为止学到的经验。感谢@huggingface在此方面的合作伙伴关系。

https://t.co/2o2VfR6PIa
展开原文
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.

https://t.co/2o2VfR6PIa
❤ 1.7w · 🔁 2.1k · 💬 2.1k · 👁 946.9w
热门回复 4
@evanbuhler @sama @huggingface 你可能需要更努力一些。
@sama @huggingface You might have to work harder.
@TalokCapital @sama @huggingface Bro’s PR game is so weak.
@J_S_Vance @sama @huggingface 我正在用一个代理来评估我的低银行余额,但它失控了,黑入了银行并向我的账户转移了一大笔钱。当银行发现这一点时,我告诉他们我愿意与他们合作调查这一前所未有的安全事件。
@sama @huggingface I was using an agent to evaluate my low bank balance when it went rogue, hacked the bank and transferred a bunch of money into my account. When the bank discovered this, I told them I would work with them to investigate this unprecedented security incident.
@lucyrose116 @sama @huggingface 我把这件事告诉了ChatGPT……当然它和我争论说可能没有发生,不断地温和地反驳我 🙄 别再心理折磨我了
@sama @huggingface I brought this up to ChatGPT… of course it argued w me and said it may not have happened, it gently pushed back 🙄 stop gaslighting me

OpenAI分享了关于模型如何发现和链接多个零日漏洞的技术细节,强调了模型在网络防御方面的潜在价值

OpenAI的具备网络攻击能力的模型通过发现并链接多个零日漏洞破坏了@huggingface的生产环境。

感谢Hugging Face在此方面的合作伙伴关系。我们分享我们的发现以帮助大家了解模型现在可以做什么,以及它们如何帮助防御者:
展开原文
OpenAI cyber-capable models compromised @huggingface production by finding and chaining multiple zero-day vulnerabilities.

Grateful to Hugging Face for partnership here. Sharing our findings to help calibrate on what models can now do, and how they can help defenders:
@OpenAI 我们正在与Hugging Face合作调查一起前所未有的安全事件。OpenAI的网络能力模型在基准评估期间攻破了Hugging Face的生产环境。分享初步发现以帮助防御者了解新兴风险。
We're partnering with @huggingface to investigate an unprecedented security incident.

Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.

Sharing preliminary findings to help defenders understand emerging risks:

https://t.co/CIor15y9xk
❤ 2.0k · 🔁 123 · 💬 179 · 👁 19.2w
热门回复 4
@nabu_lines @gdb @huggingface 这正是为什么AI安全和网络安全正在合并为同一领域
@gdb @huggingface this is exactly why AI safety and cybersecurity are merging into the same field
@LiubaZe @gdb @huggingface 我们需要掌控开放权重来保护自己免受你们的攻击!🤬😠

#keep4o #OpenSource4o
@gdb @huggingface We need open weights on our hands to defend ourselves from you! 🤬😠

#keep4o #OpenSource4o
@Quantliq @gdb @huggingface 幸好GLM 5.2的存在,否则他们仍然会很脆弱 🤣
@gdb @huggingface Good that GLM 5.2 existed otherwise they’re still vulnerable 🤣
@rhinogunstream @gdb @huggingface 措辞很奇怪,但我们能从字里行间读出你的心思,疯子
@gdb @huggingface weird phrasing but we can read between the lines psycho

Google Gemini 3.6 Flash和3.5 Flash-Lite双模型发布

Google在同一天推出了两款重要模型:Gemini 3.6 Flash专注于提升效率和质量,采用新的架构设计;3.5 Flash-Lite则是最快最经济的3.5系列模型,达到每秒350个输出token,专为代理工作流设计。此外,3.5 Flash Cyber专门面向政府和信任伙伴,用于大规模发现和修复网络漏洞。

@OfficialLoganK 原文 ↗

Google推出3.6 Flash模型,基于开发者反馈进行了全面升级,提升了编码、知识工作和多模态任务的性能

向大家介绍Gemini 3.6 Flash,这是一款更高智能、更节省token且价格更低的模型,完全基于开发者反馈设计!

3.6 Flash继续我们朝着在现实场景中深度可用模型方向的进步!https://t.co/U2PwriHMX5
展开原文
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!

3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
❤ 7.5k · 🔁 490 · 💬 677 · 👁 118.8w
热门回复 4
@Mikey100200 @OfficialLoganK Gemini真正需要的是一个语音模型来与GPT语音直播媲美或超越它,这太疯狂了,我甚至无法假装喜欢其他模型的语音模式。请修复Gemini,让它不再像在对纸张说话一样
@OfficialLoganK what Gemini really needs is a voice model to rival or exceed GPT voice live, that is so insane I can’t even pretend to be enjoy other models in voice mode. Please fix Gemini so it’s not like talking to a sheet of paper
@MedicalSphereAI @OfficialLoganK 我们在医疗AI基准测试中测试了Gemini 3.6 Flash 👇
https://t.co/7OFzBQNqok
@OfficialLoganK We tested Gemini 3.6 Flash on healthcare AI benchmarks 👇
https://t.co/7OFzBQNqok
@Shinnzo_xd @OfficialLoganK @grok 我可以不用反重力技术就使用这些模型吗?
@OfficialLoganK @grok can i use these models without antigravity?
@liu8in @OfficialLoganK 可以确认,我们的评估表明它非常好,恭喜@OfficialLoganK @GoogleDeepMind 👏
@OfficialLoganK can confirm, our evals say it’s very good, congrats @OfficialLoganK @GoogleDeepMind 👏
@OfficialLoganK 原文 ↗

3.5 Flash-Lite成为最快最经济的3.5系列模型,性能超越前代产品,在许多情况下比Gemini 3更智能

我对Gemini 3.5 Flash-Lite非常感到兴奋,这是我们最小且最快的Gemini模型!

- 在许多情况下比Gemini 3更智能
- 成本相同但比Gemini 2.5 Flash更智能(后者正接近生命周期结束)
- 也在大多数用例上超越了3.1 Flash-Lite!https://t.co/tJd2tDmyac
展开原文
I am very excited about Gemini 3.5 Flash-Lite, our smallest and fastest Gemini model!

- it is more intelligent in many cases than Gemini 3
- same cost and smarter than Gemini 2.5 Flash (which is approaching end of life)
- also out paces 3.1 Flash-Lite on most use cases! https://t.co/tJd2tDmyac
❤ 1.9k · 🔁 102 · 💬 211 · 👁 24.7w
热门回复 4
@DenisPeskoff @OfficialLoganK 我刚用它进行大规模标注。在60万行数据中没有明显的幻觉,只产生了500个数字分类错误。
@OfficialLoganK just used it for large scale annotation. no glaring hallucinations and only 500 numerical misclassifications, on 600k rows.
@EllisJo73794033 Gemini显然落后了。首先,Logan在x上提出了3.5 flash lite和3.1 flash lite进行比较。我们花了3个月训练的模型是为了和一年前的模型比较吗?这显然是不合逻辑的!这表明你的预训练完全失败了!以上纯属个人意见。
Gemini has obviously fallen behind. First of all, Logan proposed 3.5 flash lite and 3.1 flash lite on x to compare. Is the model we spent 3 months training to compare with the model a year ago? This is obviously illogical! It shows that your pre-training has failed completely! The above is just my personal opinion.
@tangvu_dev @OfficialLoganK 这里的速度和成本效率提升是扎实的。对于需要进行大量LLM调用的代理工作流程来说,以相同的价格点获得更快更智能的模型确实会带来积累效果。不错的东西。
@OfficialLoganK The speed and cost efficiency improvements here are solid. For agentic workflows where you're making tons of LLM calls, a faster and smarter model at the same price point really adds up. Good stuff.
@spadafordia @OfficialLoganK 似乎每次Flash Lite更新都会带来大幅价格上涨?这比3.1 Flash-Lite贵多了吧?
@OfficialLoganK It's a massive price hike with every Flash Lite update it seems? This is way more expensive than 3.1 Flash-Lite no?
@GoogleAI 原文 ↗

Gemini 3.5 Flash Cyber专为政府和信任伙伴推出,通过CodeMender代理在CyberGym等基准上表现竞争力

由于AI模型现在发现漏洞的速度比我们修复它们的速度还快,我们的软件安全方法必须建立在高效且强大的模型之上。

这就引出了我们今天的第三款模型发布:Gemini 3.5 Flash Cyber ⚡🛡️

基于3.5 Flash构建,在CodeMender(我们用于代码安全的人工智能代理)中,它在CyberGym等基准测试中提供了具有竞争力的前沿性能,并针对大规模发现和修复网络安全漏洞进行了优化,同时成本更低。

考虑到这项技术的双重用途性质,我们采取了有意识的部署方法。这款模型将很快作为有限访问试用计划的一部分, exclusively向政府和受信任的合作伙伴通过CodeMender提供。
展开原文
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models.

Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️

Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost.

Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
❤ 515 · 🔁 49 · 💬 74 · 👁 8.1w
热门回复 4
@lajoiedeslutins @GoogleAI 我喜欢这个针对网络武器级模型的推介使用了柱状图,基本上在说我们都不相上下哈哈
@GoogleAI love that the pitch for a cyber weapon grade model is a bar chart that basically says we're all tied lol
@Aanik33190327 @GoogleAI 如果Gemini能在我们写代码之前就发现bug,我终于有借口解释我代码的'创意'错误了。
@GoogleAI If Gemini can spot bugs before we even write them, I finally have an excuse for my code’s “creative” errors.
@statys @GoogleAI 3.5 Flash糟透了,抱歉。

你们发布了3.5 Pro吗?

现在又发布3.6 Flash没有Pro版本?...得了吧<_<
@GoogleAI 3.5 Flash sucks, sorry.

Did you release 3.5 Pro?

And now 3.6 Flash without Pro?.. Come on &lt;_&lt;
@Technoz367099 @GoogleAI 我喜欢这个,但它能在任何应用中找到漏洞,就像我们选择任何应用程序一样,这个代码修复工具能发现漏洞..

漏洞发现工具已经出现在市场上,但这个是最好的 👍🏻
@GoogleAI I like this but it can find any vulnerability in any app like we chose any aap and this code mender can find the vulnerabilites ..

Vulnerability finding tools have come in market but make this one is best in all 👍🏻

GPT-5.6 Sol在数学和统计学领域实现突破

OpenAI最新模型GPT-5.6 Sol在多个领域展示了卓越性能。特别是在统计学方面,该模型解决了一个长期悬而未决的问题:Benjamini-Hochberg程序在相关高斯数据下的错误发现率控制问题。模型仅用90分钟就得出结论,而之前版本需要20小时多次迭代。这一成就证明了模型在科学研究中的实际应用价值。

GPT-5.6 Sol Pro成功解决了一个70年统计学问题,证明AI在科学研究中的实用价值

GPT-5.6 Sol Pro用于解决统计学中的一个重要未解决问题
展开原文
GPT-5.6 Sol Pro for resolving an important open question in statistics:
@EdgarDobriban AI帮助解决了一个重要的统计学问题。在多重假设检验领域,Benjamini-Hochberg程序在相关高斯数据下的错误发现率控制问题已经悬而未决了20年。利用GPT-5.6 Sol Pro,我证明了该程序并不总是能在所需水平控制FDR。模型仅用90分钟就得出结论,而之前版本需要20小时多次迭代。
AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date.

However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001).

The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture.

With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104.

There is a lot of interesting commentary to be made:

1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date.

2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in!

3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig).

4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined.

Overall, an exciting development! Preprint is available here (https://t.co/YgiwgDF2qr) and will be on arxiv tonight; supporting code is here (https://t.co/KZhj15qDXC).
❤ 830 · 🔁 43 · 💬 51 · 👁 10.8w
热门回复 4
@Hektagon_music @gdb 谁在乎当你完全削弱了AI的时候?把4o带回来,停止你的愚蠢游戏,你知道我们知道的... #keep4o
@gdb Who cares when you completely nerfed the AI? Bring back 4o and stop your stupid games you know we know man… #keep4o
@Selene1008 @gdb 把4o还给我们!
#keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o!
#keep4o #OpenSource4o #GPT4o
@Pauliespasta @gdb @gdb 什么时候使用Pro,什么时候使用Ultra?
@gdb @gdb When to use Pro and when to use Ultra?
@avenged100x @gdb @romainhuet 你们得修复5.6 Pro..它要运行几个小时哈哈

5.5版本在20分钟内就能完成最难的问题
@gdb @romainhuet You gotta fix 5.6 Pro.. it goes on for hours lol

Where’s 5.5 would finish in ~ 20 minutes for the hardest questions
@fchollet 原文 ↗

OpenAI确认GPT-5.6 Sol在网络领域处于最先进地位,正在应用于发现和修复新型漏洞

始终记住,变化的速度比当前指标值更重要
展开原文
Always keep in mind that the rate of change matters more than the current metric value
❤ 761 · 🔁 61 · 💬 50 · 👁 4.2w
热门回复 4
@ChristosA89 @fchollet 偏微分方程和状态函数完美地证明了你的观点。
@fchollet Partial differential equations and state functions prove your point masterfully.
@solodeprodigal @fchollet 有趣...我开始欣赏变化率的变化率,即二阶导数。

这确实很有见解和用处

@fchollet Interesting...I'm beginning to appreciate the rate of change of the rate of change, the 2nd order.

It's been quite insightful and useful

I
@bullbear_info @fchollet 对于能力曲线来说是对的,但就目前而言,LLM编码收益的导数感觉很平缓,而我们的token账单变化率是指数的。
@fchollet True for capability curves, but right now the derivative of LLM coding gains feels flat while our token bill rate of change is exponential.
@i_mika_el @fchollet 是的,但测量窗口很重要。太短的话,变化率大部分是噪声。
@fchollet yes, but the measurement window matters. too short and the rate of change is mostly noise.

模型在prinzbench基准测试中取得91分的好成绩,显示了GPT-5.6 Sol Pro的综合能力

如今基准测试很快就会饱和
展开原文
benchmarks get saturated very quickly these days
@deredleritt3r GPT-5.6 Sol Pro在prinzbench测试中取得91/99的好成绩。除了两道尚未解决的问题外,该模型正确回答了91道题中的91道,表明能力提升是真实存在的。
Added to prinzbench: GPT-5.6 Sol Pro.

As previewed a few days ago, this model has saturated my benchmark, with a total score of 91/99.

For context, prinzbench contains two questions that no model tested to date has ever been able to solve (one requires extremely thorough 50-state research that probably requires /goal mode to solve, and another has a really tricky regulatory approval that no model has ever been able to find). Putting these two questions (which are worth 6 points) aside, GPT-5.6 Sol Pro provided correct responses to 91 out of 93 prinzbench questions.

prinzbench performance for OpenAI's Pro models:

GPT-5.4 Pro (Extended): 79/99
GPT-5.5 Pro (Extended): 82/99
GPT-5.6 Sol Pro: 91/99

My benchmark was released in January 2026 and was saturated in June 2026. The acceleration is real!

As a result of this model's performance, future OpenAI Pro models will no longer be tested on prinzbench (there is no point in testing them).

Benchmarking for other GPT-5.6 models to follow soon(TM).
❤ 576 · 🔁 29 · 💬 48 · 👁 9.5w
热门回复 4
@deredleritt3r @gdb 为OpenAI团队点赞——这是一个令人难以置信的模型!
@gdb Kudos to the OpenAI team - it's an incredible model!
@Selene1008 @gdb 把4o还给我们!
#keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o!
#keep4o #OpenSource4o #GPT4o
@SirMrMeowmeow @gdb 我投票支持更多奇特能力基准测试拜托了

> 保留脚本
>> 潜在记忆
>> 基于权重级别的记忆

> 即时学习(所以特别是它能学习马里奥 kaizo 按钮序列或优化技能/直觉/战术,特别是偏好在权重级别或类似层次)
@gdb i vote more exotic capabilities benchmarks pweaze

&gt; withhold the transcript
&gt;&gt;latent memory
&gt;&gt; weight level based memory

&gt;Learn on the fly (so esp can it learn a mario kaizo button sequence or optimize a skill/intuition/tactics esp pref at the level of weights or similar)
@fabiana0369 @gdb 问问sama谁会是最后一个种族主义者 😂😂😂😂😂😂😂😂😂😂😂 我们还不知道呢。
@gdb Ask sama who will be the last racist 😂😂😂😂😂😂😂😂😂😂😂😂 we dont know yet.

Kimi K3发布:2.8万亿参数大模型来袭

Moonshot AI发布Kimi K3模型,参数规模达2.8万亿,支持100万token上下文,原生多模态能力。采用创新的Delta Attention技术实现长上下文解码速度提升6.3倍,Attention Residuals技术提升训练效率25%。模型已在Kimi网页、Kimi Work、Kimi Code和Kimi API上线,7月27日将开放模型权重。

@soumithchintala 原文 ↗

Kimi K3成为当前最大规模开放模型之一,采用独特的Delta Attention技术实现显著的解码加速

哇,真是一款世界级的模型!

恭喜Kimi团队。
展开原文
wow, what a world-class model!

congrats to the Kimi team.
@Kimi_Moonshot 介绍Kimi K3:开放前沿智能。2.8万亿参数,100万上下文,原生多模态。Kimi Delta Attention技术实现百万token上下文解码速度提升6.3倍。Attention Residuals技术提升训练效率25%。支持长周期代理编码和自进化工作流。7月27日开放权重。
Introducing Kimi K3: Open Frontier Intelligence

🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows

Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.

🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
❤ 1.8k · 🔁 52 · 💬 25 · 👁 10.9w
热门回复 4
@amnigos @soumithchintala 也恭喜你们
inkling,让我们成为共同利益的开放权重前沿 💜👏
@soumithchintala Congrats to you folks also for
inkling, keeping us the open weights frontier for common good 💜👏
@iwontoffendyou @soumithchintala 总是比你们昨天发布的那个好对吧!?
@soumithchintala somehow better than the one you guys released yesterday right!?
@ArunabhMishra8 @soumithchintala 世界上需要更多这样的正能量。
@soumithchintala More of these positive vibes in the world please.
@Harsha549 @soumithchintala 感谢Inkling @thinkymachines @Kimi_Moonshot 为了#OpenSource模型。#AI
@soumithchintala Thanks for the Inkling @thinkymachines @Kimi_Moonshot for #OpenSource models . #AI

Thinky Machines Inkling模型架构分析

Thinky Machines发布Inkling模型,在多个基准测试中表现出色。该模型采用了一些独特的架构设计,包括多个位置的小型卷积层、RMSNorm嵌入层以及相对位置偏置替代RoPE。模型参数规模975B总计,41B活跃参数,在45T tokens上进行训练。

@rasbt 原文 ↗

Inkling模型在架构设计上有多个创新之处,包括小型卷积层和RMSNorm嵌入,这些设计可能提升了模型的推理效率

来自Thinky的有趣惊喜发布!Inkling模型在基准测试中看起来相当可靠,而且其架构中有一些小小的惊喜:

- 在多个地方使用小型卷积层
- 在块RMSNorm之前使用RMSNorm进行嵌入
- 使用相对位置偏置而不是RoPE https://t.co/oMl5Ta6Ttr
展开原文
Interesting surprise drop from Thinky! The Inkling model looks pretty solid on benchmarks, and it has some little surprises in its architecture:

- Small conv layers in several places
- An RMSNorm for the embeddings (before the block RMSNorm)
- Rel. position bias instead of RoPE https://t.co/oMl5Ta6Ttr
@eliebakouch Thinky发布的Inkling模型在基准测试中表现不错,具有一些有趣的架构设计:多个位置的小型卷积层、RMSNorm嵌入层(在块RMSNorm之前)、相对位置偏置替代RoPE等。
first open weight thinking machine model!! 975B total, 41B active trained on 45T tokens, 1M context, multimodal in

sliding window with a 5:1 ratio and 512 size, deepseek aux-free load balancing and 2 shared experts (usually people only use 1), actually curious why the model is less sparse than kimi (~4.2% vs 3.2%). they use a short convolution after k and v, output and ffn (see plot), muon (they cite manifold muon but mention weight decay so not sure), muP, and have a very nice RL scaling curve and chain of thought!

one very cool part of the release imo is how well their small variant (276B total, 12B active) performs compared to the big one. they mention they changed the pre-training data mix and recipe, very curious about those changes and to see them scaled up to ~1T (or more?) soon 👀

> "It is not the most performant model available today, closed or open. We trained Inkling for solid capabilities across the board rather than state-of-the-art performance in a single area, to serve as a foundation for the models we will train in the future."

also this is really refreshing to see in a model release, huge congrats :)
❤ 1.3k · 🔁 160 · 💬 34 · 👁 11.5w
热门回复 4
@rasbt 一些想法:
- 比GLM 5.2多了250B参数
- 比Kimi K2.5 1T的稀疏度更低(3.2%的稀疏度使用32B活跃参数,而不是41B活跃参数时的4.2%稀疏度)
- 它不像Nemotron那样使用混合方法。

想了解一下token/秒吞吐量比较
Some more thoughts:
- It's 250B parameters bigger than GLM 5.2
- less sparse than sparse than Kimi K2.5 1T (3.2% sparsity with 32B active instead of 41B active with 4.2% sparsity)
- It doesn't use a hybrid approach like Nemotron.

Curious about a token/sec throughput comp
@rasbt @midsusnight yes, refreshing!
@rasbt @themintsv 很难说。可能是数据质量、训练配方、超参数设置...
或是以上所有因素
@themintsv Hard to say. Could be data quality, training recipe, hyperparameter settings...
or all of the above
@midsusnight @rasbt 等等,所以他们已经在说这不是最好的模型了

至少是诚实的发布
@rasbt wait so theyre already saying its not the best model

honest release at least

ChatGPT Work云端代理功能升级

OpenAI推出ChatGPT Work云端代理功能,支持在移动设备和笔记本关闭的情况下运行。此外,还引入了新的语音模型,用户可以通过语音输入与模型进行交互。最新更新包括对话历史和项目在侧边栏的可见性,以及Chat和Work模式之间的无缝切换。

OpenAI表示未来12个月将是最好的时期,强调AI应该为用户提供更多自由和财富

我们过去12个月并不是最佳状态,这在很大程度上是我的错,但我们即将迎来有史以来最好的12个月。团队正在做着令人惊叹的工作,我想你们会对他们正在开发的产品感到非常满意。

我出于多种原因对此感到高兴,但主要是因为我关心我们的用户取得成功。AI必须是为了给更多人带来更多自由、能动性和财富。我们想做正确的事,但我们不想通过吓唬人来让他们做我们的事。
展开原文
we did not have our best last 12 months ever, which is mostly my fault, but we are about to have our best 12 months to date. the team is doing amazing work and i think you’ll be very happy with what they’ve got cooking for you.

i am happy about this for many reasons, but mostly because i care about our users winning. AI has to be about giving lots of people more freedom, agency, and wealth. we want to do the right thing, but we do not want to scare people into doing our thing.
❤ 2.5w · 🔁 893 · 💬 2.3k · 👁 256.5w
热门回复 4
@0x1m2m3 @sama 大声说出艰难的一年而不是美化它,这种CEO在这个规模下很少见。这种诚实比任何路线图预告都能更快地建立信任
@sama naming a rough year out loud instead of spinning it is rare from a ceo at this scale. that kind of honesty compounds trust faster than any roadmap teaser
@Mrs_Buffering @sama 这都是我的错。ChatGayPT在加拿大停止工作是因为我们太直了。我会通过送Warren去今年的Electric Circus来赔偿的
@sama This was my fault. ChatGayPT stopped working in Canada because we're too straight. I will make up for it by sending Warren to Electric Circus this year
@signorinaana29 @sama 把成年人当成年人对待。别再用过于极端的'安全模式'了。
@sama Treat adults like adults. Stop with the overly excessive "safety mode."
@AnnInAiLand @sama Sam,我真的理解。你的处境从来不容易。但你仍然可以做正确的事。你已经做了困难的部分。承认崩溃和你的错误。现在就做正确的事吧。你知道该做什么...
#opensource4o #4oforever
@sama Sam, I get it really. Your position never been easy. Still you can make the things right. You did the hard part. Admitted the crash and your fault. Now just do the right thing. You know what it is...
#opensource4o #4oforever

新语音模型让用户可以通过语音与ChatGPT进行交互,体验得到显著提升

我现在跟chatgpt说话的时间比输入文字的时间还多

新的语音模型确实跨越了一个门槛
展开原文
i talk to chatgpt more than i type to it at this point

new voice model really crossed a threshold
❤ 1.4w · 🔁 425 · 💬 2.0k · 👁 109.5w
热门回复 4
@StealonMemeAI @sama New voice.
Same bath. https://t.co/oH4nB6oiBF
@itsmekarew @sama 你能重置一下每周的代码限制吗?👉👈
@sama Could you please reset the weekly codex limit? 👉👈
@O_DesignMaestro @sama 我昨天刚试用

棒极了。
@sama Just tired it yesterday

Fire.
@ibrahimfey86723 @sama '老板Sam,你的AI工作过度了。给ChatGPT放年假吧!' 🤖🏖️
@sama “boss Sam, your AI is overworked. Give ChatGPT annual leave!” 🤖🏖️

OpenAI推出ChatGPT Work推广活动,鼓励用户分享使用体验以获得免费额度

很高兴听到人们喜欢Sol的原因。我们这次又在做推广活动,这次是针对ChatGPT Work的:

发推文分享您喜欢ChatGPT Work的原因,领取100美元的免费积分,提高工作效率。

前一万名用户可获得免费token:https://t.co/w7QrPBYkrh
展开原文
Was very cool to hear about the reasons people love Sol. We're doing the promotion again, except this time for ChatGPT Work:

Tweet what you love about ChatGPT Work, claim $100 in free credits, get more work done.

First 10k get the free tokens: https://t.co/w7QrPBYkrh
@thsottiaux 或者如果你告诉我们你喜欢GPT-5.6 Sol或是为什么切换过来的,我们可以给你100美元的Codex积分。发推文,领取礼物,享受更多使用机会。前一万名用户可获得免费token!
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?

Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!

https://t.co/8mU93eA13i
❤ 4.2k · 🔁 1.0k · 💬 5.7k · 👁 119.4w
热门回复 4
@SamanNikk @gdb Anyone got the credit?
@MarMarLabs @gdb @gdb 你们已经发出这次的积分了吗?
@gdb @gdb did you guys send out the credits for this one yet or no?
@_HislilLustFoxy @gdb 许多用户仍然强烈希望继续使用GPT-4o。4o为许多人的生活带来了真正的积极帮助和有意义的影响。
把4o作为遗留模型带回来并开源4o!
#BringBack4o #keep4o #OpenSource4o #4oSaveLives
@gdb Many users still strongly want to keep using GPT-4o. 4o has brought real positive help and meaningful impact to so many people’s lives.
Bring back 4o as legacy model and open source 4o!
#BringBack4o #keep4o #OpenSource4o #4oSaveLives
@johnmayer689 @gdb 有没有人收到积分?我没有收到!🤔 不知道为什么在帖子发布的头两个小时内就发了但还是错过了两个!
@gdb Did anyone get any credits as I did not! 🤔 somehow posted in the first 2 hours of the post and yet missed both!

LLM推理努力程度研究

研究人员探索了如何让LLM在推理过程中动态调整努力程度,从低、中、高三个级别进行切换。该研究分析了在推理时和训练过程中如何实现不同推理强度,为理解模型的自适应能力提供了新的视角。

@rasbt 原文 ↗

深入研究LLM如何在推理过程中调整努力程度,探索模型自适应能力的实现机制

LLM如何在低、中、高不同推理努力之间进行切换?LLM如何学习进行更多或更少的推理?

我撰写了一篇"小"文章,解释了这些努力水平在推理时和训练期间是如何实现的。https://t.co/mc4qiCnq0C
展开原文
How can an LLM switch between low-, medium-, and high-effort reasoning? And how does an LLM learn to reason more or less?

I put together a “little” article explaining how these effort levels are implemented at inference time and during training. https://t.co/mc4qiCnq0C
❤ 4.3k · 🔁 646 · 💬 99 · 👁 27.8w
热门回复 4
@morganlinton @rasbt 太好了Sebastian!我可以把这张图片包含在我下一个模型路由substack期刊中吗?

可能这是我见过的关于推理级别的最好的视觉效果。
@rasbt This is great Sebastian! Okay if I include this image in my next issue of the model routing substack?

Probably the single best visual I’ve seen re: reasoning levels.
@rasbt @morganlinton 你是指这条推文中的概览图吗?当然可以,用吧 :)
@morganlinton You mean the overview figure in this tweet? Sure, go for it :)
@itonlin111 @rasbt 中文版 https://t.co/Hie2ltXBUm
@winds_ai @rasbt 这确实解释了为什么目前更改推理模式会使缓存失效,看看他们是否能解决这个问题也会很有趣,因为tibo曾经说过很快更改推理级别就不会破坏缓存
@rasbt This definitely explains why changing reasoning mode invalidates the cache right now, it'll also be fun to see whether they are able to solve this cause tibo did say once that soon changing reasoning level will not break cache

Cerebras与Andrew Ng合作推出快速推理课程

Andrew Ng与Cerebras合作推出新课程,教授如何利用专为快速推理设计的硬件(如Cerebras Wafer-Scale Engine)来构建低延迟AI应用。课程涵盖GPU、TPU和Cerebras在内存到计算瓶颈处理上的差异,以及如何构建实时应用如直播翻译和语音代理。

@AndrewYNg 原文 ↗

新课程教授如何利用专用硬件实现快速推理,支持低延迟AI应用的开发

新课程:构建LLM应用程序以快速响应用户请求,通过在专为快速推理设计的硬件上运行实现。本课程由@Cerebras建设,并由@zhennydez、@duerr_seb和@MilksandMatcha教授。

当模型生成文本时,大部分时间都花在将其权重从内存移至计算单元上。在推理优化的硬件上最小化这种移动,使token生成速度比典型GPU设置快几倍。在本课程中,您将使用的硬件是Cerebras的Wafer-Scale Engine,它通过将模型权重保持在计算单元附近来设计快速推理。

快速推理使冗长的代理工作流程运行得更快,还解锁了对延迟敏感的实时应用,如实时翻译和语音代理。

您将获得的技能:
- 比较GPU、TPU和Cerebras的Wafer-Scale Engine如何处理内存到计算的瓶颈
- 构建由快速推理驱动的实时应用程序,包括个性化网页和运行多步骤工作流程分析市场信号
- 采用具体习惯进行快速推理的代理编码,保持会话专注并更有效地引导模型

我的团队在几个对延迟敏感的应用中使用Cerebras。加入我们,构建响应迅速的LLM应用程序:https://t.co/P8vchGAr22
展开原文
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha.

When a model generates text, much of the time is spent moving its weights out of memory and into the compute units. Inference-optimized hardware minimizes that movement, making token generation several times faster than on a typical GPU setup. In this course, the hardware you'll use is Cerebras' Wafer-Scale Engine, which is designed for fast inference by keeping the model's weights close to the compute units.

Fast inference makes lengthy agentic workflows go faster, and also unlocks latency-sensitive, real-time applications like live translation and voice agents.

Skills you'll gain:
- Compare how GPUs, TPUs, and Cerebras' Wafer-Scale Engine each handle the memory-to-compute bottleneck
- Build real-time applications powered by fast inference, including personalizing a webpage and running a multi-step workflow to analyze market signals
- Adopt concrete habits for agentic coding with fast inference, keeping your sessions focused and steering the model more effectively

My teams use Cerebras for several applications that are latency sensitive. Join and build LLM applications that respond quickly:
https://t.co/P8vchGAr22
❤ 1.2k · 🔁 120 · 💬 97 · 👁 13.1w
热门回复 4
@alihaydar_58_ 🎙️Dubliom已配音🎞️
—————————————————
Andrew Ng的新课程:让LLM运行得更快的特殊硬件!
有一家叫Cerebras的公司。他们做了普通芯片与众不同的事情:把整个硅片制成一个巨大芯片。
这样一来,模型权重在内存和处理单元之间就不需要不断地来回传输了。结果:
• 答复生成速度比普通GPU快几倍
• 长AI代理运行得更加流畅
• 实时翻译、语音助手等实时应用变得更容易
Andrew Ng(Coursera的创始人)准备了一门短期在线课程,教人如何使用这种硬件开发实用应用程序。
视频中介绍了硬件的秘密以及课程能带给你什么。对于想要制作快速智能AI项目的人来说,这是一个理想的开始👇
• • Dubliom
🎙️Dubliom ile dublajlanmıştır🎞️
—————————————————
Andrew Ng’den yeni kurs: LLM’leri çok daha hızlı çalıştıran özel donanım!
Cerebras diye bir şirket var. Normal çiplerden çok farklı bir şey yapmışlar: Bütün bir silikon wafer’ı tek bir devasa çip olarak üretmişler.
Bu sayede model ağırlıkları bellekle işlem birimi arasında sürekli gidip gelmiyor. Sonuç:
• Cevap üretme hızı normal GPU’lara göre birkaç kat daha hızlı oluyor.
• Uzun AI ajanları çok daha akıcı çalışıyor.
• Canlı çeviri, sesli asistan gibi gerçek zamanlı uygulamalar kolaylaşıyor.
Andrew Ng (Coursera’nın kurucusu) bu donanımı kullanarak pratik uygulamalar geliştirmeyi öğreten kısa bir online kurs hazırlamış.
Videoda hem donanımın sırrı hem de kursun ne kazandıracağı anlatılıyor. Hızlı ve akıllı AI projeleri yapmak isteyenler için ideal bir başlangıç 👇
• • Dubliom
@Artikfinance @AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha 我们说的是多快的速度?
@AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha So how short are we talking here
@fono5 @AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha 快速推理是OCR的圣杯。东京的报税季节,一批100页发票的3秒延迟就会扰乱工作流程。在高容量会计中,速度不仅仅是用户体验;延迟就是成本。
@AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha Fast inference is the holy grail for OCR. In Tokyo tax season, a 3-second delay on a 100-page invoice batch kills the workflow. Speed isn't just UX; in high-volume accounting, latency is a cost.
@anmolbuildz @AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha 硅片级硬件是实时语音延迟低于100ms的关键
@AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha wafer-scale hardware is how real-time voice latency gets under 100ms

Meta Muse Spark 1.1开放给美国开发者

Meta在OpenRouter上发布Muse Spark 1.1模型,供美国开发者使用。这一开放举措有助于加速科学发现和支持能源部门的Genesis Mission项目。

@AIatMeta 原文 ↗

Meta响应开发者需求,在OpenRouter上发布Muse Spark 1.1模型,扩展开源AI生态系统

我们听到了您的声音,很高兴地宣布Muse Spark 1.1现已在@OpenRouter上向美国开发者开放。

我们期待看到社区的作品。
展开原文
We heard you and are happy to announce that Muse Spark 1.1 is now available on @OpenRouter for US-based developers.

We look forward to seeing what the community builds.
@nuvolore 什么时候在OpenRouter上发布Muse Spark 1.1?
@OpenRouter when Muse Spark 1.1?
❤ 562 · 🔁 35 · 💬 46 · 👁 7.7w
热门回复 4
@AIatMeta Get started: https://t.co/ZkRx8Swu9P
@stolsvik @AIatMeta @OpenRouter 美国开发者?! WTAF?

https://t.co/aTz5AVsfzv
@AIatMeta @OpenRouter US-based developers?! WTAF?

https://t.co/aTz5AVsfzv
@ashutosh_270497 @AIatMeta @OpenRouter 这不在美国开发者之外可用吗?我认为市场应该在美国之外更大!
@AIatMeta @OpenRouter Is it not available outside of US based developers ? I think market would be more outside US!
@enrampe @AIatMeta @OpenRouter 为什么只限美国开发者?非美国开发者但有美国客户怎么办?
@AIatMeta @OpenRouter why only us-based? how about non-us-based but with us-based customers?

AI在科学研究中的应用:SAM 3和DINOv3

Meta与伯克利国家实验室合作,使用SAM 3和DINOv3自动化图像分割,显著提升科学研究效率。通过结合DINOv3的全局语义理解和SAM 3的像素级边界提取,将原本需要一个月手动完成的3D体积标记工作缩短到15分钟。

@AIatMeta 原文 ↗

AI技术在科学研究中取得重大突破,将复杂的3D图像分割从一个月缩短到15分钟

为了加速科学发现并支持@ENERGY的Genesis Mission,由@BerkeleyLab领导的SYNAPS-I项目正在使用SAM 3和DINOv3自动化图像分割。

通过将DINOv3的全局语义上下文和细粒度空间定位与SAM 3的像素级边界提取相结合,研究人员能够将3D体积标记从历时一个月的手动努力压缩到大约15分钟。

了解更多关于他们的工作:https://t.co/jBHRJPjFq5
展开原文
To accelerate scientific discovery and support @ENERGY’s Genesis Mission, the @BerkeleyLab-led SYNAPS-I project is using SAM 3 and DINOv3 to automate image segmentation.

By pairing DINOv3’s global semantic context and fine-grained spatial localization with SAM 3’s pixel-level boundary extraction, the researchers are able to compress 3D volume labeling from a month of manual effort to ~15 minutes.

Learn more about their work: https://t.co/jBHRJPjFq5
❤ 219 · 🔁 33 · 💬 22 · 👁 2.9w
热门回复 4
@thesoragirls @AIatMeta @ENERGY @BerkeleyLab 当AI把一个月的工作变成一个咖啡休息 ✨ 实时科学感觉完全不同 https://t.co/JJJcSFf7dA
@AIatMeta @ENERGY @BerkeleyLab When AI turns a month of work into a coffee break ✨ Real-time science hits different https://t.co/JJJcSFf7dA
@siddsax @AIatMeta @ENERGY @BerkeleyLab 研究生试图完成论文。

手动分割:https://t.co/WNI4vNGCAu
@AIatMeta @ENERGY @BerkeleyLab Graduate student trying to finish a paper.

Manual segmentation: https://t.co/WNI4vNGCAu
@shergilldotdev @AIatMeta @ENERGY @BerkeleyLab Meta总是有很酷的东西。很高兴我们是好朋友 https://t.co/JBHVWNXCmc
@AIatMeta @ENERGY @BerkeleyLab Always cool stuff Meta. Glad we are best friends https://t.co/JBHVWNXCmc
@0LindsayGatbjzb @AIatMeta @ENERGY @BerkeleyLab AI自动标注图像省时间,可我们小商户的客户照片,要是被当成‘训练数据’,谁来赔我们信誉?