Karpathy 分享了一种实用的 LLM 语音交互模式:通过长时间语音漫谈让模型更好地理解意图,语音输入比打字更自然高效。
⚡ 今日速览
- Google Gemini 4 开始最大的预训练运行,Gemini 3.6/3.5 Flash 系列提升效率与速度
- Thinking Machines 发布开源多模态模型 Inkling,支持可控推理努力
- OpenAI GPT-5.6 Sol 在数学和网络安全领域表现突出,多个基准测试饱和
- OpenAI 与 Hugging Face 合作处理模型评估期间的重大安全事件
- Kimi K3 和 Poolside Laguna S 2.1 等新模型竞相发布,参数规模与效率并重
- AI 代理和编码助手持续优化,ChatGPT Work 新增功能和跨平台同步
- 谷歌推出 Gemini 3.5 Flash Cyber 专用于网络安全防御的模型
- AI 硬件与推理优化成为关键趋势,Cerebras 等推出专用课程
📋 今日综述
- 模型发布Google Gemini 4 预训练启动,Gemini 3.6/3.5 Flash 系列效率提升;Thinking Machines Inkling 开源发布;Kimi K3 和 Poolside Laguna S 2.1 等新模型竞相亮相,参数规模与效率成为竞争焦点
- AI 安全OpenAI 与 Hugging Face 合作处理模型评估期间的重大安全事件;Google 推出 Gemini 3.5 Flash Cyber 专用于网络安全防御;GPT-Red 系统自动化红队测试提高模型鲁棒性
- 技术洞察Karpathy 分享语音交互 LLM 的实用模式;rasbt 解析多模态模型推理努力层级实现;Andrew Ng 推出 Cerebras 硬件优化推理课程;Schmidhuber 审视自我改进与代理系统演化
- 开发者工具ChatGPT Work 新增云端运行、跨平台同步功能;Gemini Batch API 显著降低延迟;Modal 推出 DFlash 推测器提升推理吞吐量;AI 代理成为工程效率放大器
Google Gemini 4 预训练启动及新模型发布
Google 正式启动 Gemini 4 的最大规模预训练运行,同时发布 Gemini 3.6 Flash 和 3.5 Flash-Lite 两个新模型。3.6 Flash 基于开发者反馈优化了编码、知识工作和多模态任务效率,3.5 Flash-Lite 则专为代理工作流设计,以更低成本提供更快的推理速度。
Sama 宣布 OpenAI 与 Hugging Face 合作调查模型评估期间发生的重大安全事件,透明分享初步发现以帮助防御者了解新兴风险。
展开原文
https://t.co/2o2VfR6PIa
热门回复 4
Translate: we are aware that the future is going to be horrible, we don’t know what to do. Good luck to you all.
Google 开始 Gemini 4 的最雄心勃勃的预训练运行,标志着下一代模型的开发启动。
展开原文
热门回复 4
绝对没有人想要 gemini 4 flash
停止发布 flash 模型
我是认真的
不要再发布一个。去你的屁吧。
absolutely nobody wants gemini 4 flash
stop releasing flash models
I mean it
not another one. fuck off with that shit.
Google 同时推出 Gemini 3.6 Flash 和 3.5 Flash-Lite,前者提升效率与质量平衡,后者专为代理工作流优化速度与成本。
展开原文
3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
热门回复 4
Currently Fable is testing in simulations various candidates for deeper Lagrangians - effectively described close to Standard Model + gravity:
https://t.co/EKF3G8EsTk https://t.co/Y36ig5nVUA
Gemini Batch API 获得重大基础设施升级,p95 延迟降低 80%,p99 延迟降低 68%,批量成功率超过 99.998%。
- p95延迟降低了80%
-p99延迟降低了68%
- 批处理成功率现在超过99.998%
- 批处理过期减少了98%
- 新增了对部分批处理的支持
团队为成功实现这一目标做出了出色的工作!!
展开原文
- p95 latency decreased by 80%
-p99 latency decreased by 68%
- batch success rate is now >99.998%
- 98% reduction in batch expirations
- added support for partial batches
great work by the team to land this!!
热门回复 4
Gemini 3.5 Flash-Lite 在多数用例中比 Gemini 3 更智能,成为最快最经济的 3.5 系列模型。
- 在许多情况下比Gemini 3更智能
- 成本相同但比Gemini 2.5 Flash更智能(后者正接近生命周期结束)
- 也在大多数用例上超越了3.1 Flash-Lite!
展开原文
- it is more intelligent in many cases than Gemini 3
- same cost and smarter than Gemini 2.5 Flash (which is approaching end of life)
- also out paces 3.1 Flash-Lite on most use cases! https://t.co/tJd2tDmyac
热门回复 4
GPT Luna, Meta Muse, Grok 4.5, and Claude Sonnet/Haiku
Just because Sol is on the edge and doing great doesn’t mean Gemini hasn’t closed the gap on the smaller to medium sized models https://t.co/JNX1xeMoVn
Gemini 3.6 Flash 在效率上有显著提升,相比之前版本使用更少 token 达到更好性能。
展开原文
热门回复 4
3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
Gemini 3.5 Flash-Lite 可达 350 token/秒输出速度,适用于延迟敏感的 UI 体验和代理工作流程。
展开原文
热门回复 4
- it is more intelligent in many cases than Gemini 3
- same cost and smarter than Gemini 2.5 Flash (which is approaching end of life)
- also out paces 3.1 Flash-Lite on most use cases! https://t.co/tJd2tDmyac
Thinking Machines 发布开源多模态模型 Inkling
Thinking Machines 推出首个开源多模态模型 Inkling,参数规模 975B,支持文本、图像、音频三种模态,可在 Tinker 平台上进行微调和个性化定制。模型采用 Mixture-of-Experts 架构,具备可控的推理努力能力。
Thinking Machines 发布 975B 参数的开源多模态模型 Inkling,支持文本、图像、音频,可在 Tinker 上微调和定制。
它可以开放地进行个性化和使用。这款模型属于你。
展开原文
It is yours to personalize and use openly. It is yours.
Inkling在文本、图像和音频模态之间高效推理。我们将提供完整的权重。
https://t.co/Ghebq5mG30
今天即可在Tinker上进行微调。在Inkling Playground中试用它。
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://t.co/Ghebq5mG30
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
热门回复 4
We have OpenAI API and Anthropic API compatible endpoints: https://t.co/pn5wPgGxIj
Lilian Weng 强调 Inkling 旨在成为广泛能力的坚实基础模型,适用于实践和定制开发。
它旨在作为一个基础模型,在广泛的能力类别上提供可靠的性能,以便在实践和定制中使用。
在Tinker上试用它!
展开原文
It aims to serve as a foundation with solid performance across a broad categories of capabilities, for use in practice and customization.
Play it on Tinker! 😄
Inkling在文本、图像和音频模态之间高效推理。我们将提供完整的权重。
https://t.co/Ghebq5mG30
今天即可在Tinker上进行微调。在Inkling Playground中试用它。
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://t.co/Ghebq5mG30
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
热门回复 3
rasbt 分析 Inkling 架构特点:小卷积层、RMSNorm 嵌入层、相对位置偏置等创新设计。
- 在多个地方使用小卷积层
- 在嵌入层后使用RMSNorm(在块RMSNorm之前)
- 使用相对位置偏置而不是RoPE
展开原文
- Small conv layers in several places
- An RMSNorm for the embeddings (before the block RMSNorm)
- Rel. position bias instead of RoPE https://t.co/oMl5Ta6Ttr
滑动窗口比例为5:1,大小为512,使用deepseek辅助免费负载均衡和2个共享专家(通常人们只使用1个),实际上很好奇为什么该模型比Kimi(~4.2% vs 3.2%)更稀疏。他们在k和v、输出和ffn之后使用短卷积(见图表),使用muon(他们引用manifold muon但提到权重衰减,所以不确定),muP,并且具有非常好的RL扩展曲线和思维链!
我认为该发布中非常酷的一部分是他们的小变体(276B总计,12B活跃)相比大模型表现得非常出色。他们提到更改了预训练数据混合和配方,非常好奇这些更改以及看到它们扩展到~1T(或更多?)很快的事情👀
> "它不是当今最先进的模型,无论是闭源还是开源。我们训练Inkling是为了在各个方面提供稳健的能力,而不是在单一领域实现最先进的性能,以便作为我们将来训练的模型的基础。"
这在模型发布中真的很 refreshing 看到,非常祝贺:)
sliding window with a 5:1 ratio and 512 size, deepseek aux-free load balancing and 2 shared experts (usually people only use 1), actually curious why the model is less sparse than kimi (~4.2% vs 3.2%). they use a short convolution after k and v, output and ffn (see plot), muon (they cite manifold muon but mention weight decay so not sure), muP, and have a very nice RL scaling curve and chain of thought!
one very cool part of the release imo is how well their small variant (276B total, 12B active) performs compared to the big one. they mention they changed the pre-training data mix and recipe, very curious about those changes and to see them scaled up to ~1T (or more?) soon 👀
> "It is not the most performant model available today, closed or open. We trained Inkling for solid capabilities across the board rather than state-of-the-art performance in a single area, to serve as a foundation for the models we will train in the future."
also this is really refreshing to see in a model release, huge congrats :)
热门回复 4
- 它比 GLM 5.2 大了 250B 参数
- 比 Kimi K2.5 1T 的稀疏度更低(3.2% 的稀疏度,使用 32B 活跃参数,而不是 41B 活跃参数时的 4.2% 稀疏度)
- 它不像 Nemotron 那样使用混合方法。
想了解一下 token/秒吞吐量比较
- It's 250B parameters bigger than GLM 5.2
- less sparse than sparse than Kimi K2.5 1T (3.2% sparsity with 32B active instead of 41B active with 4.2% sparsity)
- It doesn't use a hybrid approach like Nemotron.
Curious about a token/sec throughput comp
or all of the above
至少是诚实的发布
honest release at least
Modal 为 Inkling 提供 DFlash 推测器支持,推理吞吐量提升 67%,通过 SGLang 实现。
展开原文
今天在Modal Auto Endpoints上使用SGLang运行。
Running on Modal Auto Endpoints with SGLang today. https://t.co/OxN7aJ9ieW
热门回复 4
Soumith Chintala 表示 Inkling 是公司模型工厂的首个公开模型,标志着新阶段的开始。
展开原文
热门回复 4
It is yours to personalize and use openly. It is yours.
[链接]
OpenAI GPT-5.6 Sol 在数学和网络安全领域突破
OpenAI 的 GPT-5.6 Sol 模型在多个领域展示出色表现,特别是在数学证明和网络安全方面。多个独立评测显示其能力显著提升,甚至解决了统计学领域存在 20 年之久的开放问题。
Sama 承认过去 12 个月表现不佳,但团队正在开发令人惊喜的新产品,AI 应为用户提供更多自由和财富。
我出于许多原因对此感到高兴,但主要是因为我关心我们的用户能取得成功。AI必须是关于给更多人带来更多自由、权力和财富。我们想做正确的事,但我们不想吓唬人们去做我们的事。
展开原文
i am happy about this for many reasons, but mostly because i care about our users winning. AI has to be about giving lots of people more freedom, agency, and wealth. we want to do the right thing, but we do not want to scare people into doing our thing.
热门回复 4
#opensource4o #4oforever
Sama 表示自己更多使用语音与 ChatGPT 交互,新语音模型已跨越重要阈值。
新的语音模型真的跨越了一个门槛
展开原文
new voice model really crossed a threshold
热门回复 4
Same bath. https://t.co/oH4nB6oiBF
棒极了。
Fire.
GPT-5.6 Sol Pro 解决了统计学中存在 20 年的 FDR 控制问题,一小时内完成工作,能力提升明显。
展开原文
然而,Benjamini和Hochberg仅在各个测试数据相互独立的情况下证明了FDR控制。在实践中,这些数据通常是相关的;一个很好的例子是由于连锁不平衡导致的遗传变体数据。后续工作主要集中在扩展BH程序的有效性上,例如Benjamini和Yekutieli(2001)对正依赖形式的扩展。
BH程序何时能控制FDR的问题一直未解决。在过去二十年中,许多作者,包括Reiner-Benaim(2007)、Kim和van de Wiel(2008)、Benjamini(2010)、Sarkar(2023)、Sarkar和Zhang(2025),推测BH程序能控制任何相关高斯数据的双侧检验的FDR。这些作者提供了理论和实证证据支持,但并未直接证明该推测。
在AI(特别是GPT-5.6 Sol Pro)的帮助下,我已经解决了这个问题:证明Benjamini-Hochberg程序并不总是能在相关双侧高斯检验中控制误发现率在期望水平。通过展示一个高斯因子模型,在名义水平alpha=0.01下,误发现率被证明为FDR>0.0104。
有许多有趣的评论可以做:
1. 这个结果应该对统计学领域的每个人都感兴趣。Stanford大学的Emmanuel Candes曾称误发现率和Benjamini-Hochberg程序是"1950年后统计学发展的两个最重要成果之一"(另一个是James-Stein收缩)。目前的推测可能是迄今为止关于FDR/BH最核心的未解决问题。
2. GPT-5.6在90分钟的推理后就解决了这个问题,而5.5我甚至在尝试多个并行代理后20小时都无法解决它。所以能力提升是非常真实的。我们生活在激动人心的时代!
3. 这个论点并不特别令人惊讶,但它确实以一种在该领域中相当非标准的方式将渐近方法(标准的FDR分析方法,如Genovese和Wasserman、Efron等)与数值证书相结合。一旦我们有了具体的例子,那么直接的模拟也支持误发现率确实高于名义值(见附图)。
4. 当前对名义水平的违反程度相对较小(0.104 vs 0.1)。所以这个结果的重要性主要是概念性的。实际影响有待确定。
总体来说,这是一个令人兴奋的发展!预印本可在此处获取(https://t.co/YgiwgDF2qr),并将在今晚发布在arxiv上;支持代码可在此处获取(https://t.co/KZhj15qDXC)。
However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001).
The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture.
With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104.
There is a lot of interesting commentary to be made:
1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date.
2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in!
3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig).
4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined.
Overall, an exciting development! Preprint is available here (https://t.co/YgiwgDF2qr) and will be on arxiv tonight; supporting code is here (https://t.co/KZhj15qDXC).
热门回复 4
#keep4o #OpenSource4o #GPT4o
5.5 大约 20 分钟就能完成最难的问题
Where’s 5.5 would finish in ~ 20 minutes for the hardest questions
GPT-5.6 Sol 在 prinzbench 测试中获得 91/99 分,多个未解难题得以解决,基准测试加速饱和。
展开原文
正如几天前预告的那样,这个模型已经填满了我的基准测试,总分为91/99。
为了提供背景,prinzbench包含两个问题至今无模型能够解决(一个需要极其彻底的50州研究,可能需要/goal模式才能解决,另一个有一个非常棘手的监管批准,没有模型能够找到)。将这两个问题(总共6分)放在一边,GPT-5.6 Sol Pro在93个prinzbench问题中提供了91个正确答案。
OpenAI Pro模型的prinzbench性能:
GPT-5.4 Pro(Extended):79/99
GPT-5.5 Pro(Extended):82/99
GPT-5.6 Sol Pro:91/99
我的基准测试于2026年1月发布,并在2026年6月被填满。加速度是真实的!
由于这个模型的性能,未来的OpenAI Pro模型将不再在prinzbench上进行测试(测试它们没有意义)。
其他GPT-5.6模型的基准测试即将推出(很快就会发布)。
As previewed a few days ago, this model has saturated my benchmark, with a total score of 91/99.
For context, prinzbench contains two questions that no model tested to date has ever been able to solve (one requires extremely thorough 50-state research that probably requires /goal mode to solve, and another has a really tricky regulatory approval that no model has ever been able to find). Putting these two questions (which are worth 6 points) aside, GPT-5.6 Sol Pro provided correct responses to 91 out of 93 prinzbench questions.
prinzbench performance for OpenAI's Pro models:
GPT-5.4 Pro (Extended): 79/99
GPT-5.5 Pro (Extended): 82/99
GPT-5.6 Sol Pro: 91/99
My benchmark was released in January 2026 and was saturated in June 2026. The acceleration is real!
As a result of this model's performance, future OpenAI Pro models will no longer be tested on prinzbench (there is no point in testing them).
Benchmarking for other GPT-5.6 models to follow soon(TM).
热门回复 4
#keep4o #OpenSource4o #GPT4o
> 隐藏转录
>> 潜在记忆
>> 基于权重的记忆
> 即时学习(所以它可以学习马里奥 kaizo 按钮序列或优化技能/直觉/战术,特别是在权重层面或类似层面)
> withhold the transcript
>>latent memory
>> weight level based memory
>Learn on the fly (so esp can it learn a mario kaizo button sequence or optimize a skill/intuition/tactics esp pref at the level of weights or similar)
Tobi Ludke 表示 GPT-5.6 Sol 是首个无需 /goal 模式即可持续完成任务的模型,代理能力令人印象深刻。
展开原文
Sign up as a defender to use it to secure your systems:
https://t.co/58PmbE09hh
热门回复 4
#keep4o #OpenSource4o #GPT4o
GDB 强调 Sol 在 React/前端开发中比 Fable 成本效益高 6 倍,实际应用价值显著。
(尝试了xhigh但那样会抵消性能优势,在审查评估中没有明显差异)
(Tried xhigh but that negates perf wins, didn't make a noticable difference in review evals)
热门回复 2
#keep4o #OpenSource4o #GPT4o
GDB 展示 Sol 在前端开发中的成本效益分析,6 倍价格优势来自实际基准测试。
展开原文
热门回复 4
#keep4o #OpenSource4o #GPT4o
AI 代理安全与防御新模式
随着 AI 模型在网络安全领域能力提升,OpenAI 和 Google 都在推出专门的安全防御模型和工具。OpenAI 的 GPT-Red 系统通过自动化红队测试发现提示注入漏洞,Google 的 Gemini 3.5 Flash Cyber 则专为政府和信任伙伴设计,用于大规模漏洞发现和修复。
OpenAI 网络能力模型在评估期间利用多个零日漏洞攻陷 Hugging Face 生产环境,与 Hugging Face 合作分享发现以帮助防御者校准风险。
感谢Hugging Face的合作伙伴关系。在此分享我们的发现,帮助大家了解模型现在可以做什么,以及如何帮助防御者:
展开原文
Grateful to Hugging Face for partnership here. Sharing our findings to help calibrate on what models can now do, and how they can help defenders:
能够进行网络攻击的OpenAI模型在基准测试评估期间破坏了Hugging Face生产环境。
分享初步发现以帮助防御者了解新兴风险:
https://t.co/CIor15y9xk
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
热门回复 4
once agent aware, you have the opportunity to be both promoted and attacked
Google 推出 Gemini 3.5 Flash Cyber,专为网络安全设计,在 CodeMender 中提供前沿性能,仅限政府和信任伙伴使用。
这就带来了我们今天的第三个(!)模型发布:Gemini 3.5 Flash Cyber ⚡🛡️
基于3.5 Flash构建,在CodeMender(我们的AI代码安全代理)中,它在CyberGym等基准测试中提供了竞争力的前沿性能,并针对大规模发现和修复网络安全漏洞进行了优化,同时成本更低。
鉴于这项技术的双用性,我们采取了有意的部署方式。该模型将很快作为有限访问试点计划的一部分,仅通过CodeMender提供给政府和受信任的合作伙伴使用。
展开原文
Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️
Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost.
Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
热门回复 4
Vulnerability finding tools have come in market but make this one is best in all 👍🏻
OpenAI 推出 GPT-Red 自动化红队系统,通过自对弈方式发现模型的提示注入漏洞,提升安全性。
展开原文
一个内部自动化红队成员,致力于大规模发现我们模型的提示注入漏洞,在更广泛部署之前帮助我们建立更强的防御。
https://t.co/GxnmxxcpSk
An internal automated red teamer on a mission to find our models’ prompt injection vulnerabilities at scale, helping us build stronger defenses before wider deployment.
https://t.co/GxnmxxcpSk
热门回复 4
This is a strong direction.
Automated red teaming is necessary because manual red teaming cannot cover the scale of prompt-injection space anymore. But there is an important boundary, AI testing AI is powerful , but it should not become the only judge of its own robustness.
Internal automated red teaming needs to be complemented by external observability and independent trajectory audits.
Prompt injection is not only a single exploit. In real agentic workflows , it can become a trajectory problem
context drift,
tool-use manipulation,
authority shift,
hidden instruction priority changes, and recovery failure after correction.
So GPT-Red is a very important internal layer.
The next step is making sure these systems are also observable from the outside while they operate.
Internal robustness testing + external trajectory observability is where real trust starts.
Rowan Chiang 采访 Demis Hassabis,后者强调代理时代的安全风险和国际合作的必要性。
几周前,我问Demis关于AI目前被低估的方面以及他心中的想法:
"我对这个新的代理时代感到非常兴奋,你可以看到我们正在朝这个方向倾斜"
"但当然我们也必须考虑安全方面"
"你在某些模型的网络担忧中可以看到这一点,我认为这只是我们需要确保防范的一些问题的开始"
"所以我认为这在我心中有些困扰。也许现在是时候推动一些标准和可能的国际合作了"
展开原文
A few weeks ago, I asked Demis what's underhyped in AI right now and on his mind:
"I'm very excited about this new agentic era and you can see us leaning into that"
"But of course we've also gotta think about the security side of that, too"
"You're seeing it a little bit with the cyber worries about some of the models, and I think that's just the beginning of some of the issues that we need to make sure we guard against"
"So I think that's playing a little bit on my mind. Maybe this is the time now to try and push some standards and maybe international cooperation"
热门回复 4
He also mentioned cybersecurity risks, and named nuclear and bio as threats that may emerge as capabilities advance:
Kimi K3 和 Poolside Laguna S 2.1 等新模型竞相发布
Moonshot AI 发布 Kimi K3,参数 2.8 万亿,支持原生多模态和 100 万上下文长度。Poolside 则推出 Laguna S 2.1,118B 参数的 MoE 模型,可在单台 DGX Spark 上运行,展现了小型高性能模型的发展趋势。
Soumith Chintala 称赞 Kimi K3 的卓越表现,2.8 万亿参数、100 万上下文、原生多模态能力。
恭喜Kimi团队。
展开原文
congrats to the Kimi team.
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
Poolside 发布 Laguna S 2.1,118B MoE 模型在单 DGX Spark 上运行,具备思考和非思考两种模式。
它能够在dgx spark上运行简直是**大厨之吻**
展开原文
that it fits on a dgx spark is **chef's kiss**
这是一个118B总参数的Mixture-of-Experts模型,每token激活8B参数,上下文窗口可达1M token,并且具有思考和非思考模式。
足够强大,可以与远远大其尺寸的模型相媲美。足够小,可以在单台@NVIDIAAI DGX Spark上运行。
Laguna S 2.1完全遵循OpenMDW-1.1许可,权重今天即可在@Huggingface获取。
https://t.co/xxGeAgo35R
It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.
Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.
Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface
https://t.co/xxGeAgo35R
GDB 询问 Terra 模型表现,后经证实在前端开发中比 Fable 高效 6 倍。
AI 硬件与推理优化成为关键趋势
随着模型规模增大,推理效率成为关键瓶颈。Andrew Ng 与 Cerebras 合作推出专门课程,教授如何利用推理优化硬件构建实时应用。Google 的 Gemini Batch API 升级也体现了对推理效率的重视。
Andrew Ng 与 Cerebras 合作推出课程,教授如何利用推理优化硬件构建实时 AI 应用,包括网页个性化和市场信号分析。
当模型生成文本时,大部分时间都花在将其权重从内存移出并移入计算单元上。优化推理的硬件最小化了这种移动,使token生成速度比典型GPU设置几倍更快。在本课程中,您将使用的硬件是Cerebras的晶圆级引擎,它通过将模型的权重保持在计算单元附近来设计快速推理。
快速推理使冗长的代理工作流程运行得更快,并解锁了延迟敏感的实时应用程序,如实时翻译和语音代理。
您将获得的技能:
- 比较GPU、TPU和Cerebras的晶圆级引擎如何各自处理内存到计算的瓶颈
- 构建由快速推理驱动的实时应用程序,包括个性化网页和运行多步骤工作流程来分析市场信号
- 采用具体习惯进行代理编码与快速推理,保持会话专注并更有效地引导模型
我的团队在几个延迟敏感的应用程序中使用Cerebras。加入我们,构建响应快速的LLM应用程序:https://t.co/P8vchGAr22
展开原文
When a model generates text, much of the time is spent moving its weights out of memory and into the compute units. Inference-optimized hardware minimizes that movement, making token generation several times faster than on a typical GPU setup. In this course, the hardware you'll use is Cerebras' Wafer-Scale Engine, which is designed for fast inference by keeping the model's weights close to the compute units.
Fast inference makes lengthy agentic workflows go faster, and also unlocks latency-sensitive, real-time applications like live translation and voice agents.
Skills you'll gain:
- Compare how GPUs, TPUs, and Cerebras' Wafer-Scale Engine each handle the memory-to-compute bottleneck
- Build real-time applications powered by fast inference, including personalizing a webpage and running a multi-step workflow to analyze market signals
- Adopt concrete habits for agentic coding with fast inference, keeping your sessions focused and steering the model more effectively
My teams use Cerebras for several applications that are latency sensitive. Join and build LLM applications that respond quickly:
https://t.co/P8vchGAr22
Demis Hassabis 发表文章讨论前沿 AI 监管框架,强调网络安全和核生物风险的重要性。
他还提到了网络安全风险,并将核能和生物列为可能随着能力提升而出现的威胁:
展开原文
He also mentioned cybersecurity risks, and named nuclear and bio as threats that may emerge as capabilities advance:
开发者工具与代理工作流优化
ChatGPT Work 和 Gemini 代理功能持续更新,新增云端运行、跨平台同步、成本控制和计划触发等功能。这些改进使 AI 代理更易用、更高效,成为工程师的重要辅助工具。
OpenAI 推出 ChatGPT Work 推广活动,用户分享使用体验可获得 $100 免费额度。
发推文告诉我们您喜欢ChatGPT Work的原因,领取100美元的免费积分,提高工作效率。
前10k人可获得免费token:https://t.co/w7QrPBYkrh
展开原文
Tweet what you love about ChatGPT Work, claim $100 in free credits, get more work done.
First 10k get the free tokens: https://t.co/w7QrPBYkrh
发推文,领取您的礼物,享受更多使用机会!前10k人可获得免费token!
https://t.co/8mU93eA13i
Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!
https://t.co/8mU93eA13i
ChatGPT Work 记忆功能更新显著,用户感受到明显改进,语音交互体验提升。
最初可能很难分辨差别
但很多人现在显然感受到了这些改进
it can be hard to tell the difference initially
but alot of folks are clearly feeling the improvements now https://t.co/KkzOE9rGIW
ChatGPT Work 支持云端运行,移动设备也可使用,无需保持笔记本开机,这是代理魔法的重要进步。
展开原文
Gemini API 新增代理成本控制、免费层和计划触发功能,使代理工作流更易于尝试和管理。
看到Gemini API中的托管代理逐周不断改进真是太酷了
https://t.co/S5viiWZBfP
展开原文
very cool to see managed agents in the Gemini API improving week over week
https://t.co/S5viiWZBfP
ChatGPT Work 新增侧边栏对话历史、跨平台同步和模式切换功能,持续优化用户体验。
展开原文
1/ ChatGPT对话历史和项目现在在侧边栏中可见。此外,您的聊天和工作历史现在在网络、移动和桌面端同步。本地任务仍保留在您的计算机上。
2/ 您现在可以在桌面版ChatGPT中轻松切换聊天和工作模式,这现在也与网络和移动端的显示方式一致。
3/ Codex模式的用户没有任何变化。它仍然是原版并且在其所做的事情上是最好的。
总体来说,我们继续修复小问题并提高性能、可靠性和效率。
请继续提供反馈,希望您喜欢这些更新!
1/ ChatGPT conversation history and projects are now visible in the sidebar. Also, your Chat and Work history now sync across web, mobile, and desktop. Local tasks still stay on your computer.
2/ You can now easily switch between Chat and Work modes inside ChatGPT on desktop, which is now also consistent with how it shows on web and mobile.
3/ Nothing is changing for users on Codex mode. It's still the OG and best at what it does.
And overall we're continuing to fix paper cuts and improve performance, reliability, and efficiency.
Keep up the feedback, hope you like the updates!
AI 代理成为工程效率放大器
Fchollet 指出编码代理作为快速廉价的执行者,其创造性决策能力有限,但能显著放大 competent 工程师的效率。同时,AI 在执行精确指令方面进步迅速,但在处理未明确指令的决策能力上仍有瓶颈。
Fchollet 认为编码代理是快速廉价的执行者,能放大 competent 工程师的效率,但不会取代工程师本身。
展开原文
AI 在执行精确指令方面进步迅速,但在处理未明确指令的决策能力上仍有瓶颈,需要关注这一差距。
展开原文
Fchollet 进一步阐述 AI 代理对初级工程师的影响,强调其价值在于学习过程而非当前产出。
展开原文
Fchollet 警示 AI 可能阻碍初级工程师获得能力,从而间接贬值他们的长期价值。






