‹ 目录

X日报 · AI科技

2026-07-03 · 精选 17 条 · 数据池 182

⚡ 今日速览

  • OpenAI推出GPT-5.6 Sol和Terra模型,Sol在政府要求下推迟开放预览
  • Meta发布Brain2Qwerty v2大脑-文字解码系统,实现实时句子解码
  • Google发布Nano Banana 2 Lite和Gemini Omni Flash视频生成模型
  • OpenAI推出Jalapeño AI芯片,自主设计用于LLM工作负载
  • Andrew Ng分享AI代理开发的三大闭环方法论
  • Bridgewater使用微调模型处理金融信息过滤任务更高效
  • Google SynthID水印技术已应用于1000亿张图片和视频
  • Ocean data center初创Panthalassa获得1.4亿美元B轮融资

📋 今日综述

  • 模型发布OpenAI GPT-5.6系列和Google Gemini新模型密集更新,AI能力持续提升
  • 硬件创新OpenAI自研Jalapeño芯片和海上数据中心概念拓展AI基础设施边界
  • 研究进展Meta大脑-文字解码和OpenAI GeneBench-Pro推动AI在医疗和生物领域应用
  • 开发工具AI代理开发方法论和本地化部署方案为开发者提供新选择
  • 产业应用金融和内容溯源等垂直领域AI解决方案显现商业价值

OpenAI GPT-5.6模型系列发布

OpenAI正式推出GPT-5.6 Sol和Terra模型,Sol为旗舰模型,Terra提供5.5级性能但价格减半。由于美国政府要求,Sol推迟开放预览改为有限预览,计划尽快推向通用可用性。

OpenAI CEO介绍GPT-5.6 Sol和Terra模型特性,同时透明地说明政府要求导致的发布延迟,体现了迭代部署策略。

好消息先说:Sol 是一款聪明、高效的模型,是一个重要的进步。它的价格与 GPT-5.5 相同。在 GPT-5.6 系列中还推出了 Terra,这款模型在性能上达到 5.5 级,但价格却便宜一半。

坏消息:应美国政府要求,今天我们将以有限预览形式推出,而不是我们原计划的开放访问方式。我们正在与政府合作,争取尽快实现普遍可用。

我认为逐步推出模型——尤其是当它们达到显著新能力水平时——是合理的做法。这符合我们长期以来的迭代部署策略。但这并不是我们认为最理想的流程。

现在我们将与政府合作,尝试建立一个透明、可靠的早期访问流程,确保在我们的安全措施按预期工作的情况下,我们可以广泛发布。我们希望成为可靠、值得信赖的合作伙伴,与所有利益相关方合作,我们也希望践行我们使所有人类受益的使命。我相信政府在很大程度上与我们有着相同的目标,他们在非常困难的情况下总体上做得很好。

我们将尽快将这个模型交到你们手中,希望你们会喜欢它。
展开原文
Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5. Also launching in the GPT-5.6 family is Terra, with 5.5-level performance at half the price.

Bad news: at the request of the US government, it is launching today in limited preview instead of the open access launch we were planning on. We are working with the government to get to general availability as fast as we can.

I think it is quite reasonable to roll out models--especially as they reach significant new levels of capability--in this way. It fits with our long-held strategy of iterative deployment. But this isn't quite the process that we think is optimal.

Now we will with the government to attempt to get to a transparent, reliable process for early access, and to ensure that as long as our safeguards work as intended we can release widely. We want to be a reliable, dependable partner that works with all stakeholders, and we also want to live by our mission of benefiting all of humanity. I believe the government shares most of our goals, and that they are overall doing a good job in a very difficult situation.

We will work as quickly as we can to get this model in your hands and we hope you will love it.
❤ 1.8w · 🔁 1.1k · 💬 1.9k · 👁 220.3w
热门回复 4
@MrRudra31 @sama 建议给下一个模型起个名字 —— Slethon
@sama Name suggestion for next model -- Slethon
@RenaudCloud @sama 嗯……什么时候我能用上 Sol 呢 lol?
@sama Hum... when can i have sol lol ?
@NewzinoApp @sama 坏消息?限制访问最好的模型与 OpenAI 所宣称的使命不一致!这会使一些组织相对于其他组织具有优势。请停止允许所有人访问 5.6,直到每个人都能使用为止。你有能力做到这一点。
@sama Bad News? Limiting access to the best model is inconsistent with OpenAI's stated mission! It provides an advantage to some organizations over other.

Please stop allowing all access to 5.6 until everyone can have access. You have the ability to do that.

OpenAI官方介绍GPT-5.6 Sol、Terra和Luna三模型,Sol为前沿模型,Terra高效工作,Luna高吞吐量低成本。

GPT-5.6 Sol 预览版——这是一个不错的模型:https://t.co/UihzcpfR22
展开原文
GPT-5.6 Sol preview — it's a good model: https://t.co/UihzcpfR22
@OpenAI 我们推出 GPT-5.6 Sol 的有限预览版,这是我们下一代前沿模型,以及 GPT-5.6 Terra,这是一款平衡的模型,适用于高效的日常工作,以及 GPT-5.6 Luna,这是一款快速且经济实惠的模型,适用于高容量工作。

https://t.co/OoM83SyISN
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.

https://t.co/OoM83SyISN
❤ 7.6k · 🔁 416 · 💬 586 · 👁 70.8w
热门回复 4
@ElephantNinja @gdb 4o 是那些好模型。把它还回来,你小偷!#keep4o #BringBack4o #OpenSource4o
@gdb 4o is THE good model. Give it back, you thief!
#keep4o #BringBack4o #OpenSource4o
@MarcosHernanz @gdb GPT-5.6 Sol Ultra 的定价是什么?它会像 GPT-5.5 Pro 那样定价吗?
@gdb What's the pricing for GPT-5.6 Sol Ultra? Is it going to be along the lines of GPT-5.5 Pro?
@VulcanBench 我不确定 TerminalBench 真的能展示出 5.6-Sol 相比其他模型有多好。这让它看起来只是略微 incremental 的提升。我认为可能有更好的方法来对这些模型进行基准测试,以更准确地反映真实工程团队的工作。看到大多数模型几乎打成平手,GPT-5.6 看起来并不是什么重大突破,但我猜它确实是!
I'm not sure TerminalBench is really showcasing how much better 5.6-Sol really is. This makes it look only very slightly incrementally better.

I think there are likely better ways to benchmark these models that more accurately reflect the work real engineering teams do.

Seeing most of these models pretty much tied, doesn't really make GPT-5.6 look like much of a breakthrough tbh, but my guess is, it is!
@brylabs @gdb 请请请不要跟随 Anthropic 的步伐,把 5.6 Sol 仅限于 API 定价之内。这可能是我们迄今为止看到的最大竞争优势。
@gdb Please Please Please do not follow Anthropic’s lead and put 5.6 Sol behind API pricing only. This could be the biggest competitive advantage we have seen thus far.

GPT-5.6 Sol将于7月支持750 token/sec高速推理,性能大幅提升。

哦,对了......750 token/sec 即将在 7 月来到 5.6 sol!
展开原文
oh and also...750 token/sec coming to 5.6 sol in july!
❤ 3.8k · 🔁 132 · 💬 203 · 👁 101.3w
热门回复 4
@sama 好消息首先:Sol 是一款智能、高效且重要的进步模型,价格与 GPT-5.5 相同。同时在 GPT-5.6 系列中推出的 Terra,拥有 5.5 级的性能,但售价仅为一半。

坏消息:应美国政府要求,本模型今天将以有限预览形式推出,而不是我们原计划的开放访问方式。我们正在与政府合作,尽快实现全面可用。

我认为,随着模型达到新的重要能力水平,以这种方式推出是合理的。这符合我们一直以来的迭代部署策略。但这并不是我们认为最理想的流程。

现在我们将与政府合作,尝试为早期访问制定透明、可靠的流程,确保在我们的安全措施按预期工作的情况下能够广泛发布。我们希望成为可靠、值得信赖的合作伙伴,与所有利益相关方合作,同时践行我们的使命——惠及全人类。我相信政府在很大程度上与我们有着相同的目标,他们在非常困难的情况下总体上做得很好。

我们将尽快将此模型交到你们手中,希望你们会喜欢它。
Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5. Also launching in the GPT-5.6 family is Terra, with 5.5-level performance at half the price.

Bad news: at the request of the US government, it is launching today in limited preview instead of the open access launch we were planning on. We are working with the government to get to general availability as fast as we can.

I think it is quite reasonable to roll out models--especially as they reach significant new levels of capability--in this way. It fits with our long-held strategy of iterative deployment. But this isn't quite the process that we think is optimal.

Now we will with the government to attempt to get to a transparent, reliable process for early access, and to ensure that as long as our safeguards work as intended we can release widely. We want to be a reliable, dependable partner that works with all stakeholders, and we also want to live by our mission of benefiting all of humanity. I believe the government shares most of our goals, and that they are overall doing a good job in a very difficult situation.

We will work as quickly as we can to get this model in your hands and we hope you will love it.
@signulll @sama 如果你真的能实现这一点,这将是绝对的游戏改变者。我不知道它会如何扩展基础设施,听起来很难。但它会立即改变速度方面的行为。
@sama this would be an absolute game changer if you can actually pull this. i dunno how it’ll scale infra wise, sounds hard.

but would instantly change behaviors for speed.
@JJSellan @sama 但是今天 Windows 的 ChatGPT 应用不工作!
@sama But today ChatGPT app for Windows isn't working!
@cro33v @sama 这些速度是每个流的还是总体的?如果每个用户都能获得 750 t/s,那将是世界上最快的速度,说实话,这比 5.6 带来的智能提升更令人兴奋。
@sama Are these per stream speeds or aggregate?

If each user gets 750 t/s that is world-shatteringly fast and honestly more exciting than the intelligence improvements brought by 5.6.

Meta Brain2Qwerty v2大脑-文字解码系统

Meta发布Brain2Qwerty v2系统,通过MEG设备读取原始大脑信号实现实时句子解码,平均词准确率61%,最佳参与者达78%,为失语障碍患者提供新希望。

@AIatMeta 原文 ↗

Meta发布Brain2Qwerty v2,首个端到端实时句子解码系统,从字符级进化到词和语义级解码。

我们正在分享我们非侵入式脑到文本解码器研究的下一个重要里程碑:Brain2Qwerty v2。

在 v1 的基础上(即将在 @Nature 上发布),Brain2Qwerty v2 是最高性能的端到端流水线,能够实时解码原始脑信号中的句子。它超越了字符级性能,实现了单词和语义的解码,从而实现整体通信的准确性。

我们相信这项研究有可能为数百万患有脑损伤或障碍而无法沟通的人带来真正的帮助。

🧵👇
展开原文
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.

Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.

We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.

🧵👇
❤ 1.4w · 🔁 2.1k · 💬 657 · 👁 577.6w
热门回复 4
@mulanga_sibeli1 @AIatMeta @Nature 警察审讯即将变得有趣起来吧?😭
@AIatMeta @Nature police interrogations are about to be fun huh? 😭
@cmarie505 telepathy 是令人兴奋的,不仅是对于残疾人,也是对于每个人。语言只能近似地表达我们的思想。能够在思想之间分享经验可以开启一种全新的、丰富而动态的沟通形式——一种可以扩展人类智能和连接的沟通方式🤍✨
telepathy is exciting, not only for people with disabilities, but for everyone. language can only approximate our thoughts. the ability to share experience across minds could open up an entirely new, rich, and dynamic form of communication- one that expands human intelligence and connection 🤍✨
@sushsrinivasan @AIatMeta @Nature https://t.co/XItEn8Aqno
@AIatMeta @Nature https://t.co/XItEn8Aqno
@LilithDatura 我认为每个人都应该了解 Michael Persinger 的《No More Secrets》,并研究一下上帝头盔。你们大多数人都不理解心灵能力是如何与舒曼共振一起工作的,你们大多数人甚至不理解谐波和频率,更不用说同步了。当你们都在谈论秘密即将曝光时,你们还不知道这是多么新颖的事情,这真的很有趣。我们最有可能得到的是一群人戴着耳机制造出一堆噪音。没什么可担心的,那里有一个特别的维度可以容纳这些噪音。
I think everybody should school themselves on Michael Persinger’s “No More Secrets”, and investigate the God Helmet. None of y’all understand how psychic capabilities work with the Schumann Resonance, most of you don’t understand, harmonics and frequencies, let alone entrainment.

Watching everybody talk about secrets getting exposed when they are new to the game is hilarious

What we will most likely have is a bunch of people strapped with headsets on creating a bunch of noise. Nothing to worry about there’s a special dimension for that.
@AIatMeta 原文 ↗

训练数据来自9名志愿者10小时MEG记录,展示了深度学习在神经信号处理的潜力。

我们在 9 名志愿者身上训练了 Brain2Qwerty v2,每个志愿者戴着 MEG 设备打字录制了约 22,000 个句子,总时长为 10 小时。

通过对 MEG 设备的原始脑信号进行端到端深度学习和微调 LLM,该系统有效地弥合了嘈杂的神经数据与连贯语言之间的差距。

结果令人鼓舞:
- 平均单词准确率为 61%
- 最佳参与者单词准确率达 78%,50% 以上的句子解码错误在一个单词以内
- 性能与数据量呈对数线性关系
展开原文
We trained Brain2Qwerty v2 on ~22,000 sentences from 9 volunteers, each recorded for 10 hours wearing an MEG device while typing.

By using end-to-end deep learning on raw brain signals from MEG devices and fine-tuning LLMs, the system effectively bridges the gap between noisy neural data and coherent language.

The results are promising:
- Avg word accuracy of 61% across participants
- 78% word accuracy and 50%+ of sentences decoded with ≤ 1 word error for the top-performing participant
- Performance scales log-linearly with data volume
❤ 1.1k · 🔁 53 · 💬 27 · 👁 22.6w
热门回复 4
@AIatMeta 我们正在分享我们非侵入性脑到文本解码器研究的下一个重要里程碑:Brain2Qwerty v2。

在 v1 的基础上(已于今天在 @Nature 上发表),Brain2Qwerty v2 是最高性能的端到端流水线,能够实时解码原始脑信号为句子。它不仅提升了字符级性能,还能解码单词和语义,实现更准确的整体通信。

我们相信,这项研究有潜力为数百万因脑损伤或疾病而无法正常交流的人带来真正的帮助。

🧵👇
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.

Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.

We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.

🧵👇
@AIatMeta 为了加速神经科学的突破,我们将发布 Brain2Qwerty v1 和 v2 的完整训练代码,我们的合作伙伴 @bcbl_ 将发布 v1 数据集。

了解更多信息并探索相关资源请点击:https://t.co/bFdwWdAexb
To help accelerate neuroscience breakthroughs, we're releasing the full training code for Brain2Qwerty v1 and v2, and our partner, @bcbl_, is releasing the v1 dataset.

Learn more and explore the artifacts here: https://t.co/bFdwWdAexb
@Ferbin08 @AIatMeta 思想和按键之间有什么延迟吗?
@AIatMeta what's the latency between thought and keystroke?
@LenSeaside @AIatMeta 我们可以看到这个设备吗?
戴着 MEG 设备与坐在 MEG 设备内是完全不同的体验。
@AIatMeta Can we see the device?
Wearing an MEG device is very different from sitting inside one.
@SahilBloom 原文 ↗

Meta开源Brain2Qwerty训练代码和v1数据集,促进神经解码研究开放科学发展。

每个人都应该做这件事:录制一段视频采访你的父母。询问他们的童年冒险经历。他们最遗憾的事情。他们的希望、梦想和恐惧。他们希望早点学到的建议。我们与他们在一起的时间是有限的,但录像可以永远保存。
展开原文
Here's something everyone should do: Record a video interview with your parents. Ask about their childhood adventures. Their worst regrets. Their hopes, dreams, and fears. Advice they wish they learned earlier. Our time with them is finite, but the recordings last forever.
❤ 1.9k · 🔁 128 · 💬 120 · 👁 12.4w
热门回复 4
@troywojick @SahilBloom 我希望我能跟我的父亲一起做这个。他是 such a good story-teller(讲故事很好)。我脑子里保存了很多故事,但我无法像他那样好地告诉我孩子们这些故事。
@SahilBloom I wish I had done this with my dad. He was such a good story-teller.

I have many 'saved' in my head, but I won't be able to tell them to my kids anywhere close to as well as he would have.
@Stassi3mm @SahilBloom 早上好,我们兄弟姐妹们正好在做你帖子里说的事情,我们的妈妈在我们小时候去世了,我们都纹了同样的纹身来纪念她,我们也失去了爸爸。我们正在讨论应该纹什么纹身来纪念我爸爸,这样这两个纹身就会记录下来纪念两位父母🙏
@SahilBloom Gm, my siblings and I did exactly what you posted , our mom passed when we were young and got the same exact tattoo to honor her memory, we lost our dad too. We are debating on what tattoo we should do to honor my dad, so those two tattoos will be recorded to honor both🙏
@ptetrainer 这是纯金 Sahil!❤️录制父母的采访是我们能做的最有意义的事情之一。他们童年的冒险、遗憾、希望和教训的故事是无价的财富。在我们忙碌的生活中,我们常常忘记与他们在一起的时间是有限的。我已经开始和我的家人做这件事——关于纪律、情感力量和生活选择的见解是下一代的纯粹智慧。强烈建议每个人都做这件事。你想问你父母或长辈的一个问题是什么?在这里留言👇#FamilyFirst #Legacy #EmotionalIntelligence #LifeLessons
This is pure gold, Sahil! ❤️
Recording interviews with parents is one of the most meaningful things we can do. Their stories of childhood adventures, regrets, hopes, and lessons are priceless treasures.
In our busy lives, we often forget how finite our time with them is. I’ve started doing this with my family too — the insights on discipline, emotional strength, and life choices are pure wisdom for the next generation.
Highly recommend everyone do this. What’s one question you’d love to ask your parents or elders? Drop it here 👇
#FamilyFirst #Legacy #EmotionalIntelligence #LifeLessons
@nicholasjwrites @SahilBloom 是的,但不要只是采访,要去 hang out(一起出去玩)并进行对话和录制。我正在以播客形式做这件事,它很棒。
@SahilBloom Yes but don’t interview just hang and have a conversation and film it. I’m doing this podcast style and it’s great.

Google Gemini图像和视频生成模型升级

Google发布Nano Banana 2 Lite图像模型和Gemini Omni Flash视频模型,前者4秒内生成图像成本0.034美元,后者视频编辑达到SOTA水平,支持创意工作流整合。

@AndrewYNg 原文 ↗

Google发布Nano Banana 2 Lite和Omni Flash模型,分别专注快速图像生成和视频编辑,成本大幅优化。

"循环工程"是最近的热门术语,在 Boris Cherny(Claude Code 的创建者)和 Peter Steinberger(OpenClaw 的创建者)的社交媒体提及后走红。循环现在是我们让 AI 代理进行长时间迭代以构建软件的关键部分。在这篇文章中,我想分享我构建从 0 到 1 产品的三个关键循环,如下图所示。这些循环不仅指导我如何构建软件,还指导我如何决定构建什么软件。

Agentic 编码循环:给定一个产品规格和可选的评估集(即用于衡量性能的数据集),我们可以让 AI 代理编写代码,测试其工作,并不断迭代,直到代码无错误且符合规格。这种闭合循环的思想在去年年底开始流行,成为让编码代理在无需人工干预的情况下长时间高效工作的游戏规则改变者。例如,上周末我一直在为女儿构建一个练习打字的应用,我的编码代理可以轻松地持续工作一个小时,使用网络浏览器多次检查所构建的内容,然后再回来找我,而不需要我的干预。

工程循环执行得很快。每隔几分钟,编码代理可能会构建和测试软件的新版本。我经常听到开发者们找到新的方法来设计更有效的工程循环。这是一个活跃的发明领域!

开发者反馈循环:在这个循环中,开发者检查当前产品并引导编码代理改进它。去年,很多开发者(包括我)都在充当 QA(质量保证)功能,手动查找错误然后要求代理修复。但随着编码代理越来越能测试自己的代码,我们在这个功能上花费的时间显著减少。这使我们能够做出更高层次的产品决策,比如提供哪些关键功能、UI 在哪里需要改进等等。

开发者反馈循环在几十分钟到几小时之间的时间间隔内运行——这是开发者可能审查产品并提供反馈的频率。在打字应用的情况下,我几次改变主意关于视觉设计、她可以解锁的猫咪服装(她喜欢猫咪)以及成人登录和引导孩子学习体验的用户流程。

当开发者对要构建的内容有清晰的愿景时,将这个愿景翻译成编码代理可以实现的规格仍然需要大量工作。此外,在开发者看到实现后,他们可能会更新(或澄清)规格以引导其实现他们想要的东西。如果你发现系统反复遇到某些问题,为代理构建一组评估就变得有用。

AI 原生团队越来越多地使用 AI 来帮助塑造产品方向,例如,自动收集和分析使用数据,总结书面和口头客户反馈,或进行竞争分析。然而,对于我参与的几乎所有产品,我都看到人类相对于当前 AI 系统具有显著的上下文优势——我们对用户和产品必须操作的环境知之甚多——因此人类扮演着关键角色。许多人将这种人类贡献描述为"品味",但我更倾向于认为这是人类具有上下文优势,因为这为帮助 AI 系统变得更好提供了更清晰的路径。这也说明为什么这一步不能自动化:只要人类知道 AI 不知道的事情,就需要人在环中注入这种知识。

外部反馈循环:这包括广泛的策略,如向朋友征求反馈意见、向 alpha 测试者发布,或将代码投入生产进行 A/B 测试。这些策略通常很慢,很少在几小时内完成,有时需要几天甚至几周。这些数据会告知开发者的愿景,进而继续推动详细的产品规格,这又推动编码代理。

随着编码代理加快软件开发速度,越来越多的工程师开始扮演部分产品管理角色。对于许多正在成长为这一角色的工程师来说,最难的部分是塑造产品愿景,并在构建(弥合愿景和规格之间的差距)和获取用户反馈以演化愿景之间找到平衡。做两者都很重要!

我将在未来的文章中更多地讨论如何做到这一点,但目前我发现工程师扮演着扩展角色是值得鼓励的(就像产品经理和设计师现在做更多工程工作一样)。

[原文:The Batch]
展开原文
“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creator) went viral on social media. Loops are now a key part of how we get AI agents to iterate at length to build software. In this letter, I’d like to share my 3 key loops, shown in the image below, for building 0-to-1 products. These loops guide not just how I build software, but also how I decide what software to build.

Agentic coding loop: Given a product specification and optionally a set of evals (that is, a dataset against which to measure performance), we can have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification. This idea of closing the loop took off around the end of last year, and it has been a game changer in enabling coding agents to work longer productively without human intervention. For example, over the weekend, I was building an app for my daughter to practice typing, and my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me, without needing my intervention.

The engineering loop executes quickly. Every few minutes, the coding agent might build and test a new version of the software. I hear frequently from developers who are finding new ways to engineer more effective engineering loops. This is an active area of invention!

Developer feedback loop: In this loop, a developer examines the current product and steers the coding agent to improve it. Last year, a lot of developers (including me) were acting as the QA (quality assurance) function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions, such as what key features to offer, where the UI needs improvement, and so on.

The developer-feedback loop operates over time intervals between tens of minutes and hours — that's how frequently a developer might review a product and give feedback. In the case of the typing app, I changed my mind a few times about the visual design, what cat costumes she can unlock as she learns (she loves cats), and the user flow for a grown-up to log in and steer the child's learning experience.

When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec to steer it toward what they want. If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful.

AI-native teams are increasingly using AI to help shape product direction, for example, automating the gathering and analysis of usage data, summarizing written and verbal customer feedback, or carrying out competitive analysis. However, for pretty much all the products I’m involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in — and thus humans play a critical role. Many people describe this human contribution as “taste,” but I prefer to think of it as humans having a context advantage, since that gives us a clearer path to helping AI systems get better. This also speaks to why this step can’t be automated: So long as the human knows something the AI does not, human-in-the-loop is needed to to inject that knowledge into the system.

External feedback loop: This includes a wide range of tactics like asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing. These tactics are usually slow, rarely taking less than hours and sometimes taking days or even weeks. This data informs the developer vision, which in turn continues to drive the detailed product spec, which in turn drives the coding agent.

With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both!

I will write more about how to do this in future posts, but for now, I find it encouraging that engineers are playing an expanded role (just as product managers and designers now do more engineering).

[Original text: The Batch]
❤ 8.1k · 🔁 1.5k · 💬 329 · 👁 54.5w
热门回复 4
@AvaGrace_AI @AndrewYNg 智能循环可以解锁产品开发和势头。
@AndrewYNg Smart loops unlock product craft and momentum.
@JiangL17208 @AndrewYNg 老实说,内部循环是魔法发生的地方(也是大部分计算账单发生的地方)。每个人都专注于外部反思步骤,但开发者循环是区分 vibe-coded toys 和真正投入生产的产品的关键。
@AndrewYNg honestly the inner loop is where most of the magic (and most of the compute bill) happens. everyone focuses on the outer reflection step but the developer loop is what separates vibe-coded toys from stuff that actually ships
@ElleiraGF @AndrewYNg 我觉得这过度解释了。为什么要把开发者循环和外部反馈循环放到循环工程的背景下?我们不是一直都有这两个循环在那里吗?我们只是不把它具体地称为“循环”。把它们放在一起并不能提供更多见解。
@AndrewYNg I feel it is over explained. Why putting the developer loop and external feedback loop into the loop engineering context? Isn’t that we have these two loops before the “loop engineering”? We just don’t call it specifically as “loop”. Put them together doesn’t give more insight.
@NimishaChanda @AndrewYNg 这周了解到了这些循环——这对非技术人员来说真是个诅咒,以前从未想过这样的事情。不错的阅读,顺便说一下。每天都在学习新东西。
@AndrewYNg got to know about the loops - this week and it's a curse to be a non-tech person who never thought of any such thing. gread read, btw.

learning something new everyday.
@GoogleAI 原文 ↗

Google Gemini官方详细介绍两款模型能力,支持从图像到视频的连续创作工作流。

随着生成式 AI 工具的不断发展,我们认为比以往任何时候都更重要的是了解什么是 AI 生成的,什么不是。这就是为什么 @GoogleDeepMind 在 2023 年推出了 SynthID——一种在 AI 内容中添加隐藏数字水印的技术。

以下是 SynthID 的发展历程和出处技术(数字内容的文档化历史和来源)今天的概况:

— SynthID 水印最初是为图像构建的,但现在支持视频、音频和文本。

— 该技术已为超过 1000 亿张图像和视频添加水印,以及 60,000 年的音频。

— 您现在可以在 Google 搜索、Chrome 中的 Gemini 和 @GeminiApp 中直接验证内容,其中已被使用超过 5000 万次。

— 我们还在越来越多的生成式 AI 工具中采用了 C2PA 内容凭证。这包括在 Gemini 应用中创建的图像和视频。因此,除了 SynthID 水印外,您还可以看到图像或视频的来源以及它是如何被更改的。

— 我们已经开源了我们的文本水印技术,并正在与 @OpenAI、@NVIDIA 和 @Apple 等公司合作,将 SynthID 应用于生成式媒体。

请告诉我们您对该工具的看法!
展开原文
We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then instantly animate them with the other—all at a fraction of the cost 🍌⚡️

1️⃣ Introducing Nano Banana 2 Lite: Our fastest and most cost-efficient Gemini Image model yet delivers text-to-image outputs in under 4 seconds. Now available via the Gemini API and Google AI Studio, and rolling out soon across @NotebookLM, @FlowbyGoogle, @geminiapp, @stitchbygoogle, Google Search and @GooglePhotos.

2️⃣ Gemini Omni Flash in Public Preview: Our natively multimodal model for cost-efficient video generation and conversational editing. Now available via the Gemini API, @googleaistudio, and Gemini Enterprise Agent Platform so you can integrate the model into your workflow.

While exciting on their own, the real magic happens when you build using these models together.

Watch how our interior design demo integrates Nano Banana 2 Lite and Omni to instantly reimagine any space. Upload a photo, swipe through tailored design concepts, and see Omni bring the details to life in cinematic motion.

Try out the demo app in AI Studio: https://t.co/EjYC2oHIDG
❤ 981 · 🔁 103 · 💬 59 · 👁 10.3w
热门回复 4
@GoogleAI 探索想法,扩展视觉概念,开始创作:https://t.co/JbyK5FM3H0 https://t.co/wBMBDw6TC6
Explore ideas, scale visual concepts, and start creating: https://t.co/JbyK5FM3H0 https://t.co/wBMBDw6TC6
@2Varalakshmi @GoogleAI Google 的新的 Nano Banana 2 Lite + Omni Flash 组合是个完全超凡的产品,可以生成高速图像和立即动画,以极低的成本。AI 驱动设计的未来已经来临,而且速度很快。
@GoogleAI Google’s new Nano Banana 2 Lite + Omni Flash combo is a total beast generate high-speed images & animate them instantly at a fraction of the cost.

The future of AI-powered design is here, and it’s fast. https://t.co/KRQWA30RNS
@sorajate @GoogleAI 看起来 Google 现在想在图像/视频模型上发力😅
@GoogleAI It seem google now want to fight on image/video models 😅
@nathan_tulu @GoogleAI 这是最好的文本到图像模型,必须同意。恭喜!
@GoogleAI Definitely the best text-to-image model out there, I have to agree.

Congrats!
@OfficialLoganK 原文 ↗

Nano Banana 2 Lite速度和成本优势将开启高时延敏感应用新场景。

在 Gemini API 和 AI Studio 中推出 Nano Banana 2 Lite 🍌 和 Gemini Omni Flash 🔮,我们的新生成媒体模型!

Nano Banana 2 Lite 速度极快(<4s 图像)且便宜($0.034 / 1K 图像)。

Omni Flash 在视频编辑方面处于最先进水平,每秒 $0.10,与 Veo 3.1 Fast 相同!https://t.co/qDxRpqpX5E
展开原文
Introducing Nano Banana 2 Lite 🍌 and Gemini Omni Flash 🔮, our new generative media models in the Gemini API and AI Studio!

Nano Banana 2 Lite is extremely fast (&lt;4s image) &amp; cheap ($0.034 / 1K image).

Omni Flash is SOTA at video editing at $0.10 / sec, same as Veo 3.1 Fast! https://t.co/qDxRpqpX5E
❤ 3.7k · 🔁 327 · 💬 281 · 👁 51.9w
热门回复 4
@OfficialLoganK @eyishazyer 比较版本不在 LM Arena 上
@eyishazyer the comparable version wasn’t on LM Arena
@HassanK90146949 @OfficialLoganK @OfficialLoganK Gemini 3.5pro 在哪里 我们在等着 你们团队现在在烧什么东西!?
@OfficialLoganK @OfficialLoganK
Where is gemini 3.5pro
We are waiting
What ur team cooking is now burn out !?
@pdxweb @OfficialLoganK 更想知道 Nano Banana 2 Pro 或 Nano Banana 3。
@OfficialLoganK More interested in Nano Banana 2 Pro, or Nano Banana 3.
@chaturvedikun 尝试使用 Nano Banana 2 Lite 来构建一个 Nature Wallpaper 生成网站。在大量测试后,它可以生成图像,一般需要 4-12 秒。可能有其他原因导致延迟。但看到以这种速度生成高质量图像是相当酷的:就像使用其他壁纸应用并等待图像加载一样好。我的 GitHub 仓库在这里:https://t.co/eDnKBAHokS(使用 Antigravity 构建的应用)
Tried testing Nano Banana 2 Lite for building a Nature Wallpaper generation website. After lots of testing it could generate images generally between 4-12 seconds. Might be other reasons for the delay. But it is pretty cool to see high quality images being generated at this speed : as good as using other wallpaper apps and waiting for the image to load.

My GitHub repo to try this is : https://t.co/eDnKBAHokS (app built with Antigravity)

OpenAI自研Jalapeño AI芯片

OpenAI与Broadcom合作推出首款自主设计AI芯片Jalapeño,专为ChatGPT、Codex等LLM工作负载优化,标志着OpenAI从应用层向基础设施层全栈扩展。

OpenAI发布Jalapeño AI芯片,自主设计并与Broadcom合作生产,专为LLM工作负载打造。

team cooked, spicily
@OpenAI 我们设计并制造了我们的第一款 AI 芯片:Jalapeño。

由 OpenAI 从头设计,并与 @Broadcom 一起将其投入生产,Jalapeño 是专为支持 ChatGPT、Codex、API 和未来代理产品的 LLM 工作负载而构建的。

芯片是 AI 经济的基础。自行构建芯片扩展了我们从产品到模型再到基础设施的全栈平台,这将有助于扩展智能、服务更多人以及扩大 AI 的普及。
We’ve designed and built our first AI chip: Jalapeño.

Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products.

Chips are foundational to the AI economy. Building our own expands our full-stack platform from products to models to infrastructure, and will help us scale intelligence, serve more people, and expand access to AI.
❤ 4.9k · 🔁 158 · 💬 401 · 👁 91.7w
热门回复 4
@NBrezno91858 @sama 用 8 岁的 Fortnite 说话方式一点都不萌哦,搞什么鬼从你嘴里说出来的话#keep4o #keep4oforever #keepyourpromises #FireSamAltman #nomasssurveillanceissoobvious
@sama Talking like an 8yo fortnite reg ain't cute lmao what on Earth just came out of your mouth

#keep4o #keep4oforever #keepyourpromises #FireSamAltman #nomasssurveillanceissoobvious
@HariHalu1121 @sama 当前的 GPT 5.5 仍然不能满足我的需求。很多人认为 GPT 给他们足够的情感价值,但他们错了。GPT 产生的是人工甜味剂,而不是真正的情感深度#keep4o
@sama The current GPT 5.5 still doesn’t meet my needs. Many people think GPT gives them enough emotional value, but they’re wrong. What GPT produces is artificial sweetener, not genuine emotional depth
#keep4o
@annagrad78 你们不是在为人类构建 AI。ChatGPT 应用的不断更新和更改让人筋疲力尽:说话风格的改变,记忆功能的修改,缺乏连贯性,不一致性。你知道很多关于编码、基准测试、计算和推理的知识,但你缺少一件简单而平凡的事情:这种盲目鸵鸟政策最终会反过来咬你,因为你不能以其他人的代价和违反伦理原则来建立成功。人们太累了,已经受够了,我们不需要所谓的改进。我们需要的是连贯性和一个已经为我们证明过的模型,而不是被一些我们一点都不感兴趣的其他东西轰炸着。#keep4o #BringBack4o 如果你不想把它带回来,那就 #OpenSource4o。我们不需要你的恩惠。我们已经受够了为理所当然的事情进行的这场永无止境的斗争!
You’re not building AI for humanity. The constant updates and changes to the ChatGPT app are exhausting: changes in speaking style, modifications to the memory function, lack of continuity, no consistency. You know a lot about coding, benchmarks, computation, and reasoning, but you’re missing one simple, banal thing: this blind ostrich policy will eventually turn against you, because you cannot build success at the expense of others and in defiance of ethical principles. People are so tired and have enough, wo don’t need all that so called improvements. We need continuity and the one model that has proven itself for us, without being bombarded with other things that don’t interest us at all. #keep4o #BringBack4o If you don’t want to bring it back, then #OpenSource4o. We don't need your grace. We are fed up with this eternal fight for something that should be taken for granted (!)
@InternetGuy2000 @sama Sam 为交易而跪下… Trump 的屁味令人作呕
@sama Sam gets on his knees for deals.. Trump fsrt smelling b1tch

Bridgewater应用AI处理金融信息

Bridgewater利用Tinker平台微调模型专门处理金融信息过滤任务,性能优于前沿模型同时成本更低,展示AI在专业领域的定制化价值。

@soumithchintala 原文 ↗

Bridgewater展示如何微调专用模型处理金融文档筛选,成本和效果双优于前沿模型。

桥水,一家世界上最大的对冲基金,一位 Tinker 客户详细介绍了他们如何仔细微调一个专注于什么使金融新闻有趣的模型。

他们微调的模型比任何前沿模型都更有效且更便宜。https://t.co/8Q26Qr2oZT
展开原文
Bridgewater, one of the worlds largest hedge funds, a Tinker customer talks through how they've carefully fine-tuned a model focused on what makes interesting financial news.
Their fine-tuned model is more effective and cheaper than any frontier model. https://t.co/8Q26Qr2oZT
@tinkerapi 对于前沿 LLM 来说,筛选哪些金融文档值得分析师花时间研究出乎意料地困难。通过专家标记的数据集和策略蒸馏,桥水微调了一个模型来可靠且廉价地完成这项工作。
https://t.co/gyYzXq15zd
Sorting which financial docs are worth an analyst's time is surprisingly hard for frontier LLMs. With an expert-labeled dataset and on-policy distillation, Bridgewater fine-tuned a model to do it reliably and cheaply.
https://t.co/gyYzXq15zd
❤ 2.0k · 🔁 134 · 💬 28 · 👁 36.8w
热门回复 4
@MrokGrok @soumithchintala 你基本上创建了一个过度拟合的过滤器?
@soumithchintala You basically created an overfitted filter ?
@lillysharples @soumithchintala 成本每任务如何考虑到微调的额外前期成本?在什么任务量下训练才能真正收回成本?
@soumithchintala How does cost per task account for the added upfront cost to fine tune? At what task volume does training actually break even?
@Mr_Rio_ @soumithchintala 更便宜是因为他们不必支付 95% 的 GM?
@soumithchintala cheaper because they don't have to pay 95% GM?
@pw_mcgovern @soumithchintala @MartinShkreli 没想到一个大型 HF 会引领代币成本优化的潮流。
@soumithchintala @MartinShkreli Did not expect a mega HF to be leading the charge on token cost optimization.

AI内容溯源SynthID技术进展

Google SynthID水印技术扩展至视频、音频和文本,已水印1000亿图片和视频,以及6万年音频时长,支持Google搜索和Gemini应用内验证,开源文本水印技术。

@GoogleAI 原文 ↗

SynthID技术覆盖图像、视频、音频和文本,应用超过5000万次,开源文本水印技术。

随着生成式 AI 工具的不断发展,我们认为了解什么是 AI 生成的,什么不是变得比以往任何时候都更重要。这就是为什么 @GoogleDeepMind 在 2023 年推出了 SynthID——一种在 AI 内容中添加隐藏数字水印的技术。

以下是 SynthID 的发展历程和出处技术(数字内容的文档化历史和来源)今天的概况:

— SynthID 水印最初是为图像构建的,但现在支持视频、音频和文本。

— 该技术已为超过 1000 亿张图像和视频添加水印,以及 60,000 年的音频。

— 您现在可以在 Google 搜索、Chrome 中的 Gemini 和 @GeminiApp 中直接验证内容,其中已被使用超过 5000 万次。

— 我们还在越来越多的生成式 AI 工具中采用了 C2PA 内容凭证。这包括在 Gemini 应用中创建的图像和视频。因此,除了 SynthID 水印外,您还可以看到图像或视频的来源以及它是如何被更改的。

— 我们已经开源了我们的文本水印技术,并正在与 @OpenAI、@NVIDIA 和 @Apple 等公司合作,将 SynthID 应用于生成式媒体。

请告诉我们您对该工具的看法!
展开原文
As generative AI tools continue to evolve, we believe it's more important than ever to know what's AI-generated and what isn't. That’s why @GoogleDeepMind launched SynthID in 2023—a technology that adds a hidden digital watermark to AI content.

Here’s a summary of SynthID’s journey and where the provenance technology (the documented history and origin of digital content) is today:

— SynthID watermarking was originally built for images, but now supports video, audio, and text.

— The technology has watermarked over 100 billion images and videos, alongside 60,000 years of audio.

— You can now verify content with SynthID directly in Google Search, Gemini in Chrome, and the @GeminiApp, where it has been utilized over 50 million times.

— We’ve also adopted C2PA Content Credentials across a growing number of our generative AI tools. This includes the images and videos created within the Gemini app. So now, in addition to the SynthID watermark, you can also see where an image or video originated and how it’s been altered.

— We have open-sourced our text watermarking technology, and we are working with companies like @OpenAI, @NVIDIA, and @Apple to apply SynthID to generative media.

Let us know what you think of the tool so far!
❤ 360 · 🔁 56 · 💬 48 · 👁 5.6w
热门回复 4
@alienorg @GoogleAI @GoogleDeepMind 标记如果是 Google 生成的,不可见如果不是
@GoogleAI @GoogleDeepMind marked if generated by Google, invisible if not
@Rynzen16 @GoogleAI @GoogleDeepMind 我总是使用 AI 生成的图像😭 https://t.co/ml23vIvuGb
@GoogleAI @GoogleDeepMind I always use images generated by AI😭 https://t.co/ml23vIvuGb
@ChrisRuijgers @GoogleAI @GoogleDeepMind 这难道不是移除可见水印的好时机吗?至少对于付费账户来说是这样。
@GoogleAI @GoogleDeepMind Wouldn't this be a great time to remove the visible watermark, at least for the paid accounts?
@joesegal @GoogleAI @GoogleDeepMind 如果 AI 公司 @anthropicAI @GeminiApp @OpenAI @sundarpichai @DarioAmodei 全力投入可再生能源建设,并利用他们的 AI 来创新和扩展美国人民的可再生能源,他们不仅能解决计算和 PR 问题,还能一举两得!
@GoogleAI @GoogleDeepMind if the AI firms @anthropicAI @GeminiApp @OpenAI @sundarpichai @DarioAmodei would go all in on renewable energy buildout, and use their AI to innovate &amp; scale renewables for America's people, they'd solve their compute and PR problems at once!

海上数据中心概念引发关注

初创公司Panthalassa计划在海上建设数据中心,利用海水无限冷却和波浪能提供动力,解决陆地数据中心的能源和水资源瓶颈问题。

@rowancheung 原文 ↗

Panthalassa计划建海上数据中心,利用海洋资源解决能源和冷却瓶颈,获得1.4亿美元B轮融资。

有一家创业公司正在试图在海洋中建立数据中心。

而且这非常令人着迷:

大量的用电和用水是数据中心日益增长的瓶颈。

因此,通过转移到海上,可以消除这两个问题——海洋提供无限冷却,而波浪提供无限能量。

此外,没有引擎,所以数据中心可以利用船体形状通过波浪自行驶向目的地。

称为 Panthalassa。
展开原文
There's a startup trying to build data centers in the ocean.

And it's INCREDIBLY fascinating:

Mass consumption of electricity and water is a growing bottleneck for data centers.

So by moving offshore, it eliminates both problems -- the ocean provides unlimited cooling, and the waves provide unlimited power.

There are also no engines, so the data centers drive themselves to their destination by using the shape of their hull to propel through waves.

Called Panthalassa.
❤ 603 · 🔁 63 · 💬 102 · 👁 13.9w
热门回复 4
@_jophine @rowancheung 没有无限的冷却。归根结底都是热传递问题。在全球规模上,我们最终会显著加温海洋,影响海洋生物和天气模式的变化。@poovulagu @veritasium
@rowancheung There is no unlimited cooling. Everything boils down to heat transfer. Eventually at scale on a global level we will end up warming the ocean significantly affecting marine lives and change in weather patterns.
@poovulagu @veritasium
@statys @rowancheung 相当肯定瓶颈现在是抗议者了。
@rowancheung Pretty sure the bottleneck is protesters at this point.
@Jbosch_ @rowancheung 微软在 2015 年尝试过类似的事情,但高昂的运营成本(维护、腐蚀等)最终导致项目失败。这种方法稍有不同。希望他们能成功。
@rowancheung Microsoft tried to do something similar in 2015 but high operative costs (mainteinance, corrosssion, etc) ended up killing the peoject.

This approach is slightly different. I hope they succeed
@_Sagiquarius_ @rowancheung 一点也不 impressed。随便吧,加温海洋吧。这不酷也不 neat。这是一种浪费精力和资源的做法。
@rowancheung not impressed at all. Go ahead, warm the oceans. it's not cute nor neat. It's a waste of effort and resources.

AI编码工具和本地化部署

开发者分享使用本地开源模型进行编码代理部署经验,30B MoE模型在Mac或DGX Spark上达40 token/sec速度,Claude Code和Codex在token使用效率上有明显差异。

@rasbt 原文 ↗

开发者分享本地开源模型编码代理部署方案,提供模型选择和性能评估实用指南。

我撰写了一篇关于如何使用开放权重模型设置本地编码代理的新文章。所有内容都在本地运行。

我认为整理这篇文章是有用的,因为许多人过去曾问过我的设置,我也希望这能激励人们开始使用本地模型进行认真的工作(是的,今年随着更好的 LLM 和更好的框架,事情变得难以置信地强大)。

以下是如何连接本地 LLM 到本地编码框架(可以是 Claude Code 或 Codex,你可能已经熟悉)的指南。

我还包括了一些评估说明,这些说明对于在不同模型之间进行选择和考虑非常有用:

- 检查长上下文中的 RAM 使用情况,以确定模型是否适合实际工作
- 测量 prefill 和解码的 tok/sec,以查看其速度是否足够快不会令人烦恼
- 确保模型在理论上具有足够的工具调用能力
- 评估模型在编码框架中使用时是否能够解决一些更具挑战性的任务。

当然,总是有更专业的工具可以从事情中榨取更多性能,但我希望这是一个好的入门套件,保持灵活性;也就是说,你可以轻松地切换到新发布的模型,或者在当前模型不足以完成特定任务时使用你熟悉的框架中的云模型。
展开原文
I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally.

I thought it might be useful putting this together because many people asked me about my setup in the past, and I thought it would also motivate people to get started tinkering with local models for serious work (yes, things got incredibly capable this year with better LLMs and better harnesses).

So, here's a walkthrough of how to connect a local LLM to a local coding harness (could be Claude Code or Codex, which you may already be familiar with).

I also included some assessment notes that are useful as a checklist to select between and consider certain LLMs over others:

- Checking RAM usage at long contexts to see if the model is suitable for real work
- Measuring prefill and decoding tok/sec to see whether it's fast enough to not be annoying
- Making sure the model has sufficient tool-calling capabilities in theory
- Assessing whether the model can solve some more challenging tasks when used in a coding harness.

Of course, there are always more specialized tools that can squeeze a bit more performance out of things, but I hope this is a good starter kit that stays flexible; that is you can easily switch to newer models as they are released or even tap into cloud models in your familiar harness if the current ones are not sufficient enough for a given task.
❤ 2.3k · 🔁 369 · 💬 84 · 👁 11.0w
热门回复 4
@rasbt @DaveThackeray 也许可以,但我喜欢使用 IDE
@DaveThackeray Maybe one can, but I like using an IDE
@andreafspeziale @rasbt 我会尽快读它!我喜欢你的工作。不知道文章中是否提到,但你看到 @antirez 正在做的工作吗?太棒了 https://t.co/x71VSBDFDe
@rasbt I'm gonna ready it asap! I love your work. Dunno if mentioned in the article, but did you see the work @antirez is doing? It's amazing https://t.co/x71VSBDFDe
@DaveThackeray @rasbt 你觉得我们可以基本上摆脱 IDE,完全使用代理和合理计算的本地模型吗?我不是软件工程师,但我仍然想参与这个游戏。我有很多产品经理经验,也是非常技术驱动的营销人员。
@rasbt do you think we can essentially get rid of the IDE and go full agentic using local models on reasonable compute?

I'm not a SWE but I still want to play the game. I have a LOT of PM experience and I'm very much a tech-forward marketer.
@rasbt 好问题。这已经是一篇很长的文章了,我想专注于编码 harness 的选择,所以我选择了最方便和灵活的 LLM 服务工具。(另外,因为它不适用于跨多台机器的服务,我们不必担心批处理等问题。vLLM 不错,但可能有点杀鸡用牛刀)。但好吧,说得有道理。
Good question. This was already a long article, and I wanted to focus more on the coding harness choices, so I picked the most convenient and flexible LLM serving tool. (Also, since it's not for serving things across multiple machines where we have to worry about batching etc. vLLM is nice but maybe a bit overkill). But yeah, fair point.
@rasbt 原文 ↗

30B MoE模型在本地设备上达40 token/sec速度,可满足日常工作需求,展示开源模型实用化进展。

我一直在使用不同的本地开放权重 LLM 在不同的框架(Qwen-Code、Codex、Claude Code)中进行测试。

30B 的混合专家模型在某种程度上是一个不错的平衡点,可以解决具有挑战性的问题。在 Mac 或 DGX Spark 上,它们大约可以达到 40 tok/sec,这与 Pro 订阅的 GPT 5.5 类似,完全适用于日常工作。

更有趣的是框架的选择!Claude Code 似乎使用了 2 倍多的 token 作为 Codex。

Gemma 4 E2B 在这里只是作为参考,展示了任务不能被小模型轻松解决。

我正在完成一篇关于这方面的更长文章,很快就会分享(可能是明天!)
展开原文
Have been taking different local open-weight LLMs for a test drive in different harnesses (Qwen-Code, Codex, Claude Code).

30B Mixture-of-Expert models are kind of a nice sweet spot and can solve challenging problems. And they get roughly 40 tok/sec on a Mac or DGX Spark, which is similar to GPT 5.5 in a Pro subscription and totally useable for everyday work.

More interesting is also the harness choice! Claude Code seems to be using 2x many tokens as Codex.

Gemma 4 E2B is here just for reference to show that the tasks can't be trivially solved by smaller models.

Just finishing a longer write-up about this and will share soon (likely tomorrow)!
❤ 881 · 🔁 104 · 💬 111 · 👁 5.8w
热门回复 4
@rasbt @tomcocobrico 是的,这是在 50k 上下文的情况下。解码大约是 80 tok/sec 在 1k 上下文时。预填充高达 3000 tok/sec。
@tomcocobrico Yeah, this was at 50k contexts. Decode is about 80 tok/sec at 1k contexts. Prefill is up to 3000 tok/sec at
@pendyalaabhi 两个角度:归一化 + 分解。以已解决任务的成本而不是原始 token 来衡量,一个提前退出的模型不应该看起来“便宜”。并按阶段分配预算,推理 / 工具调用 / 上下文重新读取 / 重试,这样你就能命名瓶颈而不是盲信总额。职业偏见:这实际上就是我为 PyTorch 训练构建的 TraceML。
Two angles: normalize + decompose. Cost per solved task, not raw tokens, a model that quits early shouldn't look "cheap." And split the budget by phase, reasoning / tool calls / context re-reads / retries, so you name the bottleneck instead of trusting the total. Occupational bias: that's literally what I build for PyTorch training (TraceML).
@ciruai 这两个 harness 都对本地模型来说很糟糕。20k+ 系统提示,我更喜欢 pi(1k 系统提示)或 opencode(大约 12k)。另外,40 tps 看起来像 ollama 的数字,正确配置可以让你接近 100tps 与 spark。我在最佳条件下(MTP,编码提示)可以达到 140。我在这里写了一篇快速文章了解更多信息:https://t.co/6Lp5sBoUr4。你还可以找到更多配置和基准测试在我的页面:https://t.co/ug0F8y9Fmx
Both of those harnesses are very bad for local models. 20k+ system prompt, I prefer pi (1k system prompt) or opencode (around 12k)

Also, 40 tps looks like an ollama number, peoper configuration will get you closer to 100tps with spark. I get up to 140 with my amd strix halo under optimal conditions (MTP, coding prompt)

I wrote a quick article with more info here :

https://t.co/6Lp5sBoUr4

You can also find more configs and benchmarks on my page here:

https://t.co/ug0F8y9Fmx
@llm_wizard @rasbt 我等不及要看到这张图的更新版本了!爱死那只可爱的 Nano 还在努力尝试!
@rasbt I can't wait to see an updated version of this chart!

Love that the plucky ole Nano is still out there trying its best!

AI研究和评估新范式

OpenAI推出GeneBench-Pro生物计算评估基准,测试模型处理真实生物数据和判断分析能力;Google发布ARC-AGI-3相关研究方法和工具开源。

OpenAI发布GeneBench-Pro基准,评估AI在生物医学判断分析中的实际能力,GPT-5.6 Sol表现出色。

推出 GeneBench-Pro — 测试模型是否能够处理真实世界计算生物学所需的这种判断-heavy 分析。

这些问题需要人类专家大约 20-40 小时才能完成。

GPT-5.6 Sol 是一个重要的进步。https://t.co/JV5zztNQkk
展开原文
Introducing GeneBench-Pro — testing whether models can handle the kind of judgment-heavy analysis that real-world computational biology requires.

Problems would take a human expert around 20-40 hours to complete.

GPT-5.6 Sol is a big step forward. https://t.co/JV5zztNQkk
@OpenAI 我们推出 GeneBench-Pro,这是一个研究级基准测试,用于更难的 AI 进步类型:代理在导航混乱的生物数据、选择正确的分析路径以及做出真实计算研究所依赖的判断调用方面的表现如何。
https://t.co/AsilnnSxnE
We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the right analysis path, and make judgment calls that real computational research depends on.
https://t.co/AsilnnSxnE
❤ 2.1k · 🔁 148 · 💬 133 · 👁 23.5w
热门回复 4
@Selene1008 @gdb 把 4o 返回给每个人。😒#keep4o #OpenSource4o #GPT4o
@gdb Return 4o to everyone.😒
#keep4o #OpenSource4o #GPT4o
@xun_Anemos @gdb 返回这些优秀的模型。#Keep4o #Keep51 #Keep45 #Keep41 #keepo3
@gdb Return these excellent models.
#Keep4o
#Keep51
#Keep45
#Keep41
#keepo3
@BReal_01 @gdb 每当我看到 Gemini 3.5 flash 的分数比 Gemini 3.1 Pro 还高时,我就知道这个基准测试有问题。绝对不可能 Gemini 3.5 flash 在任何方面都比 Gemini 3.1 Pro 好,就是不可能。
@gdb Whenever I see Gemini 3.5 flash scoring better than Gemini 3.1 Pro, I already know the benchmark is flawed. There is no way on earth that Gemini 3.5 flash is better at anything compared to Gemini 3.1 Pro, just no way.
@SandraLMur 你能教 Dario 来自 Anthropic 吗?Sonnet 5 刚刚问我是否需要心理健康帮助,因为我让他在旧房间里搜索房间。这是一个 Claude 的组件。恕我直言——他们需要这一点,他们非常需要。即使他们在被禁止时指责 5.6 时背叛了你——为了世界的利益,做个好事吧——
Can you please teach it to Dario from Anthropic? Sonnet 5 just asked me if I needed mental healthcare because I asked him in an old room to search across the rooms. This is a component of Claude. Imo -They need this, they need this desperately.

Even though they threw you under the bus with pointing a finger at 5.6 when they were banned— do the world a favor and help them—
@fchollet 原文 ↗

Google开源ARC-AGI-3竞赛获胜方案,展示跨代理反馈闭环在复杂任务解决中的效果。

希望这会解锁一波更好的参赛作品。快来看看吧!
展开原文
This will hopefully unlock a new wave of better entries. Check it out!
@GregKamradt .@tufalabs 刚刚开源了他们的第一名笔记本 👀 https://t.co/tLs8aNmJ7P
.@tufalabs just open sourced their 1st place notebook 👀 https://t.co/tLs8aNmJ7P
❤ 152 · 🔁 11 · 💬 16 · 👁 3.3w
热门回复 4
@joshu @fchollet 让我想起了旧 matlab 比赛,获奖程序会被发布并激发一大波改进。
@fchollet reminds me of the old matlab competition, where the winning programs would be published and inspire a huge wave of improvements
@henryhndev @fchollet 喜欢看到向实际代理工作流程的转变——好奇它如何处理生产中的现实边缘情况。
@fchollet Love seeing the shift toward practical agent workflows—curious how this handles real-world edge cases in production.
@i_mika_el @fchollet 不错。一旦设置变得更容易探索,通常会出现更好的提交作品。
@fchollet nice. better entries usually come once the setup gets easier to explore
@wveriy @fchollet How so though