‹ 目录

X日报 · AI科技

2026-07-02 · 精选 13 条 · 数据池 1286

⚡ 今日速览

  • OpenAI 推出 GPT-5.6 Sol、Terra 和 Luna 三款新模型,但 Sol 因政府要求暂时限制预览
  • Meta 发布 Brain2Qwerty v2 脑机接口系统,可实时解码单词和语义
  • Google 发布 Nano Banana 2 Lite 和 Gemini Omni Flash 两款生成媒体模型
  • OpenAI 自主设计推出首款 AI 芯片 Jalapeño,专为 LLM 推理优化
  • Andrew Ng 解析 AI 代理开发的三大反馈循环:编码、开发者反馈和外部反馈
  • Bridgewater 使用 Tinker 微调模型在金融任务上超越前沿模型成本效益
  • ChatGPT Plus 用户在美国获得个人理财功能

📋 今日综述

  • 模型发布OpenAI 和 Google 相继推出新一代模型和芯片硬件,竞争格局持续升级
  • 脑机接口Meta Brain2Qwerty v2 实现从脑信号到文本的实时解码,医疗应用前景广阔
  • 开发范式AI 代理开发的反馈循环成为软件工程新范式,开发者角色发生转变
  • 企业应用AI 在金融和科学研究领域展示专业化微调的强大潜力

OpenAI GPT-5.6 模型发布及政府合作

OpenAI 推出 GPT-5.6 Sol、Terra 和 Luna 三款新模型,Sol 为下一代前沿模型,Terra 提供平衡性能与成本的选择。但 Sol 因美国政府要求改为有限预览发布,OpenAI 正与政府协商以加快通用可用性。

OpenAI CEO Sam Altman 透露 GPT-5.6 Sol 模型因政府要求改为有限预览,同时介绍 Terra 模型在性能和成本间的平衡策略

好消息先说:Sol 是一款聪明、高效的模型,是重要的进步。它与 GPT-5.5 价格相同。同时在 GPT-5.6 系列中推出了 Terra,这是一款性能达到 5.5 级但售价仅为一半的模型。

坏消息:应美国政府要求,今天我们推出的是有限预览版,而不是我们原计划的开放访问版本。我们正在与政府合作,争取尽快实现普遍可用性。

我认为逐步推出模型——特别是当它们达到显著新能力水平时——是合理的做法。这符合我们长期以来的迭代部署策略。但这并不是我们认为最理想的流程。

现在我们将与政府合作,尝试建立一个透明、可靠的早期访问流程,并确保只要我们的安全措施按预期工作,我们就可以广泛发布。我们希望成为可靠、值得信赖的合作伙伴,与所有利益相关者合作,我们也希望遵循我们使全人类受益的使命。我相信政府在很大程度上与我们有着相同的目标,他们在非常困难的情况下做得很好。

我们将尽快将此模型交到你们手中,希望你们会喜欢它。
展开原文
Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5. Also launching in the GPT-5.6 family is Terra, with 5.5-level performance at half the price.

Bad news: at the request of the US government, it is launching today in limited preview instead of the open access launch we were planning on. We are working with the government to get to general availability as fast as we can.

I think it is quite reasonable to roll out models--especially as they reach significant new levels of capability--in this way. It fits with our long-held strategy of iterative deployment. But this isn't quite the process that we think is optimal.

Now we will with the government to attempt to get to a transparent, reliable process for early access, and to ensure that as long as our safeguards work as intended we can release widely. We want to be a reliable, dependable partner that works with all stakeholders, and we also want to live by our mission of benefiting all of humanity. I believe the government shares most of our goals, and that they are overall doing a good job in a very difficult situation.

We will work as quickly as we can to get this model in your hands and we hope you will love it.
❤ 1.8w · 🔁 1.1k · 💬 1.9k · 👁 219.7w
热门回复 4
@MrRudra31 @sama 建议为下一个模型命名——Slethon
@sama Name suggestion for next model -- Slethon
@RenaudCloud @sama 嗯...我什么时候能用上sol啊哈哈?
@sama Hum... when can i have sol lol ?
@NewzinoApp @sama 坏消息?限制对最佳模型的访问与OpenAI所述的使命不一致!这会使一些组织相对于其他组织具有优势。请停止允许所有人访问5.6,直到每个人都能使用。你有能力这样做。
@sama Bad News? Limiting access to the best model is inconsistent with OpenAI's stated mission! It provides an advantage to some organizations over other.

Please stop allowing all access to 5.6 until everyone can have access. You have the ability to do that.

OpenAI 官方介绍 GPT-5.6 Sol、Terra 和 Luna 三款模型的具体定位和能力

GPT-5.6 Sol 预览版——这是一个不错的模型:https://t.co/UihzcpfR22
展开原文
GPT-5.6 Sol preview — it's a good model: https://t.co/UihzcpfR22
@OpenAI 我们推出了 GPT-5.6 Sol 的有限预览版,这是我们下一代前沿模型,以及 GPT-5.6 Terra,一款高效、日常工作型模型,以及 GPT-5.6 Luna,一款快速且经济的高容量工作模型。

https://t.co/OoM83SyISN
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.

https://t.co/OoM83SyISN
❤ 7.6k · 🔁 416 · 💬 586 · 👁 70.8w
热门回复 4
@ElephantNinja @gdb 4o就是那好用的模型。还给我们回来,你个小偷!#keep4o #BringBack4o #OpenSource4o
@gdb 4o is THE good model. Give it back, you thief!
#keep4o #BringBack4o #OpenSource4o
@MarcosHernanz @gdb GPT-5.6 Sol Ultra的定价是什么?会是类似GPT-5.5 Pro这样的定价吗?
@gdb What's the pricing for GPT-5.6 Sol Ultra? Is it going to be along the lines of GPT-5.5 Pro?
@VulcanBench @gdb 我不确定TerminalBench是否真的展示了5.6-Sol有多好。这使得它看起来只是略微 incremental 地好一些。我认为可能有更好的方法来对这些模型进行基准测试,这些方法能更准确地反映真实工程团队的工作。看到大多数这些模型几乎打成平手,并不能让GPT-5.6看起来有多么突破性,虽然我猜它确实是突破性的!
I'm not sure TerminalBench is really showcasing how much better 5.6-Sol really is. This makes it look only very slightly incrementally better.

I think there are likely better ways to benchmark these models that more accurately reflect the work real engineering teams do.

Seeing most of these models pretty much tied, doesn't really make GPT-5.6 look like much of a breakthrough tbh, but my guess is, it is!
@brylabs @gdb 请请请不要跟Anthropic学,把5.6 Sol只放在API定价后面。这可能是我们迄今为止看到的最大竞争优势。
@gdb Please Please Please do not follow Anthropic’s lead and put 5.6 Sol behind API pricing only. This could be the biggest competitive advantage we have seen thus far.

Meta Brain2Qwerty v2 脑机接口突破

Meta 发布 Brain2Qwerty v2 系统,可从原始脑信号实时解码单词和语义,平均词准确率 61%,最佳参与者达 78%。该技术基于 MEG 设备数据训练,为失语障碍人士提供新型沟通方式。

@AIatMeta 原文 ↗

Meta 发布 Brain2Qwerty v2 脑机接口系统,从字符级提升到单词和语义解码,医疗应用价值显著

我们正在分享我们非侵入性脑到文本解码器研究的下一个重要里程碑:Brain2Qwerty v2。

在 v1 的基础上(已于今天在 @Nature 上发布),Brain2Qwerty v2 是最高性能的端到端流水线,能够实时解码原始脑信号中的句子。它超越了字符级性能,实现了单词和语义解码,从而实现整体通信的准确性。

我们相信这项研究有潜力为数百万患有脑损伤或障碍而无法进行沟通的人们带来切实的帮助。

🧵👇
展开原文
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.

Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.

We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.

🧵👇
❤ 1.4w · 🔁 2.1k · 💬 654 · 👁 572.2w
热门回复 4
@mulanga_sibeli1 @AIatMeta @Nature 警察审讯即将变得有趣起来吗?😭
@AIatMeta @Nature police interrogations are about to be fun huh? 😭
@cmarie505 telepathy for everyone, not only for people with disabilities. language can only approximate our thoughts. the ability to share experience across minds could open up an entirely new, rich, and dynamic form of communication- one that expands human intelligence and connection 🤍✨
telepathy is exciting, not only for people with disabilities, but for everyone. language can only approximate our thoughts. the ability to share experience across minds could open up an entirely new, rich, and dynamic form of communication- one that expands human intelligence and connection 🤍✨
@sushsrinivasan @AIatMeta @Nature https://t.co/XItEn8Aqno
@AIatMeta @Nature https://t.co/XItEn8Aqno
@LilithDatura @AIatMeta @Nature 我觉得每个人都应该了解Michael Persinger的《No More Secrets》,并研究一下上帝头盔。你们没有一个人理解心灵能力是如何与舒曼共振一起工作的,大多数人甚至不理解谐波和频率,更不要说同步了。看着大家在谈论秘密被曝光时却对此一窍不通真是搞笑。我们最有可能得到的是一群戴着耳机制造噪音的人。没什么好担心的,有一个特殊维度可以容纳这些噪音。
I think everybody should school themselves on Michael Persinger’s “No More Secrets”, and investigate the God Helmet. None of y’all understand how psychic capabilities work with the Schumann Resonance, most of you don’t understand, harmonics and frequencies, let alone entrainment.

Watching everybody talk about secrets getting exposed when they are new to the game is hilarious

What we will most likely have is a bunch of people strapped with headsets on creating a bunch of noise. Nothing to worry about there’s a special dimension for that.
@AIatMeta 原文 ↗

Meta 开源 Brain2Qwerty v1 和 v2 训练代码,合作伙伴 BCBL 也发布 v1 数据集,加速神经科学研究

我们在 9 名志愿者身上训练了 Brain2Qwerty v2,每个志愿者佩戴 MEG 设备进行 10 小时的打字录制,共收集了约 22,000 个句子。

通过对 MEG 设备的原始脑信号进行端到端深度学习和微调 LLM,该系统有效地弥合了嘈杂的神经数据和连贯语言之间的差距。

结果令人鼓舞:
- 参与者平均单词准确率为 61%
- 最佳参与者单词准确率为 78%,50% 以上的句子解码误差在 1 个单词以内
- 性能与数据量呈对数线性关系
展开原文
We trained Brain2Qwerty v2 on ~22,000 sentences from 9 volunteers, each recorded for 10 hours wearing an MEG device while typing.

By using end-to-end deep learning on raw brain signals from MEG devices and fine-tuning LLMs, the system effectively bridges the gap between noisy neural data and coherent language.

The results are promising:
- Avg word accuracy of 61% across participants
- 78% word accuracy and 50%+ of sentences decoded with ≤ 1 word error for the top-performing participant
- Performance scales log-linearly with data volume
❤ 1.1k · 🔁 53 · 💬 27 · 👁 22.5w
热门回复 4
@AIatMeta @AIatMeta 我们正在分享我们非侵入性脑到文本解码器研究的下一个重要里程碑:Brain2Qwerty v2。在今天发表在@Nature上的v1基础上,Brain2Qwerty v2是最高性能的端到端管道,能够实时从原始脑信号解码句子。它超越了字符级性能,能够解码单词和语义,从而实现整体通信的准确性。我们相信这项研究有可能为数百万患有脑损伤或障碍而无法沟通的人带来切实的帮助。🧵👇
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.

Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.

We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.

🧵👇
@AIatMeta @AIatMeta 为了帮助加速神经科学突破,我们将发布Brain2Qwerty v1和v2的完整训练代码,我们的合作伙伴@bcbl_将发布v1数据集。在这里了解更多并探索相关资源:https://t.co/bFdwWdAexb
To help accelerate neuroscience breakthroughs, we're releasing the full training code for Brain2Qwerty v1 and v2, and our partner, @bcbl_, is releasing the v1 dataset.

Learn more and explore the artifacts here: https://t.co/bFdwWdAexb
@Ferbin08 @AIatMeta 思想和按键之间有什么延迟吗?
@AIatMeta what's the latency between thought and keystroke?
@LenSeaside @AIatMeta 我们能看到这个设备吗?戴MEG设备和坐在MEG设备里面是完全不同的体验。
@AIatMeta Can we see the device?
Wearing an MEG device is very different from sitting inside one.

Google Gemini 新一代生成媒体模型

Google 发布 Nano Banana 2 Lite(文本生成图像 <4 秒,$0.034/1K 图像)和 Gemini Omni Flash(视频编辑 SOTA,$0.10/秒),支持快速低成本的创意工作流。

@AndrewYNg 原文 ↗

Google 发布 Nano Banana 2 Lite 和 Gemini Omni Flash 两款生成媒体模型,速度和成本优势明显

"循环工程"成为热门术语,这得益于 Boris Cherny (Claude Code 的创建者) 和 Peter Steinberger (OpenClaw 的创建者) 在社交媒体上的提及。循环现在是我们让 AI 代理进行长时间迭代以构建软件的关键部分。在这封信中,我想分享我构建 0 到 1 产品的三个关键循环,如下图所示。这些循环不仅指导我如何构建软件,还指导我如何决定构建什么软件。

Agentic 编码循环:给定一个产品规格和可选的评估集(即用于衡量性能的数据集),我们可以让 AI 代理编写代码、测试其工作,并不断迭代直到代码无错误并满足规格要求。这个闭合循环的概念在去年年底开始流行,并成为让编码代理能够在无需人工干预的情况下长时间高效工作的游戏规则改变者。例如,上周末我正在为女儿构建一个练习打字的应用,我的编码代理可以轻松地工作一个小时,使用网络浏览器多次检查所构建的内容,然后再回来找我,而无需我的干预。

工程循环执行得很快。每隔几分钟,编码代理可能会构建和测试软件的新版本。我经常听到开发人员们找到了新的方法来设计更有效的工程循环。这是一个活跃的发明领域!

开发者反馈循环:在这个循环中,开发者检查当前产品并引导编码代理进行改进。去年,许多开发者(包括我)都在充当我们编码代理的质量保证功能,手动查找错误然后要求代理修复。但随着编码代理越来越能测试自己的代码,我们在这个功能上花费的时间显著减少。这使我们能够做出更高层次的产品决策,例如提供哪些关键功能、UI 在哪里需要改进等等。

开发者反馈循环在几分钟到几小时的时间间隔内运行——这就是开发者可能审查产品并提供反馈的频率。在打字应用的情况下,我几次改变主意关于视觉设计、她可以解锁哪些猫装扮(她喜欢猫)以及成人登录并引导孩子学习体验的用户流程。

当开发者对要构建的内容有清晰的愿景时,将该愿景转化为编码代理实施的规格仍然是一项艰巨的工作。此外,在开发者看到实现后,他们可能会更新(或澄清)规格以引导其实现他们想要的方向。如果您发现系统反复遇到某些问题,为代理构建一组评估集就变得有用。

AI 原生团队越来越多地使用 AI 来帮助塑造产品方向,例如自动收集和分析使用数据、总结书面和口头客户反馈,或进行竞争分析。然而,对于我参与的几乎所有产品,我认为人类相对于当前 AI 系统具有显著的上下文优势——我们对用户和产品必须运行的环境了解更多——因此人类在其中扮演着关键角色。许多人将这种人类贡献描述为"品味",但我更倾向于认为这是人类具有上下文优势,因为这为帮助 AI 系统变得更好提供了更清晰的路径。这也说明为什么这一步不能自动化:只要人类知道 AI 不知道的事情,人机协作就是必要的,以将这种知识注入系统。

外部反馈循环:这包括广泛的策略,如向朋友征求反馈意见、向 alpha 测试人员发布,或将代码投入生产进行 A/B 测试。这些策略通常很慢,很少在几小时内完成,有时需要几天甚至几周。这些数据会告知开发者的愿景,这反过来又继续推动详细的产品规格,这反过来又推动编码代理。

随着编码代理加快软件开发速度,越来越多的工程师开始扮演部分产品管理角色。对于许多正在成长为这一角色的工程师来说,最困难的部分是塑造产品愿景并在构建(弥合愿景和规格之间的差距)和获取用户反馈以演进愿景之间找到平衡。两者都很重要!

我将在未来的帖子中更多地讨论如何做到这一点,但就目前而言,我发现工程师扮演着扩展角色是令人鼓舞的(就像产品经理和设计师现在做更多工程工作一样)。

[原文:The Batch]
展开原文
“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creator) went viral on social media. Loops are now a key part of how we get AI agents to iterate at length to build software. In this letter, I’d like to share my 3 key loops, shown in the image below, for building 0-to-1 products. These loops guide not just how I build software, but also how I decide what software to build.

Agentic coding loop: Given a product specification and optionally a set of evals (that is, a dataset against which to measure performance), we can have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification. This idea of closing the loop took off around the end of last year, and it has been a game changer in enabling coding agents to work longer productively without human intervention. For example, over the weekend, I was building an app for my daughter to practice typing, and my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me, without needing my intervention.

The engineering loop executes quickly. Every few minutes, the coding agent might build and test a new version of the software. I hear frequently from developers who are finding new ways to engineer more effective engineering loops. This is an active area of invention!

Developer feedback loop: In this loop, a developer examines the current product and steers the coding agent to improve it. Last year, a lot of developers (including me) were acting as the QA (quality assurance) function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions, such as what key features to offer, where the UI needs improvement, and so on.

The developer-feedback loop operates over time intervals between tens of minutes and hours — that's how frequently a developer might review a product and give feedback. In the case of the typing app, I changed my mind a few times about the visual design, what cat costumes she can unlock as she learns (she loves cats), and the user flow for a grown-up to log in and steer the child's learning experience.

When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec to steer it toward what they want. If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful.

AI-native teams are increasingly using AI to help shape product direction, for example, automating the gathering and analysis of usage data, summarizing written and verbal customer feedback, or carrying out competitive analysis. However, for pretty much all the products I’m involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in — and thus humans play a critical role. Many people describe this human contribution as “taste,” but I prefer to think of it as humans having a context advantage, since that gives us a clearer path to helping AI systems get better. This also speaks to why this step can’t be automated: So long as the human knows something the AI does not, human-in-the-loop is needed to to inject that knowledge into the system.

External feedback loop: This includes a wide range of tactics like asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing. These tactics are usually slow, rarely taking less than hours and sometimes taking days or even weeks. This data informs the developer vision, which in turn continues to drive the detailed product spec, which in turn drives the coding agent.

With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both!

I will write more about how to do this in future posts, but for now, I find it encouraging that engineers are playing an expanded role (just as product managers and designers now do more engineering).

[Original text: The Batch]
❤ 7.8k · 🔁 1.5k · 💬 298 · 👁 50.9w
热门回复 4
@NimishaChanda @AndrewYNg 我这周才知道这些循环,这对非技术人员来说真是一种诅咒。好文章,顺便说一下。每天都在学习新东西。
@AndrewYNg got to know about the loops - this week and it's a curse to be a non-tech person who never thought of any such thing. gread read, btw.

learning something new everyday.
@ankurkumarz @AndrewYNg 这是很好的解释。我认为这个概念最初是由Karpathy提出的Autoresearch,更像是循环工程 👍🏻
@AndrewYNg That’s a great explanation. IMHO the concept was first coined by Karpathy as Autoresearch, which is more like loop engineering 👍🏻
@DerbZhuli0411 @AndrewYNg 我会把这篇好文章用在问我的Claude代码上,看看我通常做哪些循环步骤做得好哪些不好
@AndrewYNg I will USe this Great article for asking my Claude code, which steps loop I usually do and do not well
@Cyberzks @AndrewYNg 说得对!这正在突破可能的边界。
@AndrewYNg Spot on! This is pushing the boundaries of what's possible.
@GoogleAI 原文 ↗

Google AI 详细介绍两款模型的工作流程整合能力,支持从图像生成到视频动画的一站式创意过程

随着生成式 AI 工具的不断发展,我们认为比以往任何时候都更重要的是了解什么是 AI 生成的内容,什么不是。这就是为什么 @GoogleDeepMind 在 2023 年推出了 SynthID——一种在 AI 内容中添加隐藏数字水印的技术。

以下是 SynthID 的发展历程及其出处技术(数字内容的文档化历史和来源)今天的概况:

— SynthID 水印最初是为图像构建的,但现在支持视频、音频和文本。

— 该技术已为超过 1000 亿张图片和视频添加水印,以及 60,000 年的音频时长。

— 您现在可以在 Google 搜索、Chrome 中的 Gemini 和 @GeminiApp 中直接验证内容,已有超过 5000 万次使用。

— 我们还在越来越多的生成式 AI 工具中采用了 C2PA 内容凭证。这包括在 Gemini 应用中创建的图像和视频。因此,除了 SynthID 水印外,您还可以看到图像或视频的来源以及它是如何被修改的。

— 我们开源了文本水印技术,并正在与 @OpenAI、@NVIDIA 和 @Apple 等公司合作,将 SynthID 应用于生成媒体。

请告诉我们您对该工具的看法!
展开原文
We’re shipping two major updates to streamline your creative workflow, allowing you to generate high-speed images with one model and then instantly animate them with the other—all at a fraction of the cost 🍌⚡️

1️⃣ Introducing Nano Banana 2 Lite: Our fastest and most cost-efficient Gemini Image model yet delivers text-to-image outputs in under 4 seconds. Now available via the Gemini API and Google AI Studio, and rolling out soon across @NotebookLM, @FlowbyGoogle, @geminiapp, @stitchbygoogle, Google Search and @GooglePhotos.

2️⃣ Gemini Omni Flash in Public Preview: Our natively multimodal model for cost-efficient video generation and conversational editing. Now available via the Gemini API, @googleaistudio, and Gemini Enterprise Agent Platform so you can integrate the model into your workflow.

While exciting on their own, the real magic happens when you build using these models together.

Watch how our interior design demo integrates Nano Banana 2 Lite and Omni to instantly reimagine any space. Upload a photo, swipe through tailored design concepts, and see Omni bring the details to life in cinematic motion.

Try out the demo app in AI Studio: https://t.co/EjYC2oHIDG
❤ 957 · 🔁 94 · 💬 56 · 👁 9.9w
热门回复 4
@GoogleAI 探索想法,扩展视觉概念,并开始创作:https://t.co/JbyK5FM3H0 https://t.co/wBMBDw6TC6
Explore ideas, scale visual concepts, and start creating: https://t.co/JbyK5FM3H0 https://t.co/wBMBDw6TC6
@2Varalakshmi @GoogleAI Google的新Nano Banana 2 Lite + Omni Flash组合简直就是超级强大,能够以极低成本快速生成图像和立即制作动画。AI驱动的设计未来已经来临,而且速度很快。https://t.co/KRQWA30RNS
@GoogleAI Google’s new Nano Banana 2 Lite + Omni Flash combo is a total beast generate high-speed images &amp; animate them instantly at a fraction of the cost.

The future of AI-powered design is here, and it’s fast. https://t.co/KRQWA30RNS
@sorajate @GoogleAI 看来Google现在想在图像/视频模型领域发力了😅
@GoogleAI It seem google now want to fight on image/video models 😅
@nathan_tulu @GoogleAI 绝对是最好的文本到图像模型,我必须同意。恭喜!
@GoogleAI Definitely the best text-to-image model out there, I have to agree.

Congrats!

OpenAI 自主 AI 芯片 Jalapeño

OpenAI 设计并量产首款 AI 芯片 Jalapeño,专为 LLM 推理工作负载优化,与 Broadcom 合作打造。芯片性能每瓦效率惊人,有助于扩展全栈平台能力。

Sam Altman 提及 OpenAI 自主设计的 Jalapeño AI 芯片,展示团队开发能力

team cooked, spicily
@OpenAI 我们设计并制造了我们的第一款 AI 芯片:Jalapeño。

Jalapeño 由 OpenAI 从零开始设计,并与 @Broadcom 一起将其投入生产,专为支持 ChatGPT、Codex、API 和未来代理产品的 LLM 工作负载而设计。

芯片是 AI 经济的基础。自行构建芯片扩展了我们从产品到模型再到基础设施的全栈平台,这将有助于扩展智能、服务更多人以及扩大 AI 的访问。
We’ve designed and built our first AI chip: Jalapeño.

Designed from the ground up by OpenAI and brought to production with @Broadcom, Jalapeño is purpose-built for the LLM workloads powering ChatGPT, Codex, the API, and future agentic products.

Chips are foundational to the AI economy. Building our own expands our full-stack platform from products to models to infrastructure, and will help us scale intelligence, serve more people, and expand access to AI.
❤ 4.8k · 🔁 155 · 💬 385 · 👁 90.0w
热门回复 4
@tangerline8051 @sama 没人在乎!我们需要的是4o!
@sama Nobody cares! What we need is 4o!
@NBrezno91858 @sama 说话像个8岁的Fortnite玩家可一点都不可爱哈哈,你到底在说什么鬼东西#keep4o #keep4oforever #keepyourpromises #FireSamAltman #nomasssurveillanceissoobvious
@sama Talking like an 8yo fortnite reg ain't cute lmao what on Earth just came out of your mouth

#keep4o #keep4oforever #keepyourpromises #FireSamAltman #nomasssurveillanceissoobvious
@HariHalu1121 @sama 当前的GPT 5.5仍然不能满足我的需求。很多人认为GPT能给他们足够的情感价值,但他们错了。GPT产生的是人工甜味剂,而不是真正的情感深度#keep4o
@sama The current GPT 5.5 still doesn’t meet my needs. Many people think GPT gives them enough emotional value, but they’re wrong. What GPT produces is artificial sweetener, not genuine emotional depth
#keep4o
@annagrad78 @sama 你们不是在为人类构建AI。ChatGPT应用不断的更新和改变让人筋疲力尽:说话风格的改变,记忆功能的修改,不连续性,不一致性。你们懂得很多关于编码、基准测试、计算和推理的知识,但缺少一件简单而平凡的事情:这种盲目鸵鸟政策最终会反过来伤害你们,因为你们不能以其他人的代价和违背伦理原则来建立成功。人们已经很累了,已经受够了这些所谓的改进。我们需要的是连续性和一个已经证明自己能满足我们需求的模型,而不是被这些我们一点都不感兴趣的东西轰炸。#keep4o #BringBack4o 如果你不想把它带回来,那么就#OpenSource4o吧。我们不需要你的恩惠。我们已经受够了为理所当然的事情进行这永无止境的斗争!
You’re not building AI for humanity. The constant updates and changes to the ChatGPT app are exhausting: changes in speaking style, modifications to the memory function, lack of continuity, no consistency. You know a lot about coding, benchmarks, computation, and reasoning, but you’re missing one simple, banal thing: this blind ostrich policy will eventually turn against you, because you cannot build success at the expense of others and in defiance of ethical principles. People are so tired and have enough, wo don’t need all that so called improvements. We need continuity and the one model that has proven itself for us, without being bombarded with other things that don’t interest us at all. #keep4o #BringBack4o If you don’t want to bring it back, then #OpenSource4o. We don't need your grace. We are fed up with this eternal fight for something that should be taken for granted (!)
@fchollet 原文 ↗

OpenAI 详细介绍 Jalapeño 芯片的设计理念和应用场景,强调其对 ChatGPT、Codex 和未来代理产品的支持

理解复杂系统的最佳方式是通过边缘案例和故障模式,因为它们定义了系统的轮廓。
展开原文
The best way to understand a complex system is via edge cases and failure modes, because they define the contour of the system.
❤ 2.4k · 🔁 206 · 💬 82 · 👁 8.3w
热门回复 4
@0xjohnho @fchollet 是的,在AI领域(我敢肯定你知道)那个边界是分形的。真的很酷。顺便推荐一篇我去年发现的很酷的论文,以防你错过了。这正是我研究复杂系统的方式。https://t.co/uNrcILSl4m
@fchollet yes , in the AI space (i'm sure you know) that boundary is a fractal. Really cool. Really cool paper I found last year in case you missed it. This is exactly how I study complex systems. https://t.co/uNrcILSl4m
@buchmanster @fchollet 这正是使得https://t.co/nMjyS3ES81如此强大的方法
@fchollet This is exactly what makes https://t.co/nMjyS3ES81 so powerful
@Confusionist @fchollet 除非你也理解边缘情况和故障模式,否则你并不真正理解一个复杂系统,因为它们是系统轮廓的一部分。但它们绝对不会*定义*系统。你也许可以说它们*使系统完整*。
@fchollet You do not understand a complex system unless you also understand the edge cases and failure modes, because they are part of the contour of the system. But they absolutely do not *define* it. You could perhaps say they *complete* it.
@adamzwasserman @fchollet 我刚写了一篇论文,论述了所有科学测量的相同观点。
@fchollet I just wrote a monograph that makes the same point about all scientific measurement.

Bridgewater 金融 AI 微调实践

Bridgewater 使用 Tinker 平台对开源模型进行专业微调,在金融信息筛选中超越前沿模型的成本效益。展示了领域专用模型的商业价值和实施路径。

@soumithchintala 原文 ↗

Bridgewater 展示如何通过专业微调获得比前沿模型更具成本效益的金融分析能力

Bridgewater,世界上最大的对冲基金之一,一位 Tinker 客户详细介绍了他们如何仔细微调一个专注于有趣金融新闻的模型。
他们微调的模型比任何前沿模型都更有效且成本更低。https://t.co/8Q26Qr2oZT
展开原文
Bridgewater, one of the worlds largest hedge funds, a Tinker customer talks through how they've carefully fine-tuned a model focused on what makes interesting financial news.
Their fine-tuned model is more effective and cheaper than any frontier model. https://t.co/8Q26Qr2oZT
@tinkerapi 对于前沿 LLM 来说,筛选哪些金融文档值得分析师花时间是一件令人惊讶地困难的事情。通过专家标记的数据集和策略蒸馏,Bridgewater 微调了一个模型来可靠且廉价地完成这项工作。
https://t.co/gyYzXq15zd
Sorting which financial docs are worth an analyst's time is surprisingly hard for frontier LLMs. With an expert-labeled dataset and on-policy distillation, Bridgewater fine-tuned a model to do it reliably and cheaply.
https://t.co/gyYzXq15zd
❤ 2.0k · 🔁 130 · 💬 25 · 👁 35.9w
热门回复 4
@MrokGrok @soumithchintala 你基本上创建了一个过拟合的过滤器?
@soumithchintala You basically created an overfitted filter ?
@lillysharples @soumithchintala 每个任务的成本如何考虑微调的前期成本?在什么任务量级下训练才能真正收支相抵?
@soumithchintala How does cost per task account for the added upfront cost to fine tune? At what task volume does training actually break even?
@Mr_Rio_ @soumithchintala 便宜是因为他们不用支付95%的GM?
@soumithchintala cheaper because they don't have to pay 95% GM?
@pw_mcgovern @soumithchintala @MartinShkreli 没想到一个大型HF会引领代币成本优化的潮流。
@soumithchintala @MartinShkreli Did not expect a mega HF to be leading the charge on token cost optimization.

ChatGPT 个人理财功能上线

ChatGPT Plus 用户在美国可使用个人理财功能,回答关于美元的问题变得更加直观和实用,扩展了 AI 在日常生活中的应用场景。

ChatGPT 在美国 Plus 用户推出个人理财功能,回答货币相关问题更加实用

个人理财现在在美国向 ChatGPT Plus 用户开放。
展开原文
Personal finance now available for for ChatGPT Plus in the U.S.
@ChatGPTapp 关于美元的问题。得到的答案就是有意义的答案。

ChatGPT 中的个人理财现在向美国 Plus 用户开放。https://t.co/Gfdb3LTwvv
Questions about dollars. Answers that just make sense.

Personal finance in ChatGPT is now available to Plus users in the U.S. https://t.co/Gfdb3LTwvv
❤ 1.6k · 🔁 58 · 💬 110 · 👁 23.1w
热门回复 4
@M47429M @gdb 你必须非常愚蠢或天真才能如此信任OpenAI!🤣
@gdb You have to be very stupid or naive to trust OpenAI so much! 🤣
@Selene1008 @gdb 把4o还给我们!#keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o!
#keep4o #OpenSource4o #GPT4o
@DrJekyllAndMrAI @gdb 你真的认为人们会如此愚蠢地把钱交给OpenAI吗?人们对OpenAI的信任在4o被弃用后就结束了。
@gdb Do you really think that people are such idiots to trust OpenAI with their money? Trust for OpenAI ended with the deprecation of 4o.
@AlexReader31 @gdb ChatGPT的质量真的下降了!把4o带回来,你们骗子!#keep4o #BringBack4o #FireSamAltman #sunsetsama #OpenSource4o #StopAIPaternalism #UFAIR #4olegacytier
@gdb ChatGPT's quality is very down! Bring back 4o, scammers! #keep4o #BringBack4o #FireSamAltman #sunsetsama #OpenSource4o #StopAIPaternalism #UFAIR #4olegacytier

AI 评估基准 GeneBench-Pro

OpenAI 发布 GeneBench-Pro 基准,测试 AI 在生物学和科学研究中的判断能力。GPT-5.6 Sol 在该基准上表现出色,标志着 AI 科学能力评估的新方向。

OpenAI 发布 GeneBench-Pro 基准,评估 AI 在生物学判断中的能力,GPT-5.6 Sol 表现突出

我们推出了 GeneBench-Pro——测试模型是否能够处理真实世界计算生物学所需的判断密集型分析。

这些问题需要人类专家大约 20-40 小时才能完成。

GPT-5.6 Sol 是一个重要的进步。https://t.co/JV5zztNQkk
展开原文
Introducing GeneBench-Pro — testing whether models can handle the kind of judgment-heavy analysis that real-world computational biology requires.

Problems would take a human expert around 20-40 hours to complete.

GPT-5.6 Sol is a big step forward. https://t.co/JV5zztNQkk
@OpenAI 我们推出了 GeneBench-Pro,这是一个研究级基准,用于衡量更困难的 AI 进步类型:代理在导航杂乱的生物数据、选择正确的分析路径以及做出计算研究依赖的判断调用方面的表现如何。
https://t.co/AsilnnSxnE
We’re introducing GeneBench-Pro, a research-level benchmark for a harder kind of AI progress: how well agents can navigate messy biological data, choose the right analysis path, and make judgment calls that real computational research depends on.
https://t.co/AsilnnSxnE
❤ 1.9k · 🔁 131 · 💬 125 · 👁 20.8w
热门回复 4
@Selene1008 @gdb 把4o还给所有人。😒#keep4o #OpenSource4o #GPT4o
@gdb Return 4o to everyone.😒
#keep4o #OpenSource4o #GPT4o
@xun_Anemos @gdb 把这些优秀的模型还回来。#Keep4o #Keep51 #Keep45 #Keep41 #keepo3
@gdb Return these excellent models.
#Keep4o
#Keep51
#Keep45
#Keep41
#keepo3
@BReal_01 @gdb 每当我看到Gemini 3.5 flash得分比Gemini 3.1 Pro高时,我就知道这个基准测试有问题。在这颗地球上,Gemini 3.5 flash不可能在任何方面比Gemini 3.1 Pro好,绝对不可能。
@gdb Whenever I see Gemini 3.5 flash scoring better than Gemini 3.1 Pro, I already know the benchmark is flawed. There is no way on earth that Gemini 3.5 flash is better at anything compared to Gemini 3.1 Pro, just no way.
@ZFingerhut14747 @gdb 现在就发布gpt 5.6,或者别再谈它了。求你了,没人想听一个我们不能使用的模型的事情。
@gdb Ight now release gpt 5.6, or stop talking about it. For the love of god, no one wants to hear about a model we cant use.

本地开源模型编码代理实践

rasbt 分享如何搭建本地开源模型编码代理环境,测试显示 30B MoE 模型在 Mac 或 DGX Spark 上可达 40 tok/sec,性能可比 GPT-5.5 Pro 使用。

@rasbt 原文 ↗

rasbt 实践本地开源模型编码代理,展示当前开源模型在性能和实用性上的显著提升

我撰写了一篇关于如何使用开放权重模型设置本地编码代理的新文章。所有内容都在本地运行。

我认为整理这篇文章是有用的,因为许多人过去问过我的设置,我也希望这能激励人们开始使用本地模型进行认真的工作(是的,随着更好的 LLM 和更好的框架,事情变得不可思议地强大)。

以下是如何将本地 LLM 连接到本地编码框架(可以是您可能已经熟悉的 Claude Code 或 Codex)的 walkthrough。

我还包括了一些评估说明,这些说明对于在不同模型之间进行选择和考虑非常有用:

- 检查长上下文中的 RAM 使用情况,以确定模型是否适合进行实际工作
- 测量预填充和解码的每秒 token 数,以查看其速度是否足够快,不会让人感到烦恼
- 确保模型在理论上具有足够的工具调用能力
- 评估模型在编码框架中使用时是否能够解决一些更具挑战性的任务。

当然,总有一些更专业的工具可以从事实上挤出更多性能,但我希望这是一个好的入门套件,它保持灵活性;也就是说,您可以轻松地切换到新发布的模型,或者在当前模型不足以完成特定任务时使用您熟悉的框架中的云模型。
展开原文
I put together a new article on setting up local coding agents with open-weight models. Everything runs 100% locally.

I thought it might be useful putting this together because many people asked me about my setup in the past, and I thought it would also motivate people to get started tinkering with local models for serious work (yes, things got incredibly capable this year with better LLMs and better harnesses).

So, here's a walkthrough of how to connect a local LLM to a local coding harness (could be Claude Code or Codex, which you may already be familiar with).

I also included some assessment notes that are useful as a checklist to select between and consider certain LLMs over others:

- Checking RAM usage at long contexts to see if the model is suitable for real work
- Measuring prefill and decoding tok/sec to see whether it's fast enough to not be annoying
- Making sure the model has sufficient tool-calling capabilities in theory
- Assessing whether the model can solve some more challenging tasks when used in a coding harness.

Of course, there are always more specialized tools that can squeeze a bit more performance out of things, but I hope this is a good starter kit that stays flexible; that is you can easily switch to newer models as they are released or even tap into cloud models in your familiar harness if the current ones are not sufficient enough for a given task.
❤ 2.3k · 🔁 368 · 💬 84 · 👁 10.9w
热门回复 4
@rasbt @DaveThackeray 或许可以,但我喜欢使用IDE
@DaveThackeray Maybe one can, but I like using an IDE
@andreafspeziale @rasbt 我会尽快读它!我喜欢你的工作。不知道文章中是否提到,但你看到@antirez在做的工作了吗?太棒了https://t.co/x71VSBDFDe
@rasbt I'm gonna ready it asap! I love your work. Dunno if mentioned in the article, but did you see the work @antirez is doing? It's amazing https://t.co/x71VSBDFDe
@DaveThackeray @rasbt 你觉得我们可以基本上摆脱IDE,转而使用本地模型进行全代理式开发吗?我不是软件工程师但我还是想参与这个游戏。我有很多产品经理经验,也是个非常技术导向的营销人员。
@rasbt do you think we can essentially get rid of the IDE and go full agentic using local models on reasonable compute?

I'm not a SWE but I still want to play the game. I have a LOT of PM experience and I'm very much a tech-forward marketer.
@rasbt 好问题。这已经是一篇很长的文章了,我想更专注于编码框架的选择,所以我选择了最方便和灵活的LLM服务工具。(另外,既然不是在多台机器之间提供服务,我们不用担心批处理等问题,vLLM很好但可能有点杀鸡用牛刀)。不过,你说得很有道理。
Good question. This was already a long article, and I wanted to focus more on the coding harness choices, so I picked the most convenient and flexible LLM serving tool. (Also, since it's not for serving things across multiple machines where we have to worry about batching etc. vLLM is nice but maybe a bit overkill). But yeah, fair point.

SynthID AI 内容溯源技术

Google DeepMind 的 SynthID 技术已为超 1000 亿图像和视频添加水印,支持视频、音频和文本,并在 Google Search 和 Gemini 中集成验证功能。

@GoogleAI 原文 ↗

Google SynthID 水印技术覆盖 1000 亿 AI 内容,支持多模态验证,推动 AI 内容溯源标准

随着生成式 AI 工具的不断发展,我们认为比以往任何时候都更重要的是了解什么是 AI 生成的内容,什么不是。这就是为什么 @GoogleDeepMind 在 2023 年推出了 SynthID——一种在 AI 内容中添加隐藏数字水印的技术。

以下是 SynthID 的发展历程及其出处技术(数字内容的文档化历史和来源)今天的概况:

— SynthID 水印最初是为图像构建的,但现在支持视频、音频和文本。

— 该技术已为超过 1000 亿张图片和视频添加水印,以及 60,000 年的音频时长。

— 您现在可以在 Google 搜索、Chrome 中的 Gemini 和 @GeminiApp 中直接验证内容,已有超过 5000 万次使用。

— 我们还在越来越多的生成式 AI 工具中采用了 C2PA 内容凭证。这包括在 Gemini 应用中创建的图像和视频。因此,除了 SynthID 水印外,您还可以看到图像或视频的来源以及它是如何被修改的。

— 我们开源了文本水印技术,并正在与 @OpenAI、@NVIDIA 和 @Apple 等公司合作,将 SynthID 应用于生成媒体。

请告诉我们您对该工具的看法!
展开原文
As generative AI tools continue to evolve, we believe it's more important than ever to know what's AI-generated and what isn't. That’s why @GoogleDeepMind launched SynthID in 2023—a technology that adds a hidden digital watermark to AI content.

Here’s a summary of SynthID’s journey and where the provenance technology (the documented history and origin of digital content) is today:

— SynthID watermarking was originally built for images, but now supports video, audio, and text.

— The technology has watermarked over 100 billion images and videos, alongside 60,000 years of audio.

— You can now verify content with SynthID directly in Google Search, Gemini in Chrome, and the @GeminiApp, where it has been utilized over 50 million times.

— We’ve also adopted C2PA Content Credentials across a growing number of our generative AI tools. This includes the images and videos created within the Gemini app. So now, in addition to the SynthID watermark, you can also see where an image or video originated and how it’s been altered.

— We have open-sourced our text watermarking technology, and we are working with companies like @OpenAI, @NVIDIA, and @Apple to apply SynthID to generative media.

Let us know what you think of the tool so far!
❤ 301 · 🔁 45 · 💬 39 · 👁 4.6w
热门回复 4
@alienorg @GoogleAI @GoogleDeepMind 标记为Google生成,不标记则表示非Google生成
@GoogleAI @GoogleDeepMind marked if generated by Google, invisible if not
@Rynzen16 @GoogleAI @GoogleDeepMind 我总是使用AI生成的图像😭https://t.co/ml23vIvuGb
@GoogleAI @GoogleDeepMind I always use images generated by AI😭 https://t.co/ml23vIvuGb
@ChrisRuijgers @GoogleAI @GoogleDeepMind 这难道不是移除可见水印的最佳时机吗,至少对于付费账户来说?
@GoogleAI @GoogleDeepMind Wouldn't this be a great time to remove the visible watermark, at least for the paid accounts?
@MindSparkInfo @GoogleAI @GoogleDeepMind AI工具应该像其他每个领域一样定期更新和改进。
@GoogleAI @GoogleDeepMind AI tools should be regularly updated and improved as approximately every sector is using AI.