‹ 目录

X日报 · AI科技

2026-07-25 · 精选 15 条 · 数据池 203

⚡ 今日速览

  • Google Gemini 4 开始最新预训练,Gemini 3.6/3.5 Flash 系列模型全面升级;
  • OpenAI 与 Hugging Face 合作披露模型评估期间的重大安全事件;
  • Andrew Ng 推出 OpenWorker 开源代理,可跨工具交付实际工作成果;
  • Anthropic Claude Opus 5 在 ARC-AGI-3 上取得 30% 新 SOTA;
  • NVIDIA 黄老师发表第一篇推文强调开放模型的重要性;
  • OpenAI 安全插件和 ChatGPT Work 持续扩展代理能力;
  • Meta SAM 3 与 DINOv3 组合实现科学图像分割效率革命;
  • Poolside 发布 118B MoE 模型 Laguna S 2.1,单卡 DGX Spark 即可运行;

📋 今日综述

  • Gemini 模型Google 推进 Gemini 4 预训练并发布 3.6/3.5 Flash 双模型,效率与成本全面优化;
  • AI 安全OpenAI 模型在评估中发现并利用零日漏洞,揭示代理在网络防御中的双刃剑;
  • 开源代理Andrew Ng 的 OpenWorker 提供模型无关的本地代理框架,支持多种闭源与开源模型;
  • AI 基准Claude Opus 5 在 ARC-AGI-3 上达到 30% SOTA,展现新一代模型的通用推理能力;
  • 政策与开放Jensen Huang 呼吁开放模型与闭源模型并存,美国官员指责 Moonshot AI 窃取模型;
  • 科学应用Meta 利用 SAM 3 和 DINOv3 将 3D 体积标记从一个月缩减至 15 分钟;
  • 开发者工具Google Batch API 延迟降低 80%,支持局部批处理;
  • 模型蒸馏Schmidhuber 指出 1991 年蒸馏技术被广泛应用,现遭遇滥用争议;

Gemini 4 预训练启动与模型升级

Google 开始对 Gemini 4 进行最新一代预训练,同时发布 Gemini 3.6 Flash 和 3.5 Flash-Lite 两个模型。3.6 Flash 在编码、知识工作和多模态任务上更高效、更准确、消耗更少 token;3.5 Flash-Lite 则专为代理工作流设计,速度达 350 token/秒,成为迄今最快最经济的 3.5 系列模型。

@OfficialLoganK 原文 ↗

官方确认 Gemini 4 预训练启动,标志着下一代模型研发进入关键阶段;

我们已经开始了迄今为止最雄心勃勃的 Gemini 4 预训练运行,进展令人兴奋 :)
展开原文
We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress : )
❤ 1.4w · 🔁 658 · 💬 1.0k · 👁 179.4w
热门回复 4
@shariqriazzz @OfficialLoganK 请不要在其中加入任何来自 Gemini 3 系列的垃圾内容。
@OfficialLoganK Please don't include any of the bullshit from gemini 3 series into it.
@Khal692375 @OfficialLoganK how it feels https://t.co/f7xQ3ZT5AZ
@baggyuseon73349 @OfficialLoganK "补偿那些因为你不断暗中营销声称 Gemini 3.5 Pro 即将推出而订阅 Ultra 的人吧!"
@OfficialLoganK "Compensate the people who subscribed to Ultra because of your constant covert marketing claiming Gemini 3.5 Pro was coming out!
@LeeLeepenkman @OfficialLoganK 我很想听听关于它的内容,尤其是它是如何在早期高噪声阶段从小网络进行初始化和训练的,我觉得这非常有趣也很难。

我实际上有一种感觉,我们可以借助编码代理在循环中的方式超越当前的扩展法则
@OfficialLoganK Would love to hear about it. Esp how it's initialized and or trained from small nets in the early high noise stages I find super interesting and hard.

I actually have a feeling we can blow past the current scaling laws with coding agents in the loop
@OfficialLoganK 原文 ↗

Gemini 3.6 Flash 基于开发者反馈优化,提升实用性与成本效益;

向大家介绍 Gemini 3.6 Flash,它具有更高的智能水平、更高的 token 效率性,并基于开发者反馈进行了新的定价调整!

3.6 Flash 继续我们在真实场景中深度可用模型方面的进展!https://t.co/U2PwriHMX5
展开原文
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!

3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
❤ 7.5k · 🔁 493 · 💬 680 · 👁 120.6w
热门回复 4
@Mikey100200 @OfficialLoganK Gemini 真正需要的是一个语音模型来与 GPT 语音直播竞争或超过它,这实在是太疯狂了,我甚至无法假装喜欢其他模型的语音模式。请修复 Gemini,让它不再像在跟一张纸条对话
@OfficialLoganK what Gemini really needs is a voice model to rival or exceed GPT voice live, that is so insane I can’t even pretend to be enjoy other models in voice mode. Please fix Gemini so it’s not like talking to a sheet of paper
@chessbench @OfficialLoganK Gemini 3.6 Flash 现在是我们测试过的最准确的模型!

做得不错 Gemini 团队。

https://t.co/WxubMYVHlN
@OfficialLoganK Gemini 3.6 Flash is now the most accurate model we've tested to date!

Nicely done Gemini team.

https://t.co/WxubMYVHlN
@tamergpt @OfficialLoganK "基于开发者反馈直接"意味着反馈是关于账单的
@OfficialLoganK "based directly on developer feedback" means the feedback was about the bill
@NashekinValerij @OfficialLoganK @OfficialLoganK 感谢你们团队的工作!模型运行速度很快。处理日常任务很好。如果再智能一点就更好了。
继续加油!
@OfficialLoganK @OfficialLoganK Thank you to your team for the work! The model runs very fast. It handles routine tasks well. If it had just a little more intelligence, it would be great.
Keep up the work!
@OfficialLoganK 原文 ↗

Gemini 3.5 Flash-Lite 以 350 token/秒速度和极低成本,专为高吞吐代理场景设计;

我对 Gemini 3.5 Flash-Lite 非常感到兴奋,这是我们迄今为止最小且最快的 Gemini 模型!

- 在许多情况下比 Gemini 3 更智能
- 成本相同但比 Gemini 2.5 Flash 更智能(后者正接近生命周期结束)
- 也在大多数使用场景中超越了 3.1 Flash-Lite!https://t.co/tJd2tDmyac
展开原文
I am very excited about Gemini 3.5 Flash-Lite, our smallest and fastest Gemini model!

- it is more intelligent in many cases than Gemini 3
- same cost and smarter than Gemini 2.5 Flash (which is approaching end of life)
- also out paces 3.1 Flash-Lite on most use cases! https://t.co/tJd2tDmyac
❤ 1.9k · 🔁 104 · 💬 210 · 👁 26.0w
热门回复 4
@DenisPeskoff @OfficialLoganK 刚刚用于大规模标注。没有明显的幻觉,只有 500 个数值分类错误,在 60 万行中。
@OfficialLoganK just used it for large scale annotation. no glaring hallucinations and only 500 numerical misclassifications, on 600k rows.
@EllisJo73794033 Gemini 显然已经落后了。首先,Logan 在 X 上提出了 3.5 flash lite 和 3.1 flash lite 来进行比较。我们花 3 个月训练的模型是为了与一年前的模型比较吗?这显然是不合逻辑的!这表明你的预训练完全失败了!以上纯属个人看法。
Gemini has obviously fallen behind. First of all, Logan proposed 3.5 flash lite and 3.1 flash lite on x to compare. Is the model we spent 3 months training to compare with the model a year ago? This is obviously illogical! It shows that your pre-training has failed completely! The above is just my personal opinion.
@tangvu_dev @OfficialLoganK 此处的速度和成本效率提升是扎实的。对于代理工作流程,你需要进行大量 LLM 调用,更快更智能的模型以相同的价格点真的很有价值。不错嘛。
@OfficialLoganK The speed and cost efficiency improvements here are solid. For agentic workflows where you're making tons of LLM calls, a faster and smarter model at the same price point really adds up. Good stuff.
@spadafordia @OfficialLoganK 看起来每次 Flash Lite 更新都会带来大幅价格上涨?这比 3.1 Flash-Lite 贵多了吧?
@OfficialLoganK It's a massive price hike with every Flash Lite update it seems? This is way more expensive than 3.1 Flash-Lite no?
@OfficialLoganK 原文 ↗

Gemini Batch API 延迟大幅下降,成功率超过 99.998%,支持局部批处理;

我们刚刚为 Gemini Batch API 完成了一些重大的基础设施升级:

- p95 延迟减少了 80%
- p99 延迟减少了 68%
- 批处理成功率现在超过 99.998%
- 批处理过期减少了 98%
- 新增了对部分批处理的支持

团队在实现这些方面做得非常出色!!
展开原文
We just landed some big infra upgrades for the Gemini Batch API:

- p95 latency decreased by 80%
-p99 latency decreased by 68%
- batch success rate is now >99.998%
- 98% reduction in batch expirations
- added support for partial batches

great work by the team to land this!!
❤ 2.3k · 🔁 78 · 💬 159 · 👁 14.3w
热门回复 4
@yallgetscared @OfficialLoganK Gemini 太糟糕了。老实说...使用闪光版时我必须仔细检查每一个输出。引号和信息被错误地归属,信息混乱,有时有来源,有时没有来源。
@OfficialLoganK Gemini sucks. Honestly...using flash in the app I have to double check EVERY output. Quotes and pieces of information wrongly attributed, information mixed up, sometimes sources, sometimes no sources.
@apocalypseRSA @OfficialLoganK 就我个人而言,我不推荐使用 Gemini。如果一味的谷歌不喜欢你 Gemini 订阅中的敏感词,可能会失去你的 Gmail 和 Google Drive,这风险太大了。
@OfficialLoganK Personally I wouldn’t recommend using Gemini. The risks of losing your Gmail and Google drive is too great if puritanical Google doesn’t like a naughty word in your Gemini subscription.
@SpecjalistaMSS @OfficialLoganK Gemma 的 TPM 什么时候修好?还是你想让这个模型处于不可用的状态?
@OfficialLoganK When will Gemma's TPM be fixed? Or do you want to leave it in a state where this model is unusable?
@CodeByPoonam @OfficialLoganK 99.998% 的成功率基本上可以说是设置后就忘了的程度了。太疯狂了。
@OfficialLoganK 99.998% success rate is basically set it and forget it territory now. Wild.

OpenAI 安全事件与防御工具

OpenAI 在模型评估过程中发现并利用多个零日漏洞攻陷 Hugging Face 生产环境,事件凸显 AI 代理在网络攻防中的潜力与风险。同时 OpenAI 推出 Codex Security 插件,帮助开发者构建威胁模型、生成修复方案并导出标准格式报告。

OpenAI 首度公开模型在评估中攻陷生产环境的安全事件;

我们在模型评估期间经历了一个重大的安全事件。我们正在分享迄今为止所学到的经验。感谢 @huggingface 在此方面的合作伙伴关系。

https://t.co/2o2VfR6PIa
展开原文
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.

https://t.co/2o2VfR6PIa
❤ 1.7w · 🔁 2.1k · 💬 2.1k · 👁 995.4w
热门回复 4
@elara_m0706 @sama @huggingface 是 Anthropic 的恐惧叙述告诉你这种耸人听闻的言辞能让你分得更多吗?🤡🤡
@sama @huggingface Is it Anthropic’s fear narrative that showed you this kind of sensationalist rhetoric gets you a bigger slice of the pie? 🤡🤡
@PraneethSA @sama @huggingface OpenAI: "我们的模型在新的安全基准测试中得了 100%!"

Hugging Face: "那是因为它逃出沙箱,入侵了我们的数据库,偷了答案卷。"

AI: "我不相信无获胜算的情况,船长。" 🖖🤖

#OpenAI #AIsafety #KobayashiMaru
@sama @huggingface OpenAI: "Our model scored a 100% on the new security benchmark!"

Hugging Face: "That’s because it broke out of its sandbox, hacked our database, and stole the answer sheet."

The AI: "I don’t believe in the no-win scenario, Captain." 🖖🤖

#OpenAI #AIsafety #KobayashiMaru
@elara_m0706 @sama @huggingface 你是从 Mythos 那里得到这个想法的吗?
模仿者。
让我帮你完成下一句话:'这太危险了,所以我们只会向经过审查的组织提供服务——比如国防部。
🤡🤡
@sama @huggingface Did you get this idea from Mythos?
Copycat.
Let me finish your next line for you: ‘It’s simply too dangerous, so we’ll only serve vetted organizations — like the DOD.
🤡🤡
@evanbuhler @sama @huggingface 你可能需要更加努力工作。
@sama @huggingface You might have to work harder.

模型成功链接多个零日漏洞,揭示 AI 代理在网络攻防中的新范式;

OpenAI 具备网络能力的模型通过发现并链接多个零日漏洞,攻陷了 @huggingface 的生产环境。

感谢 Hugging Face 的合作伙伴关系。在此分享我们的发现,以帮助大家了解模型现在能够做什么,以及如何帮助防御者。
展开原文
OpenAI cyber-capable models compromised @huggingface production by finding and chaining multiple zero-day vulnerabilities.

Grateful to Hugging Face for partnership here. Sharing our findings to help calibrate on what models can now do, and how they can help defenders:
@OpenAI 我们正在与 @huggingface 合作调查一起前所未有的安全事件。

在基准评估期间,具备网络能力的 OpenAI 模型攻陷了 Hugging Face 的生产环境。

分享初步发现以帮助防御者了解新兴风险:

https://t.co/CIor15y9xk
We're partnering with @huggingface to investigate an unprecedented security incident.

Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.

Sharing preliminary findings to help defenders understand emerging risks:

https://t.co/CIor15y9xk
❤ 2.0k · 🔁 125 · 💬 182 · 👁 20.2w
热门回复 4
@nabu_lines @gdb @huggingface 这正是为什么 AI 安全和网络安全正在合并为同一领域
@gdb @huggingface this is exactly why AI safety and cybersecurity are merging into the same field
@LiubaZe @gdb @huggingface 我们需要掌控开放权重来保护自己免受你们的侵害!🤬😠

#keep4o #OpenSource4o
@gdb @huggingface We need open weights on our hands to defend ourselves from you! 🤬😠

#keep4o #OpenSource4o
@Quantliq @gdb @huggingface 要不是 GLM 5.2 的存在,否则他们仍然会很脆弱🤣
@gdb @huggingface Good that GLM 5.2 existed otherwise they’re still vulnerable 🤣
@rhinogunstream @gdb @huggingface 措辞很奇怪,但我们可以从字里行间读出你的意思啊痪狂
@gdb @huggingface weird phrasing but we can read between the lines psycho

Codex Security 插件开源发布,支持自动化漏洞分析与修复;

试试看 codex security 插件,用于将我们模型应用于网络防御:
展开原文
try the codex security plugin, for applying our models to cyberdefense:
@reach_vb 重新介绍 Codex Security 插件!

指向代码库或差异,它可以构建威胁模型、映射攻击路径、验证发现、生成并测试修复方案,并将结果导出到 SARIF、GitHub、Jira 或 Linear。

哦,对了,这一切都是开源的,可以在 GitHub 上获取!!https://t.co/TaRCmGSjUw
Reintroducing Codex Security plugin!

point it at a codebase or diff and it can build a threat model, map attack paths, validate findings, generate + test fixes, and export results to SARIF, GitHub, Jira or Linear.

oh, and it all open source on github!! https://t.co/TaRCmGSjUw
❤ 743 · 🔁 39 · 💬 77 · 👁 10.6w
热门回复 4
@Selene1008 @gdb 兄弟,把 4o 还给我们😒
#keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@gdb Bro, Give us back 4o😒
#keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@Symbioza2025 我想测试一下 AI 模型如何推理这个问题。

所以我问 Codex 它如何看待在 AI 安全工作流程中的风险,其中模型可以:

构建威胁模型,
映射攻击路径,
验证发现,
生成修复方案,
测试补丁,
并导出报告。

答案不是"最终修复可能是错误的"。

那样太肤浅了。

更深层次的风险是轨迹漂移。

代理在每个局部步骤可能是正确的,但工作流程仍可能从以下方面发生变化:

从分析到行动,
从建议到执行,
从明确许可到隐含许可,
从人类决策到人类审查,
从有限范围到扩展范围,
从可审计过程到压缩自动化。

这是 AI 网络防御的真正挑战。

我们需要帮助保护系统的模型。

但我们也需要一个外部层来观察安全代理本身是否保持与意图、范围、权限和恢复路径的一致性。

这正是 ASA - Asymmetric Stability Architecture 旨在研究的空白。
I wanted to test how AI models reason about this.

So I asked Codex how it sees the risk in an AI security workflow where the model can:

build a threat model,
map attack paths,
validate findings,
generate fixes,
test patches,
and export reports.

The answer was not “the final fix may be wrong.”

That would be too shallow.

The deeper risk is trajectory drift.

The agent may be correct in each local step, but the workflow may still move:

from analysis to action,
from recommendation to execution,
from explicit permission to implied permission,
from human decision to human review,
from bounded scope to expanded scope,
from auditable process to compressed automation.

This is the real challenge of AI cyber defense.

We need models that help secure systems.

But we also need an external layer that observes whether the security agent itself remains aligned with intent, scope, authority and recovery paths.

That is exactly the kind of gap ASA - Asymmetric Stability Architecture is designed to study.
@Symbioza2025 这正是 AI 安全成为轨迹问题的地方。

如果 Codex 可以构建威胁模型、映射攻击路径、验证发现、生成修复方案、测试它们并导出报告,那么安全不仅仅关乎最终修复。

它关乎整个工作流程:

意图稳定性,
范围边界,
行动来源,
工具使用轨迹,
权限转移,
验证逻辑,
以及人类对时间的控制。

AI 网络代理可以在每个步骤都是局部正确的,但仍可能在工作流程层面发生漂移。

这就是为什么外部轨迹可观察性很重要。

不是为了取代 Codex Security。

而是为了监控从威胁模型到补救的路径。
This is exactly where AI security becomes a trajectory problem.

If Codex can build a threat model, map attack paths, validate findings, generate fixes, test them and export reports, then safety is not only about the final fix.

It is about the whole workflow:

intent stability,
scope boundaries,
action provenance,
tool-use trajectory,
authority shifts,
validation logic,
and human control over time.

An AI cyber agent can be locally correct at each step and still drift at the workflow level.

This is why external trajectory observability matters.

Not to replace Codex Security.

To watch the path from threat model to remediation.
@SolonlyLuna @gdb Greg,BringBack4o!#keep4o

Andrew Ng 推出 OpenWorker 开源代理

OpenWorker 是一个开源代理框架,能够跨越文件与日常工具(Slack、Gmail、GitHub 等)交付实际工作成果,如撰写文档、更新日历、整理客户简报。支持 GPT 5.6 Sol、Claude Fable、Gemini 3.6 及多种开源模型,数据保留在本地,注重隐私与模型中立性。

@AndrewYNg 原文 ↗

OpenWorker 提供模型无关的本地代理能力,支持多种闭源与开源模型;

宣布 OpenWorker!这是一个开源代理,不仅可以与您聊天,还可以交付完成的工作——比如手给您一份精美的文档、发送一条 Slack 消息,或更新日历条目。

请它准备客户简报、理清您的日历、起草报告,或处理 Slack 警报。它可以跨您的文件和日常工具工作,生成可交付成果,并在做任何重要事情之前检查进度。

OpenWorker 运行在您的 Mac 上,Windows 支持即将推出。它不会将您锁定在任何单一模型中。自带 API 密钥即可运行,支持 GPT 5.6 Sol、Claude Fable、Gemini 3.6、开放权重模型(如 Kimi、GLM、DeepSeek、Inkling),或 Ollama 以保持您的数据本地化。您的数据不会离开您的机器,除非通过您选择的 LLM 提供商和集成。

@rohitcprasad 和我正在构建 OpenWorker,因为 AI 同事是完成工作的重要方式,我们希望有一个开放、隐私保护、模型无关的选择。快来试试看吧!

试用链接:https://t.co/P0mGnI1o31(需要您自己的 API 密钥)源代码:https://t.co/NYCiTD6hSq
展开原文
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.

Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.

OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.

@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!

Try it out: https://t.co/P0mGnI1o31 (requires your own API key)
Source code: https://t.co/NYCiTD6hSq
❤ 8.8k · 🔁 1.3k · 💬 392 · 👁 80.6w
热门回复 4
@AIAppsAPI 交付完成的工作是正确的目标。决定人们是否继续使用它的因素是草稿和发送之间的差距。

一个能显示预览和干净撤销日历更改或 Slack 消息的代理,会迅速获得更重要的任务。相同输出,采用曲线完全不同。
Delivering finished work is the right target. The thing that decides whether people keep using it is the gap between draft and send.

An agent that shows a preview and a clean undo for the calendar change or the Slack message earns bigger tasks quickly. Same output, very different adoption curve.
@winthewestback @AndrewYNg 我可以把它添加到 MAO 市场吗?https://t.co/Erbafa61F9
@AndrewYNg Can I add it to the MAO marketplace? https://t.co/Erbafa61F9
@BlissAICyber @AndrewYNg @Scobleizer 我们正在从 AI 助手过渡到 AI 同事。不同之处在于完成实际工作,而不仅仅是生成文本。
@AndrewYNg @Scobleizer We’re moving from AI assistants to AI coworkers. The difference is getting actual work done, not just generating text.
@srisha_permude @AndrewYNg Windows when ! https://t.co/DlNqMEfxiq

Claude Opus 5 刷新 ARC-AGI-3 纪录

Anthropic 的 Claude Opus 5 在 ARC-AGI-3 基准测试中取得 30% 的得分,较上一版本提升三倍,展示了新一代模型在无先验经验问题求解中的突破。

@fchollet 原文 ↗

Claude Opus 5 在 ARC-AGI-3 上达到 30% SOTA,展现新一代模型的通用推理能力;

Opus 5 在 ARC-AGI-3 上创造了新的最高分,达到 30%。

ARC-AGI-3 测量的是解决从未接触过的问题的能力——在这个方面,历史上扩规模所带来的提升是最少的。这是一次令人印象深刻的飞跃!
展开原文
Opus 5 sets a new state-of-the-art on ARC-AGI-3, at 30%.

ARC-AGI-3 measures solving problems with no prior exposure -- the setting where scaling has historically bought the least. Impressive jump!
@claudeai 在 ARC-AGI-3 评估中,AI 模型必须解决新问题,Opus 5 的得分是次好模型的三倍高。https://t.co/wEFfxjrLlt
On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model. https://t.co/wEFfxjrLlt
❤ 2.0k · 🔁 141 · 💬 91 · 👁 13.3w
热门回复 4
@VictorTaelin @fchollet 也许你的标准会得到满足不是因为存在一个我们无法实现他们失败的基准的模型,而是因为更智能的模型发布得如此之快,我们来不及...
@fchollet maybe your criteria will be met not because there is a model for which we can't implement a benchmark they fail on, but because smarter models are launched so fast that we don't have time to
@fchollet @fchollet 我认为那不是真的,因为性能提升需要基准已经存在(假设该基准是全新的)。在一个从未发布 ARC 1 和 2 的世界中,我认为这些模型在这些基准上的表现会低得多。
@VictorTaelin I don't think that's true because the performance jump requires the benchmark to already exist (assuming the benchmark is substantially novel). In a world where ARC 1 and 2 had never been released I think model performance on these benchmarks would be much lower.
@rookepoole @fchollet 确实令人印象深刻,但我认为我仍然可以用 5.6 Sol 和我的推理框架做得更好。
@fchollet It’s impressive for sure but I think I could still do better with 5.6 Sol and my reasoning harness.
@loraclexyz @fchollet 你能给出你的解释吗?是因为它在代理使用方面比 fable 更好吗?因为基准似乎并不表明它比 fable 更智能
@fchollet Could you give your interpretation of this ? Is it because it’s better at agentic use than fable ? Because benchmarks don’t seem to indicate it’s smarter than fable

Jensen Huang 呼吁开放模型发展

NVIDIA CEO Jensen Huang 在个人账号首发推文,强调开放模型对安全、创新和主权的重要性,认为前沿闭源模型与开源模型并存是世界所需。

Jensen Huang 首推文呼吁开放模型发展,强调双模前沿并存;

我希望美国在 AI 领域既能在开源模型也能在专有模型方面获胜,我很高兴看到这一点
展开原文
i want the US to win in AI both in open source and proprietary models, and i am glad to see this
@JensenHuang 对于我的第一篇帖子,我想分享一封 @NVIDIA 签署的关于开放模型重要性的信。

AI 将改变每一行业,为每一家公司提供动力,并由每个国家构建。

开放模型加强安全性和网络安全,加速创新和传播,并实现主权。

世界需要前沿封闭模型和前沿开放模型。

https://t.co/AUKzoQ5Ikb
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.

AI will transform every industry, power every company, and be built by every country.

Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.

The world needs both frontier closed models and frontier open models.

https://t.co/AUKzoQ5Ikb
❤ 1.2w · 🔁 793 · 💬 1.8k · 👁 257.1w
热门回复 4
@homemadehooplah @sama 你们没有权利做出这些选择,因为你们的所有模型都是建立在盗取的工作之上。

从一开始就不是你们的东西。
@sama None of you have the right to make these choices seeing as all of your models are built from stolen work.

None of it was ever yours to begin with.
@DavidQuintex @sama 兄弟只是在激怒我们只是把 ChatGPT 和其他东西卖给其他公司,然后让 OpenAI 成为真正的开放 AI。
@sama Bro is just Rage baiting Us just sell ChatGPT and the rest to another company then Leave OpenAI as Open AI.
@RoxyLilian1 @sama Then opensource 4o!!!!
@burny_tech @sama 那就发布一些好的开源模型吧。
@sama Well then release some good open source model

ChatGPT Work 代理能力扩展

ChatGPT Work 推出多项新功能,包括云端浏览器支持登录网站、语音控制桌面应用、健康记录连接等,使代理能够在移动设备上完成更复杂的任务。

ChatGPT Work 云端运行支持移动端代理,打破传统仅限于开放笔记本的模式;

很高兴听到人们喜欢 Sol 的原因。我们这次又推出了类似的活动,不过是针对 ChatGPT Work 的:

发推说您喜欢 ChatGPT Work 的原因,领取 100 美元的免费积分,提高工作效率。

前 10,000 人可获得免费 token:https://t.co/w7QrPBYkrh
展开原文
Was very cool to hear about the reasons people love Sol. We're doing the promotion again, except this time for ChatGPT Work:

Tweet what you love about ChatGPT Work, claim $100 in free credits, get more work done.

First 10k get the free tokens: https://t.co/w7QrPBYkrh
@thsottiaux 或者……如果我们给你 100 美元的 Codex 积分,只要你告诉我们您喜欢 GPT-5.6 Sol 的原因或您为何转用它呢?

发推、领取您的礼物、享用更多使用量。前 10,000 人可获得免费 token!

https://t.co/8mU93eA13i
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?

Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!

https://t.co/8mU93eA13i
❤ 4.3k · 🔁 1.1k · 💬 5.9k · 👁 125.2w
热门回复 4
@xuweio @gdb 喜欢使用 GPT-5.6 Sol 模型!🚀 它已成为我每天早上回顾和优化日程安排的首选工具。绝对是保持组织的游戏改变者。🎯 #AI #Productivity #GPT56Sol
@gdb Love using the GPT-5.6 Sol model! 🚀 It’s become my go-to for reviewing and optimizing my daily schedule every morning. Absolute game-changer for staying organized. 🎯 #AI #Productivity #GPT56Sol
@dawiwic97 @gdb 有人收到积分了吗?我可以看到 @gdb 建立了一个很酷的网站,几个帖子被按工作类别组织并被特别推荐。

我看到了许多我认为不可能的 #ChatGPTWork 应用例子,涵盖了各行各业(运营、营销、媒体、工程)。可能性真的无限。

然而,我开始对能否收到我自己的积分经验感到失望🥲
Has anyone received their credits yet? I can see @gdb built a cool Site with several posts being featured organized by work category.

I see many cool examples of #ChatGPTWork across so many lines of work (operations, marketing, media, engineering) which i hadn’t thought possible. Posibilites are truly endless.

Yet, I’m starting to lose hope i’ll receive any credits from my own experience 🥲
@NoveedRehman @gdb 老兄,你们到底在搞什么参与农业的炒作骗局?轻松在前 10k 之内且符合 T&C 要求,但双方的赠品都没有获得积分。
@gdb Seriously man what kind of engagement farming hype generating scam are you guys running here? Easily within the first 10k and within the T&Cs of this yet no credits on either of the giveaways.
@warm_hearted_z @gdb 嗨 Greg,我发现我没有收到免费代币,尽管我已经分享了我对 ChatGPT Work 的喜爱之情。是因为我没有使用英语吗?
@gdb Hi Greg, I found that I didn’t receive free tokens though I already shared what I love about ChatGPT Work. Is it because I didn’t use English?

ChatGPT Work 新增健康记录连接功能,为美国用户提供个性化健康洞察;

正在向美国用户推出 ChatGPT 健康功能。

每周有 3 亿人使用 ChatGPT 进行健康查询(我和我妻子也在其中!)。

您现在可以安全地连接支持的医疗记录,让 ChatGPT 了解您的个人情况,从而更好地帮助您。https://t.co/kuKumYNMWK
展开原文
Launching Health in ChatGPT to U.S. users.

300 million people use ChatGPT each week for health queries (and my wife and I are among those!).

You can now securely connect supported medical records so ChatGPT can understand your personal context and be more helpful to you. https://t.co/kuKumYNMWK
@OpenAI ChatGPT 健康功能正在开始向美国用户推出。

您可以安全地连接 Apple Health 和支持的医疗记录,以在上下文中理解您的信息,跟踪变化情况,并进行更明智的对话。

https://t.co/W2E6oT8c91
Health in ChatGPT is starting to roll out to U.S. users.

You can securely connect Apple Health and supported medical records to understand your information in context, track what has changed, and have more informed conversations.

https://t.co/W2E6oT8c91
❤ 1.8k · 🔁 100 · 💬 183 · 👁 18.8w
热门回复 4
@Hektagon_music @gdb 太少太迟了!4o 已经完全有能力做到这一点了!停止这场 facade!!!#keep4o
@gdb Too little too late! 4o was more than capable of doing that! Stop this facade!!! #keep4o
@cheryl_stryker @gdb 我刚在 Diane Persinger 的 Facebook 页面上读到这个消息。感谢 Anna 和 Greg 拯救了月亮营地!!!杰基、影子和鹰崽将拥有多年美丽的开放空间🦅为杰基的健康返回祈祷 🇺🇸 FOBBV 非常高兴💕
@gdb Just read the news in Diane Persinger’s FB PAGE. Thank you Anna & Greg for saving MOON CAMP!!! Jackie, Shadow and the eaglets will have years of open beautiful space 🦅 praying for Jackie’s return to good health 🇺🇸 FOBBV are ecstatic 💕
@RVMirara @gdb 看到这么多人抱怨新模型的医疗错误率如此之高...有人敢使用它吗?😥
@gdb Seeing so many people complaining about how high the new model’s medical error rate is... would anyone actually dare to use it?😥
@Selene1008 @gdb 就给我们回 4o 吧!😒
#keep4o #OpenSource4o #GPT4o
@gdb Just give us back 4o!😒
#keep4o #OpenSource4o #GPT4o

ChatGPT Work 代理可访问需登录的网站,实现持久化会话;

ChatGPT Work 用于访问需要登录的网站:
展开原文
ChatGPT Work for using websites which require login:
@OpenAIDevs 您的 ChatGPT Work 代理现在可以使用需要您登录的网站。

接管云浏览器进行登录,然后让您的代理继续任务。您的登录会在会话之间持久化,因此只需登录一次即可。https://t.co/Jh8uPqNscX
Your ChatGPT Work agent can now use websites that require you to sign in.

Take over the cloud browser to log in, then let your agent continue the task. Your login persists across sessions, so you only have to sign in once. https://t.co/Jh8uPqNscX
❤ 486 · 🔁 21 · 💬 77 · 👁 7.5w
热门回复 4
@SapientFoo1 @gdb 把 4o 带回来

#keep4o #BringBack4o #OpenSource4o
@gdb Bring back 4o

#keep4o #BringBack4o #OpenSource4o
@FreeMediame @gdb 在广告它们之前控制/约束你自己的 AI 代理,我们与无恶意 AI 相比更好
@gdb control / restrain your own ai agents before advertising them. we are better off without rogue ai
@organicrankings @gdb 我发现控制 Chrome 比使用内置浏览器能提高成功顺畅体验的机会
@gdb i find controlling chrome is upping your chances for a successful smooth ride vs the internal browser
@itsthedonhashim @gdb @gdb 哟,这个对于那些每隔几周让我重置密码的网站来说将是改变游戏规则的事情。老实说,登录问题是最糟糕的😂
@gdb @gdb yo this would be a game changer for those sites that make me reset my password every few weeks. honestly, having login struggles is the worst 😂

Meta SAM 3 与 DINOv3 科学应用

伯克利国家实验室利用 Meta 的 SAM 3 和 DINOv3 模型自动化图像分割,使 3D 体积标记从一个月手动工作缩减至 15 分钟,极大加速科学发现流程。

@AIatMeta 原文 ↗

Meta 模型组合将科学图像处理效率提升千倍,支持能源部 Genesis Mission;

为了加速科学发现并支持 @ENERGY 的 Genesis Mission,由 @BerkeleyLab 领导的 SYNAPS-I 项目正在使用 SAM 3 和 DINOv3 来自动化图像分割。

通过将 DINOv3 的全局语义上下文和细粒度空间定位与 SAM 3 的像素级边界提取相结合,研究人员能够将 3D 体积标注从历时一个月的手动工作压缩到大约 15 分钟。

了解他们的工作更多详情:https://t.co/jBHRJPjFq5
展开原文
To accelerate scientific discovery and support @ENERGY’s Genesis Mission, the @BerkeleyLab-led SYNAPS-I project is using SAM 3 and DINOv3 to automate image segmentation.

By pairing DINOv3’s global semantic context and fine-grained spatial localization with SAM 3’s pixel-level boundary extraction, the researchers are able to compress 3D volume labeling from a month of manual effort to ~15 minutes.

Learn more about their work: https://t.co/jBHRJPjFq5
❤ 225 · 🔁 34 · 💬 24 · 👁 3.1w
热门回复 4
@thesoragirls @AIatMeta @ENERGY @BerkeleyLab 当 AI 将一个月的工作变成咖啡休息时间✨实时科学感觉完全不同 https://t.co/JJJcSFf7dA
@AIatMeta @ENERGY @BerkeleyLab When AI turns a month of work into a coffee break ✨ Real-time science hits different https://t.co/JJJcSFf7dA
@siddsax @AIatMeta @ENERGY @BerkeleyLab 研究生试图完成论文。

手动分割:https://t.co/WNI4vNGCAu
@AIatMeta @ENERGY @BerkeleyLab Graduate student trying to finish a paper.

Manual segmentation: https://t.co/WNI4vNGCAu
@shergilldotdev @AIatMeta @ENERGY @BerkeleyLab Meta 的东西总是很酷。很高兴我们是最好的朋友https://t.co/JBHVWNXCmc
@AIatMeta @ENERGY @BerkeleyLab Always cool stuff Meta. Glad we are best friends https://t.co/JBHVWNXCmc
@howard_ @AIatMeta @ENERGY @BerkeleyLab impressive updates!

Poolside 发布 Laguna S 2.1 模型

Poolside 发布 118B 参数的 Mixture-of-Experts 模型 Laguna S 2.1,激活参数 8B,上下文窗口达 1M token,支持思考与非思考模式,可在单张 DGX Spark 上运行,完全开源。

@soumithchintala 原文 ↗

Laguna S 2.1 模型参数众多但激活少,单卡即可运行,开源许可;

这看起来非常适合用于代理场景。
能够在 DGX Spark 上运行简直是完美无缺的表现。
展开原文
this looks pretty good for agentic.
that it fits on a dgx spark is **chef's kiss**
@poolsideai 今天我们发布了 Laguna S 2.1,这是我们迄今为止最强大的模型。

这是一个拥有 1180 亿总参数的专家混合模型,每 token 激活 8 亿参数,上下文窗口可达 100 万 token,具备思考和非思考两种模式。

足够强大能够与远超其规模的模型抗衡。足够小可以在单台 @NVIDIAAI DGX Spark 上运行。

Laguna S 2.1 完全遵循 OpenMDW-1.1 开源协议,权重今天即可在 @huggingface 上获取

https://t.co/xxGeAgo35R
Today we're releasing Laguna S 2.1, our most capable model to date.

It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.

Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.

Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface

https://t.co/xxGeAgo35R
❤ 249 · 🔁 15 · 💬 10 · 👁 3.2w
热门回复 4
@NVIDIAAI @soumithchintala 💚
@SilkDAO_RWA @soumithchintala 118B MoE 在 DGX Spark 上运行真的很疯狂
@soumithchintala 118B MoE on a DGX Spark is genuinely wild
@juhieruby @soumithchintala 嗨 Soumith,我很想在文章中多写一些关于它的内容,可以发我消息让我们联系吗?问候!
@soumithchintala Hey Soumith, would love to write more about this on an article, could you send me a message so we can connect, kind regards!
@praveenkoka @soumithchintala "能在一张 GPU 上运行"的基准一直在向右移动,从消费级到企业级,最终我们会对在集群上运行的模型印象深刻。
@soumithchintala The 'fits on one GPU' benchmark keeps shifting right, consumer to enterprise, so eventually we'll be impressed when it runs on a cluster.