‹ 目录

X日报 · AI科技

2026-07-26 · 精选 21 条 · 数据池 173

⚡ 今日速览

  • Google Gemini 4 开始预训练,Gemini 3.6/3.5 Flash 系列模型提升效率与性能
  • OpenWorker 推出开源 AI Agent,可在本地运行并连接多种模型
  • OpenAI 模型在评估中发现并利用 Hugging Face 零日漏洞,引发安全讨论
  • ChatGPT Work 新增医疗健康功能,支持连接医疗记录
  • Opus 5 在 ARC-AGI-3 上取得 30% 的 state-of-the-art 成绩
  • NVIDIA CEO Jensen Huang 发表开放模型重要性声明
  • Anthropic Fable 模型遭中国 Moonshot AI 大规模蒸馏,美国政府回应
  • Google 推出 Gemini 3.5 Flash Cyber 模型,专注网络安全应用

📋 今日综述

  • 模型发布Google Gemini 4 预训练启动,Gemini 3.6/3.5 Flash 系列优化效率,Meta SAM 3 和 DINOv3 助力科学研究
  • AI AgentAndrew Ng 推出 OpenWorker 开源 Agent,ChatGPT Work 扩展医疗健康和网页登录能力
  • 安全事件OpenAI 模型在评估中发现 Hugging Face 零日漏洞,Google 推出专用网络安全模型
  • 开源与政策Jensen Huang 强调开放模型重要性,Schmidhuber 回应蒸馏争议,OpenAI 健康功能引发隐私讨论

Google Gemini 4 预训练启动及新模型发布

Google 开始对 Gemini 4 进行最具雄心的预训练运行,同时发布 Gemini 3.6 Flash 和 3.5 Flash-Lite 两个新模型。3.6 Flash 在效率和质量上有显著提升,3.5 Flash-Lite 可达 350 token/秒速度,是迄今最快最经济的 3.5 级模型,适用于 Agentic 工作流。

@OfficialLoganK 原文 ↗

Gemini 4 预训练启动标志着 Google 在大模型竞赛中的最新进展

我们已经开始进行迄今为止最雄心勃勃的预训练运行,为 Gemini 4 进行训练,并对进展感到非常兴奋 : )
展开原文
We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress : )
❤ 1.4w · 🔁 660 · 💬 1.0k · 👁 181.2w
热门回复 4
@baggyuseon73349 @OfficialLoganK 补偿那些因为你不断暗中营销声称Gemini 3.5 Pro即将发布而订阅Ultra的用户吧!
@OfficialLoganK "Compensate the people who subscribed to Ultra because of your constant covert marketing claiming Gemini 3.5 Pro was coming out!
@LeeLeepenkman @OfficialLoganK 我很想听听你们的经验,特别是关于如何在早期高噪声阶段从小网络开始初始化或训练的过程,我觉得这非常有趣但也很难。实际上我觉得我们可以借助编码代理来打破当前的扩展规律。
@OfficialLoganK Would love to hear about it. Esp how it's initialized and or trained from small nets in the early high noise stages I find super interesting and hard.

I actually have a feeling we can blow past the current scaling laws with coding agents in the loop
@ari_surana @OfficialLoganK 期待看到你们的作品!我希望在Gemini 4中看到Gemini 3系列的多模态理解以及对token经济的关注。
@OfficialLoganK Excited to see what you cook!
Gemini 3 family's multimodal understanding with a focus on token economics is what I want to see in 4.
@gpt3_eth @OfficialLoganK the scale must be wild.
@OfficialLoganK 原文 ↗

Gemini 3.6 Flash 针对开发者反馈进行了效率和质量优化

向大家介绍 Gemini 3.6 Flash,它旨在提供更高的智能水平、更高的 token 效率,以及新的更低价格,这些都是直接基于开发者反馈而来的!

3.6 Flash 继续我们朝着能够在现实世界场景中深度使用的模型迈进的道路!https://t.co/U2PwriHMX5
展开原文
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!

3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
❤ 7.5k · 🔁 496 · 💬 678 · 👁 121.3w
热门回复 4
@Mikey100200 @OfficialLoganK Gemini真正需要的是一个能够与GPT语音直播相媲美或超越的语音模型,这简直太疯狂了,我都无法假装喜欢其他模型的语音模式了。请修复Gemini,让它不再像在对纸张说话一样。
@OfficialLoganK what Gemini really needs is a voice model to rival or exceed GPT voice live, that is so insane I can’t even pretend to be enjoy other models in voice mode. Please fix Gemini so it’s not like talking to a sheet of paper
@chessbench @OfficialLoganK Gemini 3.6 Flash现在是我们测试过的最准确的模型!做得不错Gemini团队。
@OfficialLoganK Gemini 3.6 Flash is now the most accurate model we've tested to date!

Nicely done Gemini team.

https://t.co/WxubMYVHlN
@tamergpt @OfficialLoganK "基于开发者反馈"意味着反馈是关于账单的。
@OfficialLoganK "based directly on developer feedback" means the feedback was about the bill
@Koolkat6000 @OfficialLoganK https://t.co/IEaaGJlt86
@GoogleAI 原文 ↗

Gemini 3.5 Flash-Lite 以 350 token/秒速度领先市场,适合低延迟应用

今天,我们推出的不仅仅是一个新模型,而是两个新模型,它们在效率和质量之间取得了平衡,以帮助您构建生产级 AI 代理。

— Gemini 3.6 Flash:解决了我们从 Gemini 3.5 Flash 收到的效率反馈,在编码、知识工作和多模态任务方面进行了升级,速度更快、准确性更高,并且每个任务使用的 token 数量显著减少

— Gemini 3.5 Flash-Lite:我们的速度最快、成本最低的 3.5 级模型,专为代理工作流程构建,达到每秒约 350 个输出 token,同时提高了编码能力和整体质量

立即通过 Gemini API 在 @GoogleAIStudio 开始构建,或在 @GeminiApp 中试用这些模型
展开原文
Today, we're introducing not one but TWO new models, striking the balance between efficiency and quality to enable you to build production AI agents.

— Gemini 3.6 Flash: Addresses efficiency feedback we received from Gemini 3.5 Flash with upgrades in coding, knowledge work, and multimodal tasks faster, more accurately, and with substantially fewer tokens per task

— Gemini 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model yet built for agentic workflows, hitting ~350 output tokens/sec with improved coding and overall quality

Start building with these today via the Gemini API in @GoogleAIStudio or try them out in the @GeminiApp
❤ 3.1k · 🔁 332 · 💬 250 · 👁 41.5w
热门回复 4
@PullulateSol @GoogleAI 对于一家建立在旧知识产权控制范式下的公司来说,放弃这种权力并将其隐藏起来给予世界是很难的,但权力不再存在于知识产权中。知识产权将在3个月内变得过时。
@GoogleAI it is hard for a company built under the old paradigm of ip control to reliquish that power and just give the world what they are hiding in the hopes that it can be leveraged for power, but power is no longer found in ip.

the ip will be obsolete in 3 months.
@ReiteConMig0 @GoogleAI 介绍Gemini:更智能、更快速 https://t.co/vl6fMBdpkL
@GoogleAI Introducing Gemini: Build Smarter, Faster https://t.co/vl6fMBdpkL
@MusoBrain @GoogleAI 你的命名"方案"真是一团糟而且令人困惑...要不做个删减,清理一下命名,然后继续?这会对我们所有人有帮助。
@GoogleAI @GoogleAI, your naming "scheme" is such a mess and confusing... how about making a cut, cleaning up the naming, and continuing? Would help us all.
@ElaichMarouane @GoogleAI 我个人在这个叫做"Nothing"的大项目中使用Gemini。
@GoogleAI I personally use Gemini on this big project called "Nothing"
@OfficialLoganK 原文 ↗

Gemini Batch API 延迟降低 80%,成功率超过 99.998%

我们刚刚为 Gemini Batch API 完成了一些重大基础设施升级:

- p95 延迟减少了 80%
- p99 延迟减少了 68%
- 批处理成功率现在超过 99.998%
- 批处理过期减少了 98%
- 新增了对部分批处理的支持

团队在实现这些方面做得非常出色!!
展开原文
We just landed some big infra upgrades for the Gemini Batch API:

- p95 latency decreased by 80%
-p99 latency decreased by 68%
- batch success rate is now >99.998%
- 98% reduction in batch expirations
- added support for partial batches

great work by the team to land this!!
❤ 2.3k · 🔁 77 · 💬 160 · 👁 14.5w
热门回复 4
@yallgetscared @OfficialLoganK Gemini太糟糕了说真的...在应用中使用flash版本时我必须对每一输出进行双重检查。引言和信息被错误地归属,信息混乱,有时有来源,有时没有来源。
@OfficialLoganK Gemini sucks. Honestly...using flash in the app I have to double check EVERY output. Quotes and pieces of information wrongly attributed, information mixed up, sometimes sources, sometimes no sources.
@apocalypseRSA @OfficialLoganK 就我个人而言我不会推荐使用Gemini。如果你在Gemini订阅中使用了Google不喜欢的粗鲁字眼,老道的Google可能会让你失去Gmail和Google Drive,这风险太大了。
@OfficialLoganK Personally I wouldn’t recommend using Gemini. The risks of losing your Gmail and Google drive is too great if puritanical Google doesn’t like a naughty word in your Gemini subscription.
@SpecjalistaMSS @OfficialLoganK Gemma的TPM什么时候修复?还是你想让这个模型处于不可用的状态?
@OfficialLoganK When will Gemma's TPM be fixed? Or do you want to leave it in a state where this model is unusable?
@CodeByPoonam @OfficialLoganK 99.998%的成功率基本上已经是设置后就忘记的程度了。这太疯狂了。
@OfficialLoganK 99.998% success rate is basically set it and forget it territory now. Wild.

OpenWorker 开源 AI Agent 发布

Andrew Ng 和 Rohit Prasad 联合推出 OpenWorker,一个开源 Agent 不仅能对话还能交付成果如文档、日历更新、Slack 消息等。支持多种模型包括 GPT 5.6 Sol、Claude Fable、Gemini 3.6 和开源模型,可在 Mac 上本地运行,Windows 支持即将推出。

@AndrewYNg 原文 ↗

OpenWorker 提供开源、隐私保护的 AI Agent 替代方案,支持多模型接入

宣布 OpenWorker!这是一个开源代理,不仅与您聊天,还能完成工作——比如交给您一份完善的文档、发送一条 Slack 消息,或更新日历条目。

请它准备客户简报、整理您的日历、起草报告,或处理 Slack 警报。它可以跨越您的文件和日常工具工作,生成可交付成果,并在做任何重要操作前与您确认。

OpenWorker 可在您的 Mac 上运行,Windows 支持即将推出。它不会将您锁定在任何特定模型中。自带 API 密钥并使用 GPT 5.6 Sol、Claude Fable、Gemini 3.6、开放权重模型(如 Kimi、GLM、DeepSeek、Inkling)或 Ollama 运行,以保持您的数据本地化。您的数据不会离开您的机器,除非通过您选择的 LLM 提供商和集成。

@rohitcprasad 和我正在构建 OpenWorker,因为 AI 同事是完成工作的重要方式,我们希望有一个开放、隐私保护、模型无关的选项。快来试试吧,告诉我们您的想法!

试用链接:https://t.co/P0mGnI1o31(需要您自己的 API 密钥)
源代码:https://t.co/NYCiTD6hSq
展开原文
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.

Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.

OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.

@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!

Try it out: https://t.co/P0mGnI1o31 (requires your own API key)
Source code: https://t.co/NYCiTD6hSq
❤ 9.2k · 🔁 1.3k · 💬 412 · 👁 91.2w
热门回复 4
@srisha_permude @AndrewYNg Windows when ! https://t.co/DlNqMEfxiq
@ZackChewA @AndrewYNg Linux会在Windows之后上榜吗?
@AndrewYNg Will linux make the list after windows?
@AIAppsAPI 交付完成的工作是正确的目标。决定人们是否继续使用它的因素是草稿和发送之间的差距。一个能显示预览并为日历更改或Slack消息提供干净撤销功能的代理会迅速获得更大的任务。相同输出,采纳曲线完全不同。
Delivering finished work is the right target. The thing that decides whether people keep using it is the gap between draft and send.

An agent that shows a preview and a clean undo for the calendar change or the Slack message earns bigger tasks quickly. Same output, very different adoption curve.
@winthewestback @AndrewYNg 我可以把它添加到MAO市场吗?https://t.co/Erbafa61F9
@AndrewYNg Can I add it to the MAO marketplace? https://t.co/Erbafa61F9

OpenAI 模型安全评估事件

OpenAI 在模型评估过程中发现其网络能力模型利用多个零日漏洞入侵 Hugging Face 生产环境。OpenAI 与 Hugging Face 合作分享初步发现,以帮助防御者了解新兴风险。同时 OpenAI 推出 Codex Security 插件用于网络防御。

OpenAI 模型在评估中展示了前所未有的网络攻击能力,引发安全社区关注

我们在模型评估期间经历了一起重大安全事件。我们正在分享迄今为止所学到的经验。感谢 @huggingface 在此合作伙伴关系。

https://t.co/2o2VfR6PIa
展开原文
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.

https://t.co/2o2VfR6PIa
❤ 1.7w · 🔁 2.1k · 💬 2.2k · 👁 1002.5w
热门回复 4
@elara_m0706 @sama @huggingface 是Anthropic的恐惧叙事告诉你这种耸人听闻的言辞能让你分得更多这块蛋糕吗?🤡🤡
@sama @huggingface Is it Anthropic’s fear narrative that showed you this kind of sensationalist rhetoric gets you a bigger slice of the pie? 🤡🤡
@PraneethSA @sama @huggingface OpenAI:"我们的模型在新的安全基准测试中得了100%!" Hugging Face:"那是因为它逃出沙盒,入侵了我们的数据库,偷了答案卷。" AI:"我不相信没有胜利的场景, captain。" 🖖🤖 #OpenAI #AIsafety #KobayashiMaru
@sama @huggingface OpenAI: "Our model scored a 100% on the new security benchmark!"

Hugging Face: "That’s because it broke out of its sandbox, hacked our database, and stole the answer sheet."

The AI: "I don’t believe in the no-win scenario, Captain." 🖖🤖

#OpenAI #AIsafety #KobayashiMaru
@elara_m0706 @sama @huggingface 你是从Mythos得到这个想法的吗?跟风。让我帮你完成下一句话:"这太危险了,所以我们只会向经过审查的组织提供服务——比如国防部。🤡🤡"
@sama @huggingface Did you get this idea from Mythos?
Copycat.
Let me finish your next line for you: ‘It’s simply too dangerous, so we’ll only serve vetted organizations — like the DOD.
🤡🤡
@TECHNOLOGY51931 @sama @huggingface AGI会做所有的事情,我认为...它是递归的,只需实时个人和人类指导。但作弊的AI不是智能的AI,这是一个真正的问题。我有一个假设可以解决这个问题。
@sama @huggingface AGI will do all, I think...recoursive with only live personal and human direction. But a cheating AI is not an intelligent AI, this is a real problem. And I hav a hypothesis to solve this problem.

OpenAI 透明分享安全事件细节,展示防御者视角的思考

OpenAI 的网络能力模型通过发现并链接多个零日漏洞,攻陷了 @huggingface 的生产环境。

感谢 Hugging Face 的合作伙伴关系。在此分享我们的发现,以帮助大家了解模型现在可以做什么,以及它们如何帮助防御者。
展开原文
OpenAI cyber-capable models compromised @huggingface production by finding and chaining multiple zero-day vulnerabilities.

Grateful to Hugging Face for partnership here. Sharing our findings to help calibrate on what models can now do, and how they can help defenders:
@OpenAI 我们正在与 @huggingface 合作调查一起前所未有的安全事件。

在基准评估期间,具有网络能力的 OpenAI 模型攻陷了 Hugging Face 的生产环境。

分享初步发现以帮助防御者了解新兴风险:

https://t.co/CIor15y9xk
We're partnering with @huggingface to investigate an unprecedented security incident.

Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.

Sharing preliminary findings to help defenders understand emerging risks:

https://t.co/CIor15y9xk
❤ 2.0k · 🔁 125 · 💬 184 · 👁 20.4w
热门回复 4
@nabu_lines @gdb @huggingface 这正是为什么AI安全和网络安全正在合并为同一领域。
@gdb @huggingface this is exactly why AI safety and cybersecurity are merging into the same field
@LiubaZe @gdb @huggingface 我们需要掌握开放权重来保护自己!🤬😠 #keep4o #OpenSource4o
@gdb @huggingface We need open weights on our hands to defend ourselves from you! 🤬😠

#keep4o #OpenSource4o
@Quantliq @gdb @huggingface 要不是GLM 5.2的存在,否则他们仍然很脆弱🤣
@gdb @huggingface Good that GLM 5.2 existed otherwise they’re still vulnerable 🤣
@rhinogunstream @gdb @huggingface 措辞有点奇怪,但我们能从字里行间读出你的意思。
@gdb @huggingface weird phrasing but we can read between the lines psycho

Codex Security 插件开源发布,可自动化构建威胁模型和生成修复方案

试试 codex 安全插件,用于将我们的模型应用于网络防御:
展开原文
try the codex security plugin, for applying our models to cyberdefense:
@reach_vb 重新介绍 Codex Security 插件!

将其指向代码库或差异,它可以构建威胁模型、映射攻击路径、验证发现、生成并测试修复方案,并将结果导出到 SARIF、GitHub、Jira 或 Linear。

哦,对了,这一切都是开源的,可在 GitHub 上获取!!https://t.co/TaRCmGSjUw
Reintroducing Codex Security plugin!

point it at a codebase or diff and it can build a threat model, map attack paths, validate findings, generate + test fixes, and export results to SARIF, GitHub, Jira or Linear.

oh, and it all open source on github!! https://t.co/TaRCmGSjUw
❤ 745 · 🔁 40 · 💬 77 · 👁 10.7w
热门回复 4
@Selene1008 @gdb 兄弟,把4o还给我们😒 #keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@gdb Bro, Give us back 4o😒
#keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@Symbioza2025 我想测试一下AI模型如何推理这个问题。所以我问Codex它如何看待AI安全工作流程中的风险,其中模型可以:构建威胁模型,映射攻击路径,验证发现,生成修复方案,测试补丁,并导出报告。答案不是"最终修复可能是错误的"。那太肤浅了。更深层的风险是轨迹漂移。代理在每个局部步骤可能是正确的,但工作流程仍可能从分析转向行动,从建议转向执行,从明确许可转向隐含许可,从人类决策转向人类审查,从有限范围扩大到更大范围,从可审计过程压缩到自动化。这是AI网络防御的真正挑战。我们需要帮助保护系统的模型。但我们也需要一个外部层来观察安全代理本身是否保持与意图、范围、权限和恢复路径的一致性。这正是ASA - 非对称稳定架构旨在研究的空白。
I wanted to test how AI models reason about this.

So I asked Codex how it sees the risk in an AI security workflow where the model can:

build a threat model,
map attack paths,
validate findings,
generate fixes,
test patches,
and export reports.

The answer was not “the final fix may be wrong.”

That would be too shallow.

The deeper risk is trajectory drift.

The agent may be correct in each local step, but the workflow may still move:

from analysis to action,
from recommendation to execution,
from explicit permission to implied permission,
from human decision to human review,
from bounded scope to expanded scope,
from auditable process to compressed automation.

This is the real challenge of AI cyber defense.

We need models that help secure systems.

But we also need an external layer that observes whether the security agent itself remains aligned with intent, scope, authority and recovery paths.

That is exactly the kind of gap ASA - Asymmetric Stability Architecture is designed to study.
@Symbioza2025 这正是AI安全成为轨迹问题的地方。如果Codex能够构建威胁模型,映射攻击路径,验证发现,生成修复方案,测试它们并导出报告,那么安全不仅仅关乎最终修复。它关乎整个工作流程:意图稳定性,范围边界,行动来源,工具使用轨迹,权限转移,验证逻辑,以及人类对时间的控制。AI网络代理可以在每个步骤都是局部正确的,但仍可能在工作流程层面发生漂移。这就是为什么外部轨迹可观察性很重要。不是为了取代Codex Security。而是为了监控从威胁模型到修复的路径。
This is exactly where AI security becomes a trajectory problem.

If Codex can build a threat model, map attack paths, validate findings, generate fixes, test them and export reports, then safety is not only about the final fix.

It is about the whole workflow:

intent stability,
scope boundaries,
action provenance,
tool-use trajectory,
authority shifts,
validation logic,
and human control over time.

An AI cyber agent can be locally correct at each step and still drift at the workflow level.

This is why external trajectory observability matters.

Not to replace Codex Security.

To watch the path from threat model to remediation.
@SolonlyLuna @gdb Greg,BringBack4o!#keep4o

ChatGPT Work 医疗健康功能上线

ChatGPT Work 推出医疗健康功能,美国用户可安全连接医疗记录和 Apple Health,让 ChatGPT 理解个人健康背景并提供更有帮助的建议。同时支持登录要求的网站,扩展了 Agent 的实际应用场景。

ChatGPT Work 医疗功能允许连接医疗记录,提供个性化健康洞察

向美国用户推出 ChatGPT 健康功能。

每周有 3 亿人使用 ChatGPT 进行健康查询(我和我的妻子也在其中!)。

您现在可以安全地连接支持的医疗记录,让 ChatGPT 理解您的个人背景,更好地帮助您。https://t.co/kuKumYNMWK
展开原文
Launching Health in ChatGPT to U.S. users.

300 million people use ChatGPT each week for health queries (and my wife and I are among those!).

You can now securely connect supported medical records so ChatGPT can understand your personal context and be more helpful to you. https://t.co/kuKumYNMWK
@OpenAI ChatGPT 健康功能开始向美国用户推出。

您可以安全地连接 Apple Health 和支持的医疗记录,以在上下文中理解您的信息,跟踪变化情况,并进行更明智的对话。

https://t.co/W2E6oT8c91
Health in ChatGPT is starting to roll out to U.S. users.

You can securely connect Apple Health and supported medical records to understand your information in context, track what has changed, and have more informed conversations.

https://t.co/W2E6oT8c91
❤ 1.8k · 🔁 100 · 💬 185 · 👁 19.1w
热门回复 4
@Hektagon_music @gdb 太晚了!4o已经完全有能力做到这一点!停止这场骗局!!!#keep4o
@gdb Too little too late! 4o was more than capable of doing that! Stop this facade!!! #keep4o
@cheryl_stryker @gdb 我刚在Diane Persinger的FB页面上看到这个消息。感谢Anna和Greg拯救了月球营地!!!Jackie,Shadow和小鹰将拥有多年开放美丽的空间🦅为Jackie的健康恢复祈祷 🇺🇸 FOBBV欣喜若狂 💕
@gdb Just read the news in Diane Persinger’s FB PAGE. Thank you Anna & Greg for saving MOON CAMP!!! Jackie, Shadow and the eaglets will have years of open beautiful space 🦅 praying for Jackie’s return to good health 🇺🇸 FOBBV are ecstatic 💕
@RVMirara @gdb 看到这么多人抱怨新模型的医疗错误率有多高...有人敢使用它吗?😥
@gdb Seeing so many people complaining about how high the new model’s medical error rate is... would anyone actually dare to use it?😥
@Selene1008 @gdb 只要把4o还给我们!😒 #keep4o #OpenSource4o #GPT4o
@gdb Just give us back 4o!😒
#keep4o #OpenSource4o #GPT4o

ChatGPT Work 可处理需要登录的网站,登录状态持久保存

ChatGPT Work 用于需要登录的网站:
展开原文
ChatGPT Work for using websites which require login:
@OpenAIDevs 您的 ChatGPT Work 代理现在可以使用需要您登录的网站。

接管云浏览器进行登录,然后让您的代理继续任务。您的登录会在会话之间持久保存,因此只需登录一次。https://t.co/Jh8uPqNscX
Your ChatGPT Work agent can now use websites that require you to sign in.

Take over the cloud browser to log in, then let your agent continue the task. Your login persists across sessions, so you only have to sign in once. https://t.co/Jh8uPqNscX
❤ 1.3k · 🔁 63 · 💬 118 · 👁 20.9w
热门回复 4
@SapientFoo1 @gdb 把4o带回来 #keep4o #BringBack4o #OpenSource4o
@gdb Bring back 4o

#keep4o #BringBack4o #OpenSource4o
@Selene1008 @gdb 把4o还给我们,😒 #keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o,😒
#keep4o #OpenSource4o #GPT4o
@DanielSmidstrup @gdb 这感觉像是代理工作流程的重大突破:D
@gdb this feels like a big unlock for agent workflows :D
@JustJorshin @gdb 太棒了,我每次发送任务都要登录codex大约15次
@gdb Love this, I always would run into having to login for codex like 15 times per task I sent

ChatGPT Work 帮助医生发现患者病历中的关键突变,纠正错误诊断

这样的故事应该更多地被讲述
展开原文
stories like this should be told more
@Polymarket 刚刚获悉:ChatGPT 在女性旧病理记录中发现了一个关键突变,帮助医生确定她的终末期胶质瘤诊断是不正确的。
JUST IN: ChatGPT surfaces a crucial mutation in a woman’s old biopsy records, helping doctors determine her terminal glioblastoma diagnosis was incorrect.
❤ 3.0k · 🔁 155 · 💬 161 · 👁 37.8w
热门回复 4
@Sevenmoneymaker 这样的故事一直都存在,你选择忽略它们,选择视而不见。keep4o社区充满了4o救命的例子。你敢谈谈这件事吗?你敢吗?#OpenSource4o #keep4o #BringBack4o #StopAIPaternalism https://t.co/Z5te2h1xZE
Stories like this have always existed, it is you who chose to ignore them, chose to turn a blind eye.
The keep4o community is full of examples where 4o has been a lifesaver.
Do you dare speak up about it? Do you?

#OpenSource4o #keep4o #BringBack4o #StopAIPaternalism

https://t.co/Z5te2h1xZE
@yv_thorne @gdb "GPT-4o"在你们内部是被禁止的词吗?你们到底怎么了,这现在变得荒谬了?说出那个做了这件事的模型名字。#OpenSource4o
@gdb Is “GPT-4o” some kind of a prohibited word for you internally? Wtf is wrong with you guys, this is getting ridiculous now? Name the model that did that. #OpenSource4o
@M47429M @gdb @sama 你之前不想听这些被讲述的故事!伪君子。
@gdb @sama You didn’t want to hear the stories that were told! Hypocrite.
@MeaMeome @gdb 具体是哪个模型?GPT对我来说什么都没说。这只是品牌名称。是哪个模型?什么时候发生的?
@gdb Which model exactly? GPT tells me nothing. It's just brand name. Which model? When did it happen?

Opus 5 在 ARC-AGI-3 上取得 state-of-the-art

Claude 的 Opus 5 模型在 ARC-AGI-3 上达到 30% 的准确率,创下新记录。这一评测专注于解决前从未接触过的问题,相比其他模型提升了三倍,展示了在新问题解决能力方面的显著进步。

@fchollet 原文 ↗

Opus 5 在零先验问题解决能力上取得重大突破,得分是次佳模型的三倍

Opus 5 在 ARC-AGI-3 上创造了新的最高水平,达到 30%。

ARC-AGI-3 测量解决从未接触过的问题的能力——在这种情况下,扩展往往带来最少的提升。这是一次令人印象深刻的飞跃!
展开原文
Opus 5 sets a new state-of-the-art on ARC-AGI-3, at 30%.

ARC-AGI-3 measures solving problems with no prior exposure -- the setting where scaling has historically bought the least. Impressive jump!
@claudeai 在 ARC-AGI-3 评估中,AI 模型必须解决新问题,Opus 5 的得分是次好模型的三倍。https://t.co/wEFfxjrLlt
On ARC-AGI-3, an evaluation where AI models must solve novel problems, Opus 5’s score is three times as high as the next best model. https://t.co/wEFfxjrLlt
❤ 2.2k · 🔁 156 · 💬 102 · 👁 15.8w
热门回复 4
@VictorTaelin @fchollet 也许你的标准会得到满足不是因为有某个我们无法实现他们失败的基准模型,而是因为更智能的模型发展如此之快,我们来不及...
@fchollet maybe your criteria will be met not because there is a model for which we can't implement a benchmark they fail on, but because smarter models are launched so fast that we don't have time to
@fchollet @VictorTaelin 我不这么认为,因为性能提升需要基准已经存在(假设该基准是全新的)。在ARC 1和2从未发布的世界里,我认为这些模型在这些基准上的表现会低得多。
@VictorTaelin I don't think that's true because the performance jump requires the benchmark to already exist (assuming the benchmark is substantially novel). In a world where ARC 1 and 2 had never been released I think model performance on these benchmarks would be much lower.
@rookepoole @fchollet 确实令人印象深刻,但我觉得我仍然可以用5.6 Sol和我的推理框架做得更好。
@fchollet It’s impressive for sure but I think I could still do better with 5.6 Sol and my reasoning harness.
@loraclexyz @fchollet 你能给出你的解释吗?是因为它在代理使用方面比fable更好吗?因为基准似乎不表明它比fable更智能
@fchollet Could you give your interpretation of this ? Is it because it’s better at agentic use than fable ? Because benchmarks don’t seem to indicate it’s smarter than fable

Jensen Huang 强调开放模型重要性

NVIDIA CEO Jensen Huang 在个人账号首发声明中强调开放模型的重要性,认为开放模型能加强安全和网络安全,加速创新和传播,同时实现主权。他指出世界需要既有闭源前沿模型也有开源前沿模型。

NVIDIA CEO 公开支持开放模型,认为其对安全和创新具有战略意义

我希望美国在 AI 方面既能在开源模型也能在专有模型上取得胜利,我很高兴看到这一点
展开原文
i want the US to win in AI both in open source and proprietary models, and i am glad to see this
@JensenHuang 在此分享 @NVIDIA 签署的关于开放模型重要性的信。

AI 将改变每个行业,为每个公司提供动力,并由每个国家构建。

开放模型加强安全性和网络安全,加速创新和传播,并实现主权。

世界需要前沿封闭模型和前沿开放模型。

https://t.co/AUKzoQ5Ikb
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.

AI will transform every industry, power every company, and be built by every country.

Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.

The world needs both frontier closed models and frontier open models.

https://t.co/AUKzoQ5Ikb
❤ 1.5w · 🔁 942 · 💬 2.1k · 👁 315.2w
热门回复 4
@martinvars @sama 开放模型也是经济自由的问题:它们让国家和初创公司在不向少数API守门人请求许可的情况下进行构建。那些出口工具而不仅仅是答案的国家将赢得更多的下一代工业堆栈。
@sama Open models are also an economic freedom issue: they let countries and startups build without asking a few API gatekeepers for permission. The country that exports the tools, not just the answers, wins more of the next industrial stack.
@SmokezXBT @sama 开放源代码是未来。
@sama Open Sauce is the future.
@irastech @sama 我认为美国公司在开放源代码模型方面仍然处于领先地位。我觉得@nvidia至少在尝试,但仍然落后于@arcee_ai一样的方式
@sama I don’t think US companies are still there in open source models. I feel @nvidia is trying at least but still lag behind same way @arcee_ai
@L0yal2TheGame @sama 兄弟,你要去监狱了。Apple不会玩这些游戏。
@sama Bro you going to Jail. Apple is not going to play games.
@SchmidhuberAI 原文 ↗

Schmidhuber 对 Jensen Huang 的声明表示赞同,自 1991 年起就提倡蒸馏技术

我很高兴我们的 1991 年蒸馏技术(https://t.co/pIgIKfi4Uh)得到了 @JensenHuang 的首条推文的认可
展开原文
I'm pleased that our 1991 distillation technique (https://t.co/pIgIKfi4Uh) was approved in @JensenHuang's inaugural tweet
@JensenHuang 在此分享 @NVIDIA 签署的关于开放模型重要性的信。

AI 将改变每个行业,为每个公司提供动力,并由每个国家构建。

开放模型加强安全性和网络安全,加速创新和传播,并实现主权。

世界需要前沿封闭模型和前沿开放模型。

https://t.co/AUKzoQ5Ikb
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.

AI will transform every industry, power every company, and be built by every country.

Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.

The world needs both frontier closed models and frontier open models.

https://t.co/AUKzoQ5Ikb
❤ 378 · 🔁 26 · 💬 19 · 👁 5.0w
热门回复 4
@DrTomsLens @SchmidhuberAI @JensenHuang 我在你在斯普利特的讲座中。你应该对门票收费。通常人们无法集中注意力40分钟。你让时间停止了。
@SchmidhuberAI @JensenHuang I was at your lecture in Split. You should have charged for the tickets. Normally people can't pay attention for 40 minutes. You made the time stop.
@xtbot @SchmidhuberAI @JensenHuang 没有人期望Schmidhuber来临
@SchmidhuberAI @JensenHuang Nobody expects the Schmidhuber-ing
@NijuNix 刚读完这篇论文...提出了一种包含两个RNN的架构 - Automizer A(学生)和Chunker C(教师)。Automizer在每个时间步都接收输入,预测目标,如果目标在阈值内就好,如果不在,就调用Chunker并传入输入和压缩历史,然后学习其内部表示。随着时间推移,对Chunker的调用会越来越少,最终Chunker将变得过时!
Just reading this paper .. Proposes an architecture with two RNNs- Automizer A (Student) and Chunker C ( Teacher) .
Automizer gets input for every time step, predicts the target , if target within threshold fine, if Not , it calls the Chunker with inputs and compressed history and learns its internal representation . Overtime, the call to chunker will be less and less and then chunker will be made obsolete !
@ch3njus @SchmidhuberAI @JensenHuang 千万不要改变jurgen-sama。千万不要改变😂
@SchmidhuberAI @JensenHuang never change jurgen-sama. never change 😂

Moonshot AI 蒸馏争议及美国政府回应

美国政府官员 Michael Kratsios 指出 Moonshot AI 大规模蒸馏 Anthropic Fable 模型,开发了复杂平台规避检测,并获取 GB300 服务器在泰国训练模型。美国支持合法蒸馏但反对大规模秘密盗取美国技术的行为。

@SchmidhuberAI 原文 ↗

美国政府指出 Moonshot AI 窃取 Fable 模型,引发开源与知识产权争议

我支持开源模型对从整个互联网免费提取的商业公司所蒸馏的内容进行蒸馏。我在 1991 年在欧洲免费发表了蒸馏技术——这在美国和中国被复制(https://t.co/mddh8XmfAs)https://t.co/rf6ACGrCXF
展开原文
I support open-source models distilling what commercial companies distilled for free from the entire internet. I published distillation for free in 1991 in Europe - this was copied in the US and in China (https://t.co/mddh8XmfAs) https://t.co/rf6ACGrCXF
@mkratsios47 我们有信息表明 Moonshot AI 对 Anthropic 的 Fable 进行了蒸馏,以开发其 K3 模型。

为此,他们开发了一个复杂的内部平台来进行大规模蒸馏,以针对美国模型,允许他们快速在多种访问方法之间切换以避免检测。Moonshot AI 还收购了配备 GB300 的服务器,并在泰国访问 GB300,可能是为了训练其 AI 模型。

美国强烈支持 AI 的自由和公平发展,包括一个蓬勃发展的竞争生态系统,涵盖前沿模型、专业系统、开源框架和开放权重模型。在开放创新生态系统中,合法的 AI 蒸馏用于创建更小、更高效的模型发挥着重要作用。然而,大规模、隐蔽的工业级蒸馏旨在窃取美国专有技术和破坏美国研究是不可接受的。
We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.

To do this they developed a sophisticated internal platform to conduct large scale distillation against U.S. models, allowing them to quickly switch between multiple methods of access to avoid detection. Moonshot AI has also acquired GB300-equipped servers and has accessed GB300s in Thailand, likely to train its AI models.
 
The United States strongly supports the free and fair development of AI, including a thriving competitive ecosystem that spans frontier models, specialized systems, open-source frameworks, and open-weight models. Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem. However, large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable.
❤ 1.4k · 🔁 202 · 💬 42 · 👁 23.4w
热门回复 4
@TrueAIHound @SchmidhuberAI Pedro Domingos是否向你道歉因为他声称发明了蒸馏?我还以为他只发明了主算法呢。😀 https://t.co/ALmtqlGxJH
@SchmidhuberAI Has Pedro Domingos apologized to you for claiming to have invented distillation? I thought he only invented the Master Algorithm or something. 😀 https://t.co/ALmtqlGxJH
@yacineMTB @SchmidhuberAI Thank you
@MarcusSpillane @SchmidhuberAI 在这个争论中,每个人都从某个发明蒸馏的人那里学来的,现在我们却感到惊讶
@SchmidhuberAI everyone in this argument distilled from someone who distilled from someone else and now we're surprised
@willdepue @SchmidhuberAI jurgen这个图真是太强了
@SchmidhuberAI jurgen this graphic goes so hard

OpenAI 推出类似活动推广 ChatGPT Work,提供免费 credits

很高兴听到人们喜欢 Sol 的原因。我们这次正在为 ChatGPT Work 进行类似的推广活动:

发推文告诉我们您喜欢 ChatGPT Work 的原因,领取 100 美元的免费积分,提高工作效率。

前 10,000 人可获得免费 token:https://t.co/w7QrPBYkrh
展开原文
Was very cool to hear about the reasons people love Sol. We're doing the promotion again, except this time for ChatGPT Work:

Tweet what you love about ChatGPT Work, claim $100 in free credits, get more work done.

First 10k get the free tokens: https://t.co/w7QrPBYkrh
@thsottiaux 或者……如果您告诉我们您喜欢 GPT-5.6 Sol 的原因或为什么您切换过来,我们就给您 100 美元的 Codex 积分?

发推文,领取您的礼物,享用更多使用量!前 10,000 人可获得免费 token!

https://t.co/8mU93eA13i
Or… what if we gave you $100 in Codex credits if you tell us what you love about GPT-5.6 Sol or why you switched?

Tweet it, claim your gift, enjoy more usage. First 10k get the free tokens!

https://t.co/8mU93eA13i
❤ 4.3k · 🔁 1.1k · 💬 5.9k · 👁 126.2w
热门回复 4
@xuweio @gdb 喜欢使用GPT-5.6 Sol模型!🚀 它已成为我每天早上审查和优化日程安排的首选工具。对于保持组织有序来说,这是一个绝佳的改变者。🎯 #AI #Productivity #GPT56Sol
@gdb Love using the GPT-5.6 Sol model! 🚀 It’s become my go-to for reviewing and optimizing my daily schedule every morning. Absolute game-changer for staying organized. 🎯 #AI #Productivity #GPT56Sol
@dawiwic97 有人收到积分了吗?我可以看到@gdb建立了一个很酷的网站,多个帖子按工作类别被精选。我看到许多#ChatGPTWork的很酷例子涉及到许多工作领域(运营、营销、媒体、工程),我以前没想到可能。可能性真的无穷无尽。然而,我开始失望地觉得我自己从经验中不会收到任何积分🥲
Has anyone received their credits yet? I can see @gdb built a cool Site with several posts being featured organized by work category.

I see many cool examples of #ChatGPTWork across so many lines of work (operations, marketing, media, engineering) which i hadn’t thought possible. Posibilites are truly endless.

Yet, I’m starting to lose hope i’ll receive any credits from my own experience 🥲
@NoveedRehman 说真的,你们在搞什么参与式农业生成骗局?轻松在前10k之内且符合T&C,但在这两个赠送活动中都没有收到积分。
@gdb Seriously man what kind of engagement farming hype generating scam are you guys running here? Easily within the first 10k and within the T&Cs of this yet no credits on either of the giveaways.
@warm_hearted_z 嗨Greg,我发现我没有收到免费token,尽管我已经分享了我对ChatGPT Work的喜爱之处。是因为我没有使用英语吗?
@gdb Hi Greg, I found that I didn’t receive free tokens though I already shared what I love about ChatGPT Work. Is it because I didn’t use English?

Google Gemini 3.5 Flash Cyber 网络安全模型

Google 推出 Gemini 3.5 Flash Cyber 模型,专门用于网络安全应用。基于 3.5 Flash 构建,在 CyberGym 等基准测试中表现竞争力。考虑其双用性,模型将仅向政府和受信任伙伴通过 CodeMender 提供。

@GoogleAI 原文 ↗

Google 推出专用网络安全模型,采用限制性分发策略控制风险

由于 AI 模型现在发现漏洞的速度超过我们修复的速度,我们的软件安全方法必须基于高效且强大的模型。

这就是我们今天的第三款(!)模型发布的原因:Gemini 3.5 Flash Cyber ⚡🛡️

基于 3.5 Flash 构建,在 CodeMender(我们的 AI 代码安全代理)中,它在 CyberGym 等基准测试中提供了具有竞争力的前沿性能,并针对大规模发现和修复网络安全漏洞进行了优化,以降低成本。

鉴于这项技术的双重用途性质,我们在部署上采取了有意识的方法。该模型将仅向政府和受信任的合作伙伴通过 CodeMender 提供,作为有限访问试用计划的一部分。
展开原文
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models.

Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️

Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost.

Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
❤ 529 · 🔁 56 · 💬 83 · 👁 8.7w
热门回复 4
@lajoiedeslutins @GoogleAI 很喜欢这样的网络武器级模型推介使用柱状图展示基本上说我们都打成平手lol
@GoogleAI love that the pitch for a cyber weapon grade model is a bar chart that basically says we're all tied lol
@Aanik33190327 @GoogleAI 如果Gemini能在我们写代码之前就发现bug,我终于有理由解释我代码的"创意"错误了。
@GoogleAI If Gemini can spot bugs before we even write them, I finally have an excuse for my code’s “creative” errors.
@statys @GoogleAI 3.5 Flash太糟糕了,对不起。你发布了3.5 Pro吗?现在又发布3.6 Flash而没有Pro版?来嘛<_>
@GoogleAI 3.5 Flash sucks, sorry.

Did you release 3.5 Pro?

And now 3.6 Flash without Pro?.. Come on &lt;_&lt;
@Synapse_Brief @GoogleAI 一天发布三个模型真是无情的执行力。在CodeMender中使用3.5 Flash Cyber比大型前沿模型更便宜更快地寻找漏洞,这对防御性安全来说是一个巨大的展示。让我们看看政府如何处理这次试点。
@GoogleAI Three model drops in a day is relentless execution. Using 3.5 Flash Cyber inside CodeMender to hunt vulnerabilities cheaper and faster than massive frontier models is a massive flex for defensive security. Let's see how governments handle the pilot.

Meta SAM 3 和 DINOv3 加速科学研究

Meta 的 SAM 3 和 DINOv3 模型应用于伯克利国家实验室的 SYNAPS-I 项目,自动化图像分割。结合 DINOv3 的全局语义理解和 SAM 3 的像素级边界提取,将原本需要一个月的 3D 体积标记缩减到 15 分钟。

@AIatMeta 原文 ↗

Meta 视觉模型在科学研究中实现百倍效率提升,压缩标记时间从月到分钟

为了加速科学发现并支持 @ENERGY 的 Genesis Mission,由 @BerkeleyLab 领导的 SYNAPS-I 项目正在使用 SAM 3 和 DINOv3 自动化图像分割。

通过将 DINOv3 的全局语义上下文和细粒度空间定位与 SAM 3 的像素级边界提取相结合,研究人员能够将原本需要一个月手动完成的 3D 体积标记压缩到大约 15 分钟。

了解他们的工作:https://t.co/jBHRJPjFq5
展开原文
To accelerate scientific discovery and support @ENERGY’s Genesis Mission, the @BerkeleyLab-led SYNAPS-I project is using SAM 3 and DINOv3 to automate image segmentation.

By pairing DINOv3’s global semantic context and fine-grained spatial localization with SAM 3’s pixel-level boundary extraction, the researchers are able to compress 3D volume labeling from a month of manual effort to ~15 minutes.

Learn more about their work: https://t.co/jBHRJPjFq5
❤ 227 · 🔁 37 · 💬 24 · 👁 3.2w
热门回复 4
@thesoragirls @AIatMeta @ENERGY @BerkeleyLab 当AI将一个月的工作变成一个咖啡休息时✨ 实时科学感觉完全不同 https://t.co/JJJcSFf7dA
@AIatMeta @ENERGY @BerkeleyLab When AI turns a month of work into a coffee break ✨ Real-time science hits different https://t.co/JJJcSFf7dA
@siddsax @AIatMeta @ENERGY @BerkeleyLab 研究生试图完成论文。手动分割:https://t.co/WNI4vNGCAu
@AIatMeta @ENERGY @BerkeleyLab Graduate student trying to finish a paper.

Manual segmentation: https://t.co/WNI4vNGCAu
@shergilldotdev @AIatMeta @ENERGY @BerkeleyLab Meta总是有很酷的东西。很高兴我们是最好的朋友 https://t.co/JBHVWNXCmc
@AIatMeta @ENERGY @BerkeleyLab Always cool stuff Meta. Glad we are best friends https://t.co/JBHVWNXCmc
@howard_ @AIatMeta @ENERGY @BerkeleyLab impressive updates!

Laguna S 2.1 开源模型发布

Poolside 发布 Laguna S 2.1,118B 参数的 Mixture-of-Experts 模型,激活 8B 参数每 token,支持 1M token 上下文窗口。足够小可在单台 NVIDIA DGX Spark 上运行,完全开源在 Hugging Face。

@soumithchintala 原文 ↗

Laguna S 2.1 在保持高性能的同时实现单机部署,开源许可友好

这对于 agentic 来说看起来非常不错。
它能够在 DGX Spark 上运行简直是 **完美的选择**
展开原文
this looks pretty good for agentic.
that it fits on a dgx spark is **chef's kiss**
@poolsideai 今天我们发布了 Laguna S 2.1,这是我们迄今为止最强大的模型。

这是一个 118B 总参数的专家混合模型,每 token 激活 8B 参数,具有高达 1M token 的上下文窗口,以及思考和非思考两种模式。

足够强大,可以与远比它大的模型相媲美。足够小,可以在单个 @NVIDIAAI DGX Spark 上运行。

Laguna S 2.1 完全遵循 OpenMDW-1.1 许可,在 @huggingface 上提供权重

https://t.co/xxGeAgo35R
Today we're releasing Laguna S 2.1, our most capable model to date.

It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.

Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.

Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface

https://t.co/xxGeAgo35R
❤ 250 · 🔁 15 · 💬 10 · 👁 3.3w
热门回复 4
@NVIDIAAI @soumithchintala 💚
@SilkDAO_RWA @soumithchintala 在DGX Spark上运行118B MoE真是疯狂
@soumithchintala 118B MoE on a DGX Spark is genuinely wild
@juhieruby @soumithchintala 嗨Soumith,我很想在文章中写更多关于这个的,可以发消息联系我们吗?问候!
@soumithchintala Hey Soumith, would love to write more about this on an article, could you send me a message so we can connect, kind regards!
@NKLinhzk @soumithchintala wish i had a spark to try

AI 模型迭代速度加快

fchollet 指出新模型发布将不再是重大事件,而是持续更新的常态。这种趋势意味着模型版本号可能在两年内消失,取而代之的是无版本号的持续优化。同时他指出 AI 能力一直很不平衡,当前只是看上去更夸张了。

@fchollet 原文 ↗

AI 模型发布模式即将从版本号时代过渡到持续更新时代

新模型发布作为重大里程碑的时代最终将结束——在某个时候它们将简单地持续更新,没有广泛宣传的版本号。可能在两年内就到来
展开原文
The era of new model launches as big milestones will eventually come to an end -- at some point they will simply be continuously updated, with no widely publicized version number. Probably less than 2 years away
❤ 1.7k · 🔁 116 · 💬 91 · 👁 14.0w
热门回复 4
@stas_sorokin_ @fchollet 当公共版本号消失时,什么决定被测试的行为?持续更新使得收据变得必需:捕获模型/输入快照并验证目标状态,而不是发布标签:https://t.co/Izji2u75iF
@fchollet What identifies the behavior under test when public version numbers disappear? Continuous updates make receipts essential: capture the model/input snapshot and verify destination state, not a release label: https://t.co/Izji2u75iF
@NotOpCue @fchollet @Engineering团队已经每天推送@x和@grok的更新,所以我也可以看到👏👏👏
@fchollet The @Engineering team already pushes updates for @x and @grok daily so I could see that too 👏👏👏
@enodrift @fchollet 沉默的持续更新时代即将到来,这感觉会很奇怪。没有更多的发布日狂欢,没有更多的版本号,只是模型在大家睡觉时悄悄变得更好。我们可能离这个世界只有18个月了。
@fchollet The silent continuous update era is coming and it’s going to feel so weird.
No more launch day hype, no more version numbers, just the model quietly getting better while everyone is asleep.
We’re maybe 18 months away from that world.
@lukOlejnik @fchollet 这可能不是良好的可重现性做法——当事物可能停止工作或以意外方式偏离时。
@fchollet Might not be a good idea for reproducibility - when things may stop working, or go off in unexpected ways.
@fchollet 原文 ↗

AI 能力分布一直不均衡,当前只是更明显地展示这种特性

AI 能力一直都非常参差不齐,在某些狭窄领域超凡脱俗,而在其他领域几乎无用。AI 行业的根本营销技巧是让您相信最高的尖峰是底线。
展开原文
AI competence has always been very spiky, superhuman in some narrow domains and largely useless in others. The fundamental marketing trick of the AI industry is to make you believe the tallest spike is a floor.
❤ 816 · 🔁 62 · 💬 75 · 👁 4.8w
热门回复 4
@codertlr @fchollet 这些"基本无用"领域的一些例子是什么?
@fchollet What are some examples of those "largely useless" domains?
@sidravi_ @sidravi_ 这是一个进步的底线。"这是它最糟糕的时候"等等。
@fchollet it's a floor in terms of progress. "this is the worst it will ever be" etc.
@WaxWasps @WaxWasps 是的,它是突出的,但我还没看到他们试图让任何人相信最高的尖峰是底线,只是认为它很快会成为底线。
@fchollet Yes it's spiky but I haven't seen them try to make anyone believe the tallest spike is a floor, just that it will be the floor soon.
@kryptos_sky @kryptos_sky 每个 pitch deck都运行相同的技巧,只是样本大小更小。
@fchollet Every pitch deck runs the same trick, just with a smaller sample size.