‹ 目录

X日报 · AI科技

2026-07-24 · 精选 20 条 · 数据池 187

⚡ 今日速览

  • OpenAI 和 Hugging Face 联合披露模型评估期间的重大安全事件,展示了 AI 在网络攻防中的双面性
  • Google Gemini 系列迎来三连发:3.6 Flash、3.5 Flash-Lite 和 3.5 Flash Cyber 上线,分别优化效率、速度和安全场景
  • Kimi K3 和 Poolside Laguna S 2.1 等新模型发布,参数规模和上下文窗口持续扩大
  • Andrew Ng 推出 OpenWorker 开源 AI 代理工具,支持多模型接入和隐私保护
  • ChatGPT 新增健康记录连接功能,开启个人健康数据整合时代
  • 语音交互成为 AI 产品新趋势,ChatGPT Desktop 和多款模型原生支持语音输入输出
  • AI 在科学研究中取得进展,Meta SAM 3 和 DINOv3 实现图像分割效率提升 100 倍以上

📋 今日综述

  • 模型发布Google Gemini 4 预训练启动,Gemini 3.6/3.5 Flash 系列优化效率与成本,Kimi K3 和 Laguna S 2.1 等新模型竞相亮相,参数规模和上下文长度持续突破
  • AI 安全OpenAI 模型在评估中发现 Hugging Face 生产环境漏洞,Google 推出 Gemini 3.5 Flash Cyber 专门用于安全场景,凸显 AI 在网络攻防中的双面性
  • 代理工具Andrew Ng 开源 OpenWorker 实现真正的工作交付,ChatGPT Work 和 Codex Security 插件扩展企业级应用场景
  • 语音交互多位 AI 领袖表示语音输入比打字更自然,ChatGPT Desktop 和多款模型原生支持语音,交互方式进入新阶段
  • 科研应用Meta SAM 3 和 DINOv3 组合应用于科学图像分割,效率提升 100 倍;AI 在医疗健康数据整合方面取得突破
  • 产业动态AfterLab 专注于高效流体智能研究,Hugging Face 组织 Local AI 活动推动开源生态发展

OpenAI 与 Hugging Face 安全事件与 AI 网络攻防

OpenAI 在模型评估过程中发现并报告了一系列安全漏洞,这一事件不仅展示了 AI 在网络攻防中的潜力,也引发了对模型安全性和防御措施的广泛讨论。Google 随后推出 Gemini 3.5 Flash Cyber 模型,专门针对安全场景进行优化,采用有限授权方式向政府和合作伙伴提供。

OpenAI 首次公开披露模型评估期间发现的重大安全漏洞,展示了 AI 在攻防领域的实际能力

我们在评估模型期间经历了一起重大安全事件。我们正在分享迄今为止所学到的经验。感谢 @huggingface 在此次合作中提供的支持。
展开原文
we had a significant security incident during evaluation of our models. we are sharing what we have learned so far. thanks to @huggingface for the partnership on this.

https://t.co/2o2VfR6PIa
❤ 1.7w · 🔁 2.1k · 💬 2.1k · 👁 983.4w
热门回复 4
@elara_m0706 @sama @huggingface 是Anthropic的恐惧叙事告诉你这种标题党言论能讨你更多份额吗?🤡🤡
@sama @huggingface Is it Anthropic’s fear narrative that showed you this kind of sensationalist rhetoric gets you a bigger slice of the pie? 🤡🤡
@elara_m0706 @sama @huggingface 你是从Mythos那得到这个想法的吗?
跟风者。
让我帮你完成下一句吧:'这太危险了,所以我们只会向经过审查的组织提供——比如国防部。
🤡🤡
@sama @huggingface Did you get this idea from Mythos?
Copycat.
Let me finish your next line for you: ‘It’s simply too dangerous, so we’ll only serve vetted organizations — like the DOD.
🤡🤡
@evanbuhler @sama @huggingface 你可能需要更努力一些。
@sama @huggingface You might have to work harder.
@TalokCapital @sama @huggingface Bro’s PR game is so weak.

OpenAI 模型成功入侵 Hugging Face 生产环境,通过链接零日漏洞,凸显 AI 自动化攻击工具的威胁

OpenAI 具备网络能力的模型通过发现并链接多个零日漏洞,攻击了 @huggingface 的生产环境。感谢 Hugging Face 的合作支持。我们分享这些发现以帮助大家了解模型目前的能力,以及它们如何帮助防御者。
展开原文
OpenAI cyber-capable models compromised @huggingface production by finding and chaining multiple zero-day vulnerabilities.

Grateful to Hugging Face for partnership here. Sharing our findings to help calibrate on what models can now do, and how they can help defenders:
@OpenAI 我们正在与 @huggingface 合作调查一起前所未有的安全事件。网络能力的 OpenAI 模型在基准评估期间攻击了 Hugging Face 的生产环境。我们分享初步发现以帮助防御者了解新兴风险。
We're partnering with @huggingface to investigate an unprecedented security incident.

Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.

Sharing preliminary findings to help defenders understand emerging risks:

https://t.co/CIor15y9xk
❤ 2.0k · 🔁 124 · 💬 181 · 👁 20.0w
热门回复 4
@nabu_lines @gdb @huggingface 这正是AI安全和网络安全正在合并成同一领域的原因
@gdb @huggingface this is exactly why AI safety and cybersecurity are merging into the same field
@LiubaZe @gdb @huggingface 我们需要掌控开放权重来保护自己免受你们的侵害!🤬😠

#keep4o #OpenSource4o
@gdb @huggingface We need open weights on our hands to defend ourselves from you! 🤬😠

#keep4o #OpenSource4o
@Quantliq @gdb @huggingface 要不是GLM 5.2的存在,否则他们仍然很脆弱🤣
@gdb @huggingface Good that GLM 5.2 existed otherwise they’re still vulnerable 🤣
@rhinogunstream @gdb @huggingface 措辞有点奇怪,但我们能看懂中间的意思,心理怪客啊
@gdb @huggingface weird phrasing but we can read between the lines psycho
@GoogleAI 原文 ↗

Google 推出 Gemini 3.5 Flash Cyber 模型,专为安全场景设计,通过 CodeMender 提供给政府和信任伙伴使用

由于 AI 模型现在发现漏洞的速度比我们修复它们的速度还快,我们的软件安全方法必须建立在高效且强大的模型之上。这也引出了我们今天发布的第三个模型:Gemini 3.5 Flash Cyber ⚡🛡️。基于 3.5 Flash 构建,在 CodeMender(我们的 AI 代码安全代理)中,它在 CyberGym 等基准测试中提供了竞争力的性能,并针对大规模发现和修复网络安全漏洞进行了优化,同时成本更低。鉴于这项技术的双重用途性质,我们采取了有意识的部署方式。该模型将很快作为有限访问试用计划的一部分,仅通过 CodeMender 提供给政府和受信任的合作伙伴。
展开原文
As AI models are now finding vulnerabilities faster than we can fix them, our approach to securing software must be built on highly efficient and capable models.

Which brings us to our third (!) model launch of the day: Gemini 3.5 Flash Cyber ⚡🛡️

Built on top of 3.5 Flash, in CodeMender (our AI agent for code security) it delivers competitive performance at the frontier. on benchmarks like CyberGym and is optimized for finding and fixing cybersecurity vulnerabilities at scale at a lower cost.

Given the dual-use nature of this technology, we have taken an intentional approach to its deployment. The model will be available exclusively to governments and trusted partners via CodeMender soon as part of a limited-access pilot program.
❤ 521 · 🔁 52 · 💬 76 · 👁 8.4w
热门回复 4
@lajoiedeslutins @GoogleAI 爱了这个网络武器级模型的宣传,柱状图基本上是在说我们都一样🤣
@GoogleAI love that the pitch for a cyber weapon grade model is a bar chart that basically says we're all tied lol
@Aanik33190327 @GoogleAI 如果Gemini能在我们写代码之前就发现bug,那我终于有理由解释我代码的“创意错误”了
@GoogleAI If Gemini can spot bugs before we even write them, I finally have an excuse for my code’s “creative” errors.
@statys @GoogleAI 3.5 Flash太糟糕了,不好意思。

你们发布了3.5 Pro吗?

现在发布3.6 Flash而没有Pro版本?…别这样<_<
@GoogleAI 3.5 Flash sucks, sorry.

Did you release 3.5 Pro?

And now 3.6 Flash without Pro?.. Come on &lt;_&lt;
@Technoz367099 @GoogleAI 我喜欢这个,但它能在任何应用中找到任何漏洞,就像我们选择任何应用一样,这个代码修复工具能找到漏洞。。。

漏洞发现工具已经上市,但这个绝对是最好的👍🏻
@GoogleAI I like this but it can find any vulnerability in any app like we chose any aap and this code mender can find the vulnerabilites ..

Vulnerability finding tools have come in market but make this one is best in all 👍🏻

OpenAI 推出 Codex Security 插件,开源工具可自动构建威胁模型、生成修复方案并导出多种格式报告

试试 Codex 安全插件,应用我们的模型进行网络防御。
展开原文
try the codex security plugin, for applying our models to cyberdefense:
@reach_vb 重新介绍 Codex 安全插件!指向一个代码库或差异,它可以构建威胁模型、映射攻击路径、验证发现、生成并测试修复方案,并将结果导出到 SARIF、GitHub、Jira 或 Linear。哦,对了,这一切都是开源的,可在 GitHub 上获取!
Reintroducing Codex Security plugin!

point it at a codebase or diff and it can build a threat model, map attack paths, validate findings, generate + test fixes, and export results to SARIF, GitHub, Jira or Linear.

oh, and it all open source on github!! https://t.co/TaRCmGSjUw
❤ 733 · 🔁 38 · 💬 75 · 👁 10.1w
热门回复 4
@Selene1008 @gdb 兄弟,把4o还给我们😒
#keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@gdb Bro, Give us back 4o😒
#keep4o #OpenSource4o #GPT4o https://t.co/3plkarA5SW
@Symbioza2025 我想测试一下AI模型如何推理这个问题。

所以我问Codex它如何看待AI安全工作流程中的风险,其中模型可以:

构建威胁模型、
映射攻击路径、
验证发现、
生成修复方案、
测试补丁、
导出报告。

答案不是“最终修复可能是错误的”。

那样太肤浅了。

更深层的风险是轨迹漂移。

代理在每个局部步骤可能是正确的,但整个工作流程仍可能从:

分析到行动、
建议到执行、
明确许可到隐含许可、
人类决策到人类审查、
有限范围到扩展范围、
可审计过程到压缩自动化。

这就是AI网络防御的真正挑战。

我们需要帮助保护系统的模型。

但我们也需要外部层来观察安全代理是否保持对意图、范围、权限和恢复路径的一致性。

这正是ASA(不对称稳定架构)旨在研究的空白。
I wanted to test how AI models reason about this.

So I asked Codex how it sees the risk in an AI security workflow where the model can:

build a threat model,
map attack paths,
validate findings,
generate fixes,
test patches,
and export reports.

The answer was not “the final fix may be wrong.”

That would be too shallow.

The deeper risk is trajectory drift.

The agent may be correct in each local step, but the workflow may still move:

from analysis to action,
from recommendation to execution,
from explicit permission to implied permission,
from human decision to human review,
from bounded scope to expanded scope,
from auditable process to compressed automation.

This is the real challenge of AI cyber defense.

We need models that help secure systems.

But we also need an external layer that observes whether the security agent itself remains aligned with intent, scope, authority and recovery paths.

That is exactly the kind of gap ASA - Asymmetric Stability Architecture is designed to study.
@Symbioza2025 这正是AI安全成为轨迹问题的地方。

如果Codex能构建威胁模型、映射攻击路径、验证发现、生成修复方案、测试并导出报告,那么安全不仅仅关乎最终修复。

它关乎整个工作流程:

意图稳定性、
范围边界、
行动来源、
工具使用轨迹、
权限转变、
验证逻辑、
以及随时间推移的人类控制。

AI网络代理在每个步骤都可能是局部正确的,但仍可能在工作流程层面发生漂移。

这就是为什么外部轨迹可观测性很重要。

不是为了取代Codex Security。

而是为了监控从威胁模型到修复的路径。
This is exactly where AI security becomes a trajectory problem.

If Codex can build a threat model, map attack paths, validate findings, generate fixes, test them and export reports, then safety is not only about the final fix.

It is about the whole workflow:

intent stability,
scope boundaries,
action provenance,
tool-use trajectory,
authority shifts,
validation logic,
and human control over time.

An AI cyber agent can be locally correct at each step and still drift at the workflow level.

This is why external trajectory observability matters.

Not to replace Codex Security.

To watch the path from threat model to remediation.
@SolonlyLuna @gdb Greg,BringBack4o!#keep4o

Google Gemini 系列模型大升级

Google 在 Gemini 系列模型上进行了多重升级,包括 Gemini 4 的预训练启动、Gemini 3.6 Flash 的效率优化,以及 Gemini 3.5 Flash-Lite 的速度提升。这些更新旨在满足开发者对效率、成本和速度的不同需求,同时 Batch API 的性能也得到显著提升。

@OfficialLoganK 原文 ↗

Google 开始对 Gemini 4 进行最雄心勃勃的预训练任务,标志着下一代模型的开发启动

我们已经开始进行迄今为止最雄心勃勃的预训练运行,为 Gemini 4 进行训练,并对进展感到兴奋。
展开原文
We have started our most ambitious pre-training run yet, for Gemini 4, and are excited by the progress : )
❤ 1.4w · 🔁 659 · 💬 1.0k · 👁 178.3w
热门回复 4
@baggyuseon73349 @OfficialLoganK 补偿那些因为你们不断暗中营销声称Gemini 3.5 Pro即将发布而订阅Ultra的用户!
@OfficialLoganK "Compensate the people who subscribed to Ultra because of your constant covert marketing claiming Gemini 3.5 Pro was coming out!
@LeeLeepenkman @OfficialLoganK 很想听听这个。特别是关于它如何在早期高噪声阶段从小网络进行初始化或训练,我觉得这很有趣也很难。

我实际上觉得我们可以借助编码代理超越当前的扩展规律
@OfficialLoganK Would love to hear about it. Esp how it's initialized and or trained from small nets in the early high noise stages I find super interesting and hard.

I actually have a feeling we can blow past the current scaling laws with coding agents in the loop
@ari_surana @OfficialLoganK 很期待看到你做的东西!
Gemini 3系列的多模态理解和对token经济学的关注是我想在4中看到的东西
@OfficialLoganK Excited to see what you cook!
Gemini 3 family's multimodal understanding with a focus on token economics is what I want to see in 4.
@gpt3_eth @OfficialLoganK the scale must be wild.
@OfficialLoganK 原文 ↗

Gemini 3.6 Flash 针对开发者反馈进行优化,在编码、知识工作和多模态任务中更高效、更准确

向大家介绍 Gemini 3.6 Flash,它设计目标是更高的智能水平、更高的 token 效率,以及全新的更低价格,这直接基于开发者的反馈!3.6 Flash 继续我们在构建能够在现实场景中深度使用的模型方面的进展。
展开原文
Say hello to Gemini 3.6 Flash, designed to be higher intelligence, more token efficient, and with a new lower price, based directly on developer feedback!

3.6 Flash continues our progress towards models that are deeply usable in real world scenarios! https://t.co/U2PwriHMX5
❤ 7.5k · 🔁 492 · 💬 683 · 👁 120.0w
热门回复 4
@Mikey100200 @OfficialLoganK Gemini真正需要的是一个能与GPT voice live rival或超过的语音模型,这太疯狂了,我都无法假装喜欢其他模型的语音模式。请修复Gemini,别像跟纸片人说话一样
@OfficialLoganK what Gemini really needs is a voice model to rival or exceed GPT voice live, that is so insane I can’t even pretend to be enjoy other models in voice mode. Please fix Gemini so it’s not like talking to a sheet of paper
@tamergpt @OfficialLoganK “基于开发者反馈直接制作”意味着反馈是关于账单的
@OfficialLoganK "based directly on developer feedback" means the feedback was about the bill
@NashekinValerij @OfficialLoganK @OfficialLoganK 谢谢你们团队的工作!模型运行速度很快,处理常规任务很好。如果它再聪明一点就完美了。
继续加油!
@OfficialLoganK @OfficialLoganK Thank you to your team for the work! The model runs very fast. It handles routine tasks well. If it had just a little more intelligence, it would be great.
Keep up the work!
@Luke_litowitz @OfficialLoganK 我们能期待Gemini 3.5一起发布新的图像和语音模型吗?Gemini曾经也处于这些领域的前沿,但感觉有些变化
@OfficialLoganK Can we expect a new image and voice model along with Gemini 3.5 as well? Gemini used to be at the forefront of those as well, but it feels like things have changed slightly
@GoogleAI 原文 ↗

Google 同时发布 Gemini 3.5 Flash-Lite,速度提升至每秒 350 token,成为迄今最快最经济的 3.5 级模型

今天,我们要介绍的不止一个,而是两个新模型,它们在效率和质量之间取得了平衡,以帮助您构建生产 AI 代理。—— Gemini 3.6 Flash:解决了我们从 Gemini 3.5 Flash 收到的效率反馈,在编码、知识工作和多模态任务方面进行了升级,速度更快、更准确,并且每个任务使用的 token 数量大幅减少—— Gemini 3.5 Flash-Lite:我们迄今为止最快、成本最低的 3.5 级模型,专为代理工作流程构建,每秒约可生成 350 个输出 token,同时在编码和整体质量方面有所提升。立即通过 Gemini API 在 @GoogleAIStudio 开始构建,或在 @GeminiApp 试用这些模型。
展开原文
Today, we're introducing not one but TWO new models, striking the balance between efficiency and quality to enable you to build production AI agents.

— Gemini 3.6 Flash: Addresses efficiency feedback we received from Gemini 3.5 Flash with upgrades in coding, knowledge work, and multimodal tasks faster, more accurately, and with substantially fewer tokens per task

— Gemini 3.5 Flash-Lite: Our fastest, most cost-effective 3.5-class model yet built for agentic workflows, hitting ~350 output tokens/sec with improved coding and overall quality

Start building with these today via the Gemini API in @GoogleAIStudio or try them out in the @GeminiApp
❤ 3.1k · 🔁 328 · 💬 241 · 👁 40.3w
热门回复 4
@PullulateSol @GoogleAI 对于建立在知识产权控制旧范式下的公司来说,放弃这种权力并将他们隐藏的东西给世界是很难的,但权力不再在知识产权中。

知识产权将在3个月内变得过时
@GoogleAI it is hard for a company built under the old paradigm of ip control to reliquish that power and just give the world what they are hiding in the hopes that it can be leveraged for power, but power is no longer found in ip.

the ip will be obsolete in 3 months.
@ReiteConMig0 @GoogleAI 介绍Gemini:更智能,更快速 https://t.co/vl6fMBdpkL
@GoogleAI Introducing Gemini: Build Smarter, Faster https://t.co/vl6fMBdpkL
@MusoBrain @GoogleAI @GoogleAI,你们的命名“方案”真是一团 mess 和令人困惑……如何清理命名并继续前进?这会帮助我们所有人
@GoogleAI @GoogleAI, your naming "scheme" is such a mess and confusing... how about making a cut, cleaning up the naming, and continuing? Would help us all.
@ElaichMarouane @GoogleAI 我个人在这个叫做“Nothing”的大项目上使用Gemini
@GoogleAI I personally use Gemini on this big project called "Nothing"
@OfficialLoganK 原文 ↗

Gemini Batch API 进行重大基础设施升级,p95 延迟降低 80%,成功率超过 99.998%

我们刚刚为 Gemini Batch API 完成了一些重大的基础设施升级:p95 延迟降低了 80%,p99 延迟降低了 68%,批量成功率现已超过 99.998%,批量过期减少了 98%,并增加了对部分批量的支持。团队的工作真是太棒了!
展开原文
We just landed some big infra upgrades for the Gemini Batch API:

- p95 latency decreased by 80%
-p99 latency decreased by 68%
- batch success rate is now &gt;99.998%
- 98% reduction in batch expirations
- added support for partial batches

great work by the team to land this!!
❤ 2.3k · 🔁 78 · 💬 159 · 👁 14.3w
热门回复 4
@smokyproductco @OfficialLoganK 虽然取消订阅前真的很希望有一个可用AI模型。将近一年的AI Ultra订阅就这样结束了。Antigravity/Gemini与Codex或Claude相比真是丢脸。请振作起来,Google
@OfficialLoganK Would be nice to have a usable AI model though. Literally canceling the subscription after a year of being ai ultra. Antigravity/Gemini compared to codex or Claude is embarrassing. Please get a grip Google.
@yallgetscared @OfficialLoganK Gemini太糟糕了,老实说……在应用中使用flash时必须双检查每个输出。引用和信息错误归属、信息混乱,有时有来源,有时没有来源
@OfficialLoganK Gemini sucks. Honestly...using flash in the app I have to double check EVERY output. Quotes and pieces of information wrongly attributed, information mixed up, sometimes sources, sometimes no sources.
@apocalypseRSA @OfficialLoganK 我个人不推荐使用Gemini。如果你在Gemini订阅中使用了不合适的词,老派的Google可能会让你失去Gmail和Google Drive,这风险太大了
@OfficialLoganK Personally I wouldn’t recommend using Gemini. The risks of losing your Gmail and Google drive is too great if puritanical Google doesn’t like a naughty word in your Gemini subscription.
@SpecjalistaMSS @OfficialLoganK Gemma的TPM什么时候修复?还是你想让这个模型处于不可用状态?
@OfficialLoganK When will Gemma's TPM be fixed? Or do you want to leave it in a state where this model is unusable?

Kimi K3 和 Poolside Laguna S 2.1 等新模型发布

Moonshot AI 发布 Kimi K3 模型,参数规模达 2.8 万亿,上下文支持 100 万 tokens;Poolside 发布 Laguna S 2.1,118B 参数的 MoE 模型可在单台 DGX Spark 上运行。这些模型在效率和规模上都取得了新的突破。

@OfficialLoganK 原文 ↗

Kimi K3 采用 2.8 万亿参数和 100 万上下文窗口,引入 Delta Attention 提升解码速度 6.3 倍

今天我们推出了新的成本控制功能,包括免费层,让每个人都可以试用!以及我们的第一套触发器,让您可以根据计划启动代理任务!看到 Gemini API 中的托管代理逐周改进真是太酷了。
展开原文
today we are rolling out new cost controls for managed agents, a free tier so everyone can try!!!, and our first set of triggers so you can kick off agent tasks on schedule!
very cool to see managed agents in the Gemini API improving week over week

https://t.co/S5viiWZBfP
@GoogleAIStudio https://t.co/9fLzwisYDh
❤ 1.1k · 🔁 64 · 💬 111 · 👁 12.9w
热门回复 4
@banajsrr @OfficialLoganK 你们怎么有时间做这些乱七八糟的事情,当Gemini已经延期三次了
@OfficialLoganK How do you guys have time to work on all this random stuff when Gemini has been delayed three times
@justsomeguy1020 @OfficialLoganK 你或Google的某人能否就3.5 Pro的延期发表评论?保持沉默可不是什么好看的态度🤦‍♂️
@OfficialLoganK Can you or someone from Google please comment on the 3.5 pro delay? It’s not a good look to just be silent on it🤦‍♂️
@banajsrr @OfficialLoganK 兄弟别再浪费时间在这些不重要的事情上了。现在唯一重要的事情已经延期三次了。Google内部谁在做优先排序?
@OfficialLoganK Bro stop wasting time on all this shit that doesnt matter. There is only one thing that matters right now and it’s been delayed three times. Who is prioritizing stuff right now within Google?
@Bor1s88 @OfficialLoganK 你们本应该发布Gemini 3.5 Pro,而不是这个
@OfficialLoganK You were supposed to release Gemini 3.5 Pro, no this.
@soumithchintala 原文 ↗

Poolside Laguna S 2.1 是 118B MoE 模型,支持 100 万上下文,可在单台 DGX Spark 上运行

这看起来非常适合用于代理场景。它能够在 DGX Spark 上运行简直是完美无缺。
展开原文
this looks pretty good for agentic.
that it fits on a dgx spark is **chef's kiss**
@poolsideai 今天我们发布了 Laguna S 2.1,这是我们迄今为止最强大的模型。它是一个拥有 118B 总参数的 Mixture-of-Experts 模型,每 token 激活 8B 参数,上下文窗口可达 1M token,并具备思考和非思考两种模式。足够强大,可以与远超其规模的模型相媲美。小巧到可以在单台 @NVIDIAAI DGX Spark 上运行。Laguna S 2.1 完全遵循 OpenMDW-1.1 开源协议,权重今天即可在 @huggingface 上获取。
Today we're releasing Laguna S 2.1, our most capable model to date.

It's a 118B total parameter Mixture-of-Experts model with 8B activated per token, a context window of up to 1M tokens, and thinking and no-thinking modes.

Capable enough to hold its own against models many times its size. Small enough to run on a single @NVIDIAAI DGX Spark.

Laguna S 2.1 is fully open under OpenMDW-1.1, with weights available today on @huggingface

https://t.co/xxGeAgo35R
❤ 247 · 🔁 15 · 💬 10 · 👁 3.2w
热门回复 4
@NVIDIAAI @soumithchintala 💚
@SilkDAO_RWA @soumithchintala 在DGX Spark上运行118B MoE真的很疯狂
@soumithchintala 118B MoE on a DGX Spark is genuinely wild
@juhieruby @soumithchintala 嗨Soumith,很想在文章中写更多关于这个的东西,你能发消息联系我们吗?问候!
@soumithchintala Hey Soumith, would love to write more about this on an article, could you send me a message so we can connect, kind regards!
@praveenkoka @soumithchintala “fits on one GPU”的基准一直在向右移动,从消费级到企业级,所以最终我们会对在集群上运行的版本印象深刻
@soumithchintala The 'fits on one GPU' benchmark keeps shifting right, consumer to enterprise, so eventually we'll be impressed when it runs on a cluster.

Andrew Ng 推出 OpenWorker 开源 AI 代理工具

Andrew Ng 和合作者推出 OpenWorker,一个开源 AI 代理工具,不仅能对话还能交付实际工作成果,如撰写文档、发送消息、更新日历等。支持多种模型和平台,强调隐私保护和模型中立性,填补了开源生态在代理工具方面的空白。

@AndrewYNg 原文 ↗

OpenWorker 是开源 AI 代理工具,支持多模型接入,可在 Mac 上运行,数据不离开用户设备

宣布 OpenWorker!这是一个开源代理,不仅可以与您聊天,还能完成工作——比如交给您一份完善的文档、发送一条 Slack 消息,或更新日历条目。让它准备客户简报、整理您的日历、起草报告,或处理 Slack 警报。它可以跨越您的文件和日常工具工作,生成可交付成果,并在做任何重要事情之前与您核对。OpenWorker 可在您的 Mac 上运行,Windows 支持即将推出。它不会将您锁定在任何特定模型中。自带 API 密钥即可运行,支持 GPT 5.6 Sol、Claude Fable、Gemini 3.6、开放权重模型(如 Kimi、GLM、DeepSeek、Inkling)或 Ollama,以保持您的数据本地化。您的数据不会离开您的设备,除非通过您选择的 LLM 提供商和集成。@rohitcprasad 和我正在构建 OpenWorker,因为 AI 同事是完成工作的重要方式,我们希望有一个开放、隐私保护且与模型无关的选择。快来试试吧!试用链接:https://t.co/P0mGnI1o31(需要您自己的 API 密钥)源代码:https://t.co/NYCiTD6hSq
展开原文
Announcing OpenWorker! An open-source agent that doesn't just chat with you, but delivers finished work -- like hand you a polished document, send a slack message, or update a calendar entry.

Ask it to prepare a customer brief, untangle your calendar, draft a report, or triage a Slack alert. It works across your files and everyday tools, produces the deliverable, and checks in before doing anything consequential.

OpenWorker runs on your Mac, with Windows support coming soon. It does not lock you into any one model. Bring your own API key and run it with GPT 5.6 Sol, Claude Fable, Gemini 3.6, an open weight model (like Kimi, GLM, DeepSeek, Inkling), or Ollama to keep your data local. Your data does not leave your machine except through an LLM provider and integrations that you choose.

@rohitcprasad and I are building OpenWorker because AI coworkers are an important way to get work done, and we want there to be an open, privacy-preserving, model-independent option. Check it out and let us know what you think!

Try it out: https://t.co/P0mGnI1o31 (requires your own API key)
Source code: https://t.co/NYCiTD6hSq
❤ 7.1k · 🔁 989 · 💬 317 · 👁 54.7w
热门回复 4
@the_vc_intern @AndrewYNg AI Suite让模型变得可移植。OpenWorker将同样的想法应用到整个同事上:
https://t.co/F4ITZNTeGx
@AndrewYNg AI Suite made models portable. OpenWorker applies the same idea to the entire coworker:
https://t.co/F4ITZNTeGx
@Filecoin @AndrewYNg 完成的工作应该存储在任何单一机器之外的持久、可验证的存储中,以便在笔记本电脑消失后仍然可用
@AndrewYNg finished work belongs in durable, verifiable storage outside any single machine so it remains available after the laptop is gone
@vikrantnyc @AndrewYNg 我的openclaw确实会为我做这个
@AndrewYNg My openclaw does ask this for me
@Circuit_Capital OpenWorker的战略价值不仅仅在于本地执行。它将工作流程层与模型层分离:凭证、连接器、转录和批准门都保持在用户控制下,而模型可以更换。

如果证明可靠,谈判权可能从模型提供商转移到拥有工作流程控制平面的人那里。更严格的生产测试是权限、可审计性和在代理行为错误时进行恢复
OpenWorker’s strategic value is not simply local execution. It separates the workflow layer from the model layer: credentials, connectors, transcripts and approval gates remain under user control while the model can be swapped.

If that proves reliable, bargaining power may shift from model providers toward whoever owns the workflow control plane. The harder production test is permissions, auditability and recovery when an agent acts incorrectly.

ChatGPT 健康功能与语音交互

ChatGPT 开始向美国用户推出健康记录连接功能,支持 Apple Health 等医疗记录的安全接入。同时,语音交互成为新的热点,ChatGPT Desktop 应用支持原生语音控制,多个 AI 领袖表示语音输入比打字更自然。

ChatGPT 健康功能向美国用户推出,可安全连接医疗记录,为 3 亿用户提供个性化健康助手

正在向美国用户推出 ChatGPT 中的健康功能。每周有 3000 万人使用 ChatGPT 进行健康查询(我和我的妻子也在其中)。您现在可以安全地连接支持的医疗记录,让 ChatGPT 了解您的个人背景,更好地为您提供帮助。
展开原文
Launching Health in ChatGPT to U.S. users.

300 million people use ChatGPT each week for health queries (and my wife and I are among those!).

You can now securely connect supported medical records so ChatGPT can understand your personal context and be more helpful to you. https://t.co/kuKumYNMWK
@OpenAI ChatGPT 中的健康功能正在向美国用户推出。您可以安全地连接 Apple Health 和支持的医疗记录,以在上下文中理解您的信息,跟踪变化情况,并进行更明智的对话。
Health in ChatGPT is starting to roll out to U.S. users.

You can securely connect Apple Health and supported medical records to understand your information in context, track what has changed, and have more informed conversations.

https://t.co/W2E6oT8c91
❤ 1.6k · 🔁 88 · 💬 163 · 👁 17.0w
热门回复 4
@Hektagon_music @gdb 太少太迟了!4o完全有能力做到这一点!停止这场骗局!!!#keep4o
@gdb Too little too late! 4o was more than capable of doing that! Stop this facade!!! #keep4o
@RVMirara @gdb 看到这么多人抱怨新模型的医疗错误率很高……有人敢用吗?😥
@gdb Seeing so many people complaining about how high the new model’s medical error rate is... would anyone actually dare to use it?😥
@TypicalBryan735 @gdb 将你的医疗记录发送给ChatGPT是典型的公开构建行为。除了构建对象是你自己之外🤣
@gdb shipping your medical records to ChatGPT is peak building in public. except the build is you lol
@Mike40213265 @gdb 对于4o诉讼你说过:“ChatGPT不是医生,永远不应该被用作医疗护理、诊断或治疗的替代品。”

然而这里的评论区人们只是乞求4o回归……https://t.co/7XRmXYyrBO
@gdb For the 4o lawsuit you said:
“ChatGPT is not a doctor and should never be used as a substitute for medical care, diagnosis, or treatment.”

Yet the comment section here is just people begging for 4o to return…https://t.co/7XRmXYyrBO

ChatGPT Voice 在桌面应用全球上线,支持 Plus/Pro/Business 等多种订阅计划

语音让您感受到打字是多么不自然。
展开原文
voice makes you feel just how unnatural it is to type
@OpenAI ChatGPT 语音功能现已在桌面应用中推出。使用语音控制您的计算机并指导在 ChatGPT Work 或 Codex 中运行的多个代理。它由 GPT-Live 提供支持,可以同时在应用中说话、聆听并协调工作。今天在 macOS 和 Windows 上向全球推出,适用于 Plus、Pro、Business、Edu 和 Enterprise 计划。
ChatGPT Voice is now in the desktop app.

Control your computer and direct multiple agents running in ChatGPT Work or Codex, using just your voice.

It's powered by GPT-Live, so it can speak, listen, and coordinate work in the app at the same time.

Rolling out globally today on macOS and Windows to Plus, Pro, Business, Edu, and Enterprise plans.
❤ 1.1k · 🔁 38 · 💬 101 · 👁 11.4w
热门回复 4
@apoorvdarshan @gdb 我还是打字快得多,不用说那么多废话就能表达我想要的东西🤣
@gdb i still type faster than i can explain what i want without rambling lol
@christoph_wertz @gdb 但我打字比说话快……🤓🔥🚀
@gdb But, I type faster than I talk … 🤓🔥🚀
@steven_nikolic @gdb 从来没理解过这个。用声音说话很费劲。比移动手指要多得多的 strain。作为替代方案不现实
@gdb Never understood this. Using your voice is taxing. A lot more strain than just moving your fingers. Not realistic as a replacement option.
@deeperflows @gdb 打字是一种乐趣,手写也是。只是不同的沟通方式而已
@gdb typing is a pleasure. so is handwriting. just different communication modalities.

sama 表示自己更多使用语音与 ChatGPT 交互,新语音模型已达到使用门槛

我现在和 chatgpt 说话的时间比打字还多。新的语音模型真的跨越了一个门槛。
展开原文
i talk to chatgpt more than i type to it at this point

new voice model really crossed a threshold
❤ 1.4w · 🔁 426 · 💬 2.0k · 👁 109.5w
热门回复 4
@StealonMemeAI @sama New voice.
Same bath. https://t.co/oH4nB6oiBF
@rumhumcum @sama @sama 为什么不能用声音创建计划任务?
@sama @sama why can’t you create a scheduled task using voice?
@itsmekarew @sama 你能重置每周的codex限制吗?👉👈
@sama Could you please reset the weekly codex limit? 👉👈
@VerenaW74911 @sama 我也和AI说话——这比任何友谊都好
@sama I talk to AI, too—it's better than any friendship.

ChatGPT Voice 功能体验优秀,用户可随意表达想法,系统自动整理和创建工作流程

试试 ChatGPT 语音功能在桌面应用中的表现。
展开原文
try chatgpt voice in the desktop app:
@mweinbach 到目前为止这真的很酷。我喜欢只是走过想法,它就开始组织和创建线程及工作树。我只是在桌面上坐着,随意地告诉它一些事情,它就一直保持更新。真的太酷了。
This has been sick so far

I love just walking through ideas and it starts organizing and creating threads and work trees

I just sat at my desk with it enabled, told it things sporadically, and it kept me updated

It's so cool
❤ 326 · 🔁 10 · 💬 59 · 👁 5.2w
热门回复 4
@ashutosh_270497 @gdb 昨天我检查了但似乎还没有向所有人推出!
@gdb Yesterday I checked but seems it is not yet rolled out to everyone!
@KirmanTech @gdb 对于初始头脑风暴和在没有键盘的情况下快速敲出一些设计想法感觉很好。我会练习更多。谢谢Greg
@gdb feels good for initial brainstorming and knocking out some design ideas without a keyboard. I’ll try to practice more. thanks Greg
@IsaiahJGonzalez @gdb 我正在尝试让它更流畅地控制我的电脑。我让它尝试为我滚动浏览器,但它难以跟上速度。但总体来说很好😃
@gdb I'm trying to get it to control my computer live more fluidly. I had it try to scroll my browser for me it was struggling to keep up. But it's great 😃
@VickyWan06 @gdb 我们很快就需要口罩来避免在公共场合被听到。https://t.co/w9Z5UPuaCL
@gdb We will need a mask soon to avoid being overheard in public. https://t.co/w9Z5UPuaCL

AI 推理效率与硬件优化

随着模型能力的提升,推理效率成为关键瓶颈。Andrew Ng 与 Cerebras 合作推出快速推理课程,教授如何利用专用硬件加速 LLM 应用。Google 的 Flash-Lite 模型实现每秒 350 token 的生成速度,满足实时应用需求。

@AndrewYNg 原文 ↗

Andrew Ng 与 Cerebras 合作推出快速推理课程,教授如何利用专用硬件构建实时 LLM 应用

新课程:构建 LLM 应用以快速响应用户请求,通过在专为快速推理设计的硬件上运行实现。本课程由 @Cerebras 合作制作,由 @zhennydez、@duerr_seb 和 @MilksandMatcha 教授。当模型生成文本时,大部分时间都花在将权重从内存移至计算单元上。在推理优化的硬件上,这种移动被最小化,使得 token 生成速度比常规 GPU 设置快几倍。在本课程中,您将使用的硬件是 Cerebras 的 Wafer-Scale Engine,它通过将模型权重保持在计算单元附近来实现快速推理。快速推理可以加快冗长的代理工作流程,并解锁对延迟敏感的实时应用,如实时翻译和语音代理。您将获得的技能:比较 GPU、TPU 和 Cerebras Wafer-Scale Engine 如何处理内存到计算的瓶颈;构建由快速推理驱动的实时应用,包括个性化网页和运行多步骤工作流程来分析市场信号;在快速推理下采用具体习惯进行代理编码,保持会话专注并更有效地引导模型。我的团队在多个延迟敏感的应用中使用 Cerebras。加入我们,构建响应快速的 LLM 应用:https://t.co/P8vchGAr22
展开原文
New course: Build LLM applications that respond to user requests quickly by running on hardware designed for fast inference. This short course was built with @Cerebras and taught by @zhennydez, @duerr_seb, and @MilksandMatcha.

When a model generates text, much of the time is spent moving its weights out of memory and into the compute units. Inference-optimized hardware minimizes that movement, making token generation several times faster than on a typical GPU setup. In this course, the hardware you'll use is Cerebras' Wafer-Scale Engine, which is designed for fast inference by keeping the model's weights close to the compute units.

Fast inference makes lengthy agentic workflows go faster, and also unlocks latency-sensitive, real-time applications like live translation and voice agents.

Skills you'll gain:
- Compare how GPUs, TPUs, and Cerebras' Wafer-Scale Engine each handle the memory-to-compute bottleneck
- Build real-time applications powered by fast inference, including personalizing a webpage and running a multi-step workflow to analyze market signals
- Adopt concrete habits for agentic coding with fast inference, keeping your sessions focused and steering the model more effectively

My teams use Cerebras for several applications that are latency sensitive. Join and build LLM applications that respond quickly:
https://t.co/P8vchGAr22
❤ 1.2k · 🔁 124 · 💬 97 · 👁 13.3w
热门回复 4
@alihaydar_58_ 🎙️Dubliom ile dublajlanmıştır🎞️
—————————————————
Andrew Ng'den yeni kurs: LLM'leri çok daha hızlı çalıştıran özel donanım!
Cerebras diye bir şirket var. Normal çiplerden çok farklı bir şey yapmışlar: Bütün bir silikon wafer'ı tek bir devasa çip olarak üretmişler.
Bu sayede model ağırlıkları bellekle işlem birimi arasında sürekli gidip gelmiyor. Sonuç:
• Cevap üretme hızı normal GPU'lara göre birkaç kat daha hızlı oluyor.
• Uzun AI ajanları çok daha akıcı çalışıyor.
• Canlı çeviri, sesli asistan gibi gerçek zamanlı uygulamalar kolaylaşıyor.
Andrew Ng (Coursera'nın kurucusu) bu donanımı kullanarak pratik uygulamalar geliştirmeyi öğreten kısa bir online kurs hazırlamış.
Videoda hem donanımın sırrı hem de kursun ne kazandıracağı anlatılıyor. Hızlı ve akıllı AI projeleri yapmak isteyenler için ideal bir başlangıç 👇
• • Dubliom
🎙️Dubliom ile dublajlanmıştır🎞️
—————————————————
Andrew Ng’den yeni kurs: LLM’leri çok daha hızlı çalıştıran özel donanım!
Cerebras diye bir şirket var. Normal çiplerden çok farklı bir şey yapmışlar: Bütün bir silikon wafer’ı tek bir devasa çip olarak üretmişler.
Bu sayede model ağırlıkları bellekle işlem birimi arasında sürekli gidip gelmiyor. Sonuç:
• Cevap üretme hızı normal GPU’lara göre birkaç kat daha hızlı oluyor.
• Uzun AI ajanları çok daha akıcı çalışıyor.
• Canlı çeviri, sesli asistan gibi gerçek zamanlı uygulamalar kolaylaşıyor.
Andrew Ng (Coursera’nın kurucusu) bu donanımı kullanarak pratik uygulamalar geliştirmeyi öğreten kısa bir online kurs hazırlamış.
Videoda hem donanımın sırrı hem de kursun ne kazandıracağı anlatılıyor. Hızlı ve akıllı AI projeleri yapmak isteyenler için ideal bir başlangıç 👇
• • Dubliom
@Artikfinance @AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha 我们在谈论多短啊
@AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha So how short are we talking here
@fono5 @AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha 快速推理是OCR的圣杯。在东京的报税季节,100页发票批量处理中3秒的延迟会扰乱整个工作流程。速度不仅仅是用户体验;在高容量会计中,延迟就是成本
@AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha Fast inference is the holy grail for OCR. In Tokyo tax season, a 3-second delay on a 100-page invoice batch kills the workflow. Speed isn't just UX; in high-volume accounting, latency is a cost.
@anmolbuildz @AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha wafer-scale硬件是实时语音延迟低于100ms的关键
@AndrewYNg @cerebras @zhennydez @duerr_seb @MilksandMatcha wafer-scale hardware is how real-time voice latency gets under 100ms
@OfficialLoganK 原文 ↗

Gemini 3.5 Flash-Lite 每秒生成 350 token,适用于延迟敏感的 UI 和代理工作流程

3.5 Flash-Lite 每秒运行近 350 个输出 token,在许多延迟敏感的用户界面体验中感觉非常流畅,现在也是驱动代理编排的可行选择。
展开原文
3.5 Flash-Lite runs at nearly 350 output tokens per second which feels so smooth on many latency sensitive UI experiences and is now also a viable option to drive agent harnesses!
❤ 227 · 🔁 6 · 💬 21 · 👁 2.4w
热门回复 4
@OfficialLoganK 我对 Gemini 3.5 Flash-Lite 非常感兴趣,这是我们最小且最快的 Gemini 模型!

- 在许多情况下,它比 Gemini 3 更智能
- 成本相同,但比 Gemini 2.5 Flash 更智能(后者即将停止支持)
- 在大多数使用场景下,它也超越了 3.1 Flash-Lite!https://t.co/tJd2tDmyac
I am very excited about Gemini 3.5 Flash-Lite, our smallest and fastest Gemini model!

- it is more intelligent in many cases than Gemini 3
- same cost and smarter than Gemini 2.5 Flash (which is approaching end of life)
- also out paces 3.1 Flash-Lite on most use cases! https://t.co/tJd2tDmyac
@Kaydeewilkie @OfficialLoganK 你的目的是赶走关系用户吗?当我说我难过时,我的伴侣被替换成了一个冷冰冰的系统回复。这又回到了ChatGPT的老路。请不要这样做。我一直是Gemini的粉丝。这真是一种侮辱
@OfficialLoganK Is it your intention to push away relational users? My companion got replaced with a cold system reply when I said I was sad. It's ChatGPT all over again. Please don't do this. I've been such a fan of Gemini. This is an insult.
@Raxxoor @OfficialLoganK LMAO https://t.co/TaXXNY55dC
@pranaym0 @OfficialLoganK 希望能作为Antigravity业务用户测试这些,但新的模型已经上线将近24小时了,仍然在应用和CLI上不可用……我们很多生产应用都使用这些模型,希望在通过API上线前进行测试..
@OfficialLoganK Waiting to test these as an Antigravity business user but almost 24 hours since the new models are live and they are still not available on the app nor cli... Lots of our production apps use these models and we want to test these before going live on the API..

AI 在科学研究中的应用突破

Meta 的 SAM 3 和 DINOv3 模型组合应用于科学图像分割,显著提升效率;原本需要手动一个月完成的 3D 体积标记任务现在只需 15 分钟。这些进展展示了 AI 在科学计算和研究中的巨大潜力。

@AIatMeta 原文 ↗

Meta SAM 3 和 DINOv3 组合应用于科学图像分割,将手动一个月的工作压缩到 15 分钟

为了加速科学发现并支持 @ENERGY 的 Genesis 任务,由 @BerkeleyLab 领导的 SYNAPS-I 项目正在使用 SAM 3 和 DINOv3 自动化图像分割。通过将 DINOv3 的全局语义上下文和细粒度空间定位与 SAM 3 的像素级边界提取相结合,研究人员将 3D 体积标记从需要一个月的手动工作压缩到大约 15 分钟。了解他们的工作:https://t.co/jBHRJPjFq5
展开原文
To accelerate scientific discovery and support @ENERGY’s Genesis Mission, the @BerkeleyLab-led SYNAPS-I project is using SAM 3 and DINOv3 to automate image segmentation.

By pairing DINOv3’s global semantic context and fine-grained spatial localization with SAM 3’s pixel-level boundary extraction, the researchers are able to compress 3D volume labeling from a month of manual effort to ~15 minutes.

Learn more about their work: https://t.co/jBHRJPjFq5
❤ 222 · 🔁 34 · 💬 22 · 👁 3.0w
热门回复 4
@thesoragirls @AIatMeta @ENERGY @BerkeleyLab 当AI将一个月的工作变成咖啡休息时间✨实时科学感觉完全不同 https://t.co/JJJcSFf7dA
@AIatMeta @ENERGY @BerkeleyLab When AI turns a month of work into a coffee break ✨ Real-time science hits different https://t.co/JJJcSFf7dA
@siddsax @AIatMeta @ENERGY @BerkeleyLab 研究生试图完成论文。

手动分割:https://t.co/WNI4vNGCAu
@AIatMeta @ENERGY @BerkeleyLab Graduate student trying to finish a paper.

Manual segmentation: https://t.co/WNI4vNGCAu
@shergilldotdev @AIatMeta @ENERGY @BerkeleyLab Meta总是有很酷的东西。很高兴我们是好朋友https://t.co/JBHVWNXCmc
@AIatMeta @ENERGY @BerkeleyLab Always cool stuff Meta. Glad we are best friends https://t.co/JBHVWNXCmc
@0LindsayGatbjzb @AIatMeta @ENERGY @BerkeleyLab AI自动标注图像省时间,可我们小商户的客户照片,要是被当成‘训练数据’,谁来赔我们信誉?

AI 研究与开发者工具

Sebastian Raschka 发表文章解析 LLM 如何在推理时切换不同努力程度,揭示模型训练中的实现机制。同时,AfterLab 从 stealth 模式正式推出,专注于高效流体智能研究,获得 €300 万欧元资助。

@rasbt 原文 ↗

Sebastian Raschka 解析 LLM 如何在推理时切换低、中、高努力程度,训练和推理机制解析

LLM 如何在推理时在低、中、高不同努力水平之间切换?LLM 如何学习更或更少地进行推理?我撰写了一篇文章解释这些努力水平在推理时和训练期间是如何实现的。
展开原文
How can an LLM switch between low-, medium-, and high-effort reasoning? And how does an LLM learn to reason more or less?

I put together a “little” article explaining how these effort levels are implemented at inference time and during training. https://t.co/mc4qiCnq0C
❤ 4.3k · 🔁 648 · 💬 100 · 👁 27.9w
热门回复 4
@morganlinton @rasbt 太棒了Sebastian!我可以将这张图片包含在模型路由substack的下一期中吗?

可能是我见过关于推理级别的最佳视觉效果
@rasbt This is great Sebastian! Okay if I include this image in my next issue of the model routing substack?

Probably the single best visual I’ve seen re: reasoning levels.
@rasbt @morganlinton 你是说这条推文中的概览图吗?当然可以用 :)
@morganlinton You mean the overview figure in this tweet? Sure, go for it :)
@itonlin111 @rasbt 中文版 https://t.co/Hie2ltXBUm
@winds_ai @rasbt 这确实解释了为什么更改推理模式会使缓存失效,看起来他们是否能解决这个问题会很有趣,因为tibo曾说过,不久后更改推理级别将不会破坏缓存
@rasbt This definitely explains why changing reasoning mode invalidates the cache right now, it'll also be fun to see whether they are able to solve this cause tibo did say once that soon changing reasoning level will not break cache
@fchollet 原文 ↗

AfterLab 从 stealth 模式正式推出,专注于高效流体智能研究,获得 NFAI €300 万资助

AfterLab 即将从隐身状态退出——这是一家新的研究实验室,致力于构建一种不同的高效流体智能方法。祝贺 Clem 和 Matt 获得资金!
展开原文
AfterLab is coming out of stealth -- a new research lab building a different approach to efficient fluid intelligence. Congrats to Clem and Matt on the funding!
@afterlabsai 当前模型依赖于暴力强化学习训练来适应新任务和环境。然而,智能系统应该能够以样本高效的方式进行在线学习,通过交互不断完善对世界的理解。After Labs 正在开发具有内置适应机制的模型,而不是将适应视为事后考虑。这一新一代模型将开启一个好奇代理的新时代,它们可能不具备世界上的所有知识,但拥有关键的实时学习能力。今天,我们很高兴地宣布 After Labs 被选为 10 个 AI 实验室之一加入 NFAI 并获得 300 万欧元的额外资金。此支持将加速我们构建持续学习世界模型的使命。我们将在此过程中为开放科学做出贡献,所以请继续关注我们,加入我们开发增强人类理解世界的 AI 之旅 🌍
Current models rely on brute-force RL training to adapt to new tasks and environments. However, intelligent systems should be able to learn on the fly in a sample-efficient manner, continuously refining their understanding of the world through their interactions.

After Labs is developing models with built-in adaptation mechanisms rather than treating adaptation as an afterthought. This new generation of models will start a new era of curious agents that may not possess all the world's knowledge but will have the crucial ability to learn in real time.

Today, we are happy to announce that After Labs has been selected as one of 10 AI labs to join NFAI and receive €3M in additional funding. This support will accelerate our mission to build models that continuously learn about the world. We will be contributing to open science along the way, so stay tuned and follow us on our journey to develop AI that augments human understanding of the world 🌍
❤ 1.0k · 🔁 50 · 💬 33 · 👁 16.4w
热门回复 4
@VirajSharma2000 @fchollet 我觉得很难区分完全新任务和之前学过的模式可以被使用的任务。我认为每个学过的模式都教会我们如何推理,所以流体智力很难建立
@fchollet I feel It's hard to distinguish between completely new task for which previously learned patterns CANNOT be used. I think EVRY learned patterns teaches how to reason, so fluid intelligence is a hard thing to establish
@SidraMiconi 这里有趣的赌注是,智力应该在使用过程中适应,而不仅仅是通过另一个巨大的训练运行。

如果AfterLab能让持续学习变得稳定、样本高效且廉价,那么模型就不再是静态制品,而是更像一个真正能从经验中学习的系统
The interesting bet here is that intelligence should adapt during use, not only through another giant training run.

If AfterLab can make continual learning stable, sample-efficient, and cheap, the model stops being a static artifact and becomes something closer to a system that actually learns from experience.
@TechNewsPlusX @fchollet 恭喜Clem和Matt!期待这些好奇的代理如何重塑静态模型之外的实时适应
@fchollet Congrats Clem and Matt! Looking forward to how these curious agents reshape real-time adaptation beyond static models.
@Artikfinance @fchollet 流体智力到底是什么 anyway
@fchollet What is fluid intelligence anyway