‹ 目录

X日报 · AI科技

2026-07-20 · 精选 3 条 · 数据池 168

⚡ 今日速览

  • OpenAI GPT-5.6 Sol 性能突增,价格效率提升6倍,引领AI编码与数学应用
  • Thinking Machines发布首个开源多模态模型Inkling,975B参数可供微调
  • Kimi K3震撼发布:2.8万亿参数、100万上下文长度,开源权重即将开放
  • Meta AI模型在亚洲物理奥林匹克理论考中取得满分30/30
  • OpenAI推出GPT-Red自动红队系统,提升模型安全性
  • Google AI Studio推出管理代理新功能,新增免费层和触发机制
  • AI在网络安全领域应用取得显著进展,Sol模型发现新型漏洞
  • Cerebras推出专为快速推理设计的硬件课程,支持实时AI应用

📋 今日综述

  • OpenAI GPT-5.6 Sol性能与成本效率双提升引领行业变革
  • 开源模型竞赛白热化Inkling与Kimi K3加速多模态与长上下文能力
  • AI代理与编码工具成熟ChatGPT Work与Codex推动开发者生产力革命
  • AI安全与鲁棒性成为关键议题GPT-Red与持续学习研究值得关注

AI在学术竞赛与评测中的突破

Meta AI模型在亚洲物理奥林匹克理论考试中取得了满分成绩,与顶尖学生并列第一。这一成绩展示了AI在物理学等理科领域的强大推理能力。同时,多个基准测试显示新模型正在快速超越旧标准。

@AIatMeta 原文 ↗

Meta AI模型在亚洲物理奥林匹克理论考中获得满分30/30,与顶3学生并列第一名

为了展示 Meta AI 的高级推理和多模态能力,我们提交了一个模型参加亚洲物理奥林匹克竞赛的理论考试。我们很高兴地宣布,我们的模型取得了 30/30 的满分成绩,与前三名的学生选手并列第一。

感谢 APhO 委员会允许我们的模型参加比赛:https://t.co/dpyMyST2n4
展开原文
To demonstrate Meta AI's advanced reasoning and multimodal capabilities, we submitted a model to participate in the Asian Physics Olympiad’s theoretical exam. We’re happy to share that our model achieved a perfect score of 30/30, tying with the top 3 student contestants.

We appreciate the APhO committee for letting our model participate in the competition: https://t.co/dpyMyST2n4
❤ 388 · 🔁 57 · 💬 41 · 👁 17.0w
热门回复 4
@owenthcarey @AIatMeta 前三名的学生选手看着一台服务器机架滚上领奖台去领取它的奖牌。https://t.co/IRjZNWA0gN
@AIatMeta The top three student contestants watching a server rack roll up to the podium to claim its medal. https://t.co/IRjZNWA0gN
@leo_pe2 @AIatMeta 这是什么鬼东西,AI生成的证书吗?
@AIatMeta What in the AI generated certificate is that
@dirtydiamondss_ @AIatMeta 假的。你伪造了你的学生。你的AI甚至无法区分汽车和牛拉手推车。在Instagram上删除人类的账号,称他们为机器人,而BOT账号却 thriving。你的AI就是一坨狗屎
@AIatMeta FAKE. You forged your students. Your AI can't even differentiate between a Car and Bullock Kart. Deleting accounts of Humans on Insta, calling them Bot, while BOT accounts thrive. Your AI is a shithole
@epochster @AIatMeta 没有什么比一个万亿美元的公司跟十几岁的孩子打成平手更能说明这是个历史性里程碑了。
@AIatMeta Nothing says historic milestone like a trillion dollar company tying with teenagers.

prinzbench测试显示GPT-5.6 Sol Pro得分91/99,几乎完全饱和了该测试标准

如今,基准测试很快就会被填满
展开原文
benchmarks get saturated very quickly these days
@deredleritt3r 已添加到 prinzbench:GPT-5.6 Sol Pro。

正如几天前所预览的,这款模型已经填满了我的基准测试,总得分为 91/99。

需要说明的是,prinzbench 包含两道题目,到目前为止没有任何模型能够解决(其中一道需要极其彻底的 50 个州的研究,可能需要 /goal 模式才能解决,另一道则涉及非常棘手的监管审批,目前没有任何模型能够找到答案)。如果将这两道题目(总分 6 分)排除,GPT-5.6 Sol Pro 在 prinzbench 的 93 道题目中正确回答了 91 道。

OpenAI Pro 模型在 prinzbench 上的表现:

GPT-5.4 Pro(扩展版):79/99
GPT-5.5 Pro(扩展版):82/99
GPT-5.6 Sol Pro:91/99

我的基准测试于 2026 年 1 月发布,并在 2026 年 6 月被填满。加速是真实的!

由于这款模型的表现,未来的 OpenAI Pro 模型将不再在 prinzbench 上进行测试(没有必要再测试了)。

其他 GPT-5.6 模型的基准测试将很快发布(敬请期待)。
Added to prinzbench: GPT-5.6 Sol Pro.

As previewed a few days ago, this model has saturated my benchmark, with a total score of 91/99.

For context, prinzbench contains two questions that no model tested to date has ever been able to solve (one requires extremely thorough 50-state research that probably requires /goal mode to solve, and another has a really tricky regulatory approval that no model has ever been able to find). Putting these two questions (which are worth 6 points) aside, GPT-5.6 Sol Pro provided correct responses to 91 out of 93 prinzbench questions.

prinzbench performance for OpenAI's Pro models:

GPT-5.4 Pro (Extended): 79/99
GPT-5.5 Pro (Extended): 82/99
GPT-5.6 Sol Pro: 91/99

My benchmark was released in January 2026 and was saturated in June 2026. The acceleration is real!

As a result of this model's performance, future OpenAI Pro models will no longer be tested on prinzbench (there is no point in testing them).

Benchmarking for other GPT-5.6 models to follow soon(TM).
❤ 576 · 🔁 30 · 💬 47 · 👁 9.3w
热门回复 4
@deredleritt3r @gdb 向OpenAI团队致敬——这是一个令人难以置信的模型!
@gdb Kudos to the OpenAI team - it's an incredible model!
@Selene1008 @gdb 给我们把4o还回来!
#keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o!
#keep4o #OpenSource4o #GPT4o
@SirMrMeowmeow @gdb 我投票支持更多奇特的能力基准测试拜托了

>保留转录
>>潜在记忆
>>基于权重的记忆

>即时学习(例如它能学习马里奥卡佐按钮序列或优化技能/直觉/战术,尤其是在权重层面或类似层面)
@gdb i vote more exotic capabilities benchmarks pweaze

> withhold the transcript
>>latent memory
>> weight level based memory

>Learn on the fly (so esp can it learn a mario kaizo button sequence or optimize a skill/intuition/tactics esp pref at the level of weights or similar)
@fabiana0369 @gdb 问问sama谁会是最后一个种族主义者😂😂😂😂😂😂😂😂😂😂😂😂我们还不知道呢。
@gdb Ask sama who will be the last racist 😂😂😂😂😂😂😂😂😂😂😂😂 we dont know yet.
@fchollet 原文 ↗

After Labs获得欧洲NFAI资助,将专注于开发具备持续学习能力的AI模型

AfterLab 即将从隐身状态退出——这是一家新的研究实验室,致力于构建一种不同的方法来实现高效的流体智能。祝贺 Clem 和 Matt 获得资金!
展开原文
AfterLab is coming out of stealth -- a new research lab building a different approach to efficient fluid intelligence. Congrats to Clem and Matt on the funding!
@afterlabsai 当前模型依赖于暴力强化学习训练来适应新任务和环境。然而,智能系统应该能够以样本高效的方式进行实时学习,通过交互不断完善对世界的理解。

After Labs 正在开发具有内置适应机制的模型,而不是将适应视为事后的想法。这种新一代模型将开启一个好奇代理的新时代,它们可能不掌握世界的所有知识,但拥有关键的实时学习能力。

今天,我们很高兴地宣布 After Labs 被选为 10 个 AI 实验室之一加入 NFAI 并获得 300 万欧元的额外资金。此项支持将加速我们构建持续学习世界模型的使命。我们将在此过程中为开放科学做出贡献,所以请继续关注我们的旅程,共同开发能够增强人类对世界理解的 AI 🌍
Current models rely on brute-force RL training to adapt to new tasks and environments. However, intelligent systems should be able to learn on the fly in a sample-efficient manner, continuously refining their understanding of the world through their interactions.

After Labs is developing models with built-in adaptation mechanisms rather than treating adaptation as an afterthought. This new generation of models will start a new era of curious agents that may not possess all the world's knowledge but will have the crucial ability to learn in real time.

Today, we are happy to announce that After Labs has been selected as one of 10 AI labs to join NFAI and receive €3M in additional funding. This support will accelerate our mission to build models that continuously learn about the world. We will be contributing to open science along the way, so stay tuned and follow us on our journey to develop AI that augments human understanding of the world 🌍
❤ 1.0k · 🔁 48 · 💬 33 · 👁 16.2w
热门回复 4
@VirajSharma2000 @fchollet 我觉得完全新任务,无法使用之前学到的模式,是很难区分的。我认为每一个学到的模式都教会了我们如何推理,因此流体智力是很难建立的
@fchollet I feel It's hard to distinguish between completely new task for which previously learned patterns CANNOT be used. I think EVRY learned patterns teaches how to reason, so fluid intelligence is a hard thing to establish
@SidraMiconi @fchollet 这里有趣的赌注是,智力应该在使用过程中适应,而不仅仅是通过另一个巨大的训练运行。如果AfterLab能让持续学习变得稳定、样本高效且廉价,那么模型就不再是静态的工件,而是变成了一个更接近于真正从经验中学习的系统
The interesting bet here is that intelligence should adapt during use, not only through another giant training run.

If AfterLab can make continual learning stable, sample-efficient, and cheap, the model stops being a static artifact and becomes something closer to a system that actually learns from experience.
@alexanderbenz @fchollet "高效流体智力"在实践中意味着什么,与仅仅高效变换器相比呢
@fchollet what does "efficient fluid intelligence" mean in practice vs just efficient transformers
@TechNewsPlusX @fchollet 恭喜Clem和Matt!期待这些好奇的代理如何重塑静态模型之外的实时适应。
@fchollet Congrats Clem and Matt! Looking forward to how these curious agents reshape real-time adaptation beyond static models.