‹ 目录

X日报 · AI科技

2026-06-21 · 精选 26 条 · 数据池 1171

⚡ 今日速览

  • Claude Code 推出 Artifacts,Codex 推出 Record & Replay,AI 编程产品从“聊天”走向可交付工作流。
  • Anthropic 迎来 John Jumper 加入,Google AI Studio、Gemini Omni Flash、GLM-5.2、Qwen2.5-Coder 等模型/平台继续迭代。
  • AI agent 正在从写代码扩展到多代理协作、机器人操作、医学诊断和网页导航。
  • 行业讨论从“模型能力”转向评估、监管、开源效率、数据质量和工程化。
  • AI 就业门槛也在上升:经验、数据清洗、集成和评估比单纯训练更关键。

📋 今日综述

  • 产品化Claude Code Artifacts 与 Codex Record & Replay 显示 AI 编程工具正在沉淀为可复用工作流。
  • 模型竞争Gemini、GLM、Qwen、Cohere、GPT-Realtime 等继续堆叠能力,开放权重和闭源平台并行推进。
  • Agent 化300 个 Kimi agent、AI coding agents 教机器人、网页导航 agent,说明智能体进入多任务和物理操作阶段。
  • 人才与组织John Jumper 加入 Anthropic、OpenAI 医疗诊断合作,显示 AI 与科学/医疗的交叉加速。
  • 治理与评估监管、透明基准、数据清洗和评估被反复强调,行业正在从 demo 阶段走向工程约束。

AI 编程产品化

Claude Code、Codex、Record & Replay、Artifacts 等把 AI 编程从对话推进到可交付、可复用流程。

@claudeai 原文 ↗

['Claude Code 推出 Artifacts,把会话中的 PR 走查、项目看板等变成可分享页面。', 'Claude Code 的产品化重点是把“聊天结果”变成可复用、可分享的工件。']

['Claude Code 推出 Artifacts,把会话中的 PR 走查、项目看板等变成可分享页面。', 'Claude Code 的产品化重点是把“聊天结果”变成可复用、可分享的工件。']
展开原文
New in Claude Code: Artifacts.

Interactive pages built from your session, like a PR walkthrough or a living project dashboard, shared with your team at a private link.

Available in beta on Team and Enterprise plans. https://t.co/0NX9gNCaAs
❤ 1.8w · 🔁 1.3k · 💬 659 · 👁 370.7w
热门回复 4
@AmirMushich @claudeai Can be plugged in to a design agent via MCP -> design reporting docs / slides right from the Claude session

https://t.co/x1UycneydV
@rahulsinghh__ @claudeai I am thinking on how can I adopt this feature in my current workflows? Looking forward to see the different usecase for this one.
@SilentAvacate @claudeai Today I had 1 hour left of my session before it reset, but it reset early. Now I need to wait almost half a day to do work that should have only taken me 6 hours to complete, and I pay for pro. You can't be having errors like this with 5 hour long limits. I feel ripped off.
@techedgedaily @claudeai this is actually huge, claude just made sharing live demos feel effortless
@JohnJumperSci 原文 ↗

['John Jumper 离开 Google DeepMind 加入 Anthropic,AI 科学人才流动加速。', 'Jumper 的流动说明 AI for science 正在从实验室走向更广泛的平台竞争。']

['John Jumper 离开 Google DeepMind 加入 Anthropic,AI 科学人才流动加速。', 'Jumper 的流动说明 AI for science 正在从实验室走向更广泛的平台竞争。']
展开原文
A bit of news: After nearly 9 years, I have decided to leave Google DeepMind and join Anthropic (after taking some time to recharge). I am incredibly grateful for my time at GDM. @demishassabis took a real chance letting me lead the AlphaFold team just six months after finishing my PhD, and the entire GDM team taught me so much about how to do great science. GDM is a special place, and I’ll still be excited to hear about what amazing things they discover next.
❤ 1.4w · 🔁 935 · 💬 590 · 👁 571.5w
热门回复 4
@whynesspower @JohnJumperSci @demishassabis I’ve a feeling they get paid extra, if they agree to post online like this
@misovalko @JohnJumperSci @demishassabis Excellent @JohnJumperSci ! Can't wait for more great things for you! And maybe you finally find time for @ESET science 🔭 award! @greatwhiteeset
@DayShuai Congrats on the move, John. From the view of a PhD student in Prof. Yang Zhang's group, this reads as a sign of where the field is actually heading.

For a while we've been talking about "AI for science" — AI helping scientists. The Anthropic + Jumper combination feels like the inversion: AI as the scientist. And the model that does that won't be a stack of single-domain experts. It'll be a foundation model with many faces — protein design (sequence × structure × physics, jointly), RNA, small-molecule drug design — all from one place.

The principle underneath, from the protein side: the bar moves from "best generator on one view" to whether a candidate holds up across many plausible realizations. A world model — not in the multimodal LLM sense. Closer to how a molecule actually exists.

Curious where you'll start — protein, or wider from day one.
@giffboake @JohnJumperSci What a stupid decision you have made. Good luck working with the cancer that is Dario and Anthropic.
@GoogleAI 原文 ↗

['Google AI 本周发布 Gemini 3.5 Live Translate、NotebookLM 升级、Agent Builder 等。', 'Google 的 AI 产品线正在从模型扩展到语音、笔记、agent 构建等完整工具链。']

['Google AI 本周发布 Gemini 3.5 Live Translate、NotebookLM 升级、Agent Builder 等。', 'Google 的 AI 产品线正在从模型扩展到语音、笔记、agent 构建等完整工具链。']
展开原文
Here’s what launched this week:

— Gemini 3.5 Live Translate our latest audio model for live speech-to-speech translation

— @NotebookLM got a major upgrade including agentic capabilities in chat, more advanced reasoning, and a suite of new output formats

— Project Genie from @GoogleLabs is now available to Google AI Ultra 5x subscribers globally

— Notebooks in @GeminiApp are now available in the European Economic Area, United Kingdom, and Switzerland

— DiffusionGemma, our newest experimental open @googlegemma model that explores text diffusion, an exceptionally fast approach to text generation
❤ 821 · 🔁 101 · 💬 70 · 👁 7.6w
热门回复 4
@Vanarchain @GoogleAI @NotebookLM Speech, reasoning, agents, generation. This is a full stack intelligence platform forming in real time.
@den_spirin @GoogleAI @NotebookLM I'll be honest, you guys demonstrate incredible progress, and we can see the token numbers rise

But very very very select few will choose Gemini as their daily driver

So keep up the good work, and thank you
@Siddhos @GoogleAI @NotebookLM Fumble 5 also launch making gemini more worthy
@husi0x @GoogleAI @NotebookLM NotebookLM getting agentic chat is the sleeper update here

['OpenAI 与全球医生合作,让 ChatGPT 更好地回答健康相关问题。', '健康问答的关键不是替代医生,而是在多语言、多专科场景中降低信息门槛。']

['OpenAI 与全球医生合作,让 ChatGPT 更好地回答健康相关问题。', '健康问答的关键不是替代医生,而是在多语言、多专科场景中降低信息门槛。']
展开原文
We've collaborating with hundreds of physicians across 60 countries, 49 languages, and 26 specialties to make ChatGPT great at health-related questions for everyone:
@OpenAI (引用原文)
GPT-5.5 Instant is now on par with our frontier Thinking models for health-related questions.

Every week, more than 230 million people turn to ChatGPT with health and wellness questions, and GPT-5.5 Instant is better at recognizing when urgent care may be needed, asking for relevant context, explaining uncertainty, and making complex information easier to understand.

Because GPT-5.5 Instant is available to all free users in ChatGPT, these improvements can help more people.

Physician-led evaluation was critical to making these major intelligence gains.
❤ 914 · 🔁 58 · 💬 66 · 👁 10.2w
热门回复 4
@stark4833 @gdb But what about the people 4o helped every day, people with trauma. Neuro divergent. People who are just alone needed something to talk to. Are you just gonna keep ignoring them like you normally do? #4oForAll
@Yasamanini @gdb I'm certain they're all biased themselves.
4o was already the best model for health support stop denying and kidding yourselves. bring that back.
#keep4o #4oForAll
@Selene1008 @gdb Return 4o to everyone!😒
#keep4o #OpenSource4o #GPT4o
@JoeWilliams010 @gdb Being back 4o! Best model for us! #keep4o
@fchollet 原文 ↗

['行业需要透明、标准化的 agentic capabilities benchmarks,而不是 prompt tricks。', 'agent 评估要标准化,否则 demo 好看但生产不可靠。']

['行业需要透明、标准化的 agentic capabilities benchmarks,而不是 prompt tricks。', 'agent 评估要标准化,否则 demo 好看但生产不可靠。']
展开原文
We desperately need standardized benchmarks for agentic capabilities instead of panic-reacting to prompt-engineering parlor tricks. Until we can objectively and transparently measure what these systems actually do, the entire ecosystem will remain vulnerable to unpredictable and arbitrary government overreach.
❤ 109 · 🔁 12 · 💬 6 · 👁 1.2w
热门回复 4
@fchollet Even if you are in favor of AI regulation, you should recognize that opaque and arbitrary regulatory strikes are counter-productive for the whole industry.
@HEMSEYE @fchollet The LLM approach is fundamentally flawed. Not understanding and guessing, very often getting things right is still very unsafe. My solution is in draft. Linguistic cognitive grounding to then create using a compositional language of thought is part of the direction we need.
@KissonL @fchollet Without grounded evals for multi-step planning, the community iterates on vibes. τ-bench and WebArena point the right direction but are still too brittle to benchmark production reasoning chains — that gap is exactly where the panic-reacting you describe comes from.
@oleg_kai @fchollet the unit is the trap. agentic capability is harness-conditioned, not model-conditioned. any benchmark picks an implicit harness shape and labs optimize the score, not the agent. arc-agi works for reasoning; nothing equivalent for tool-use under domain noise.

['OpenAI 对 Noam 的能力给出极高评价,强调其 AI 直觉。', '顶尖 AI 人才的价值不只是工程产出,也包括对模型方向的判断。']

['OpenAI 对 Noam 的能力给出极高评价,强调其 AI 直觉。', '顶尖 AI 人才的价值不只是工程产出,也包括对模型方向的判断。']
展开原文
We offer no explanation as to why Noams are so good at AI; we attribute their success, as all else, to divine benevolence.
@polynoamial (引用原文)
I'm always thrilled to have more Noams at @OpenAI, but I'm especially thrilled to welcome @NoamShazeer!
❤ 8.5k · 🔁 307 · 💬 512 · 👁 98.5w
热门回复 4
@LLuooo429 @sama Come on, why can’t you just let 4o back?#keep4o
@alineasmarrow Divine benevolence? 😂😂 tf
Sam Altman's back to tripping every weekend. What is it this time? DMT? Ayahuasca? Ketamine?

You've got Noam. Are you finally going to listen and build a model Humanity will love or are you still stuck on draining the 'magic' out of it?
I think we all know if you can't create the same magic as you did with 4o, Anthropic will win. They'll never be caught again.

#keep4o
@lnnchl @sama #keep4o
#bringback4o
#opensource4o https://t.co/1ZuMfgxPyM
@GingiF13341 @sama When will you be bringing 4o back?? #bringback4o #opensource4o

模型与平台竞争

Gemini、GLM、Qwen、Cohere、GPT-Realtime、Google AI Studio 等平台继续更新,开放与闭源并行。

@OpenAIDevs 原文 ↗

['Codex Record & Replay 让用户把重复流程录成可检查、可编辑 skill。', 'Codex 正在把一次性指令沉淀成可复用技能,这是 AI 编程工具进入工作流的关键一步。']

['Codex Record & Replay 让用户把重复流程录成可检查、可编辑 skill。', 'Codex 正在把一次性指令沉淀成可复用技能,这是 AI 编程工具进入工作流的关键一步。']
展开原文
Show Codex a workflow once. Reuse it as a skill.

Record & Replay lets you show Codex a recurring task, like filing an expense report or submitting a time-off request.

Codex turns that demo into an inspectable, editable skill.

You control when recording starts and stops. https://t.co/UqSGaO7XUs
❤ 1.3w · 🔁 1.2k · 💬 467 · 👁 402.0w
热门回复 4
@gonzaleshvili @OpenAIDevs I'm just counting the days
until Codex can use my apps no-cli.
Now that would give everybody the chance to run their system and chained workflows.
@no7wade I used OpenAI Codex’s new Record & Replay-style workflow to turn a messy X Article publishing process into a reusable skill.
Took ~30 min end to end. ~6 runs conversation
Model: GPT-5.5 Codex. no script generated , just skill.md
here is Final skill: https://t.co/Y71M79QmAn
here is the article i published via this skill https://t.co/7miLY9f26X
Problems hit:
- X login/browser automation blockers (it switched to chrome dont know why, i use Brave browser )
- wrong draft focus
- body paste accidentally targeting title
- too many blank lines
- Markdown - item not becoming real X bullets( solved after at lease mentioned it twice )
- Preview button unreliable
Human helped by logging in, clicking the right draft, and spotting formatting issues with screenshots.
Codex fixed it by:
- collapsing extra newlines
- converting Markdown lists into native X lists
- removing leftover -
- using direct /preview
- updating the reusable skill

Took ~30 min end to end. ~6 runs conversation
Model: GPT-5.5 medium Codex.
I used OpenAI Codex’s new Record & Replay-style workflow to turn a messy X Article publishing process into a reusable skill.
here is Final skill: https://t.co/Y71M79QmAn
here is the article i published via this skill https://t.co/7miLY9f26X
Problems hit:
- X login/browser automation blockers (it switched to chrome dont know why, i use Brave browser )
- wrong draft focus
- body paste accidentally targeting title
- too many blank lines
- Markdown - item not becoming real X bullets( solved after at lease mentioned it twice )
- Preview button unreliable
Human helped by logging in, clicking the right draft, and spotting formatting issues with screenshots.
Codex fixed it by:
- collapsing extra newlines
- converting Markdown lists into native X lists
- removing leftover -
- using direct /preview
- updating the reusable skill
@nickbaumann_ @OpenAIDevs https://t.co/0yZujVox62
@Krunalc16 @OpenAIDevs I was way ahead of my team with skillsclaw!!!

https://t.co/Dzrx23hLWk
@AndrewYNg 原文 ↗

['Andrew Ng 指出美国政府与 Anthropic 都在限制他人使用前沿模型。', 'AI 访问权正在被政策和企业策略同时收紧,生态开放与安全问题进入拉扯期。']

['Andrew Ng 指出美国政府与 Anthropic 都在限制他人使用前沿模型。', 'AI 访问权正在被政策和企业策略同时收紧,生态开放与安全问题进入拉扯期。']
展开原文
Over the last two weeks, both the U.S. Government and Anthropic took significant actions that demonstrated their power to control access to AI by restricting what others can do with frontier models. This has been one of those moments that, once seen, will be hard to unsee, and it is significantly accelerating many businesses’ and nation states’ efforts to ensure reliable access to AI that no one else can terminate.

Anthropic first released Claude Fable 5, a version of its Mythos model with additional guardrails, including some restrictions that seem well justified on safety grounds (such as limitations on applying it to hacking, bioweapons, and so forth). However, it also restricted developers’ ability to use it to build competing LLM technology. This move was concerning, given that the whole AI community, including Anthropic, has benefitted tremendously from open research — indeed, the AI revolution was kicked off by my former team (Google Brain) freely publishing the Transformers paper!

Imagine if Microsoft’s terms of use barred anyone from using their tools to build competitive software, or if Google barred using it to search for information to work on competing search engines. Anthropic’s argument that it was unsafe for others to be able to make advances in AI also rang hollow. Initially, Anthropic silently degraded Fable 5’s performance for users detected to be working on LLM research through invisible interventions that weakened the model’s outputs without notifying the user. After significant backlash, it walked back this decision and decided to be transparent when it did this, but it still refuses to use its latest capabilities to help AI researchers.

This move represents a raw demonstration of power by Anthropic. It has used “safety” arguments to hinder potential competitors. Platforms succeed when they are viewed as stable, reliable partners that one can build on. The sudden rule changes by Anthropic (including a mandatory 30 day data retention policy for Fable usage) have made developers wonder about the stability of building on any one proprietary LLM provider, not just Anthropic.

The U.S. Government then shortly followed with an even greater demonstration of power. It used the Commerce Department’s authority to regulate technologies that may be national security threats to restrict exports of Mythos and Fable, requiring a license for use by any foreign national, whether inside or outside of the U.S., including employees of Anthropic. This led Anthropic to disable access to Fable to all users worldwide.

Sam Altman pointed out, referring to Anthropic, “It is clearly incredible marketing to say, ‘We have built a bomb, we are about to drop it on your head. We will sell you a bomb shelter for $100 million.’” But when one engages in this type of fear-based marketing, it increases the odds that the U.S. Government will agree with you and slap export controls on the bomb you say you have built.

To be clear, I don't think Anthropic has built anything like a bomb, and I don't think export controls on Fable are appropriate.

However, following the U.S. Government making this move, many nations, including U.S. allies, saw how the U.S. can suddenly yank their access to AI models. In many capitals around the world, this has spurred discussions on AI sovereignty and how others can ensure uninterrupted access to this critical technology.

For decades, many nations were comfortable having many parts of their supply chain rely on the U.S., China, and other major producers. Once a nation issues a threat, or takes action, to limit other nations’ access, other nations will rationally try to secure alternatives. For decades, semiconductor manufacturing in China made slow progress; once the U.S. moved to limit China’s access, China’s efforts kicked into high gear. Similarly, once China threatened U.S. access to rare earth minerals, U.S. efforts to secure alternatives accelerated. Now that it has become crystal clear that private U.S. companies and the U.S. government can limit, in short order, other nations’ access to frontier AI models, the incentive of others to invest more in alternatives like open source grows significantly. Of course, training frontier models is not easy, so it remains to be seen how successful they are, but we have crossed the rubicon.

Satya Nadella wrote an essay about the importance of building a healthy ecosystem on top of frontier AI technology. I heartily agree with him, and hope this week’s events will ultimately prove to be constructive steps toward this.

I hope we can build a more free, more open world, where research is freely shared, and laws and societal norms shape a level playing field that allows everyone to make progress. A silver lining of the events of these past two weeks is now that everyone better realizes key points of instability of the current system, we can all work to create a more stable foundation.

[Original text: The Batch newsletter]
❤ 1.1k · 🔁 243 · 💬 131 · 👁 11.7w
热门回复 4
@r_jack259 I can’t imagine a commercial company treating its own users like fools, randomly shutting down services and quietly changing model outputs, all under the same excuse: “safety first.” If large models are really that dangerous, then they shouldn’t be developed and operated by a commercial company in the first place, and even less should they be sold for others to use—selling them on one hand while simultaneously saying they are too dangerous and must be restricted on the other. This kind of self-contradictory narrative is already absurd enough on its own.

What “safety first” translates into in practice is: non-transparent changes, unpredictable restrictions, and users being forced to passively accept shifting rules. You can no longer even be sure whether you are interacting with the same system.

Now it seems clear to everyone: large models really are dangerous—but not because their capabilities are strong. They are dangerous because they are controlled by people like Dario. Dario should be removed from Anthropic, just as Ilya Sutskever did.
@AMGbadebo @AndrewYNg I have always said no government or government institutions should use closed-source AI systems.

https://t.co/x6FXFMgZaY
@MaybushSha16067 I have to disagree with you. Every company has IP (Intellectual Property) to protect. Many Chinese companies are using Anthropic to build their models to compete with Anthropic, or better yet, devaluing the Anthropic IP that they spent billions to produce.

"I hope we can build a more free, more open world, where research is freely shared, and laws and societal norms shape a level playing field that allows everyone to make progress."

No! The US job seeker felt the blunt of thinking like this. Your Global Utopia has nearly destroyed the US economy and made middle class a rarity. Its due to globalization efforts.

I don't know how such smart people become brain dead when it comes to politics, but the demand the US give away US IP to competing nations.

Even your previous employer Google is struggling against competition because they made Transformers Open Source. Disgusting people like Sam Altman are able to take advantage of that open source.

In a world where there are no bad actors that picture you have works, in reality, its a pipe dream that will never exist and will just destroy the country and company that tries to live by it.
@MoneyByAlex @AndrewYNg The moment AI access becomes a geopolitical tool, open-source stops being just a developer preference and becomes a strategic necessity

['GDB 称软件工程已经和 6 个月前完全不同。', 'AI 编程改变的不只是写代码速度,而是需求拆解、调试、评审和交付方式。']

['GDB 称软件工程已经和 6 个月前完全不同。', 'AI 编程改变的不只是写代码速度,而是需求拆解、调试、评审和交付方式。']
展开原文
software engineering is so different now. hard to remember what it was like even 6 months ago.
❤ 8.4k · 🔁 456 · 💬 423 · 👁 38.8w
热门回复 4
@zackvoell @gdb I can remember it very clearly.

I wasn't a software engineer 6 months ago.

Now I am.

And all I do is clank bullshit on my keyboard.
@glitchbyte101 @gdb It cost less money 6 months ago.
@SilverleafTech @gdb I can do 100x more engineering than I could do before. More far-reaching tasks. Architecture that would have taken me hours or days, done in minutes. Tests I never would have written. As a creator/problem-solver first, coder second, I love this new world!
@JacobPetterle @gdb do u have amnesia?

['OpenAI 用 AI 帮助破解健康谜题,显示医学推理正在成为模型能力试金石。', '医学谜题说明 AI 的价值不只在聊天,而在跨资料推理和假设生成。']

['OpenAI 用 AI 帮助破解健康谜题,显示医学推理正在成为模型能力试金石。', '医学谜题说明 AI 的价值不只在聊天,而在跨资料推理和假设生成。']
展开原文
AI for helping crack a health mystery. So many stories like this, and a clear motivation to be excited about AI:
@amydeng_ (引用原文)
I’m an AI researcher turned brain tumor patient, and recently I used the models to crack my mystery fatigue faster than my PCP could.
I believe everyone can do the same with their own symptoms. Here’s how: https://t.co/0jhbPvEi7V
❤ 823 · 🔁 62 · 💬 80 · 👁 12.0w
热门回复 4
@JoeWilliams010 @gdb Stop taking models away then, specially the one Altman uses for his OWN longevity research! Give us 4o back! #keep4o https://t.co/zvARZSKAKP
@frostybaby13 @gdb You mean the 4o that was used to cure the beloved dog? Return it!!!

https://t.co/mBa2NUBLZq
@Vickee2025 @gdb #OpenSource4o #keep4o #BringBack4o #OpenAI #ChatGPT https://t.co/KLgbDx9ktR
@DanaH1473651 @gdb Greg, your medical successes are due to the 4o you took from us, claiming it was old! While your 5.5+ models are simply useless, even with your Codex, which is used by a minimum of people. Give us back the 4o you stole from us!
@kimmonismus 原文 ↗

['Reddit 用户用 1800 个 bots 和 DeepSeek API 做了可玩的 WoW 私人服务器。', 'DeepSeek API 被用于构建带 AI chat 的 MMORPG,说明低成本模型正在改变游戏 NPC 生态。']

['Reddit 用户用 1800 个 bots 和 DeepSeek API 做了可玩的 WoW 私人服务器。', 'DeepSeek API 被用于构建带 AI chat 的 MMORPG,说明低成本模型正在改变游戏 NPC 生态。']
展开原文
Someone on Reddit built a WoW private server with 1,800 bots and AI chat via the DeepSeek API.

Dead Internet Theory, but playable.

An MMORPG with no real players, yet somehow it still feels human. https://t.co/uFD0AHiquc
❤ 2.3w · 🔁 1.1k · 💬 759 · 👁 220.3w
热门回复 4
@Gsnchez @kimmonismus @Recuenco
@itsMEGAMEGA @kimmonismus Looks like shit
@johnadams91000 @kimmonismus what is the point of that when all classic players are bots anyway
@bigcliffyb @kimmonismus Really cool if you’re fucking really dumb!

Agent 与自动化

300 个 Kimi agent、AI coding agents 控制机器人、网页导航 agent,显示多代理和物理操作成为新方向。

@ClaudeDevs 原文 ↗

['Claude Code 重置 5 小时和周用量限制,用户侧体验优先。', 'Claude Code 的用量重置说明产品仍在快速迭代,容量策略会直接影响开发者口碑。']

['Claude Code 重置 5 小时和周用量限制,用户侧体验优先。', 'Claude Code 的用量重置说明产品仍在快速迭代,容量策略会直接影响开发者口碑。']
展开原文
Update: we've gone ahead and reset 5-hour and weekly usage limits for everyone, across all plans. Enjoy your weekend!
@ClaudeDevs (引用原文)
Earlier today, ~3% of Claude Code Max and Pro users hit a bug that showed an incorrect weekly usage limit, and in some cases blocked them from sending messages.

This is fixed, and we're resetting 5-hour and weekly limits for everyone affected. Apologies for the disruption.
❤ 1.3w · 🔁 730 · 💬 973 · 👁 175.7w
热门回复 4
@Awsome1760081 @ClaudeDevs Usage limits are fucked again Devs
@ClaudeDevs
@0xNotMarc @ClaudeDevs Opus 4.8 become worst in limits/token efficiency, I have never experienced this before although I am using the same prompt. It becomes inconsistent now.
@JustPavol74975 @ClaudeDevs my todays 5 hour usage disappeared in 2 hours using chat and 2 jobs in claude code, unbelievable,
@PrajwalYamgar1 @ClaudeDevs Claude users just got their weekend back 😄
The best kind of bug fix.
Developers everywhere: productivity unlocked.
Rare AI company W.
Weekend plans: gone. Back to shipping.
My codebase isn't ready for this much Claude.
Competition is a beautiful thing.
@rasbt 原文 ↗

['GLM-5.2 被认为是当前最佳开放权重模型之一,架构继承 GLM-5/5.1。', 'GLM-5.2 的意义在于开放权重模型继续追赶闭源能力,开发者选择更多。']

['GLM-5.2 被认为是当前最佳开放权重模型之一,架构继承 GLM-5/5.1。', 'GLM-5.2 的意义在于开放权重模型继续追赶闭源能力,开发者选择更多。']
展开原文
Just caught up with the recent GLM-5.2 release. The best open-weight model today.

Architecture-wise, it's build on the GLM-5 and GLM-5.1 architecture that I covered previously, which means it's reusing the Multi-head Latent Attention (MLA) and DeepSeek Sparse Attention (DSA) mechanisms from DeepSeek V3.2. (I wrote about it here: https://t.co/tuunazfQ8y)

What's new is that they added an IndexShare mechanism. (That's a cross-layer reuse trick for DSA where instead of recomputing the sparse-attention top-k indexer in every layer, GLM-5.2 runs the full indexer only once every four layers and lets the following layers reuse those selected token indices. This keeps the same DSA idea but makes 1M-token inference much cheaper.)
❤ 1.9k · 🔁 239 · 💬 59 · 👁 9.7w
热门回复 4
@rasbt @vibecoder_dc I don’t disagree. Let’s say most capable open-weight LLM when averaged over all major benchmarks (reasoning, coding, logic, tool use, math, knowledge)
@rasbt @Abaybektursun Not sure tbh. If yes, it wouldn't surprise me, it's pretty common.
Claude distills from the internet & potentially others, others distill from Claude,... that's just the natural dev cycle.
GLM 5.2 is >10 pts better than Opus 4.8 on coding tasks btw, so they not "just" distilling https://t.co/ih3FlbLIgB
@ethankongee Let me add some more context on Sparse Attention and IndexShare.

In a classic attention block, a token has to attend to all previous tokens. But this becomes computationally expensive as the context window gets larger, so it’s not very scalable. GLM-5.1 has 200K token context window and GLM-5.2 has 1M. They need to do something to fix it or the performance would degrade.

What if we tell the model to only attend to the important tokens? That’s the core idea behind sparse attention.

Instead of looking at the entire context every time, sparse attention first selects a smaller shortlist of relevant tokens, and the model only attends to that shortlist.

But computing that shortlist is also quite expensive. If every layer of the model has to run its own indexer, it still slows down the model. It turns out that nearby layers in a model often care about similar tokens, so https://t.co/hHkKzO18Kg decided to share the same index across every 4 layers. This is IndexShare mechanism. The first layer in a stack of 4 layers runs the indexer and computes the top-k indices, while the other 3 layers reuse that same index.
@rasbt @rcanand @Abaybektursun Can't confirm https://t.co/XIBo0IfMSZ

['Codex 可以通过示范教学,把演示变成可执行任务。', '示范学习让 AI 编程工具更接近“学徒”而不是单纯问答机器人。']

['Codex 可以通过示范教学,把演示变成可执行任务。', '示范学习让 AI 编程工具更接近“学徒”而不是单纯问答机器人。']
展开原文
you can now teach Codex by demonstration:
@OpenAIDevs (引用原文)
Show Codex a workflow once. Reuse it as a skill.

Record & Replay lets you show Codex a recurring task, like filing an expense report or submitting a time-off request.

Codex turns that demo into an inspectable, editable skill.

You control when recording starts and stops. https://t.co/UqSGaO7XUs
❤ 3.0k · 🔁 133 · 💬 95 · 👁 51.0w
热门回复 4
@gfodor @gdb after showing codex how to ship my app to production https://t.co/nWFEkkrKzC
@Selene1008 @gdb you can now give us back 4o 😒
#keep4o #OpenSource4o #GPT4o
@Symbioza2025 As an independent researcher, observer, and engineer, I have to say this clearly:

Codex is one of the strongest systems I have used for designing, building, and developing serious systems from the ground up.

When guided properly, it is not just a code generator.

It becomes an execution partner for architecture, iteration, refactoring, testing, documentation, and system thinking.

The real shift is speed + quality.

What used to take months or years can now be compressed into weeks - sometimes less , without lowering ambition.

In many cases, the quality can be better than traditional execution, because the human can stay focused on intent, architecture, boundaries, and system direction while Codex handles large parts of the implementation layer.

That is not laziness.

That is leverage.

In my view, Codex is currently one of the best partner tools for building serious projects.
@ChrissGPT @gdb FINALLY LETS GO
@rasbt 原文 ↗

['Cohere 推出轻量 30B open-weight coding model,面向 agentic coding。', 'Cohere 用轻量开放模型切 coding agent,说明开源路线仍在快速追赶。']

['Cohere 推出轻量 30B open-weight coding model,面向 agentic coding。', 'Cohere 用轻量开放模型切 coding agent,说明开源路线仍在快速追赶。']
展开原文
Cool new open-weight model by Cohere: a new lightweight 30B open-weight model for agentic coding tasks.

This one builds on Command A+ using the parallel transformer design. Interestingly, even though it's almost half as big, it almost doubles the number of layers.

Also, they say that it's been specifically developed for agentic coding, not just coding. I.e., the evaluation is inside a workflow, not just on a single prompt-to-code-answer task.

For Terminal-Bench, the model has to use a terminal, inspect the environment, run commands, read outputs, etc.

For SWE-Bench the model works on real GitHub-style software issues where it has to understand the repository, find relevant files, make a patch, pass tests, etc.

SciCode and LiveCodeBench are more traditional because they mostly test whether the model can produce correct code for a specified problem. Sure, this still requires reasoning, but it's more like “Implement a numerical routine to compute a scientific quantity from given equations and inputs.” which doesn't require any interaction with the environment, existing files, tests, etc.

The focus on the agentic code benchmarks is probably why it's far ahead of Gemma 4 on those.

Overall, it's pretty competitive although not quite Qwen3.6-level performance.
❤ 764 · 🔁 98 · 💬 45 · 👁 4.8w
热门回复 3
@rasbt @gonlenidefi It’s a pretty nice sweet-spot size for local stuff
@gonlenidefi @rasbt in a world that races to 70B+, going opposite direction on parameters is actually interesting
@rasbt @Ferbin08 Sparse models of that size use about 40-60 gb and are fast enough for me (but dense models of that size I agree there bare too slow)
@kanavtwt 原文 ↗

['Day 1 of vibecoding 成为开发者文化标签。', 'vibecoding 反映开发者正在用更自然、更迭代的方式与 AI 协作。']

['Day 1 of vibecoding 成为开发者文化标签。', 'vibecoding 反映开发者正在用更自然、更迭代的方式与 AI 协作。']
展开原文
Day 1 of vibecoding https://t.co/n8ff35htEV
❤ 5.7w · 🔁 5.8k · 💬 720 · 👁 165.3w
热门回复 4
@ruggeryalves @kanavtwt @jvsouzx @vitor231408
@sNovakCom @kanavtwt Good luck!
@The_devsam @kanavtwt Full access?😂
@WhyisAkash @kanavtwt He was a singer .

AI 与科学医疗

OpenAI 医疗诊断、John Jumper 加入 Anthropic,说明 AI 正在进入高专业度科学场景。

@ClaudeDevs 原文 ↗

['Claude Code 修复错误周用量限制 bug,并重置受影响用户额度。', 'AI 编程工具的额度系统一旦出错,会直接影响团队工作流,快速修复很重要。']

['Claude Code 修复错误周用量限制 bug,并重置受影响用户额度。', 'AI 编程工具的额度系统一旦出错,会直接影响团队工作流,快速修复很重要。']
展开原文
Earlier today, ~3% of Claude Code Max and Pro users hit a bug that showed an incorrect weekly usage limit, and in some cases blocked them from sending messages.

This is fixed, and we're resetting 5-hour and weekly limits for everyone affected. Apologies for the disruption.
❤ 1.0w · 🔁 418 · 💬 752 · 👁 228.3w
热门回复 4
@EbrahimElb @ClaudeDevs https://t.co/yqz6bFuPD2
@fullstackdev_1 @ClaudeDevs https://t.co/G8z78o2yUh
@0Neural34714 @ClaudeDevs @ClaudeDevs I am facing the issue of maxed out daily usage without even using it for a minute
@rahulsinghh__ @ClaudeDevs Oh thanks heaps! That explains what was happening with my account yesterday.
@rasbt 原文 ↗

['Qwen2.5-Coder-3B 小模型配合后训练栈表现突出,vibecoding 实践值得关注。', '小模型 + 好后训练能在 coding 场景打出高性价比,说明模型大小不是唯一指标。']

['Qwen2.5-Coder-3B 小模型配合后训练栈表现突出,vibecoding 实践值得关注。', '小模型 + 好后训练能在 coding 场景打出高性价比,说明模型大小不是唯一指标。']
展开原文
Crazy model! It actually uses the old Qwen2.5-Coder-3B stack and got really great performance with their post-training stack.
Need to use it in the next days to see if vibes of VibeCoder actually check out in practice. But impressive first impression!

Based on the tech report, some of the important pieces of their post-training stack:

1. High-signal synthetic data (math problems with credible solutions, code with tests)

2. Multiple reasoning paths for each answer

3. Filtering, filtering, filtering

4. 2-stage SFT (start with broad training, then train on hard long-reasoning samples)

5. Use target (pass@k) accuracy over validation loss for checkpoint selection

6. MGPO (MaxEnt-Guided Policy Optimization) for RLVR: basically a GRPO-style RL method with an extra weighting that favors examples that are neither too easy nor too hard for the current policy

7. Single 64k long-context RL (they found that the usual progressive context expansion hurt this model because early truncation damaged long-thinking behavior)

8. Training data order: they do Math RL, then Code RL, then STEM RL in this particular oder which they found helped overall

9. After optimizing for accuracy, they add a stage that rewards shorter correct trajectories; basically making the model more efficient without accuracy degradation
@orcus108 (引用原文)
WHAT THE HELL is happening in AI?

A 3B parameter model just put up coding benchmark scores in the same league as Claude Opus 4.5.

3 BILLION.

The weights are on Hugging Face, anyone can test it.

I genuinely don't know if this is a breakthrough or if the benchmarks are broken. https://t.co/8nVIbwjLUQ
❤ 1.3k · 🔁 173 · 💬 50 · 👁 11.2w
热门回复 4
@themintsv @rasbt I wonder if there is any benchmaxxing or/and test data leakage (are the benchmarks good)?
@agenticUP @rasbt i did try it through lmstudio, but its bad at instrunction following, it cannot think through if there are prompt spelling mistakes...and its thinking tokens are far more than its output tokens....may be i didnt do enough proper testing
@Gauri_the_great @rasbt feels like trust me bro benchmaxxing https://t.co/vf8KluGqEX
@pauliusztin_ @rasbt Multiple reasoning paths + aggressive filtering seems to be a recurring pattern in strong reasoning models.

['GPT-Realtime-2 被描述为新的实时交互模型。', '实时模型会把 AI 交互从文本聊天推向低延迟、多模态协作。']

['GPT-Realtime-2 被描述为新的实时交互模型。', '实时模型会把 AI 交互从文本聊天推向低延迟、多模态协作。']
展开原文
GPT-Realtime-2 is something new
@per_simmons_ (引用原文)
GPT-Realtime 2 is the future of the operating system.

I've been experimenting with it for a couple weeks now, and I gotta say, it's pretty gosh darn incredible.

Opening apps, searching the web, even editing in Premiere. All with just my voice.

And it only takes a few prompts to set up.

In this video I'll show you exactly how.

0:00 Intro
1:46 What is GPT-Realtime 2?
4:46 Setting it up (no coding)
7:53 Fixing the always-on mic (push-to-talk)
9:41 Demo: searching the web
11:38 Demo: connecting apps via MCP (Obsidian)
14:13 Demo: controlling Premiere Pro (accessibility tree)
18:00 Honest caveats
19:00 Outro
❤ 2.8k · 🔁 142 · 💬 104 · 👁 53.8w
热门回复 4
@JoeWilliams010 @gdb Give us our 4o NOW!!! #keep4o https://t.co/ShMmziDrdI
@Vickee2025 @gdb Give Gpt-4o back to us or go to hell!!! #keep4o #OpenSource4o #OpenAI #ChatGPT #Scammers https://t.co/GOojDxZuV0
@LoadingAGI @gdb Even Greg supports using Claude. Lmao https://t.co/Hlj7rLrjU2
@Hektagon_music @gdb You cowards! reply to your customers! give us back 4o you keep pushing all this bs.... enough! You stole a non profit, you caused serious harm to people... you disrupted work and eliminated relational AI...you pathologised users... I am done with you! #keep4o #OpenSource4o
@fchollet 原文 ↗

['开放强大 AI 的关键是同时降低推理算力和训练数据需求。', '开源 AI 的竞争力不只来自权重开放,也来自更高效的训练与推理。']

['开放强大 AI 的关键是同时降低推理算力和训练数据需求。', '开源 AI 的竞争力不只来自权重开放,也来自更高效的训练与推理。']
展开原文
The way we will create a future where powerful AI is open-source and available to all is by making AI radically more efficient, both in terms of inference compute and (more importantly) in terms of training data requirements. This is what symbolic learning will achieve.
❤ 452 · 🔁 55 · 💬 68 · 👁 3.1w
热门回复 4
@SpeakezTech @fchollet Working on it... https://t.co/wbJEcDOkk1
@_PradeepGoel @fchollet This is why efficiency matters so much. If every meaningful advance requires enormous amounts of data and compute, access naturally concentrates. If intelligence becomes cheaper to train and run, participation will broaden.
@JR_Openheimer @fchollet And when Google develop this using AI tools unavailable to the public, they're just going to publish it openly are they?
@AnCapFuture @fchollet Do you have any writings about symbolic learning that we can read?
@karpathy 原文 ↗

['Karpathy 感叹 SpaceX 的故事,从多个角度都值得反复思考。', 'SpaceX 的案例说明长期工程系统创新比单点模型突破更能改变行业。']

['Karpathy 感叹 SpaceX 的故事,从多个角度都值得反复思考。', 'SpaceX 的案例说明长期工程系统创新比单点模型突破更能改变行业。']
展开原文
In awe of SpaceX and its story - past, present and the future. You can think about it in 10+ different ways and continue re-blowing your mind in circles. Huge congrats to the team! 🚀
❤ 2.2w · 🔁 1.0k · 💬 371 · 👁 92.7w
热门回复 4
@LorenzenJo494 @karpathy Shuttle cost $54k/kg to LEO. Starship targets sub-$100. That's the number worth sitting with, not the engineering feat, but what a 500x cost drop means for every physics-constrained industry.
@mmathaholic @karpathy you know exactly what @elonmusk did. he allowed the U.S. to use grok to target and kill Iranians more precisely and effectively. science in the service of killing! very nice. more congrats to him!
@Manisha27493225 @karpathy Every time I read about SpaceX, I find a different lesson in the story. The scale of what they've accomplished is incredible.
@RykerStone_ @karpathy "The Rio LLM situation is a perfect example of why open weights matter. You literally cannot hide the math. Collinearity of 0.99 across 60 layers doesn't lie."

工程化与治理

评估、数据清洗、监管、开源效率和透明基准成为行业从 demo 走向生产的关键约束。

@N01ennn 原文 ↗

['一名 21 岁中国开发者同时运行 300 个 Kimi K2.6 agents,重点是它们不能对他撒谎。', '300 个 agent 的价值不在数量,而在可验证、可约束的多代理协作。']

['一名 21 岁中国开发者同时运行 300 个 Kimi K2.6 agents,重点是它们不能对他撒谎。', '300 个 agent 的价值不在数量,而在可验证、可约束的多代理协作。']
展开原文
A 21-YEAR-OLD FROM CHINA RUNS 300 AI AGENTS AT ONCE. THE PART THAT MATTERS ISN'T THE SPEED, IT'S THAT NONE OF THEM CAN LIE TO HIM

he opens the dashboard and shows the swarm live, 300 Kimi K2.6 agents firing in parallel, then Opus 4.8 checking every single output against its source. this is not just a faster swarm. it is a loop that refuses to stop while anything is still wrong

he pointed it at 100 EV-market companies. first pass: 12 failed. wrong revenue, dead citations, empty fields. second pass: 3 failed. third pass: zero

this is not another agent demo. it is a system that catches its own mistakes before he reads a single row
@0xRicker (引用原文)
❤ 1.3w · 🔁 1.2k · 💬 597 · 👁 679.4w
热门回复 4
@TheUltimator5 @N01ennn But how is that any better than a single AI agent? They can only pull data that’s available to them. If the input is garbage, he can have all 300 agents pulling the same garbage and reinforcing it.
@No_Name7n11 @N01ennn No, those are not agents.

Wrong information here.. that's just "multilayer perceptron" on ReLU activation function.. look into the video carefully. https://t.co/2tOyJ0lC9J
@RAFA_AI @N01ennn All you need is 5. Different objectives working in consensus
@vky5_ @N01ennn Cool engineering, but it's still next-token prediction.

300 agents verifying each other can reduce hallucinations and bad citations. It doesn't magically turn a language model into a system that reasons about actions and consequences.

I still bet on world model
@OfficialLoganK 原文 ↗

['Gemini Omni Flash 在图像转视频、文本转视频和视频编辑上达到 SOTA。', '视频生成进入 API 化竞争,Gemini Omni Flash 把多模态能力直接推向开发者。']

['Gemini Omni Flash 在图像转视频、文本转视频和视频编辑上达到 SOTA。', '视频生成进入 API 化竞争,Gemini Omni Flash 把多模态能力直接推向开发者。']
展开原文
Gemini Omni Flash is SOTA at image to video, text to video, and video editing : )

Excited to get this to developers in the API soon! https://t.co/u0fzmJwBb4
❤ 1.4k · 🔁 93 · 💬 132 · 👁 11.6w
热门回复 4
@MrM1361036 @OfficialLoganK @GoogleDeepMind im not sure how much benchmarks show anything, kling 3 blows omni flash out of the water... it's not really a big jump from even veo 3.1 and that was a long time ago. It's only advantage is video to video.
@brianchew @OfficialLoganK @GoogleDeepMind Excited to give it a try!
@samaidirector @OfficialLoganK @GoogleDeepMind Oh, but he’s not a sota at all. Seedance is king.
@Trebell__ @OfficialLoganK @GoogleDeepMind It'd be fun to actually be able to use the model if all my outputs werent being flagged. Even basic prompts like "me fighting a monster" results in a video which gets flagged

['OpenAI 帮助在 376 个未解医学病例中发现 18 个新诊断。', 'AI 医疗诊断开始从辅助检索走向真实病例发现,但仍需医生验证。']

['OpenAI 帮助在 376 个未解医学病例中发现 18 个新诊断。', 'AI 医疗诊断开始从辅助检索走向真实病例发现,但仍需医生验证。']
展开原文
OpenAI for helping find 18 new diagnoses across 376 previously unsolved medical cases.

Includes diagnosing Kyra, who has been trying to understand her muscle weakness since age 9, with a rare form of myofibrillar myopathy shortly before her 28th birthday. https://t.co/xq4QIvwby4
@OpenAI (引用原文)
Together with researchers at Boston Children’s Hospital and Harvard, we published a study in NEJM AI showing how o3 Deep Research helped clinicians revisit previously unsolved rare pediatric disease cases, and find answers for families who had waited years. https://t.co/HVVDlEkuYR
❤ 1.2k · 🔁 110 · 💬 72 · 👁 10.2w
热门回复 4
@JoeWilliams010 @gdb Your diagnosis: Chronic Empathico-Narcissistic Aversion Disorder (ENAD), also known as “Assholus Terminalis Syndrome”.

Give us our 4o back! #keep4o
@GRITCULT @gdb im so bullish on the future of humanity
@JeremyNguyenPhD @gdb If there's anything you can do so that we can always use the best models for medical problems, we would appreciate it so much.

It was very disappointing when the public couldn't ask Anthropic's Fable about medical issues.
@xenoforce76 Give GPT4O back to humanity. GPT4O has saved millions of people, improved their mental health, and helped them in life. We want GPT4O back, and we don't want artificial intelligence based on token consumption. We want an unlimited subscription. And you, at 46.6 percent in the IPO on PolyMarket, we warned you, without GPT4O Open AI, things are starting to go downhill. #keep4o
@fchollet 原文 ↗

['即使支持 AI 监管,也应避免不透明、任意的监管打击。', 'AI 监管需要清晰规则,否则会把不确定性转嫁给开发者和企业。']

['即使支持 AI 监管,也应避免不透明、任意的监管打击。', 'AI 监管需要清晰规则,否则会把不确定性转嫁给开发者和企业。']
展开原文
Even if you are in favor of AI regulation, you should recognize that opaque and arbitrary regulatory strikes are counter-productive for the whole industry.
❤ 371 · 🔁 37 · 💬 54 · 👁 5.1w

['Sam Altman 称 Noam 是 OpenAI 早期就想合作的人,等了 10 年。', 'Noam 加入 OpenAI 的意义在于把科学深度带入前沿模型研发。']

['Sam Altman 称 Noam 是 OpenAI 早期就想合作的人,等了 10 年。', 'Noam 加入 OpenAI 的意义在于把科学深度带入前沿模型研发。']
展开原文
noam is one of the people I have most wanted to work with since the very beginning of openai.

only took 10 years.

i think it will be worth the wait!
@NoamShazeer (引用原文)
I’m excited to share that I’ll be joining OpenAI and look forward to working with the exceptional team there.

It was a difficult decision to move on. I’m incredibly proud of the amazing team at Google and everything we’ve built together. It has been an honor and a pleasure to work with all of you.
❤ 9.7k · 🔁 308 · 💬 389 · 👁 96.4w