Introducing Claude Sonnet 5, our most agentic Sonnet yet.
It makes plans, uses tools like browsers and terminals, and runs autonomously at a level that just a few months ago required larger and more expensive models. https://t.co/UKK8G7ww5h
@claudeai Actually a beneficial add on. Hope to see more great papers coming out of this and not slop.
@nicomusitu @claudeai 想象一下用这首音乐配插片。科学突破应该配更好的发布视频音乐。老式企业销售幻灯片音乐加 AI 元素会让科学显得无聊。
是时候让科学成为最鼓舞人心和令人惊叹的事物了。
@claudeai imagine interstellar with this music. scientific breakthrough deserves better launch video music. old corporate sales slides music with an ai twist makes science feel boring.
time to make science the most inspiring and awe-provoking thing ever.
Claude Fable 5 will be available again globally tomorrow.
After a series of productive conversations with the US government, we're redeploying the model with a new set of classifiers to target and block more cybersecurity tasks. In the near term, some routine tasks like coding and debugging will fall back to Opus 4.8. We’ll continue to refine these classifiers over the coming weeks to reduce false positives and better distinguish genuine misuse from legitimate requests.
We’ve also begun drafting a consensus framework—with Amazon, Microsoft, Google, and other Glasswing partners—for assessing the severity of AI jailbreaks and how AI developers should respond to them. We invite other industry partners and model providers to join us in this effort.
Finally, we’re scaling up our collaboration with the US government on model testing and safeguards. This will include pre-release access to models and safeguards for evaluation, information sharing on jailbreaks and misuse, and dedicated resources for joint research.
Thank you to our users for your patience, and to our partners across the government, industry, and the research community who worked alongside us to make Fable 5 available again.
Wow… I’ve been really really enjoying the Claude code setup for months but..
- fable was already heavily labotomized during the 3 days I used it (I never once asked anything malicious) - my $200 a month plan will no longer cover the top model… - and during this 6 day release it may not even be able to perform coding tasks?? 🤣
I hope SOL gets released soon. I might have to move back to OpenAI Codex.
@VolumeOnMaxx @AnthropicAI 你现在可以花 2 倍的钱来使用 Opus 4.8 伪装成 Fable 吗?
@AnthropicAI You can now pay 2x to use Opus 4.8 pretending to be Fable?
Good new first: Sol is a smart, efficient, and a significant step forward. It is the same price as GPT-5.5. Also launching in the GPT-5.6 family is Terra, with 5.5-level performance at half the price.
Bad news: at the request of the US government, it is launching today in limited preview instead of the open access launch we were planning on. We are working with the government to get to general availability as fast as we can.
I think it is quite reasonable to roll out models--especially as they reach significant new levels of capability--in this way. It fits with our long-held strategy of iterative deployment. But this isn't quite the process that we think is optimal.
Now we will with the government to attempt to get to a transparent, reliable process for early access, and to ensure that as long as our safeguards work as intended we can release widely. We want to be a reliable, dependable partner that works with all stakeholders, and we also want to live by our mission of benefiting all of humanity. I believe the government shares most of our goals, and that they are overall doing a good job in a very difficult situation.
We will work as quickly as we can to get this model in your hands and we hope you will love it.
@sama Bad News? Limiting access to the best model is inconsistent with OpenAI's stated mission! It provides an advantage to some organizations over other.
Please stop allowing all access to 5.6 until everyone can have access. You have the ability to do that.
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.
@gdb Curious how it holds up on chained tool calls. We kept hitting state drift after 3-4 steps and solved it with explicit state summaries between calls. How does the consistency look in your tests?
Google Gemma 4和多模态模型更新
Google发布Gemma 4模型,强调设备端AI能力;同时推出Nano Banana 2 Lite(<4秒图像生成)和Gemini Omni Flash(视频编辑SOTA)两款多模态模型。这些更新丰富了Google在AI模型生态中的产品线,特别是在移动和创意应用场景。
@OfficialLoganK Seriously you should rebrand it. Nano Banana.. Does sound google, doesn't sound serious. In fact most of the time i'm asking something to Nano Banana it refuses/bugs and never give the reason. Grok imagine is much faster, better, and always output in this area.
@tts23665 @OfficialLoganK 说真的,Logan?
你们公司说 Gemini 3.5 Pro 会在一个月内发布。但截止日期已经过去了,仍然没有解释、更新或透明度。
你们公司到底在做什么?
@OfficialLoganK Seriously, Logan?
Your company said Gemini 3.5 Pro would be released within a month. That deadline has already passed, and there's still no explanation, no update, and no transparency.
The dura is the brain's armor: a membrane so tough that a surgeon normally cuts through it with a scalpel. For the first time in our clinical trials, we inserted the electrode threads of our implant straight through the dura and into the cortex, keeping the dura intact.
Here's how we did it 🧵
❤ 2.0w · 🔁 3.7k · 💬 3.0k · 👁 303.9w
热门回复 3
@Symbioza2025 这是令人难以置信的进步。
但当技术开始直接与大脑接口时,问题就不再仅仅是技术性的。
我们如何确保人类在他们所创造的奇迹中不会失去自我意识?
不是恐惧。不是阻挡进步。
人类必须处于中心位置。 安全不能只是一个承诺——它必须是架构。
This is an incredible step forward.
But when technology begins to interface directly with the brain, the question is no longer only technical.
How do we make sure humans do not lose their own awareness in the wonder of what they have created?
Not fear. Not blocking progress.
Humanity must remain at the center. Safety cannot be only a promise - it must be architecture.
@elonmusk Happy belated birthday, Elon! 🫡 (P.S. My birthday is also on June 27—almost the same day as the boss’s) --- I didn’t know you were a Cancer too... cool... and thanks for everything you do... I admire you and your work... because it just makes sense... thanks, Elon
@futur3funk @elonmusk 然而——地球上最贫穷的人现在正因你而死亡......
@elonmusk and yet - the poorest people on the planet are now dying thanks to you......
The dura is the brain's armor: a membrane so tough that a surgeon normally cuts through it with a scalpel. For the first time in our clinical trials, we inserted the electrode threads of our implant straight through the dura and into the cortex, keeping the dura intact.
@neuralink If CAPTAIN JEAN- LUC PICARD could have an .... Observation about this..... GENTLEMEN, this is ' WRONG ANSWER '. The ' BORG COLLECTIVE '.... HAS NO SOUL! Your LAUGHTER, won't PREVENT...... ASSIMILATION! Welcome to ' The MACHINE '!
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.
Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.
We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.
telepathy is exciting, not only for people with disabilities, but for everyone. language can only approximate our thoughts. the ability to share experience across minds could open up an entirely new, rich, and dynamic form of communication- one that expands human intelligence and connection 🤍✨
To help accelerate neuroscience breakthroughs, we're releasing the full training code for Brain2Qwerty v1 and v2, and our partner, @bcbl_, is releasing the v1 dataset.
We’re sharing the next major milestone in our non-invasive brain-to-text decoder research: Brain2Qwerty v2.
Building on v1, which was published today in @Nature, Brain2Qwerty v2 is the highest-performing end-to-end pipeline capable of real-time sentence decoding from raw brain signals. It advances beyond character-level performance to decoding words and semantics, enabling accuracy for overall communication.
We believe this research has the potential to make a real difference for the millions of people who suffer from brain lesions or disorders that prevent them from communicating.
We trained Brain2Qwerty v2 on ~22,000 sentences from 9 volunteers, each recorded for 10 hours wearing an MEG device while typing.
By using end-to-end deep learning on raw brain signals from MEG devices and fine-tuning LLMs, the system effectively bridges the gap between noisy neural data and coherent language.
The results are promising: - Avg word accuracy of 61% across participants - 78% word accuracy and 50%+ of sentences decoded with ≤ 1 word error for the top-performing participant - Performance scales log-linearly with data volume
As engineering, product, design, DS, etc. melt into a new kind of role, I was reflecting on what roles might look like in the future. For example, when I look at the Claude Code team I see what I think is five archetypes:
1. Prototyper: comes up with brand new ideas; churns out many ideas, most of which don't ship 2. Builder: quickly turns a prototype/idea into production-grade product/infra 3. Sweeper: cleans up the UI, simplifies the code and system, unships, optimizes performance 4. Grower: takes a product that has been built and iterates on it to improve Product-Market Fit 5. Maintainer: owns a mature system to make it secure, reliable, fast, and efficient as it scales
Many people span across 2 roles, and sometimes 3 roles. I also notice that these roles are not really tied to job function -- eg. across Anthropic, some designers match category 1, some 2, some 3; same for engineers, PM, DS.
A healthy team needs a mix of these, depending on the product:
- A product that is new and pre-PMF needs people that are strong at 1+2+3 - A product that is growing and has found PMF needs 2+3+4 and some 5 - A product that has strong PMF needs 3+4+5 and some 2
Maybe product roles of the future will look more like this, and less like the domain-specific roles of today?
What dies is the mediocre middle layer. Not designers. Not engineers. The person who existed to bridge the gap between them. I believe. A dev can now prototype without waiting on a designer. A designer can validate without waiting on an engineer. The connective-tissue jobs are gone. A builder can build without asking for help of any. What's left? Every role gets harder. You survive by being genuinely good at your domain, not just good enough to justify being in the room. The bar to be worth hiring just got a lot higher. And more software ships because of it, not less.
"循环工程" 是最近的热门流行语,在 Boris Cherny(Claude Code 的创建者)和 Peter Steinberger(OpenClaw 的创建者)的社交媒体提及后走红。循环现在是我们让 AI 代理长时间迭代以构建软件的关键部分。在这篇文章中,我想分享我构建 0 到 1 产品的三个关键循环,如图所示。这些循环不仅指导我如何构建软件,还指导我如何决定构建什么软件。
Agentic 编码循环:给定一个产品规格说明和可选的一组评估标准(即用于衡量性能的数据集),我们可以让 AI 代理编写代码,测试其工作成果,并不断迭代,直到代码无错误且满足规格要求。这个闭合循环的概念在去年年底开始流行,成为让编码代理能够长时间高效工作而无需人工干预的游戏规则改变者。例如,上周末我一直在为女儿构建一个练习打字的应用,我的编码代理可以轻松地连续工作一个小时,使用网络浏览器多次检查所构建的内容,然后再回报给我,而无需我的干预。
AI 原生团队越来越多地使用 AI 来帮助塑造产品方向,例如自动收集和分析使用数据、总结书面和口头客户反馈,或进行竞争分析。然而,对于我参与的几乎所有产品,我都看到人类在用户和产品运营环境方面拥有当前 AI 系统无法比拟的丰富背景知识——因此人类在其中扮演着关键角色。许多人将这种人类贡献描述为"品味",但我更倾向于认为这是人类拥有背景优势,因为这为帮助 AI 系统变得更好指出了更清晰的路径。这也说明了为什么这一步无法自动化:只要人类知道 AI 不知道的事情,就需要人在环中注入这些知识。
“Loop engineering” is a hot buzzphrase after mentions of it by Boris Cherny (Claude Code’s creator) and Peter Steinberger (OpenClaw's creator) went viral on social media. Loops are now a key part of how we get AI agents to iterate at length to build software. In this letter, I’d like to share my 3 key loops, shown in the image below, for building 0-to-1 products. These loops guide not just how I build software, but also how I decide what software to build.
Agentic coding loop: Given a product specification and optionally a set of evals (that is, a dataset against which to measure performance), we can have an AI agent write code, test its work, and keep iterating until the code is bug-free and meets its specification. This idea of closing the loop took off around the end of last year, and it has been a game changer in enabling coding agents to work longer productively without human intervention. For example, over the weekend, I was building an app for my daughter to practice typing, and my coding agent could easily work for around an hour, using a web browser to check what it had built multiple times before getting back to me, without needing my intervention.
The engineering loop executes quickly. Every few minutes, the coding agent might build and test a new version of the software. I hear frequently from developers who are finding new ways to engineer more effective engineering loops. This is an active area of invention!
Developer feedback loop: In this loop, a developer examines the current product and steers the coding agent to improve it. Last year, a lot of developers (including me) were acting as the QA (quality assurance) function for our coding agents, manually finding bugs and then asking the agent to fix them. But with coding agents much more able to test their own code, the amount of time we need to spend on this function has decreased significantly. This allows us to make higher-level product decisions, such as what key features to offer, where the UI needs improvement, and so on.
The developer-feedback loop operates over time intervals between tens of minutes and hours — that's how frequently a developer might review a product and give feedback. In the case of the typing app, I changed my mind a few times about the visual design, what cat costumes she can unlock as she learns (she loves cats), and the user flow for a grown-up to log in and steer the child's learning experience.
When a developer has a clear vision for what to build, it is still a lot of work to translate that vision into a specification for a coding agent to implement. Further, after the developer has seen an implementation, they might update (or perhaps clarify) the spec to steer it toward what they want. If you find that the system repeatedly runs into certain problems, building a set of evals for the agent becomes useful.
AI-native teams are increasingly using AI to help shape product direction, for example, automating the gathering and analysis of usage data, summarizing written and verbal customer feedback, or carrying out competitive analysis. However, for pretty much all the products I’m involved in, I see humans as having a significant context advantage over current AI systems — we know a lot more than the AI system about the users and the context the product has to operate in — and thus humans play a critical role. Many people describe this human contribution as “taste,” but I prefer to think of it as humans having a context advantage, since that gives us a clearer path to helping AI systems get better. This also speaks to why this step can’t be automated: So long as the human knows something the AI does not, human-in-the-loop is needed to to inject that knowledge into the system.
External feedback loop: This includes a wide range of tactics like asking a few friends for feedback, launching to alpha testers, or putting the code into production with A/B testing. These tactics are usually slow, rarely taking less than hours and sometimes taking days or even weeks. This data informs the developer vision, which in turn continues to drive the detailed product spec, which in turn drives the coding agent.
With coding agents speeding up software development, more engineers are starting to play a partial product management role. For many engineers who are growing into this role, the hardest part is shaping the product vision and striking a balance between building (bridging the gap between vision and spec) and getting user feedback to evolve the vision. It is important to do both!
I will write more about how to do this in future posts, but for now, I find it encouraging that engineers are playing an expanded role (just as product managers and designers now do more engineering).
[Original text: The Batch]
❤ 5.7k · 🔁 1.1k · 💬 219 · 👁 32.3w
热门回复 4
@Oldnoob007 @AndrewYNg "品味"对我来说总感觉像是敷衍的词 Andrew 的重新定义更好:"上下文优势"
@AndrewYNg "taste" always felt like a cop-out word to me andrew's reframe is better: "Context Advantage"
it means the developer who spends more time with users , builds better products than the developer who doesn't that's always been true. AI just made it the only remaining edge.
@AndrewYNg The key risk is measuring the loop too loosely: if the eval set is static or vague, agents can overfit. I’d add a small failure-case review every few iterations.
The hardest part of growing into a product role as an engineer is not the product thinking. It is tolerating the slower external feedback loop after spending years optimizing for the fast engineering one. Two completely different relationships with time and uncertainty running in parallel.
@AndrewYNg this matches what i'm seeing building with agents daily... the coding loop is basically solved enough at this point. The actual scarce skill now is writing specs precise enough that "iterate until it matches" means something. spec writing is the new debugging tbh
@rowancheung There is no unlimited cooling. Everything boils down to heat transfer. Eventually at scale on a global level we will end up warming the ocean significantly affecting marine lives and change in weather patterns. @poovulagu @veritasium
@statys @rowancheung 相当确定瓶颈现在是抗议者了。
@rowancheung Pretty sure the bottleneck is protesters at this point.