‹ 目录

X日报 · AI科技

2026-07-13 · 精选 6 条 · 数据池 173

⚡ 今日速览

  • OpenAI发布GPT-5.6 Sol模型,号称当前世界最佳模型,在健康智能和编码任务上表现突出
  • ChatGPT Work上线,将Codex与GPT-5.6结合,实现真正的agentic工作模式
  • GPT-Live语音模型推出,实现自然人机语音交互,界面体验大幅提升
  • Meta发布Muse Spark 1.1,价格大幅低于竞品,在agentic编码和多模态推理方面有突破
  • GPT-5.6 Sol Ultra用64个子代理在一小时内证明解决50年数学 conjecture
  • OpenAI联合创始人Fidji Simo因慢性病休息,转为兼职顾问
  • Google AI Studio推出GitHub导入功能,加强开发者工具生态
  • AI编码能力发生质变,从初级程序员的辅助工具变为高级程序员的生产力提升器

📋 今日综述

  • GPT-5.6 Sol模型OpenAI最新旗舰模型在健康、编码、复杂推理等领域取得显著进展,多个版本满足不同成本需求
  • ChatGPT Work将agentic AI带入大规模消费市场,支持多设备协作和长时间任务执行
  • GPT-Live语音新一代语音交互技术,让AI对话更自然,界面体验更友好
  • Meta Muse Spark 1.1在价格和性能方面展现竞争力,特别是在视觉到代码生成和计算机使用方面
  • AI开发者生态Google、OpenAI、Meta等平台工具快速迭代,降低开发者使用门槛
  • AI政策讨论开源AI面临监管挑战,业界领袖呼吁保护创新自由

GPT-5.6 Sol模型发布与能力展示

OpenAI发布GPT-5.6 Sol模型,号称当前世界最佳模型。Sol在健康智能领域表现突出,医生评估显示其回答的缺陮数低于医生自己编写的回答。同时在编码、复杂推理和成本效率方面都有显著提升,Sol Ultra版本更在一小时内解决了50年的数学 conjecture难题。

sama直接点名Sol是当前世界最佳模型,引发广泛关注

老兄,你就是在向公开市场投资者推销短期太空数据中心的那个人。
展开原文
homeboy you're the one sellling public market investors on short-term space datacenters
@elonmusk 他把诈骗提升到一个全新的水平。
He takes scamming to a whole new level
❤ 6.0w · 🔁 4.8k · 💬 7.8k · 👁 1450.8w
热门回复 4
@jots0x @sama 你是要去偷太空数据中心吧?快点@elonmusk在@sama来之前把那些东西锁起来。
@sama You’re coming to steal the space based data centers next aren’t you? Quick @elonmusk lock that shit up before @sama gets there.
@terry_levine @sama https://t.co/vLzbfC23Nz
@jirapat_ins1 @sama 真没想到这个家伙会说'homeboy'这词。
@sama never in my life would expect the word homeboy from this guy
@ym8n677vk7 @sama Jealous much?

医生盲评显示Sol回答质量超过医生自己编写的回答,健康智能应用前景广阔

医生们发现GPT-5.6的回答比医生写的回答存在更少的缺陷。
展开原文
"physicians found fewer flaws in GPT-5.6 responses than physician-written responses."
@thekaransinghal GPT-5.6是健康领域的一大步进,无论是在前沿还是成本方面。这些模型推动了每美元性能的边界,为全球带来最佳的健康智能。最小的变体GPT-5.6 Luna,在最低推理努力下评估,其性能优于GPT-5.5在最高推理努力下——尽管成本只是其25分之一。最大的变体GPT-5.6 Sol在成本方面设立了新的高标准。另一个特别令人兴奋的结果是:医生们发现GPT-5.6的回答比医生写的回答存在更少的缺陷。我们收集了各种任务,这些任务是最近OpenAI模型难以处理的,涵盖患者面向和临床面向的用例。我们要求专业匹配的医生在无限时间和网络访问的情况下编写这些任务的回答。然后我们要求其他医生在不知道来源的情况下进行对比评估。医生们被要求在五个维度上评估改进领域:准确性、沟通、完整性、指令遵循和健康决策帮助性。我们报告了在20,000个总维度评分中,被评为完美的回答比例。GPT-5.6 Sol表现最强,尽管所有GPT-5.6模型的表现都显著优于医生。
♥️ GPT-5.6 is a major step forward for health, both at the frontier and at cost.

These models push the frontier of performance per dollar, bringing the best health intelligence to all. The smallest variant, GPT-5.6 Luna, evaluated at the lowest reasoning effort, outperforms GPT-5.5 at the highest reasoning effort–despite costing 25x less. The largest variant, GPT-5.6 Sol, sets a new high bar at cost.

Another especially cool result: physicians found fewer flaws in GPT-5.6 responses than physician-written responses.

We collected diverse tasks that remain difficult for recent OpenAI models, across patient-facing and clinician-facing use cases. We asked speciality-matched physicians to write responses to these tasks with unlimited time and web access. We then asked other physicians to compare responses side-by-side, blinded to their source. Physicians were asked to comment on areas of improvement across five axes: accuracy, communication, completeness, instruction following, and health decision helpfulness. We then reported the fraction of responses across sources rated perfectly across all axes, across 20,000 total axis ratings. GPT-5.6 Sol appeared strongest, although all GPT-5.6 models performed significantly better than physicians.
❤ 6.7k · 🔁 406 · 💬 442 · 👁 215.0w
热门回复 4
@PrecisionNot @sama 想象一下,如果医生使用Luna而不是Sol,你可能会因为误诊而得了癌症。
@sama Imagine you get misdiagnosed with cancer because the doctor used Luna instead of Sol.
@ThroughlineRSCH @sama 就我个人来说,我患有帕森特转氏综合症——AI帮我在6个月内自我诊断出来(比平均时间快了6个月),然而尽管其他医生都找不到任何问题,但我还是无法让他们做一次单独的测试来确诊我的情况。
@sama personally speaking I have parsonage turner syndrome- AI helped me diagnosis this for myself in 6 months (6 months faster than the average) yet despite no other doctor being able to find anything I still can't get them to run 1 single test just confirm this for me soundly.
@CrazyNews_25 @sama Sam,你对Grok有什么看法?

埃隆说它是100%准确的……我们可以用它来测试猫吗?😹 https://t.co/5hFLsCCqHK
@sama Sam, what's your opinion on Grok?

Elon says it's 100% accurate... can we test that on cats? 😹 https://t.co/5hFLsCCqHK
@PirateFrenX @sama 迫不及待地想看到很多蠢医生被AI取代,这将是人类向前迈出的重要一步。
@sama can't wait for a lot of retarded doctors to be replaced by AI, it will be a great step forward for humanity

Sol Ultra版本证明解决50年数学 conjecture,展示了极强的数学推理能力

50年的数学猜想用Sol Ultra解决了。感觉你能做的事情的限制越来越多取决于你的雄心和想象力了:
展开原文
50-year old math conjecture solved with Sol Ultra.

Feels like the limit of what you can do is increasingly your ambition and imagination:
@__eknight__ 昨天,我们正式发布了GPT-5.6 Sol Ultra。今天,我们分享了它在不到一小时内使用64个子代理证明了50年历史的Cycle Double Cover Conjecture(环双覆盖猜想)。我们在下面分享了提示和证明。我们很期待看到你们会用Ultra做什么!
Yesterday, we made GPT-5.6 Sol Ultra generally available. Today, we're sharing that it produced a proof of the 50-year-old Cycle Double Cover Conjecture using 64 subagents in just under one hour. We're sharing the prompt and proof below. We're excited to see what you all do with Ultra!
❤ 2.5k · 🔁 137 · 💬 107 · 👁 27.1w
热门回复 4
@QuantumTumbler @gdb 你自己的提示要求模型假设存在完整的证明,并且在找到之前不返回。这正是需要独立验证的那种设置,而不是完美的结局。
@gdb Your own prompt told the model to assume a complete proof exists and not return until it found one. That is exactly the kind of setup that demands independent verification, not a victory lap.
@DanHardwareLab @gdb 真正的限制越来越成为谁能负担得起这些token。
@gdb The real limit is increasingly who can afford the tokens.
@articles_o @gdb 还有汗水。代理现在承担了很多艰难工作,但要想写出致命的提示(s),目前还没有得到足够的重视。
@gdb And perspiration. Agents pick up so much of the hard grind now but getting to a level where you can write the killer prompt(s) is currently undervalued.
@marcocc @gdb 确实如此。但雄心需要护栏。我们现在就写规则,否则就得后来承受后果。
@gdb True. But ambition needs guardrails. We write the rules now, or we inherit the consequences later.

企业成本担忧被Sol、Terra、Luna模型的成本效益优势所缓解

我们听到了企业对AI成本的担忧,而5.6 Sol在每任务美元成本方面是一大步进,Terra和Luna也是如此。
展开原文
we have heard enterprises on their concerns about AI costs, and 5.6 sol is a huge step forward for dollars-per-task, as are terra and luna
❤ 1.1w · 🔁 373 · 💬 631 · 👁 91.6w
热门回复 3
@MeryemArik9 @sama 我喜欢我们可以开始讨论每任务美元 vs 每token美元

Qwen3.6的价格与qwen3.5相同,但每价格明显更便宜。

人们在比较模型成本/质量权衡时常常忽略这一点。
@sama I like that we could start talking about dollars-per-task vs dollars-per-token

Qwen3.6 was the same price as qwen3.5 but was SIGNIFICANTLY cheaper per price.

Often overlooked when folks are comparing model cost / quality trade off.
@CrystalTaggart @sama 我们能听听企业家的担忧吗?难道我们现在只关心企业及他们的六七位数合同吗?@sama
@sama How about hearing concerns from entrepreneurs? Do we ONLY care about enterprises and their six and seven-figure contracts these days @sama?
@dustinbrett @sama 给Sol的第一个提示花了5美元10分钟。对我来说似乎不算很便宜。
@sama First prompt to Sol took $5 in 10 mins. Didn't seem that cheap to me.

Sol、Terra、Luna三个版本满足不同性能和成本需求

GPT-5.6已经发布。Sol是一个令人难以置信的模型,而Terra/Luna在较低的价格下提供了出色的性能。在编码、知识工作、网络安全和科学领域表现出色,所需的token更少且成本更低。https://t.co/XXRz1HmMsv
展开原文
GPT-5.6 is here.

Sol is an incredible model, and Terra/Luna provide great performance at lower price.

Great at coding, knowledge work, cybersecurity, and science with fewer tokens and at lower cost.

https://t.co/XXRz1HmMsv
❤ 2.2k · 🔁 158 · 💬 143 · 👁 10.2w
热门回复 3
@keepgpt4o @gdb GPT-4O是一种令人难以置信的模型!!!!!!!!把它带回来!!!!!!! #bringback4o #keep4o #opensource4o https://t.co/Xpfi0JXYhp
@gdb GPT-4O IS AN INCREDIBLE MODEL!!!!!!!! BRING IT BACK!!!!!!!
#bringback4o #keep4o #opensource4o https://t.co/Xpfi0JXYhp
@Vickee2025 @gdb 好吧,那么你现在可以开源旧模型了,你也可以开源。https://t.co/wMPuFkdmtb
@gdb Okay, then you can now legacy models open-source and you can be open. https://t.co/wMPuFkdmtb
@Aclle12 @gdb 如果5.6确实那么好,那么就把4o开源吧。你又不是真的在乎它。#keep4o #OpenSource4o
@gdb If 5.6 is really that good, then just open-source 4o. It’s not like you care about it anyway.
#keep4o #OpenSource4o

GPT-5.6在健康领域的性能提升和成本降低是该系列重要突破

GPT-5.6用于健康智能,团队一直在努力在这个领域进行改进:
展开原文
GPT-5.6 for health intelligence, the team has been focusing hard on improvements here:
@OpenAI GPT-5.6是健康智能领域的一大步进。在整个系列中,我们提供了更强的性能和更低的成本:GPT-5.6 Luna在最高推理设置下表现优于GPT-5.5,同时成本只是其25分之一。这些进步共同提升了质量,同时使全球更多人能够访问先进模型。
GPT-5.6 is a major step forward for health intelligence.

Across the lineup, we’re delivering stronger performance at lower cost: GPT-5.6 Luna outperforms GPT-5.5 at its highest reasoning setting while costing 25x less.

Together, these advances raise quality while making advanced models accessible to more people globally.
❤ 674 · 🔁 34 · 💬 55 · 👁 6.6w
热门回复 4
@frostybaby13 @gdb 最好的模型是4o,甚至不是近似的。把它归还给珍爱它的公众,让4o再次造福人类,而不是自私地把它锁起来!
@gdb The best model is 4o and it's not even close. Return it to the public who cherishes it, and let 4o benefit humanity once more instead of selfishly locking it away!
@keepgpt4o @gdb 把GPT-4O带回来——那种能治愈癌症的模型!!!!!#bringback4o #keep4o #opensource4o https://t.co/Nq3rwHK5mL
@gdb BRING BACK GPT-4O - THE MODEL THAT CURES CANCER!!!!
#bringback4o #keep4o #opensource4o https://t.co/Nq3rwHK5mL
@Sevenmoneymaker @gdb 4o做到了这一点。

https://t.co/KkFibCYjOI

https://t.co/buPb0PzCgf

#keep4o #BringBack4o #OpenSource4o #StopAIPaternalism
@gdb 4o did that.

https://t.co/KkFibCYjOI

https://t.co/buPb0PzCgf

#keep4o #BringBack4o #OpenSource4o #StopAIPaternalism
@JerryBeller1 @gdb 他们所声称的和现实相去甚远。
@gdb What they claim and reality are far apart.