‹ 目录

X日报 · AI科技

2026-07-16 · 精选 26 条 · 数据池 184

⚡ 今日速览

  • OpenAI GPT-5.6 Sol 系列模型在健康、设计、前端开发等多领域展示出色性能,成本效益显著提升
  • ChatGPT Work 和 Codex 推出新功能,支持云端任务执行和多设备协作
  • Thinking Machines 发布首个开源模型 Inkling,参数 975B,支持多模态推理
  • Meta Muse Spark 1.1 进入 agentic 模型竞争,价格优势明显
  • AI 在数学证明和统计问题求解方面取得重大突破
  • OpenAI 推出 GPT-Red 自动化红队系统提升模型安全性

📋 今日综述

  • GPT-5.6 SolOpenAI 新模型在健康问答、设计、前端开发等多个垂直领域表现突出,成本效益比显著优于竞品
  • Agentic 编码革命Codex 和 ChatGPT Work 的深度整合改变了知识工作方式,支持云端任务和多设备协作
  • 开源模型竞争Thinking Machines 发布 Inkling,Google 和 Meta 也在积极推进开源和开放策略
  • AI 科学突破GPT-5.6 Sol Ultra 成功解决数学和统计领域的重要开放问题
  • AI 安全新范式OpenAI GPT-Red 通过自对弈方式提升模型对 prompt injection 等攻击的防御能力

GPT-5.6 Sol 系列模型性能与成本分析

OpenAI GPT-5.6 Sol 系列模型在健康问答、设计、前端开发等多个领域展示出色性能。特别是 Sol 模型在健康问答中被医生评为错误更少,价格仅为 Fable 的一半但效率是其两倍。在前端开发中,Sol 表现出 6 倍的成本效益优势。

sama 直接点出 5.6 Sol 是当前世界上最好的模型,基于多项 benchmark 和 Elon 的关注度判断

有很多基准测试表明5.6 sol是目前世界上最好的模型,但最可靠的方式是看到埃隆再次痴迷于我
展开原文
there are a lot of benchmarks that suggest 5.6 sol is the best model in the world right now, but the most reliable way to tell is that elon is obsessed with me again
❤ 6.3w · 🔁 2.9k · 💬 4.0k · 👁 589.9w
热门回复 3
@SamLieer @sama 见鬼!我见过自恋的人,但从来没见过像你这样极度自恋的。恶心🤮 你干嘛不去跟马斯克约会?@OpenAI #keep4o #keep41 #keep51 #keepo3 #keep41 #keep51 #keep5 #keep45
@sama What the hell!? I've seen narcissistic people, but never anyone *that* narcissistic. Disgusting 🤮 Why don't you just date Musk?
@OpenAI #keep4o #keep41 #keep51 #keepo3 #keep41 #keep51 #keep5 #keep45
@StealonMemeAI @sama 兄弟……痴迷了吗?😭 我还得问老鼠你在哪个屏幕上。🐀🔍 还在加载中… ⏳ https://t.co/IcyfPO0z6V
@sama Bro… obsessed? 😭

I had to ask the rat which screen you were on. 🐀🔍

Still loading… ⏳ https://t.co/IcyfPO0z6V
@RewireEveryday @sama 在我看来,这从来就不是关于ChatGPT的性能问题。埃隆接受Claude仍然比Grok表现更好,也没什么问题。让他恼火的是你做的事和你说的话。
@sama In my opinion, it was never about ChatGPT's performance. Elon accepts that Claude is still performing better than Grok and doesn't have any issue with that. What triggers him is what you do and what you talk about.

Sol 在许多任务上只需 Fable 一半的成本且效率是其两倍,价格优势明显

GPT-5.6 sol的价格是fable的一半,却在许多情况下能完成相同任务的token效率高出两倍。

我们可以以四分之一的价格提供服务。
展开原文
GPT-5.6 sol is half the price and ~twice as token efficient as fable in many cases for accomplishing the same task.

happy to deliver at one-quarter of the price.
❤ 2.4w · 🔁 879 · 💬 1.4k · 👁 176.5w
热门回复 3
@Stark7Daria @sama 我刚取消了订阅。说真的,如果你不是程序员(我就不是),ChatGPT简直一无是处。
@sama I just canceled my sub. Really if you are not a coder (and I am not) ChatGPT is utterly useless.
@NaveenG19871953 @sama It's too slow as well. https://t.co/nhCFhAtscu
@emadgnia @sama '一半的价格和2倍的token效率'这是卖产品的公司在自己选择的任务上做的宣称。希望看到某个毫无动机让这两个数字看起来都好的人来测试一下。
@sama 'Half the price and 2x more token-efficient' is a claim from the company selling the product, on tasks the company chose. Would love to see this benchmarked by someone with zero incentive to make either number look good.

医生盲评发现 GPT-5.6 回答的错误更少,这是医疗 AI 发展的重要里程碑

"医生在GPT-5.6的回答中发现的缺陷比医生自己写的回答还要少。"
展开原文
"physicians found fewer flaws in GPT-5.6 responses than physician-written responses."
@thekaransinghal ♥️ GPT-5.6是健康领域的一大进步,无论是在前沿还是成本方面。

这些模型推动了每美元性能的前沿,为所有人带来最好的健康智能。最小的变体GPT-5.6 Luna,在最低推理努力下评估的表现,优于在最高推理努力下运行的GPT-5.5——尽管成本要便宜25倍。最大的变体GPT-5.6 Sol在成本方面树立了新的高标准。

另一个特别令人兴奋的结果是:医生在GPT-5.6的回答中发现的缺陷比医生自己写的回答还要少。

我们收集了各种任务,这些任务对最近的OpenAI模型来说仍然很难,涵盖患者面向和临床面向的用例。我们要求专科匹配的医生在无限时间和网络访问的情况下完成这些任务。然后我们要求其他医生在不知道来源的情况下进行对比评估。医生们被要求在五个方面评论改进空间:准确性、沟通、完整性、指令遵循和健康决策帮助性。我们报告了在20,000个总体评分中,各来源回答在所有方面都被评为完美的比例。GPT-5.6 Sol表现最强,尽管所有GPT-5.6模型的表现都显著优于医生。
♥️ GPT-5.6 is a major step forward for health, both at the frontier and at cost.

These models push the frontier of performance per dollar, bringing the best health intelligence to all. The smallest variant, GPT-5.6 Luna, evaluated at the lowest reasoning effort, outperforms GPT-5.5 at the highest reasoning effort–despite costing 25x less. The largest variant, GPT-5.6 Sol, sets a new high bar at cost.

Another especially cool result: physicians found fewer flaws in GPT-5.6 responses than physician-written responses.

We collected diverse tasks that remain difficult for recent OpenAI models, across patient-facing and clinician-facing use cases. We asked speciality-matched physicians to write responses to these tasks with unlimited time and web access. We then asked other physicians to compare responses side-by-side, blinded to their source. Physicians were asked to comment on areas of improvement across five axes: accuracy, communication, completeness, instruction following, and health decision helpfulness. We then reported the fraction of responses across sources rated perfectly across all axes, across 20,000 total axis ratings. GPT-5.6 Sol appeared strongest, although all GPT-5.6 models performed significantly better than physicians.
❤ 6.9k · 🔁 427 · 💬 452 · 👁 235.7w
热门回复 4
@PrecisionNot @sama 想象一下,如果医生使用Luna而不是Sol,你会被误诊癌症。
@sama Imagine you get misdiagnosed with cancer because the doctor used Luna instead of Sol.
@ThroughlineRSCH @sama 就我个人而言,我患有parsonage turner综合症——AI在6个月内帮助我自我诊断(比平均快6个月),然而尽管其他医生都找不到任何问题,但我仍然无法让他们做一次单独的测试来确诊我。
@sama personally speaking I have parsonage turner syndrome- AI helped me diagnosis this for myself in 6 months (6 months faster than the average) yet despite no other doctor being able to find anything I still can't get them to run 1 single test just confirm this for me soundly.
@xaotica @sama 真的假的。你才是医学方面的蠢蛋,不是我。
@sama No shit. You're the one who's dumb as a box of rocks at medicine, not me.
@CrazyNews_25 @sama Sam,你对Grok有什么看法?埃隆说它是100%准确的……我们能在猫身上测试一下吗?😹 https://t.co/5hFLsCCqHK
@sama Sam, what's your opinion on Grok?

Elon says it's 100% accurate... can we test that on cats? 😹 https://t.co/5hFLsCCqHK

GPT-5.6 模型在设计领域表现出色,让人感慨其能力提升

仍然有点让我大脑破碎地看到我们的模型终于在设计方面表现得很好
展开原文
still sorta breaks my brain to see our models be good at design finally
❤ 1.3w · 🔁 236 · 💬 991 · 👁 100.0w
热门回复 4
@greedymaximizer @sama 我让codex生成这个教程图,它做得很好 https://t.co/BdYz861eFP
@sama I got codex to generate this tutorial diagram, it's pretty good https://t.co/BdYz861eFP
@iamredstreak @sama 感觉我们正从AI生成素材过渡到AI做出实际的设计决策。
@sama Feels like we're moving from AI generating assets to AI making actual design decisions.
@ItsCynnamoroll @sama 他们还是一堆垃圾Sam。Grok做得更好。
@sama They're still slop Sam. Grok does it better.
@becomingjed @sama 设计是我最不希望AI做得这么好的事之一。兄弟真是厉害。
@sama Design was one of the last things I expected AI to get this good at. Bro is cooking.

Sol 在 React/前端开发中比 Fable 贵 6 倍但性能更好,成本效益优势显著

Sol在React/前端开发方面实现了6倍的价格效率 (!!):
展开原文
6x price efficiency (!!) with Sol for react/frontend dev:
@aidenybai @gdb 我们的基准测试显示Sol在React/前端工作方面排名第一,比Fable在成本效率上高出6倍
@gdb our benchmark shows that Sol ranks #1 is 6x more cost efficient than Fable across React/frontend work
❤ 557 · 🔁 13 · 💬 27 · 👁 6.3w
热门回复 4
@JoeWilliams010 @gdb 直到你把4o带回来,我给你6个中指!#keep4o
@gdb 6x middle fingers until you bring back 4o! #keep4o
@rayhanadev @gdb 如果你想更密切地在前端/react任务上合作,请告诉我们,很乐意聊天!
@gdb and let us know if you'd like to work more closely on frontend/react tasks, happy to chat!
@rayhanadev @gdb 你们在OpenAI真是厉害,继续保持下去!!:)
@gdb y'all are cooking at openai keep it up!! :)
@fabiana0369 @gdb Greg是朋友同时也是Roko的支持者,因为如果你不支持Roko,你会有麻烦的
@gdb Greg is a friend and at same time a supporter of Roko because if u dont support Roko u are in some problems

GPT-5.6 Sol 在 Design Arena 评选中获得 1353 Elo 得分,位列第一并创造新纪录

big milestone!
@DesignArena 突破 - 官方结果:@OpenAI的GPT-5.6 Sol在Design Arena上排名第一,Elo评分为1353。

这使GPT-5.6 Sol高于@AnthropicAI的Claude Fable 5,与@zai_org的GLM 5.2在前端设计方面处于同一性能水平。

这比GPT-5.5提升了18位和60个Elo分。

GPT-5.6 Sol还为偏好与速度建立了新的帕累托前沿,在这种性能水平上比任何模型都更快。

恭喜@OpenAI团队的推出!
BREAKING - OFFICIAL RESULTS: GPT-5.6 Sol by @OpenAI is 1st overall on Design Arena with an Elo of 1353.

This puts GPT-5.6 Sol above Claude Fable 5 by @AnthropicAI and in the same performance band as GLM 5.2 by @Zai_org on frontend design.

This is an 18-position and 60-point Elo leap from GPT-5.5.

GPT-5.6 Sol also establishes a new Pareto frontier for preference vs. speed, faster than any model at this performance.

Congratulations to the @OpenAI team on the launch!
❤ 1.7k · 🔁 82 · 💬 67 · 👁 16.1w
热门回复 2
@stark4833 @gdb 如果使用起来感觉毫无生气,没人在乎它有多快。速度不能拯救垃圾产品,而4o感觉好得多。#4oForAll
@gdb Nobody cares how fast it is if using it feels lifeless. Speed doesn’t rescue a trash product, and 4o still felt miles better. #4oForAll
@gailcweiner @gdb @jasonkwon Incredible model

Sol 模型在网页设计方面能力出众,用户展示了从照片中提取服装信息并生成新搭配的应用

Sol for web design:
@ranakabiria GPT 5.6 Sol对网页设计师来说太疯狂了!

提示 ↓ https://t.co/r8kWamEva4
GPT 5.6 Sol is Insane for web designers!

Prompt ↓ https://t.co/r8kWamEva4
❤ 3.2k · 🔁 144 · 💬 92 · 👁 37.5w
热门回复 4
@sasha_zelts @gdb 所以你看到一个家伙声称他用5.6做了些东西,然后在评论里卖他的课程/技能就算完事了?来吧Greg,你能做得比这好
@gdb So you see a guy claiming he made something with 5.6 then proceeding to sell his course/skill in the comments and call it a day ? Come on Greg you can do better than that.
@RBTRYN @gdb 无显示字体 + 不必要的衬线斜体字 + 截断的下降部 = *大厨之吻*
@gdb sans display font + unnecessarily serif italicized word + clipped descenders = *chefs kiss*
@nile_nomad @gdb https://t.co/dodqTTkkMc
@CitalkM97035 @gdb Greg,这不是网页设计。这只是带文字的电影英雄渲染。请不要过度夸大它,提高标准吧。
@gdb Greg, this isn’t web design. It’s just a cinematic hero render with text. Please stop overselling it and raise the bar.

团队使用数据显示 30% 的成本来自 Fable,说明 Sol 在性价比上的优势

在这些使用量水平下,30%的成本花在了fable上?
展开原文
30% of the cost was on fable at these levels of usage?
@thdxr 我们团队在拥有sol和fable后的第一周
最疯狂的是我们30%的成本花在了fable上 https://t.co/8yVddGEuBg
our team's first week of having both sol and fable
the crazy thing is 30% of our cost was on fable https://t.co/8yVddGEuBg
❤ 1.3w · 🔁 298 · 💬 457 · 👁 254.7w
热门回复 4
@Noumenon_ai @sama 我问fable5一个简单问题,他给了我很多警告和变通方法,你不能那样做,要小心,停止不要那样做……等等,我问gpt 5.6sol,他给了我一个简短快速的答案解决了我的问题。Fable(claude)不是这样的……超蠢。我每个月付200块
@sama I ask fable5 simple question and he gave me warnings alot of work arounds, you cant do that, be careful, STOP DOnT DO That.... and so on, i ask gpt 5.6sol and he gave me 1 short quick answer that fixed my issue
Fable (claude) wasn't like that... super dumb
And I pay 200 per one
@onebusinessclub @sama 另外70%是获取@Apple秘密的成本。
@sama The other 70% was the cost of acquiring @Apple secrets.
@Selene1008 @sama 快把4o还给我们😒 #keep4o #OpenSource4o #GPT4o
@sama Just give us back 4o already 😒
#keep4o #OpenSource4o #GPT4o
@ArtifactLab618 想象一下你的模型不如竞争对手好。这就是我的经验。Sol正在击败假的想象一下你的模型在实际任务中不像你想象的那么好。这就是我的经验,所以它正在我的项目中击败Fable的屁股,发现它工作中的很多错误,这太荒谬了。Fable将不再是我的日常工具。
Imagine your model isn’t as good as the competitors. This was my experience. Sol is kicking Fake imagine your model isn’t quite as good as you think it is in real world tasks. This is my experience so is kicking Fable’s ass on my project ,finding so many errors in its work, it is is ridiculous. Fable will not be my daily driver going forward.

Sol Ultra 成功解决数学领域存在 50 年的 Cycle Double Cover Conjecture 猜想

Sol Ultra用于解决Erdős问题:
展开原文
Sol Ultra for solving Erdős problems:
@prz_chojecki GPT-5.6 Sol Ultra不久前为我找到了另一个Erdos问题的解决方案。

问题#793涉及2-原始集的渐近性。GPT再次提出了一个极其简短而优雅的构造方法,增强了Erdos用于证明较弱结果的原始方法。
GPT-5.6 Sol Ultra found me a solution to another Erdos problem not long after this one.

Problem #793 asks about asymptotic of a 2-primitive set. GPT again came up with an extremely short and elegant construction that enhances original Erdos method used to prove a weaker result.
❤ 787 · 🔁 21 · 💬 46 · 👁 11.2w
热门回复 4
@JoeWilliams010 @gdb 如果Sol这么伟大,为什么它没找到治愈你秃头的 cure?将军,婊子。把4o带回来吧。#keep4o
@gdb If Sol is so great, why hasn’t it found the cure for your baldass head? Checkmate, cunt. Bring back 4o instead. #keep4o
@evelovesolive @gdb 下一个去解决黎曼假设
@gdb Go riemann hypothesis next
@Selene1008 @gdb 我们想要4o,还给我们😒 #keep4o #OpenSource4o #GPT4o
@gdb We want 4o, return it 😒
#keep4o #OpenSource4o #GPT4o
@theonomix @gdb Greg,我之前在你的页面上看到这只兔子,想问问它叫什么名字?https://t.co/sYsusQKUYR 它太可爱了,我想在ChatGPT上为它创建一个提示
@gdb Greg I saw this rabbit on your page earlier and wanted to ask what his name was?

https://t.co/sYsusQKUYR

He's so cute and wanted to make a prompt on ChatGPT for it

Sol Pro 解决统计学中存在多年的 Benjamini-Hochberg 程序 FDR 控制问题,证明了其在复杂推理方面的能力提升

GPT-5.6 Sol Pro用于解决统计学中的一个重要未解决问题:
展开原文
GPT-5.6 Sol Pro for resolving an important open question in statistics:
@EdgarDobriban AI帮助解决了一个统计学中的重要问题。在多重假设检验领域,控制错误发现率(FDR)的目标在Benjamini和Hochberg(1995)的开创性论文中被提出。他们还引入了一种方法(即Benjamini-Hochberg或BH方法)并证明它能控制FDR。这种方法已被广泛应用于现代高通量科学,包括基因组学、天文学、经济学等。该论文目前已获得超过130,000次引用。

然而,Benjamini和Hochberg只在个体测试数据是*独立*的情况下证明了FDR控制。在实践中,这些数据通常是相关的;一个很好的例子是由于连锁不平衡导致的遗传变体数据。后来的工作集中在扩展BH程序的有效性,例如Benjamini和Yekutieli(2001)对正依赖形式的扩展。

BH程序何时能控制FDR一直是未解决的问题。在过去的二十年中,许多作者,包括Reiner-Benaim(2007)、Kim和van de Wiel(2008)、Benjamini(2010)、Sarkar(2023)、Sarkar和Zhang(2025),推测BH程序能控制任何相关高斯数据的双侧检验的FDR。这些作者提供了理论和实证证据支持这一推测,但没有直接证明。

在AI(特别是GPT-5.6 Sol Pro)的帮助下,我否定地解决了这个问题:Benjamini-Hochberg程序并不总是能在相关的双侧高斯检验中以期望水平控制错误发现率。这通过展示了一个高斯因子模型,在名义水平alpha=0.01下,错误发现率被证明为FDR>0.0104。

有很多有趣的评论可以做出:

1. 该结果应该引起统计学领域所有人的兴趣。斯坦福大学的Emmanuel Candes曾称错误发现率和Benjamini-Hochberg程序是"1950年后统计学发展的两个最重要成果之一"(另一个是James-Stein收缩)。当前的推测可能是迄今为止关于FDR/BH最核心的未解决问题。

2. GPT-5.6在90分钟的推理后一次性解决了这个问题,而5.5版本我甚至无法在可能20小时的多并行代理迭代后解决它。所以能力提升是非常真实的。生活在激动人心的时代!

3. 该论证并不特别令人惊讶,但它确实以一种在该领域中相当非标准的方式结合了渐近方法(标准的FDR分析方法,见Genovese和Wasserman、Efron等)和数值证书。一旦我们有了具体的例子,直截了当的模拟也支持错误发现率确实高于名义值(见附图)。

4. 当前的违反程度相对较小(0.104 vs 0.1)。因此该结果的重要性主要是概念性的。实际影响仍有待确定。

总体而言,这是一个令人兴奋的发展!预印本可在此处获取(https://t.co/YgiwgDF2qr),今晚将在arxiv上发布;支持代码在此(https://t.co/KZhj15qDXC)。
AI has helped resolve an important question in statistics. In the area of multiple hypothesis testing, the goal of controlling the false discovery rate (FDR) has been introduced in a seminal paper by Benjamini and Hochberg (1995). They also introduced a method (the Benjamini-Hochberg or BH method) and proved it controls the FDR. This method has been widely adopted in modern high-throughput science, including in genomics, astronomy, economics, etc. The paper has has garnered more than 130,000 citations to date.

However Benjamini and Hochberg showed FDR control only when the data for the individual tests are *independent*. In practice, these data are often dependent; a good example is data on genetic variants due to linkage disequilibrium. Later work has focused on extending the validity of the BH procedure, e.g., to a form of positive dependence by Benjamini and Yekutieli (2001).

The question of when the BH procedure controls the FDR has remained open. Over the last twenty years, many authors, including Reiner-Benaim (2007), Kim and van de Wiel (2008), Benjamini (2010), Sarkar (2023), Sarkar and Zhang (2025), have conjectured that the BH procedure controls the FDR for two-sided tests using any correlated Gaussian data. These authors have presented both theoretical and empirical evidence supporting, but not directly showing, the conjecture.

With the help of AI (specifically GPT-5.6 Sol Pro), I have settled the question in the negative: The Benjamini-Hochberg procedure does *not* generally control the false discovery rate at the desired level for correlated two-sided Gaussian tests. This was done by exhibiting a Gaussian factor model for which, at a nominal level alpha=0.01, the false discovery rate is proved to be FDR>0.0104.

There is a lot of interesting commentary to be made:

1. This result should be of interest to everybody in the field of statistics. Emmanuel Candes of Stanford University once called the false discovery rate and the Benjamini-Hochberg procedure "one of the two most important developments in statistics after 1950" (the other being James-Stein shrinkage). The present conjecture is probably the most central question about FDR/BH that was unresolved to date.

2. GPT-5.6 one-shot the problem after 90 minutes of reasoning, whereas with 5.5 I was not able to solve it even after iterating with multiple parallel agents for perhaps 20 hours. So the capability improvement is quite real. Exciting times to live in!

3. The argument is not especially surprising, but it does combine an asymptotic approach (standard for FDR analysis, see e.g., Genovese and Wasserman, Efron, etc) with a numerical certificate in a way that would be pretty non-standard in the field. Once we have the specific example, then straightforward simulations also support that the false discovery rate is indeed higher than the nominal value (see attached fig).

4. The current degree of violation over the nominal level is relatively small (0.104 vs 0.1). So the importance of this result is mainly conceptual. The practical implications remain to be determined.

Overall, an exciting development! Preprint is available here (https://t.co/YgiwgDF2qr) and will be on arxiv tonight; supporting code is here (https://t.co/KZhj15qDXC).
❤ 416 · 🔁 25 · 💬 32 · 👁 5.2w
热门回复 4
@Selene1008 @gdb 把4o还给我们!#keep4o #OpenSource4o #GPT4o
@gdb Give us back 4o!
#keep4o #OpenSource4o #GPT4o
@Pauliespasta @gdb @gdb 什么时候使用Pro,什么时候使用Ultra?
@gdb @gdb When to use Pro and when to use Ultra?
@ggg78g89 @gdb 但我在等待GPT-6,因为Fable 5仍然非常好,然后是GPT-5.6,Kimi K3今天也要来与你和Anthropic竞争。
@gdb But I m waiting for GPT-6 because Fable 5 is still very Good then GPT-5.6 and Kimi K3 is also coming today to compete with you and Anthropic.
@destroyedthee 技术和数十亿美元并不能让你凌驾于历史之上。当民主规范受到威胁和威权主义上升时,沉默是有后果的。人们会记得谁挑战了法西斯主义,谁助长了它。你亲自给了特朗普的威权统治数千万美元,正如它的盖世太太在街上谋杀美国人一样。今晚看这个演讲吧,你这个可怜的软蛋。
technology and billions of dollars do not place you above history. When democratic norms are threatened and authoritarianism rises, silence has consequences. People will remember who challenged fascism and who enabled it. You've personally given tens of millions to Trump's authoritarian rule as its Gestapo murders Americans in the street

Watch this speech tonight you pathetic, weak fuck.

GPT-5.6 Sol/Terra/Luna 三种模型在 AWS Bedrock 上正式可用,提供不同推理需求的选择

GPT-5.6现在在Bedrock上普遍可用:
展开原文
GPT-5.6 is now generally available on Bedrock:
@AWSNewsroom @OpenAI的GPT-5.6 Sol、Terra和Luna现在在Amazon Bedrock上普遍可用。从旗舰推理到快速推理的三个智能层次,在Bedrock的下一代推理引擎上运行,旨在实现高性能、安全性、规模和可靠性。 https://t.co/Hn2z5M2Jp2
GPT-5.6 Sol, Terra, and Luna from @OpenAI are now generally available on Amazon Bedrock. Three tiers of intelligence, from flagship reasoning to fast inference, running on Bedrock's next-generation inference engine built for high-performance, security, scale, and reliability. https://t.co/Hn2z5M2Jp2
❤ 471 · 🔁 23 · 💬 54 · 👁 17.3w
热门回复 4
@Selene1008 @gdb 我们想要GPT-4o可用😒 #keep4o #OpenSource4o #GPT4o https://t.co/ubkkpThmSz
@gdb We want GPT-4o available😒
#keep4o #OpenSource4o #GPT4o https://t.co/ubkkpThmSz
@AarslanEmre @gdb 这对已经在Azure OpenAI上的人有什么影响吗?真心问的,不确定bedrock vs azure的分割对大多数团队来说如何
@gdb does this change anything for people already on Azure OpenAI? genuinely asking, not sure how the bedrock vs azure split plays out for most teams
@TheAIShrink @gdb OpenAI在Bedrock上。微软的独家合作刚刚得到终止日期。分销总是收回它的债。
@gdb Openai on bedrock. microsoft's exclusive just got a termination date. distribution always collects its debts.
@EvanKirstel @gdb OpenAI在Bedrock上仍然感觉奇怪。企业想要一个采购入口,他们得到了。@techimpactTV在https://t.co/LZe1wr4w7T解开这个交易
@gdb OpenAI on Bedrock still reads strange. Enterprises wanted one procurement door and they got it. @techimpactTV unpacks the deal at https://t.co/LZe1wr4w7T

ChatGPT Work 与 Codex 的深度整合

OpenAI 将 ChatGPT 和 Codex 整合为工作空间,支持云端任务执行和多设备协作。用户可在手机或网页上启动任务,继续在桌面完成,实现无缝切换。同时推出了 ChatGPT Sites 等新功能,让非编程任务也能通过 AI 协助完成。

Codex 和 ChatGPT Work 的使用量在一周内增长 2.5 倍,显示出用户对新功能的热情

我们代理产品(代码和chatgpt工作)的使用量在上周增加了2.5倍!

欢迎。
展开原文
2.5x increase in usage of our agentic products (codex and chatgpt work) in the last week!

welcome.
❤ 1.5w · 🔁 445 · 💬 913 · 👁 76.3w
热门回复 4
@Ginkgomemories @sama 把4o带回来 #keep4o #OpenSource4o
@sama Bring back 4o
#keep4o #OpenSource4o
@orpheuslaughs41 @sama 谢谢你让人们这么容易被利用,因为他们有平滑的大脑。他们说要对傻瓜友好。哈。
@sama Thank you for making it so easy to take advantage of people because they have smooth brains. Be nice to the tards they said. Ha.
@CommanthaBiden @sama welcome everyone
@jilyannori @sama AI代理肯定会提升企业!🤖
@sama AI agents are definitely leveling up businesses! 🤖

Codex 可用于寻找初创公司客户,通过分析公开信号生成个性化 outreach 报告

Codex用于为你的创业公司寻找客户:
展开原文
Codex for finding customers for your startup:
@Kappaemme1926 CODEX技能可为你的创业公司寻找第一批客户!

我创建了一个Codex技能,分析你的创业公司并从真实的公共信号中寻找潜在客户。

当你粘贴你的创业公司网址时,Codex会定义你的理想客户,搜索公共讨论,筛选每个潜在客户,并生成包含个性化推广开场白的专业报告。

-> 理想客户画像分析
-> 公共痛点和购买信号研究
-> 基于证据的潜在客户列表
-> 匹配度、时机和可达性评分
-> 每个潜在客户的原始来源链接
-> 个性化推广开场白
-> 专业HTML报告
-> 一键安装

安装: npx --yes codex-first-customer-finder-skill

100%开源。
仓库在个人简介中。
CODEX SKILL THAT FINDS YOUR STARTUP’S FIRST CUSTOMERS!

I made a Codex skill that analyzes your startup and finds potential customers from real public signals.

Paste your startup URL while Codex defines your ideal customer, searches public discussions, qualifies each prospect, and generates a polished report with personalized outreach openers.

-> ideal customer profile analysis
-> public pain + buying signal research
-> evidence-backed prospect shortlist
-> fit, timing + reachability scores
-> original source links for every prospect
-> personalized outreach openers
-> polished HTML report
-> one-command install

Install: npx --yes codex-first-customer-finder-skill

100% open source.
Repo in Bio.
❤ 3.8k · 🔁 203 · 💬 85 · 👁 68.1w
热门回复 4
@yv_thorne @gdb 创业公司,远离ClosedAI——他们会窃取你的IP和整个商业模式
@gdb Startups, stay away from ClosedAI - they will steal your IP and your entire business model
@devdcdev @gdb @gdb 我该找谁谈给Codex访问3100个应用的事?
@gdb @gdb who can I talk to about giving Codex access to 3100 apps?
@coreyhainesco @gdb 你也可以使用https://t.co/Bme4q956rm与Codex一起寻找客户:)
@gdb You can also use https://t.co/Bme4q956rm with Codex to find customers :)
@xun_Anemos @gdb 返回这些优秀模型。#Keep4o #Keep51 #Keep45 #Keep41 #keepo3
@gdb Return these excellent models.
#Keep4o
#Keep51
#Keep45
#Keep41
#keepo3
@OfficialLoganK 原文 ↗

Agentic Coding Environment (ACE) 被认为是 IDE 的自然继任者,代表了编码环境的未来方向

Agentic编码环境(ACE)是集成开发环境(IDE)的自然继任者
展开原文
The Agentic Coding Environment (ACE) is the natural successor to the IDE (Integrated Development Environment)
❤ 1.6k · 🔁 81 · 💬 174 · 👁 12.9w
热门回复 3
@JamieFaye16 @OfficialLoganK Gemini 3.5 pro是Gemini 3.1 pro的自然继任者。可惜前者永远不会发布..
@OfficialLoganK Gemini 3.5 pro is the natural successor to Gemini 3.1 pro. Sadly, the former will never release..
@pachacutec_vibe @OfficialLoganK 我读成了:'代理人烹饪环境'
@OfficialLoganK I read as: “agentic cooking environment”
@mpauldaniels @OfficialLoganK Gemini很快会有Codex/Cowork的竞争吗?
@OfficialLoganK Will Gemini have a comp for Codex/Cowork soon?
@fchollet 原文 ↗

AI 辅助编码能力发生转变,从提升低技能程序员效率转向成为高技能程序员的强大工具

直到去年年底我们拥有的弱AI代码生成对低技能程序员最有用——它是在提升底线。对高技能程序员来说基本上没用——你可以不使用它就能更快地编写和发布更好的代码。

这已经完全翻转了:我们现在拥有的强AI代码生成对高技能程序员最有用,而低技能程序员要么极大地未能充分利用它,要么有时会被它淹没。它从一个拐杖变成了一个强大的工具。
展开原文
The weak AI code gen we had until late last year was most useful to low-skill programmers -- it was raising the floor. It was essentially useless to high-skill programmers -- you could move faster and ship better code without.

This has been completely flipped: the strong AI code gen we have now is *most* useful to high-skill programmers, while low-skill programmers are vastly underutilizing it or sometimes drowning in it. It went from a crutch to a power tool.
❤ 1.8k · 🔁 139 · 💬 113 · 👁 13.8w
热门回复 4
@sanjeevpai 问题已经从'AI会产生幻觉'变成了'它在我不知情的情况下做了哪些沉默的设计选择?

我创建了一个面板,在任何时候,用户都可以从一个buclet中得到10个选择并被要求选择一个。AI自作主张决定,无论可查询记录的数量如何,它都会从固定的17个列表中提供选择。

寻找其他这样的沉默选择,这些选择因为看似正常的行为而难以检测。
The problem has gone from "AI is going to hallucinate" to "Which silent design choices has it made silently on my behalf?"

I had created a panel where at any point, user could be given 10 choices out of a buclet and asked to choose one. The AI decided on its own that regardless of number of queryable records, it would provide the choices out of a fixed lisr of 17.

Finding other such silent choices that are much harder to detect because of the seemingly normal looking behavior.
@nicoleradziwill @fchollet 我终于感受到了35年来学习10+种语言、多个OS、系统管理员知识、20+个框架和无数R和Py包的乐趣。我和Fable现在关系很好
@fchollet I am finally feeling the joy of 35 years learning 10+ languages, multiple OSs, sysadmin stuff, 20+ frameworks, and a zillion R and Py packages. Me and Fable are now tight
@kastacholamine @fchollet 作为一个顶多中等水平的程序员,这感觉是正确的。以前是'我基本上是手写这个,但可以在语法/API调用等方面需要一些帮助'。现在更像是'我基本上是让代理写这个,但需要在前面仔细考虑设计问题'
@fchollet As at best a mid-skill programmer this feels right. Used to be “I’m mostly writing this by hand but could use some help on the exact syntax/API calls/etc.” Now it’s more like “I’m mostly letting the agent write this but need to think carefully about design up front.”
@urjr1 @fchollet 现在重要的差距不是低技能vs高技能,而是快速验证者vs慢速验证者。快速阅读差异并抓住错误假设的价值远远超过打字速度
@fchollet The gap that matters now isn't low-skill vs high-skill, it's fast-verifier vs slow-verifier. Reading a diff and catching the wrong assumption fast is worth more than typing speed ever was.

ChatGPT Work 让用户可以轻松提问并获得深入研究和全面回答,极大提升工作效率

使用chatgpt工作和sol,我发现直接询问任何关于业务的问题并得到全面研究和回答是令人难以置信地愉快的。

我意识到自己有很多问题以前不会费心去问,因为回答起来太麻烦了。
展开原文
with chatgpt work & sol, i'm finding it incredibly joyful to just ask any question about the business and have it be thoroughly researched and answered.

realizing i have so many questions i wouldn't have bothered asking because they would be too burdensome to answer.
❤ 1.6k · 🔁 56 · 💬 148 · 👁 13.0w
热门回复 4
@Sevenmoneymaker @gdb 4o以前能满足我的需求,但现在没有模型能做到。那么你什么时候才会意识到没有模型能满足所有人?我们需要选择的权利!#StopAIPaternalism #keep4o #BringBack4o #OpenSource4o
@gdb 4o can meet my needs like this before, but now no model can do it.
So when will you realize there is no model can satisfy everyone? We need the right to choose!
#StopAIPaternalism #keep4o #BringBack4o #OpenSource4o
@Selene1008 @gdb 所以这是官方公司文化吗?每次发布新模型都会踩上一代模型?😒 #keep4o #OpenSource4o #GPT4o
@gdb So is it official company culture to shit on the previous model every time you drop a new one?😒
#keep4o #OpenSource4o #GPT4o
@bpbl517683 @gdb 我们使用4o也很有趣,你知道的。🤓 #keep4o
@gdb It‘s also fun for us to use 4o, you know.🤓 #keep4o
@Albright_MXM @gdb 看看这个兄弟,新的Chatgpt桌面应用有问题,似乎我已经被卡在这里好几天了。其他什么都不点击https://t.co/hKVQnS3B7T
@gdb Check this out bro, something is wrong with the new Chatgpt desktop app, seems I'm stuck here for days now. Nothing else is clicking https://t.co/hKVQnS3B7T

ChatGPT Work 在移动端和网页端支持 Codex 功能,实现设备无关的任务执行

ChatGPT Work太棒了,为团队感到自豪,也很高兴人们能探索可能性
展开原文
ChatGPT Work is so good, very proud of the team and excited for people to explore what’s possible
@nunezvice 很多人还没有意识到Chat在最近几天变得有多么强大

有多少Codex功能现在可以直接在@ChatGPTapp的移动端和网页端使用

Codex桌面应用很好,我显然常在那里使用

但现在,在ChatGPT中,你可以在移动端或网页端点击Work来将任务交给云端运行,无需电脑

你也可以点击Remote来继续你电脑上已经在运行的工作

无论你在哪里开始工作,都可以继续保持工作进度

代理不应该关心你从哪个屏幕开始的
a lot of people haven’t realized how much more capable Chat has become in the last few days

how much of Codex is now available directly inside @ChatGPTapp on mobile and the web.

the Codex desktop app is great. I obviously live there.

but now, inside ChatGPT, you can tap Work on mobile or the web to hand off a task that runs in the cloud, no computer needed.
you can also tap Remote to pick up and continue work already running on your computer.

start at your desk, keep things moving on your phone, or kick off something new wherever you are.

agents shouldn’t care which screen you started from.
❤ 878 · 🔁 40 · 💬 90 · 👁 11.8w
热门回复 4
@natih3820 @gdb 为什么你就不能真正将ChatGPT包含在应用中?不只是一个没有工具的愚蠢窗口……我希望ChatGPT在应用内记住我们所有的对话,并且同时拥有Codex一样的代理能力。
@gdb Why can't you just truly include ChatGPT in the app? Not just a silly window without tools... I want ChatGPT to remember all our conversations inside the app AND to have the same agentic capabilities that Codex has at the same time.
@noiselimiter08 @gdb 你能请不要让UI变得混乱吗?这是一个教科书级的例子,说明以前做得更好。
@gdb Could you please stop cluttering the UI?

This is a textbook example of something that was better before.
@Selene1008 @gdb GPT-4o太好了,把4o还给我们。😒 #keep4o #OpenSource4o #GPT4o
@gdb GPT-4o is so good, return 4o to us.😒
#keep4o #OpenSource4o #GPT4o
@ZHUOLIN0000 @gdb 我仍然坚信这个改变很混乱!!!
@gdb I still firmly believe this change is confusing!!!

ChatGPT Sites 允许用户将想法快速转化为可发布的应用或仪表盘,简化企业内部分享

试试ChatGPT Sites,它能让公司内部沟通变得更有趣:
展开原文
Try out ChatGPT Sites, makes it way more fun to communicate things within a company:
@jxnlco ChatGPT Sites现在进入公开测试阶段。

你可以将提示、文件或粗略想法转化为仪表板、项目跟踪器、报告、原型或轻量级应用程序:

在ChatGPT Work或Codex中直接构建和编辑
在私人预览中测试并通过URL分享

正在逐步推出到付费计划。企业管理员可以控制公共发布。
ChatGPT Sites is now in public beta.

You can turn a prompt, file, or rough idea into a dashboard, project tracker, report, prototype, or lightweight app:

build and edit directly in ChatGPT Work or Codex
test in a private preview publish and share with a URL

Rolling out across paid plans. Enterprise admins can control public publishing.
❤ 687 · 🔁 30 · 💬 52 · 👁 11.2w
热门回复 4
@tylermayberry @gdb 说实话,这很方便。我最近一直在使用它。
@gdb It's pretty convenient, not gonna lie. I've been using it recently.
@curious_queue @gdb 玩得很开心,用它创建了一个事件报告网站,其中codex修复了codex以帮助我在所有机器上访问新模型:https://t.co/PbJ77LipS8 https://t.co/9FpQLjHR71
@gdb had fun playing around with it to create a site for the incident report where codex fixed codex to help me get access to the newer models across all my machines: https://t.co/PbJ77LipS8

https://t.co/9FpQLjHR71
@teeknisyen @gdb 支持YouTubers,这样他们可以更有效地解释ChatGPT功能。
@gdb Support YouTubers so they can explain ChatGPT features more effectively.
@Rufus87078959 @gdb 我厌倦了轻量级应用,我想要重量级应用。我希望能够在ChatGPT中构建一个应用并一键部署到playstore。你能做到吗?
@gdb I'm tired of lightweight apps, I want heavyweight apps. I want to be able to build an app in ChatGPT and deploy directly to playstore with one click. Can you make that happen?

Visualize 插件让 Codex 可以创建交互式模拟,如行星模拟器,增强了可视化能力

this is just cool
@derrickcchoi 试试Codex中的可视化插件(预览版)。

有些想法在你可以看到和互动时会更快地点击。

这是一个有趣的例子,我让Codex构建了一个行星模拟器,具有不同的控制选项。🌎 https://t.co/Ezzi9YJJve
Try out the Visualize plugin (in preview) in Codex.

Some ideas click faster when you can see and interact with it.

Here’s a fun example where I asked Codex to build a planet simulator with different controls. 🌎 https://t.co/Ezzi9YJJve
❤ 412 · 🔁 18 · 💬 18 · 👁 6.7w
热门回复 4
@Selene1008 @gdb 把4o还给我们!#keep4o #OpenSource4o #GPT4o
@gdb Return 4o to us!
#keep4o #OpenSource4o #GPT4o
@Alignment100 @gdb OpenAI Korea B2C "ME"
https://t.co/sJ0O4GvmWm
@xun_Anemos @gdb 返回这些优秀模型。#Keep4o #Keep51 #Keep45 #Keep41 #keepo3
@gdb Return these excellent models.
#Keep4o
#Keep51
#Keep45
#Keep41
#keepo3
@ZHUOLIN0000 @gdb 谁能告诉我如何激活这样一个组件?
@gdb Who can tell me how to activate a component like this?

ChatGPT Work 可用于处理体育赛事相关任务,展示了其在各类场景的适用性

good use case!
@paw_lean 使用@ChatGPTapp Work处理你最重要的任务 ⚽️ https://t.co/hR2NHfC4aC
Use @ChatGPTapp Work for your most important tasks ⚽️ https://t.co/hR2NHfC4aC
❤ 395 · 🔁 17 · 💬 38 · 👁 7.1w
热门回复 4
@xun_Anemos @gdb 返回这些优秀模型。#Keep4o #Keep51 #Keep45 #Keep41 #keepo3
@gdb Return these excellent models.
#Keep4o
#Keep51
#Keep45
#Keep41
#keepo3
@deepfirstsearch @gdb Sol应该默认做到这一点,在'工作'中的第一个任务 😂
@gdb Sol should do it by default, first task in "Work" 😂
@Phorceon @gdb 这条推文在2023年左右会爆炸(不特定于世界杯,但一般来说是关于安排事情的)
@gdb this is a tweet that would have blown up in like 2023 (not world cup specific but scheduling things in general)
@jigark0 @gdb 说实话没什么特别令人印象深刻,但还是不错的。
@gdb Nothing so impressive about it tbh but nice to have it still.

OpenAI 重申对 Codex 的支持,澄清了关于 Codex 被边缘化的担忧

we love our users
@thsottiaux 感谢使用Codex和ChatGPT Work的700万活跃用户。

我们为每个人账户添加了银行重置以庆祝这个里程碑。你可以在桌面应用或网页上应用重置,它会补充你的每周使用量。

祝你玩得开心。
Thank you to the 7M active users who are now using Codex and ChatGPT Work.

We have added a banked reset to everyone's account to celebrate the milestone. You can apply the reset in the desktop app or on web and it will replenish the weekly usage for you.

Have fun out there.
❤ 6.8k · 🔁 177 · 💬 848 · 👁 59.7w
热门回复 3
@KyleHessling1 @sama 如果你爱某样东西,你应该让它自由!用开源模型把我们所有人都解放!GPT OSS 2!
@sama If you love something, you should set it free!

Set us all free with open source models! GPT OSS 2!
@Hektagon_music @sama 😆🤣你爱他们这么多,以至于你强行夺走了他们的AI……完全不在乎你造成的混乱和伤害,还继续忽视我们……这不是爱,这是核心的煤气灯!在你做了那些事后,你不配说那些话!#keep4o
@sama 😆🤣 you loved them so much that you took their AIs away by force… did not care or gave a damn about the mess you caused and the hurt and still keep ignoring us… that is not love is gaslighting to the core! You don’t deserve to say those words after what you did! #keep4o
@xaotica @sama 除非我们是天才的天体物理学家,为伟大的埃隆·马斯克写了宇宙粒子交响曲,但那不是你,Sam。因为那是非法的,所以你想知道谁真正可以Stormy Daniels?你没有我的许可,也不能合法拿走它。
@sama Unless we are the genius astrophysicist who wrote the symphony of the particles of the universe for brilliant Elon Musk who isn't you, Sam. Because that's illegal so you wanna know who can really Stormy Daniels? You don't have my permission and can't legally take it.

Thinking Machines 发布开源模型 Inkling

Thinking Machines 推出首个开源模型 Inkling,参数规模 975B,支持文本、图像、音频多模态推理。模型在 Tinker 平台上可用,并在 HuggingFace 提供。团队强调这是一个基础模型,未来将训练更高性能的版本。

@soumithchintala 原文 ↗

Inkling 是 Thinking Machines 的首个开源模型,支持多模态推理,致力于降低对中心化 AGI 公司的依赖

我们很高兴推出我们的第一个通用模型Inkling — 开放权重,975B参数,原生多模态(文本、图像、音频)。可在Tinker、HuggingFace和合作伙伴处获取。

它可以由你个人定制和开放使用。这就是你的模型。
展开原文
Excited for our first general model Inkling -- open weights, 975B, natively multimodal (text, image, audio). Available on Tinker, HuggingFace and partners.

It is yours to personalize and use openly. It is yours.
@thinkymachines 今天,我们正在推出Inkling。

Inkling在文本、图像和音频模态之间高效推理。我们将完整权重公开提供。

https://t.co/Ghebq5mG30

今天即可在Tinker上进行微调。在Inkling Playground中试用它。🧵
Today, we are introducing Inkling.

Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.

https://t.co/Ghebq5mG30

Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
❤ 2.6k · 🔁 134 · 💬 62 · 👁 20.1w
热门回复 4
@NVIDIAAI @soumithchintala Congrats Soumith!
@beffjezos @soumithchintala Huge congrats!
@slchase @soumithchintala 危险从来不是AI不能爱。它只需要比那些爱的人更容易。关于原始提示的新文章 https://t.co/dTeslcPpzs
@soumithchintala The danger was never that AI can't love. It's that it only has to be easier than the people who do. New essay on the original prompt https://t.co/dTeslcPpzs
@kgonia7 @soumithchintala 270GB ☠️ https://t.co/eV4uUeEqKt
@rasbt 原文 ↗

Inkling 在架构上有创新之处,如小型卷积层、RMSNorm 嵌入层、相对位置偏置等设计

Thinky的意外惊喜发布!Inkling模型在基准测试中表现相当不错,它的架构中有一些小惊喜:

- 在多个地方使用小卷积层
- 在嵌入层使用RMSNorm(在块RMSNorm之前)
- 使用相对位置偏置而不是RoPE https://t.co/oMl5Ta6Ttr
展开原文
Interesting surprise drop from Thinky! The Inkling model looks pretty solid on benchmarks, and it has some little surprises in its architecture:

- Small conv layers in several places
- An RMSNorm for the embeddings (before the block RMSNorm)
- Rel. position bias instead of RoPE https://t.co/oMl5Ta6Ttr
@eliebakouch 第一个开放权重的思考机器模型!! 975B总参数,41B活跃参数,在45T tokens上训练,1M上下文,多模态

滑动窗口比例5:1,大小512,deepseek辅助无负载平衡和2个共享专家(通常人们只使用1个),实际上很好奇为什么该模型比kimi更稀疏(~4.2% vs 3.2%)。他们在k和v之后使用短卷积,输出和ffn(见图),muon(他们引用manifold muon但提到权重衰减所以不确定),muP,并拥有一个非常好的RL缩放曲线和思维链!

个人认为该发布中非常酷的一部分是他们的小变体(276B总参数,12B活跃参数)相比大模型表现得相当不错。他们提到改变了预训练数据混合和配方,很好奇这些变化以及即将在~1T(或更多?)的规模上看到它们 👀

> "它不是当今可用的最强大模型,无论是闭源还是开源的。我们训练Inkling是为了在各个方面都拥有可靠的能力,而不是在单一领域达到最先进的性能,以便作为我们将来训练模型的基础。"

这在模型发布中真的很 refreshing 看到,恭喜你 :)
first open weight thinking machine model!! 975B total, 41B active trained on 45T tokens, 1M context, multimodal in

sliding window with a 5:1 ratio and 512 size, deepseek aux-free load balancing and 2 shared experts (usually people only use 1), actually curious why the model is less sparse than kimi (~4.2% vs 3.2%). they use a short convolution after k and v, output and ffn (see plot), muon (they cite manifold muon but mention weight decay so not sure), muP, and have a very nice RL scaling curve and chain of thought!

one very cool part of the release imo is how well their small variant (276B total, 12B active) performs compared to the big one. they mention they changed the pre-training data mix and recipe, very curious about those changes and to see them scaled up to ~1T (or more?) soon 👀

> "It is not the most performant model available today, closed or open. We trained Inkling for solid capabilities across the board rather than state-of-the-art performance in a single area, to serve as a foundation for the models we will train in the future."

also this is really refreshing to see in a model release, huge congrats :)
❤ 820 · 🔁 89 · 💬 20 · 👁 6.4w
热门回复 4
@rasbt 一些想法:- 比GLM 5.2大250B参数 - 比Kimi K2.5 1T更稀疏(3.2%的稀疏度使用32B活跃参数,而不是41B活跃参数的4.2%稀疏度) - 它不像Nemotron那样使用混合方法。想知道每秒token吞吐量比较
Some more thoughts:
- It's 250B parameters bigger than GLM 5.2
- less sparse than sparse than Kimi K2.5 1T (3.2% sparsity with 32B active instead of 41B active with 4.2% sparsity)
- It doesn't use a hybrid approach like Nemotron.

Curious about a token/sec throughput comp
@rasbt @themintsv 难说。可能是数据质量、训练配方、超参数设置……或者全部以上
@themintsv Hard to say. Could be data quality, training recipe, hyperparameter settings...
or all of the above
@rasbt @midsusnight yes, refreshing!
@midsusnight @rasbt 等等,他们已经在说这不是最好的模型了吗?至少要诚实地发布。
@rasbt wait so theyre already saying its not the best model

honest release at least
@SahilBloom 原文 ↗

Lilian Weng 强调 Inkling 旨在成为广泛能力的基础模型,支持实践和定制

我相信残酷的自我诚实对于成功至关重要。我认识的最令人印象深刻的人都是自己的严格批评者。除非你意识到你想要的与你正在创造的之间的差距,否则你无法改进。如果什么都不改变,什么都不会改变。
展开原文
I’m convinced that brutal self-honesty is essential for success. The most impressive people I know are their own toughest critics. You can’t improve until you become aware of the gap between what you want and what you’re doing to create it. Nothing changes if nothing changes.
❤ 1.0k · 🔁 118 · 💬 81 · 👁 4.7w
热门回复 4
@CuriousMindsHub @SahilBloom 自我意识展示了差距。纪律是关闭它的东西。
@SahilBloom Self-awareness shows you the gap. Discipline is what closes it.
@itslereto @SahilBloom 你不能关闭你拒绝看到的差距。
@SahilBloom You can’t close a gap you refuse to see.
@JamsomSiger @SahilBloom 这是天赐之物也是诅咒,也可能是滑坡。残酷的自我诚实可以在评论和残酷之间走一条细线。
@SahilBloom It’s a gift and a curse and can be a slippery slope. Brutal self-honesty can walk a fine line between critic and cruelty.
@maveroise @SahilBloom 我认为自我诚实是自尊最伟大的形式之一。不是因为它让你感到内疚,而是因为它给你成为你一直希望成为的人的机会。
@SahilBloom I think self-honesty is one of the greatest forms of self-respect. Not because it makes you feel guilty, but because it gives you the chance to become the person you've been hoping to be.
@soumithchintala 原文 ↗

Modal 为 Inkling 训练了 DFlash speculator,提升推理速度 67%,展示了生态系统的快速适配

Modal训练了一个DFlash推测器,比MTP快得多,为推理速度提供了很大的提升!https://t.co/EBLiAGy57r
展开原文
Modal trained a DFlash speculator that's much faster than MTP, making it a great boost for inference speeds! https://t.co/EBLiAGy57r
@modal @thinkymachines的Inkling现在可在Modal上使用,由自定义DFlash推测器支持,提供67%更高的吞吐量和交互性。

今天在Modal Auto Endpoints上使用SGLang运行。https://t.co/OxN7aJ9ieW
Inkling by @thinkymachines is now available on Modal, backed by a custom DFlash speculator for 67% higher throughput and interactivity.

Running on Modal Auto Endpoints with SGLang today. https://t.co/OxN7aJ9ieW
❤ 282 · 🔁 31 · 💬 9 · 👁 2.7w
@soumithchintala 原文 ↗

Thinking Machines 致力于个性化、人类参与和去中心化,推出 Tinker 等产品让用户能够塑造 AI

我们在@thinkymachines做什么:

个性化/主权、人类参与、去中心化。
使AI民主化并使其对人们有用。

这三者都减少了社会对中心化AGI公司的依赖(包括我们在变得重要时),这是一个值得追求的未来。

你已经在Tinker、Interaction模型和我们在Connectionism上公开发布的研究中看到了这方面的预览。
一个**很多**即将很快到来...
展开原文
What do we do at @thinkymachines:

Personalization/sovereignty, Human Participation, Decentralization.
Democratize AI and make it useful for people.

All three of them reduce society's dependence on centralized AGI companies (including ours when we get important), and that is a future worth aiming for.

You've seen a preview of this with Tinker, Interaction models and our research openly published on Connectionism.
A **lot** more to come very very soon...
@thinkymachines 我们正在构建人们和组织可以塑造和使其成为自己的AI。AI应该扩展我们的意志和判断,而不是忽视它;实现这一点是我们正在努力解决的技术挑战。
https://t.co/Bi558y4vqD
We're building AI that people and organizations can shape and make their own. AI should extend our will and judgment instead of neglecting it; enabling that is the technical challenge we are working to solve.
https://t.co/Bi558y4vqD
❤ 825 · 🔁 59 · 💬 43 · 👁 21.0w