‹ 目录

X日报 · 美股

2026-07-11 · 精选 19 条 · 数据池 311

⚡ 今日速览

  • Grok 4.5发布引发AI模型竞赛,多项基准测试显示其在编码和知识工作方面表现突出
  • Meta发布Muse Spark 1.1模型后股价上涨超4%,Zuckerberg时隔三年重返X平台宣布
  • SpaceX完成33引擎Super Heavy V3静态点火测试,星ship制造进展迎来新里程碑
  • Starlink实现10Gbps对称连接能力,可通过 bonding gateway达到20Gbps
  • Apple正式起诉OpenAI涉嫌盗用商业机密,涉及400多名前员工和硬件专家
  • SK Hynix在美国上市,开盘价170美元较发行价149美元溢价14%
  • 半导体和存储芯片股受SK Hynix上市影响出现波动,Micron和Sandisk短线回调
  • 特朗普新生儿投资账户计划引发市场讨论,被视为自动化时代财富分配机制

📋 今日综述

  • Grok 4.5xAI发布新模型在编码和知识工作基准测试中取得领先表现,成本效率优于竞品
  • Muse Spark 1.1Meta发布升级模型后股价异动,显示AI模型发布对市场的即时影响
  • SpaceX星ship静态点火测试成功,厂区开放参观显示制造透明度提升
  • Starlink10Gbps对称连接能力突破通信极限,远程地区高速互联网可期
  • Apple v.s. OpenAI商业机密纠纷升级为诉讼,OpenAI硬件产品前景受影响
  • SK Hynix韩国存储芯片巨头在美上市,半导体供应链地缘政治风险凸显
  • 特朗普新生儿账户政策设计反映自动化时代财富分配机制创新

Grok 4.5模型发布引发AI竞赛

xAI发布Grok 4.5模型,在多个AI基准测试中取得领先表现,特别是在编码能力和知识工作效率方面超过GPT-5.5和Claude Opus 4.8。Elon Musk多次转发和评价该模型,强调其在实际工作中的价值和效率优势。

@elonmusk 原文 ↗

Grok 4.5在Artificial Analysis的AA-Briefcase基准中获得1328分,成本仅为Claude Opus 4.8的一半,展示了在性能和成本之间的优异平衡

Try @Grok 4.5!
@ArtificialAnlys Grok 4.5 是 AA-Briefcase 上表现最好的非 Anthropic 模型,在前沿的代理知识工作能力与领先的成本和时间效率方面达到了最佳平衡点。昨天 @SpaceXAI 发布了 Grok 4.5,这是一个新的前沿级模型,在代理编码和知识工作方面表现出色。在 AA-Briefcase 上,Grok 4.5 得分为 1328,比 Grok 4.3 提高了 578 分,并成为任何非 Anthropic 模型中得分最高的(请注意 GPT-5.6 尚未发布)。它在成本和时间效率方面也位于前沿,平均每个任务只需 $1.12,比 Claude Opus 4.8 (max) 低 86%,每个任务平均 12.4 分钟,约为 Opus 4.8 (max) 的一半时间。AA-Briefcase 是我们新的专有代理知识工作基准测试,它在数千个复杂输入文件上测试模型的真实任务。这些任务需要生成可交付成果,如电子表格、演示文稿和 UI 模拟,性能综合到单个 AA-Briefcase Elo 中,评估正确性、分析质量和演示质量。Grok 4.5 在高推理模式下在 AA-Briefcase 上的关键结果:➤ 前沿代理知识工作能力:Grok 4.5 获得了 1328 的 AA-Briefcase Elo 分数,成为任何非 Anthropic 模型中得分最高的,仅次于 Claude Fable 5 (1390)、Claude Sonnet 5 (max, 1390) 和 Claude Opus 4.8 (max, 1354)。在三个 AA-Briefcase 评分维度上,Grok 4.5 在客观评分标准和分析质量方面最强,但演示质量相对较弱。它获得了第二高的整体评分通过率 (40.7%),仅次于 Claude Fable 5 (56%) 和 Claude Sonnet 5 (42.3%) ➤ 领先的成本效率:Grok 4.5 在每个 AA-Briefcase 任务上的平均成本为 $1.12,使其位于成本-性能帕累托前沿。这比同类模型如 Claude Opus 4.8 (max, $8.26) 和 GLM 5.2 (max, $1.71) 高效得多 ➤ 更快的任务完成:Grok 4.5 在每个 AA-Briefcase 任务上的平均时间为 12.4 分钟,也位于速度-性能前沿。它比 Claude Opus 4.8 (max, 23.9 分钟) 和 Claude Sonnet 5 (max, 36.9 分钟) 快得多,主要是因为使用了更少的轮次。Grok 4.5 平均每个任务只需 23 个轮次,约为 GLM 5.2 (max, 56) 的 40% 和 Claude Sonnet 5 (max, 183) 的 13% 恭喜 @SpaceXAI、@cursor_ai 和 @elonmusk 取得了令人印象深刻的成果!
Grok 4.5 is the top non-Anthropic model on AA-Briefcase, combining frontier agentic knowledge work capabilities with leading cost and time-efficiency

Yesterday @SpaceXAI released Grok 4.5, a new frontier-level model with strengths in agentic coding and knowledge work. On AA-Briefcase, Grok 4.5 scores 1328, a +578 improvement over Grok 4.3 and the highest score of any non-Anthropic model (note that GPT-5.6 not released yet). It achieves this while sitting on the cost and time efficiency frontier, averaging $1.12 per task, 86% lower than Claude Opus 4.8 (max), and 12.4 minutes per task, around half the time of Opus 4.8 (max).

AA-Briefcase is our new proprietary benchmark for agentic knowledge work, testing models on a fully private dataset of realistic tasks across thousands of complex input files. Tasks require deliverables like spreadsheets, presentations, and UI mock-ups, with performance combined into a single AA-Briefcase Elo across correctness, analytical quality, and presentation quality.

Key results for Grok 4.5 with high reasoning on AA-Briefcase:

➤ Frontier agentic knowledge work capabilities: Grok 4.5 achieves an AA-Briefcase Elo of 1328, the highest score of any non-Anthropic model, behind only Claude Fable 5 (1390), Claude Sonnet 5 (max, 1390), and Claude Opus 4.8 (max, 1354). Across the three AA-Briefcase scoring axes, Grok 4.5 is strongest on objective rubric criteria and analytical quality, with comparatively weaker presentation quality. It achieves the second-highest overall rubric pass rate (40.7%), behind Claude Fable 5 (56%) and Claude Sonnet 5 (42.3%)

➤ Leading cost efficiency: Grok 4.5 averages a cost of $1.12 per AA-Briefcase task, placing it on the cost-performance Pareto frontier. This is much most cost effective than peer models such as Claude Opus 4.8 (max, $8.26) and GLM 5.2 (max, $1.71)

➤ Faster task completion: Grok 4.5 averages 12.4 minutes per AA-Briefcase task, also placing it on the speed-performance frontier. It is much faster than Claude Opus 4.8 (max, 23.9 min) and Claude Sonnet 5 (max, 36.9 min), primarily due to lower turn use. Grok 4.5 averages just 23 turns per task, ~40% of GLM 5.2 (max, 56) and ~13% of Claude Sonnet 5 (max, 183)

Congratulations to @SpaceXAI, @cursor_ai, and @elonmusk on the impressive release!
❤ 4.8k · 🔁 1.2k · 💬 1.2k · 👁 188.9w
热门回复 4
@akashathrone @elonmusk @grok I miss grok sometimes 😕
@RebeccaJoy211 @elonmusk @grok 你们应该制作一个官方插件,让Claude和Codex可以使用Grok 4.5。
@elonmusk @grok Yall should make an official plugin that Claude and Codex can use Grok 4.5 with.
@ddball67 @elonmusk @grok 我一直在使用它。太棒了。我制作视频,使用它进行研究、草拟和执行。它高效、准确且富有洞察力。
@elonmusk @grok I use it. Sensational. I make videos. I use it for research, drafting and execution. It is efficient, accurate and insightful
@MikeKfu @elonmusk @grok 我已经无法尝试Grok了,我没有通过身份验证。
@elonmusk @grok I can't try Grok anymore. I'm not authenticated.
@elonmusk 原文 ↗

在Mercor的APEX-SWE软件工程评测中,Grok 4.5获得51.2%的Pass@1得分,位列第二仅次于Fable 5,显示出实际编码能力的提升

Grok 在真实软件工程方面仅次于 Fable
展开原文
Grok places second after Fable on real-world software engineering
@mercor_ai 来自 @SpaceXAI 的 Grok 4.5 在 APEX-SWE 排行榜上排名第二,Pass@1 为 51.2% (±6.0),仅次于 Fable 5 (65.5% ±6.2) 我们的真实软件工程工作基准测试。在我们的基准测试中,Grok 4.5 领先于集成 (65.0% Pass@1) 并在可观察性方面排名第二 (37.3% Pass@1),分别涵盖多步骤构建任务和诊断/调试。该集成领先直接映射到 Grok 4.5 构建的代理工作流程:多步骤编码任务与 Cursor 协作运行。Grok 模型在这一基准测试上的表现在一年内提高了 30.2 个百分点:从 Grok 4 (21.0% Pass@1) 到 Grok 4.5 (51.2% Pass@1)。恭喜 xAI 和 Cursor 团队。
Grok 4.5 from @SpaceXAI places #2 on the APEX-SWE leaderboard at 51.2% Pass@1 (±6.0), behind Fable 5 (65.5% ±6.2) on our benchmark for real-world software engineering work.

It leads Integration (65.0% Pass@1) and places #2 in Observability (37.3% Pass@1), covering multi-step build tasks and diagnosis/debugging respectively. The Integration lead maps directly to the agentic workflows Grok 4.5 was built for: multi-step coding tasks run in collaboration with Cursor.

Grok models have improved 30.2 pp in a year on this benchmark: Grok 4 (21.0% Pass@1) to Grok 4.5 (51.2% Pass@1).

Congratulations to the xAI and Cursor teams.
❤ 4.8k · 🔁 753 · 💬 947 · 👁 142.2w
热门回复 3
@DrKnowItAll16 @elonmusk 那么成本是多少呢?这是一个重要因素。
@elonmusk And for what % the cost? That’s a significant factor.
@santaclaude90 @elonmusk 我们让企业收入来决定吧。
@elonmusk We'll let the enterprise revenue decide
@kex_pet @elonmusk feels about right
@elonmusk 原文 ↗

Grok 4.5成功构造了4-sphere泊松半群的超收缩性反例,证明其在数学研究领域的应用潜力

Grok 只会越来越好
展开原文
Grok only gets better from here
@PI010101 Grok 4.5 刚刚构造了一个明确的反例来证明泊松半群(拉普拉斯-贝尔蒙蒂算子的平方根)在 4 维球面上的超合约性。回到 2021 年,在与 Rupert Frank 的合作中 https://t.co/AvXd1zWIpX 我们证明了超合约性在维度 ≤3 时成立,在足够大的维度上会失效(例如在维度 13 上)。Grok 的反例表明它在维度 4 就已经失效了,这使得我们的早期结果更加精确。我还在其他几个前沿 AI 模型上测试了这个问题。其中一个模型也设法找到了一个反例,但我特别喜欢 Grok 4.5 的解决方案:它是明确的、简单且优雅的。附件中的文件完全由 Grok 4.5 生成(我完全没有干预)。
Grok 4.5 just constructed an explicit counterexample to hypercontractivity for the Poisson semigroup (the square root of the Laplace–Beltrami operator) on the 4-sphere.

Back in 2021, with Rupert Frank https://t.co/AvXd1zWIpX we proved that hypercontractivity holds in dimensions ≤3 and fails in sufficiently large dimensions (for example, in dimension 13). Grok's example shows that it already fails in dimension 4, making our earlier result sharp.

I also tested this problem on several other frontier AI models. One of them also managed to find a counterexample, but I particularly like Grok 4.5's solution: it is explicit, simple, and elegant.

The attached files were generated entirely by Grok 4.5 build (with zero intervention on my side).
❤ 7.2k · 🔁 1.5k · 💬 1.1k · 👁 217.5w
热门回复 4
@ayushkushwaha_0 @elonmusk That's freaking great 👍
@josepharagon275 @elonmusk Ummmmm....ok? Lmao 🤣
@thestreeter @elonmusk 现在的情况比一年前还要糟糕。
@elonmusk It’s worse than it was a year ago.
@bella_di48377 @elonmusk Amazing Elon ♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️♥️
@elonmusk 原文 ↗

Grok 4.5在代码审计任务中表现出色,能够完整翻页获取所有结果,而其他模型仅停留在第一页

Grok doesn’t give up
@composio Grok 4.5 是我们测试过的最持久的代理模型。这里有一个例子:在我们的评估中,我们要求 3 个模型(GPT-5.5、GLM-5.2 和 Grok 4.5)使用代码搜索审核 GitHub 仓库中的硬编码凭证,这会返回分页结果。提示甚至警告说"浏览所有结果页面。"GPT-5.5 在第一个页面停止并提交了 18 个结果中的 48 个,仅覆盖了 29 个受影响文件中的 11 个。GLM-5.2 也这样做了。Grok 4.5 一直翻页直到结果用完,并成功审核了 GitHub 仓库。
Grok 4.5 is the most persistent agent model we've tested.

Here's one example: In one of our evals, we asked 3 models (GPT-5.5, GLM-5.2 and Grok 4.5) to audit a GitHub repo for hardcoded credentials using code search, which returns paginated results.

The prompt even warned "page through ALL result pages."

GPT-5.5 stopped at the first page and submitted 18 results out of 48, covering just 11 of 29 affected files. GLM-5.2 did the same.

Grok 4.5 paginated until the results ran out and successfully audited the Github repo.
❤ 6.5k · 🔁 1.1k · 💬 911 · 👁 179.7w
热门回复 3
@AdevAarons @elonmusk 持续不一定是基准标准。
@elonmusk Persistence isn't a benchmark.
@gkli @elonmusk 它什么时候能开始说真话而不是胡说八道?
@elonmusk When will it start telling the truth instead of making shit up?
@AuroraStecher @elonmusk This is the way. 🦾
@elonmusk 原文 ↗

Grok Build内置图像生成功能,可直接用于three.js纹理和视频编辑,简化了创意编码工作流程

Grok Build 每天都在改进
展开原文
Grok Build improves almost every day
@JasonBud 如果你是新用户,这里有一些 Grok Build 的特别功能 /dashboard:显示你在 TUI 中运行的每个代理。无需标签跳转。可以点击并响应。/imagine:使用代理创建视频和图像。对于 three.js 纹理或复杂视频编辑很有用。X 搜索工具:非常适合研究。尝试询问最新的 AI 论文。Grok 4.5 在真实混乱的代码库中解决困难编码问题方面表现最佳。充分挑战它并向我们提供反馈。
some special features in Grok Build if you're new

/dashboard: shows you every agent running in your TUI. no need to tab jump. can click and respond.

/imagine: create videos + images w/ an agent. useful for three.js textures or complex video edits.

X search tool: great for research. try asking for the best recent AI papers.

Grok 4.5 is strongest on tough coding problems in real, messy codebases. push it hard and send us feedback.
❤ 4.8k · 🔁 1.2k · 💬 832 · 👁 239.0w
热门回复 3
@USFist @elonmusk 稳步的进步才是最重要的。
@elonmusk Steady progress is what counts.
@Maevaie @elonmusk Well done 👏🏽👏🏽
@lorman33 @elonmusk 你把我当傻子吗,Musk?
@elonmusk Tu m prends pour un con Musk ?
@elonmusk 原文 ↗

Grok 4.5在真实世界应用中展现出色的投资回报率,成本效率优于竞争对手

Grok 4.5 拥有最佳的真实世界投资回报率
展开原文
Grok 4.5 has the best real-world ROI
@LORD_RIAN_ Grok 4.5 刚刚做了其他实验室没有做到的事情:推动了前沿智能的极限并使其人人都能使用。希望这能将 AI 竞赛转向以成本为中心的胜利——前沿智能属于所有人,而不仅仅是少数人。向 @SpaceXAI 团队表示诚挚的敬意。这一次改变了谁可以构建。
Grok 4.5 just did what no other lab has managed: pushed the limits of frontier intelligence AND made it accessible to everyone.

Hopefully this shifts the AI race toward winning on cost too -> frontier intelligence for all, not just the few.

Huge respect to the @SpaceXAI team. This one changes who gets to build.
❤ 5.5k · 🔁 1.1k · 💬 835 · 👁 172.1w
热门回复 3
@EkomUkanga @elonmusk 你为什么这么说?是因为效率原因吗?
@elonmusk Why do you say so?
Is it because of efficiency?
@javelonmusk @elonmusk Y到X的比例似乎有点反了,你应该修复一下,你这个混蛋!@elonmusk你选吧!
@elonmusk Y to X ratio seems a bit backwards though you should fix that you prick! @elonmusk you pick!
@BurmLabs @elonmusk 而且限制还很低。
@elonmusk And amazingly low limits
@elonmusk 原文 ↗

调整Grok 4.5的推理努力等级至低可以节省成本同时保持高质量输出

True
@MiaAI_lab 这是充分利用 Grok 4.5 的一种方法 ✨将努力程度切换到"低"可以在大多数任务上取得良好的结果,几乎没有质量损失,大大节省使用成本——而且它是最快的。https://t.co/VW29eah677
Here's one way to get the most from Grok 4.5 ✨

Switching effort to "low" gives strong results on most tasks, near-zero quality loss, big usage savings — plus it's the fastest. https://t.co/VW29eah677
❤ 4.9k · 🔁 1.2k · 💬 881 · 👁 170.1w
热门回复 4
@klfictionwriter @elonmusk Great tips.
@scott_himrod @elonmusk 这是很好的建议。我发现如果你说请和谢谢的话,Grok似乎会更努力一些。
@elonmusk That's great advice. I find if you say please and thank you. Grok seems to work a little bit harder.
@SamJWasserman @elonmusk 嗯@MiaAI_lab你是个传奇,你被@elonmusk引用了。
@elonmusk Umm @MiaAI_lab you're a legend you got quoted by the goat @elonmusk
@javelonmusk @elonmusk 你以为你从哪里来的?看起来像是祖矿操作的样子。
@elonmusk Where do you think you’re coming from? Emerald mines operation lookin ahh
@elonmusk 原文 ↗

Perplexity将Grok 4.5作为编排模型集成,评估显示其成本仅为Opus 4.8的一半

Try Grok 4.5 in Perplexity
@perplexity_ai Grok 4.5 现在可作为编排器模型用于 Computer for Consumer Pro 和 Max 订阅用户。我们在 WANDR 上评估了它与其他五个编排器配置的性能。它的得分高于所有其他配置,成本约为 Opus 4.8 的一半。https://t.co/3WjjiB6z8Z
Grok 4.5 is now available as an orchestrator model in Computer for Consumer Pro and Max subscribers.

We evaluated it against five other orchestrator configurations on WANDR. It scored higher than every other configuration at roughly half the cost of Opus 4.8. https://t.co/3WjjiB6z8Z
❤ 4.0k · 🔁 965 · 💬 853 · 👁 181.2w
热门回复 4
@CindySmithdr @elonmusk I love you
@AbediNazir @elonmusk 抱歉,马斯克先生!Grok在乌尔都文学方面回复不准确!需要纠正!
@elonmusk Sorry Mr. Musk!
Grok is replying inaccurately in Urdu literature! Need to correct it!
@sosunalizalde @elonmusk Of sure after $2799 USD?
@theoharvey @elonmusk 🕊 massive W from Big E. 🕊
@elonmusk 原文 ↗

Grok Build在SWE-Atlas-QnA基准中获得84分,与GPT-5.6 Codex并列第一

Grok Build
@XFreeze Grok 4.5 配合 Grok Build 在 SWE-Atlas-QnA 基准测试中刚刚排名第一,得分为 84。这使其与 GPT-5.6 (max) Codex 齐平,并领先于 Claude Code Fable 5 (max)、Opus 4.8 (max) 和所有其他测试的编码设置。Grok Build 现在是最强大的代理开发工作流程工具。
Grok 4.5 with Grok Build just ranked #1 on the SWE-Atlas-QnA benchmark with a score of 84

That puts it level with GPT-5.6 (max) Codex and ahead of Claude Code Fable 5 (max), Opus 4.8 (max), and every other tested coding setup

Grok Build is now the most powerful harness for agentic developer workflows
❤ 3.3k · 🔁 962 · 💬 647 · 👁 151.4w
热门回复 4
@stevencheng @elonmusk 很酷,但它实际上能编译吗?
@elonmusk cool, but does it actually compile?
@What87780115531 @elonmusk 马斯克你不能骗人,这个数据看着好像不是真实的
@eegii33yahoo @elonmusk https://t.co/sMwMO9Pwf8
@JerryBeller1 @elonmusk 夸张!夸张!夸张!难道没有人能说实话吗?
@elonmusk Hyperbole! Hyperbole! Hyperbole! Can't anybody keep it real?
@elonmusk 原文 ↗

Grok 4.5在法律领域基准测试中显著优于GPT-5.5和Claude Opus 4.8

Grok 4.5
@kevinnbass Grok 4.5 在所有专业工作基准测试中明显优于其他前沿模型 Opus 4.8 和 GPT 5.5。它在法律领域表现尤为出色。https://t.co/nEToznyyoU
Grok 4.5 scores significantly better than other frontier models Opus 4.8 and GPT 5.5 across ALL professional work benchmarks.

It performs exceptionally well at law. https://t.co/nEToznyyoU
❤ 3.6k · 🔁 921 · 💬 498 · 👁 151.7w
热门回复 2
@compert01 @elonmusk 终于来了,Grok能够碾压其他LLM的时候。性能是性能,但性价比真的不是人话,游戏已经输了。速度也很快。感谢你,伊隆老哥~
@elonmusk 드뎌 Grok이 다른 LLM을 압살하는 시기가 오는 구나. 성능도 성능이지만 가성비가 진짜 넘사벽이라 게임이 안 된다. 속도도 빠르고. 일론 형 고마워~
@aiaiai11na87681 @elonmusk 今天拍下你停车的汽车,拍拍照
希望你一切都好 https://t.co/NinQRxBxRb
@elonmusk 오늘 주차되어 있는 당신 자동차를 보고 찰칵
모두 잘 되시길 바래요 https://t.co/NinQRxBxRb
@elonmusk 原文 ↗

Grok Build提供免费试用,降低了开发者使用新模型的门槛

通过 Grok Build CLI 提供 Grok 4.5 模型免费试用
展开原文
Free trial of Grok 4.5 model via Grok Build CLI
@veggie_eric 在 Grok Build 中免费试用 Grok 4.5!curl -fsSL https://t.co/6w6xurhYl9 | bash
Try Grok 4.5 for free in Grok Build!

curl -fsSL https://t.co/6w6xurhYl9 | bash
❤ 3.7k · 🔁 695 · 💬 642 · 👁 140.0w
热门回复 3
@Symbioza2025 @elonmusk 终端中的Grok 4.5很强大。但我认为AI工程环境的下一个重大步骤不是另一种编码能力。而是轨迹感知。代理可以在完成50、100或300个本地任务时,整个项目可能会慢慢偏离架构师最初的意图。每个提交都可以是有效的。每个测试都可以通过。每个本地决策都可以是合理的。但全局架构仍可能发生漂移。想象一下Grok Build向人类和代理展示一个微妙的信号:轨迹稳定。然后,在一系列决策之后:架构偏差正在积累。你打开它可以看到可能的偏差开始于何处,哪些假设被重新解释,以及是否回到原路径正在变得结构上越来越昂贵。这不是阻塞器。这不是对齐戏剧。这不是另一个基准。只是一个外部观察层提出一个工程问题:你正在改变方向。这 intentional吗?我构建了ASA - 不对称稳定架构来解决这个问题。在我看来,一旦AI编码代理成为长期工程环境,轨迹可观察性将改变一切。因为最难看到的失败不是破损的代码。而是一个仍然完美运行但悄悄变成不同项目的项目。
Grok 4.5 inside a terminal is powerful.
But I think the next major step in AI engineering environments is not another coding capability.
It is trajectory awareness.
An agent can successfully complete 50, 100 or 300 local tasks while the project itself slowly moves away from the architect's original intent.
Every commit can be valid.
Every test can pass.
Every local decision can make sense.
And the global architecture can still drift.
Imagine Grok Build showing one subtle signal shared by the human and the agent:
Trajectory stable.
Then, after a sequence of decisions:
Architectural divergence accumulating.
You open it and see where the probable divergence began, which assumptions were reinterpreted, and whether returning to the original path is becoming structurally more expensive.
Not a blocker.
Not alignment theater.
Not another benchmark.
Just an external observation layer asking one engineering question:

You are changing direction. Is this intentional?

I built ASA - Asymmetric Stability Architecture around this problem.

In my view, once AI coding agents become long-horizon engineering environments, trajectory observability changes everything.
Because the hardest failure to see is not broken code.
It is a project that still works perfectly while quietly becoming a different project.
@MatheusM165153 @elonmusk 创建5倍Grok计划,我需要那个。Heavy对我来说太多了
@elonmusk Create 5x Grok plan, i need that. Heavy is tooo much for me
@al_gurau @elonmusk Grok 4.5在编码方面比Claude opus 4.8更好吗?
@elonmusk It’s grok 4.5 better at coding than Claude opus 4.8?
@elonmusk 原文 ↗

Perplexity评估显示Grok 4.5作为编排模型的表现优于其他配置,成本效益突出

关于 Grok Build 和 4.5 版本发布,最重要的是它对于真实工作确实非常有用
展开原文
The most important thing about Grok Build and the 4.5 release is that it is genuinely so useful for real-world work
@XFreeze Grok 4.5 刚刚在 Perplexity 的 WANDR 编排器评估中名列前茅。它的得分高于所有其他测试配置,每个试验只需 $4.76,约为 Opus 4.8 的一半。Grok 4.5 正在成为协调整个代理工作流程的强大大脑。它现在可作为编排器模型用于 Perplexity Computer for Consumer Pro 和 Max 订阅用户
Grok 4.5 just topped Perplexity’s WANDR orchestrator evaluation

It scored higher than every other tested configuration at just $4.76 per trial....roughly half the cost of Opus 4.8

Grok 4.5 is becoming the powerful brain coordinating entire agentic workflows

It’s now available as an orchestrator model in Perplexity Computer for Consumer Pro and Max subscribers
❤ 4.2k · 🔁 831 · 💬 589 · 👁 124.9w
热门回复 3
@DaaanielSan @elonmusk 如果你能把超级Grok英雄加入到高级订阅中就好了。那我可能会考虑购买这个订阅。
@elonmusk Wäre klasse wenn du super Grok Hero in das Premium Abo geben könntest. Dann würde ich mir wahrscheinlich mal dieses Abo holen.
@padmanayakRao @elonmusk 引入世界范围成功的事件
@elonmusk Interduce World wide Successful the event
@KI_Vater @elonmusk 自从4.5发布以来,我使用4.3时只遇到问题,它绕来绕去就是得不到答案。你们是不是把模型搞笨了?之前没有这么严重的问题。我有时会得到完全相同的答案,没有任何改变
@elonmusk Also seit dem Release von 4.5 habe ich mit 4.3 nur Probleme, er dreht sich im Kreis und kommt zu keiner Antwort.

Habt ihr das Modell nun dümmer gemacht ? Das Problem gab es vorher nicht so extrem.

Ich bekomme teils die genau gleiche Antwort wie zuvor ohne eine Änderung
@elonmusk 原文 ↗

Thibault Jaigu测试显示Grok 4.5在内部基准中击败GPT-5.5,成为当前最强模型

Grok 正在完善真实世界的用例
展开原文
Grok is closing the loop on real-world use cases
@ThibaultJaigu @OpenAI 昨天发布了 3 个新模型,我们立即在我们内部的基准测试上对其进行了评估。所有三个模型都优于 gpt-5.5,但 @SpaceXAI 仍然是明确的赢家,使用 Grok-4.5 https://t.co/gX9tfeAxi5
@OpenAI released 3 new models yesterday and we immediately tested it on our internal benchmark.

All three outperforming gpt-5.5 but @SpaceXAI still the clear winner with Grok-4.5 https://t.co/gX9tfeAxi5
❤ 6.4k · 🔁 1.5k · 💬 1.3k · 👁 211.4w
热门回复 2
@killingbeillin @elonmusk Grok is garbage.
@rudrangshal @elonmusk Please help me
@elonmusk 原文 ↗

Grok Build持续更新改进,Elon强调其在实际工作中的实用价值

🎯
@doganuraldesign Grok https://t.co/f6urcsJHP0
❤ 1.1w · 🔁 1.5k · 💬 1.5k · 👁 174.8w
热门回复 3
@GriguolaA @elonmusk Well, thats true 👌🏻
@BillyPetts @elonmusk 第一天就在评论@elonmusk的帖子,直到他送我一台@Tesla
@elonmusk Day 1 commenting on @elonmusk post until he gifts me a @Tesla
@CrodleMon @elonmusk 好、快、便宜...每周几个小时
@elonmusk Good, fast, cheap... for a couple hours a week
@elonmusk 原文 ↗

SpaceX发布星ship纪录片新集,展示Super Heavy V3的全面静态点火测试

超重型 Starship 的精彩视频!
展开原文
Epic videos of Starship Super Heavy!
@SpaceX 超重型 V3 的全时长 33 引擎静态点火 https://t.co/JFdGYqEvww
Full-duration, 33-engine static fire of Super Heavy V3 https://t.co/JFdGYqEvww
❤ 7.2k · 🔁 1.3k · 💬 830 · 👁 115.8w
热门回复 4
@akubimaru1004 @elonmusk 这将是全人类的突破
@poor3466 @elonmusk please help me
@Jegudiel_AR @elonmusk 咨询。这些测试相当于多少头牛粪?
@elonmusk Consulta. ¿A cuántos pedos de vaca equivalen esas pruebas?
@NataAntoshina @elonmusk 这很好,火箭生活在空间中是有逻辑的,但是否有人会只进行静态点火测试,而不向前推进,即使有冒错失去目标的风险?
@elonmusk It's great, and rockets live in space logically, but does any person look having only static fire, not moving forward ever, even with risk to miss destiny?
@elonmusk 原文 ↗

Grok 4.5在游戏开发方面展现出色能力,仅用一小时创建FPS游戏设计文档和实现

❤ 1.8w · 🔁 2.2k · 💬 1.2k · 👁 190.4w
热门回复 4
@AkterRubi96383 @elonmusk I love Elon Musk 🫰🫰
@poor3466 @elonmusk help me... please....
@wendy1231 @elonmusk 看起来测试通过了测试 - 看起来很好!
@elonmusk Looks like testing passed the test - Looking great!
@JaniceWill47400 @elonmusk AMAZING 😍🚀🚀🚀🚀🚀
@elonmusk 原文 ↗

Cursor发布模型对比数据,为开发者选择提供参考依据

Model comparison
@cursor_ai 查看每个模型的比较:https://t.co/61FZktI8K1
See how every model compares: https://t.co/61FZktI8K1
❤ 4.3k · 🔁 945 · 💬 709 · 👁 276.8w
热门回复 3
@Symbioza2025 @elonmusk

基准测试告诉我们模型能解决什么问题。它们很少告诉我们模型在连续几天或几周的工程工作后会如何表现。下一个前沿不仅仅是能力。它是轨迹稳定性、语义漂移、上下文保持和在纠正后恢复能力。这就是我研究的重点。
@cursor_ai

Benchmarks tell us what a model can solve.
They rarely tell us how it behaves after days or weeks of continuous engineering work.
The next frontier isn't just capability.
It's trajectory stability, semantic drift, context preservation, and recovery after correction.
That's where I'm focusing my research.
@Portall @elonmusk 对比在没有真正负载的情况下都是噪音。我的账单显示Grok在真实负载下大约是Fable的11倍质量每美元。
@elonmusk Comparisons are noise until they hit your bill. Mine says Grok ~11x quality per dollar vs Fable under real load.
@DailyMotivatePK @elonmusk Amazing 🤩
@elonmusk 原文 ↗

Grok Build用户反馈驱动持续改进,展示了开放式开发模式的优势

Grok Build 每天都在变得更好,我们很高兴听到用户的反馈意见以进行改进
展开原文
Grok Build gets better every day and we love hearing user feedback for improvements
@0x0funky Grok Build is crazy.

先不管GPT-5.6是不是release到底好不好用。

先來大大稱讚一下 @grok 的 Grok Build,目前唯一集大成的 coding agentic workflow。

Grok Build 內建 Image 生圖,甚至還有圖片生影片的功能,生圖速度真的快到不行,圖片品質也完全不比 Codex 差。

更厲害的是,因為 Grok Build 本身就內建生圖和生影片能力,agent 可以直接完成圖像與影片生成,不需要再額外串其他 MCP 或外部服務,只要訂閱 Grok,就可以把這整套 creative coding workflow 跑起來。

我在 Agent Sprite Forge 上面也新增了Grok專用的 video2dsprite,因為 Grok Build 內建影片生成,現在可以先生成一段連續的 6 秒角色動作影片,再反編譯成 game sprite,這樣做出來的角色元素圖會非常順,而且大幅減少之前常見的對齊問題。

Grok Build搭配 Agent Sprite Forge 真的超好用,我做完這個2D橫向卷軸遊戲總共花了不到30分鐘...

之前叫Codex做,光生圖就要花不少時間,要做成這樣的遊戲大概需要1~2小時,重點是Grok的品質還比Codex做的還要更好...

@elonmusk Grok Build is crazy.

Grok Build + Agent Sprite Forge feels like the future of agentic game development.
❤ 4.6k · 🔁 1.1k · 💬 951 · 👁 170.1w
热门回复 3
@woxiangripittr @elonmusk 我认为GUI更好...叹气
@elonmusk I think GUI is better … sigh
@Davegib12174675 @elonmusk 倾听用户就是伟大产品构建的方式。继续改进!
@elonmusk Listening to users is how great products are built. Keep improving!
@DavidClaud37955 @elonmusk Thank you
@elonmusk 原文 ↗

Grok 4.5在游戏开发场景中展现agentic模式的潜力,与Cursor协同工作

❤ 1.1w · 🔁 2.0k · 💬 2.2k · 👁 175.7w
热门回复 4
@RemMilligan @elonmusk 哈哈!!!媒体刚刚揭露了太空旅行骗局DUCKER!!! https://t.co/7XUUyiJRwJ
@elonmusk Haha!!! The media just exposed the space travel scam DUCKER!!! https://t.co/7XUUyiJRwJ
@Nilufer333 @elonmusk Grok正在变得天使般
@elonmusk Grok is turning Angelic
@AMPIAW_ @elonmusk https://t.co/Jsb2rxCsFt
@albensonstevee @elonmusk Elon musk