Grok 4.5在Long-Horizon Terminal-Bench上取得领先表现,任务完成数从4.2的0.080跃升至13个
展开原文
Grok 4.5 现在达到了13个,Fable 5达到了12个。单个任务耗费约990万个token、231个episode和85分钟的挂钟时间。这意味着代理在整个过程中都保持计划并完成任务,这种能力在两个月内几乎翻倍。
SpaceXAI 位居榜首,他们宣称4.2倍的输出token效率,这还低估了实际效果。每百万个token投入2美元,产出6美元。在一个任务消耗1000万个token的基准测试中,账单主要由输入重放构成,他们表示Grok 4.5在不到一半的步骤内解决任务,因此每次调用时需要重新发送的累积上下文更少。效率在输入端复合,这也是成本高的方面。
Fable 5落后于Fable一个任务。他们自己的发布图表显示Fable输给Fable 1.1 17分,Grok 4.20在同一排行榜上排名0.080且没有完成任何任务,因此4.5版本的飞跃不是家族特征。
我的理解是4.5版本的飞跃来自于与Cursor一起训练,这是一种实时代理编辑轨迹的数据流,其他人在那个体量上没有这样的数据,而反证据中的任何内容都不能反驳这种能力会累积到下一个检查点。
Grok 4.5 is now at 13, and Fable 5 is at 12. A single task costs around 9.9M tokens, 231 episodes, and 85 minutes of wall clock time. That means agents are holding a plan across all of it and finishing, and that capability nearly doubled in two months.
SpaceXAI is on top, and they marketed the 4.2x output token efficiency, which undersells it. Two dollars in, six out, per million. On a benchmark where one task burns ten million tokens, the bill is dominated by input replay, and they say Grok 4.5 solves tasks in under half the number of steps, so there is less accumulated context to resend on every call. The efficiency compounds on the input side, which is the side that costs money.
Fable 5 is one task behind. Their own launch chart has them losing DeepSWE 1.1 to Fable by 17 points, and Grok 4.20 sits on this same board at 0.080 with zero completions, so whatever happened in 4.5 is not a family trait.
My read is that the 4.5 jump came out of training alongside Cursor, which is a stream of real agentic edit trajectories nobody else has at that volume, and nothing in the counterevidence argues against it compounding into the next checkpoint.
















