@XFreeze Grok 4.5 现在在 Long-Horizon Terminal-Bench 的二元通过率排名中位居第一,明显领先 Claude Fable 5、Claude Opus 4.8 和 GPT-5.6-sol。在最严格的评分标准下——只有在完全解决任务且奖励完美、无错误时才算通过——Grok 4.5 明显领先所有其他模型。这对于实际的编码、自动化和工程工作至关重要,因为长时程终端任务需要的不仅仅是取得部分进展。模型必须保持上下文、从错误中恢复并在数百个步骤跨整个工作流成功完成。
Grok 4.5 now also ranked #1 on the Long-Horizon Terminal-Bench by binary pass rate, outperforming Claude Fable 5, Claude Opus 4.8 and GPT-5.6-sol
Under the strictest scoring metric - where a task counts only if it is fully solved with a perfect reward and zero errors......Grok 4.5 finished clearly ahead of every other tested model
This matters for real-world coding, automation and difficult engineering work because long-horizon terminal tasks require much more than making partial progress
The model has to maintain context, recover from mistakes and successfully complete an entire workflow across hundreds of steps
Grok 4.5 is showing serious strength on complex agentic tasks over extended periods of time
@elonmusk 09 11 01 is a plan by Obama/the Dems Admin who controlled our Congress to send George Bush to a WAR in IRAQ. To steal more of our taxpayers money. Those Muslims Terrorist are sponsored:funded by Dems Admin members to bring them here as a Collage students from Saudi Arabia.
@elonmusk This is really cool. Long-horizon tasks are where a lot of models still struggle, so seeing Grok 4.5 lead on the strict binary metric is impressive. Going to try Grok Build this week 💜
In Snorkel AI’s benchmark, Grok 4.5 combined with Grok Build was tested against GPT 5.5 and Claude Opus 4.8 across nearly 2,000 expert-created workplace tasks involving real documents, spreadsheets, presentations and professional analysis
Grok outperformed both GPT 5.5 and Claude Opus 4.8 overall, while leading by even wider margins across several high-judgment fields:
It also recorded the lowest failure rate across every error category Snorkel measured, including missing analysis, incorrect recommendations, poor structure and missing sources
The most important result is not simply that Grok completed more tasks
It produced better professional deliverables, made fewer critical mistakes and provided more specific, actionable recommendations where competing models often returned generic work
Grok 4.5 is proving that real-world usefulness matters more than benchmark scores alone
BREAKING: Grok 4.5 leads VulcanBench’s new coding benchmark. 🔥
Grok scored 91.3%, solving 21 of 23 real-world software tasks across five languages, beating Claude Fable 5 and GPT-5.6 Sol while also owning the cost-efficiency frontier.
@elonmusk Interesting stance from Haley Stevens on unconditional aid how does this square with oversight on every other foreign program? https://t.co/2rtU1nMYl2
@leanaloving04 @elonmusk Grok目前是最好的选择
@elonmusk Grok is currently the best
@ivanutor @elonmusk 好啊,我等它对学生免费的时候再试用一下,先生。
@elonmusk Yeah, I will try it when it is free for students, sir.
Starlink’s progress in under seven years is totally insane
2019 — The first dedicated Starlink launch carried 60 satellites 2020 — Wider public beta service began 2024 — The first Direct to Cell satellites launched, followed by the first commercial satellite-texting rollout 2025 — Direct to Cell expanded across more carriers, devices and markets 2026 — Starlink reached approximately 10.3 million subscribers across 164 countries, territories and other markets
By March 31, 2026, SpaceX had approximately 9,600 Starlink broadband and mobile satellites in low-Earth orbit
Direct to Cell was already providing satellite-to-mobile texting and voice services to approximately 7.4 million monthly unique devices across roughly 30 countries
What began with a single launch of 60 satellites is becoming a global communications layer connecting homes, businesses, aircraft, ships, government operations, rescue missions and ordinary smartphones
SpaceX is building connectivity infrastructure for the entire planet
❤ 1.0w · 🔁 1.6k · 💬 932 · 👁 184.6w
热门回复 4
@EvasTeslaSPlaid @elonmusk @SpaceX Yes
@mattkennedy777 @elonmusk @SpaceX 我有一个想法,可以扩展 Tesla 的充电网络/电网,同时与教育合作伙伴建立关系,这种方式成本极低且互利互惠。我已经担任学校校长 19 年了。@elonmusk 感谢阅读这条消息。
@elonmusk @SpaceX I have an idea that expands Tesla's charging imprint/grid while leveraging a relationship with the educational partners that's of minimal cost and mutually beneficial. I've been a school principal for 19 years. @elonmusk thanks for reading this.
@elonmusk @SpaceX Mr Elon we had such trouble with Star Link in Mississippi. Wonder why. It did not work well with us. We followed all the directions. UGH.
@elonmusk Analyzing the economic impact of tariffs helps policymakers assess their utility. Data-driven adjustments ensure that trade policies support national growth.
@Ai5670670574913 @elonmusk Come on Man, Let’s go!
IREN AI 云服务需求旺盛,ARR 指引上调
IREN 签约 28 亿美元多年 AI 云服务合同,将 2026 年末 AI 云 ARR 目标从 37 亿提升至 40 亿以上。公司从去年 3MW 自建算力扩容至 480MW,2027 年目标 1.2GW,客户包括超大型科技公司和 AI 开发者。
3.4B ARR is already under contract as of today🤯 you are not bullish enough $IREN will have 12+ Billion ARR by end of 2027 and currently has a marketcap of 12bil. Repeat after me, Undervalued.
@danroberts0101 IREN 与领先的 AI 开发者签署了 28 亿美元的新的多年 AI 云服务合同,并将年末 2026 年 AI 云 ARR 目标从 37 亿提升至 40 亿以上。我们的垂直整合 AI 云平台正在快速扩大。在过去 12 个月中,我们将自建 AI 云容量从大约 3MW 扩展到 480MW,计划在 2027 年达到 1.2GW,客户群扩展到超大型公司、企业和 AI 开发者。我们很自豪能够支持在设计、物理 AI 和机器人、生成媒体、AI 搜索和模型开发方面构建前沿应用的领先公司。
12 months ago we had ~3MW of self-built AI Cloud capacity. Today: 480MW being delivered this year, $2.8bn in new contracts signed, and our 2026 ARR target raised to $4bn+ with ~85% already under contract.
Demand continues to exceed everything we can build. Recent contracts include customer prepayments covering ~45% of the associated GPU capex, with weighted average contract terms of ~4 years across the portfolio.
Data centers. Compute. Software. The three-layer thesis, executing as written: https://t.co/VP7KeVpV3F
$IREN has 7.6bil in cash with a 13bil marketcap and starting Q1 2027 we will have quarterly rev of 1 bil which will become 3 bil quarterly by Q1 2028. Talk about undervalued. Oh and might I add that we are getting 45% prepayments.
@EliteOptions2 My brother in Christ. You pushed $1500 7/17 near ATH and overbought and now you are pushing $700 on oversold with PEG being negative by then a super expensive contracts. I can see a gap fill to $780 for sure but I think you trade the stock not the chart on this one
数据中心版税收入同比增长逾倍,而且现在所有超大型数据中心 CPU 都运行在 Arm 上:Graviton、Axion、Cobalt、Grace,每一个都向 $ARM 支付费用。再加上 AGI CPU 即将在 FY27 年第四季度开始产生第一笔收入,以及 150 亿美元的长期目标,这一切都已经写好了剧本。
At least that’s what the market is pricing after a 46% crash in three weeks. There’s just one problem: the chip everyone is selling hasn’t shipped a single unit yet.
453 to 243 in four weeks. Downgrades, valuation calls, ETF outflows, every bear argument hit at once. But look where the selling stopped: $ARM is backtesting its breakout zone just above 189, the level that capped this stock for two years before the breakout. Price has held it so far.
Growth Catalyst The core business is compounding while the new one loads. $ARM closed FY26 with $4.92B in revenue, up 23%, its third straight year of 20%+ growth. This was capped by a record $1.49B quarter in May.
Data center royalty revenue more than doubled year over year, and roughly half of all hyperscaler CPUs now run on Arm: Graviton, Axion, Cobalt, Grace, every one of them paying $ARM a toll. Layer the AGI CPU on top, with first production revenue hitting in Q4 FY27 and a $15B long-term target, and the setup writes itself.
Two catalysts on deck: $INTC reports July 23 and $ARM wins either way. Strong Intel means CPU demand is hot, weak Intel means the market share is already sitting in $ARM’s pocket. Then $ARM prints July 29.
When a stock retests the level that launched an explosive move, the bounce off it can be just as violent as the breakout. Most traders wait until buying feels comfortable again. By then the discount is gone. Pullbacks like this don’t come often.
@EliteOptions2 RISC-V isn't fantasy: Open-source architecture is attracting massive capital from Meta, Google & Qualcomm to bypass Arm’s licensing long-term.