Sol Pro在prinzbench上91/99分,几乎饱和该难度评测,显示出极强的推理能力
展开原文
正如几天前预览的那样,此模型已填满我的基准测试,总分为 91/99。
需要说明的是,prinzbench 包含两个至今无模型能够解决的问题(其中一个需要极其彻底的 50 州研究,可能需要 /goal 模式来解决,另一个则有一个非常棘手的监管审批,无模型曾成功找到)。如果将这两个问题(总计 6 分)放在一边,GPT-5.6 Sol Pro 在 93 个 prinzbench 问题中提供了 91 个正确答案。
OpenAI Pro 模型在 prinzbench 上的表现:
GPT-5.4 Pro(扩展版):79/99
GPT-5.5 Pro(扩展版):82/99
GPT-5.6 Sol Pro:91/99
我的基准测试于 2026 年 1 月发布,并在 2026 年 6 月被填满。这种加速度是真实存在的!
由于此模型的表现,未来 OpenAI Pro 模型将不再在 prinzbench 上进行测试(测试它们没有意义)。
其他 GPT-5.6 模型的基准测试即将推出(很快就会发布)。
As previewed a few days ago, this model has saturated my benchmark, with a total score of 91/99.
For context, prinzbench contains two questions that no model tested to date has ever been able to solve (one requires extremely thorough 50-state research that probably requires /goal mode to solve, and another has a really tricky regulatory approval that no model has ever been able to find). Putting these two questions (which are worth 6 points) aside, GPT-5.6 Sol Pro provided correct responses to 91 out of 93 prinzbench questions.
prinzbench performance for OpenAI's Pro models:
GPT-5.4 Pro (Extended): 79/99
GPT-5.5 Pro (Extended): 82/99
GPT-5.6 Sol Pro: 91/99
My benchmark was released in January 2026 and was saturated in June 2026. The acceleration is real!
As a result of this model's performance, future OpenAI Pro models will no longer be tested on prinzbench (there is no point in testing them).
Benchmarking for other GPT-5.6 models to follow soon(TM).
热门回复 4
#keep4o #OpenSource4o #GPT4o
> withhold the transcript
>>latent memory
>> weight level based memory
>Learn on the fly (so esp can it learn a mario kaizo button sequence or optimize a skill/intuition/tactics esp pref at the level of weights or similar)

