医生盲检评分显示GPT-5.6 Sol在健康回复质量上超过人类医生,标志着AI在医疗领域的重要突破
展开原文
These models push the frontier of performance per dollar, bringing the best health intelligence to all. The smallest variant, GPT-5.6 Luna, evaluated at the lowest reasoning effort, outperforms GPT-5.5 at the highest reasoning effort–despite costing 25x less. The largest variant, GPT-5.6 Sol, sets a new high bar at cost.
Another especially cool result: physicians found fewer flaws in GPT-5.6 responses than physician-written responses.
We collected diverse tasks that remain difficult for recent OpenAI models, across patient-facing and clinician-facing use cases. We asked speciality-matched physicians to write responses to these tasks with unlimited time and web access. We then asked other physicians to compare responses side-by-side, blinded to their source. Physicians were asked to comment on areas of improvement across five axes: accuracy, communication, completeness, instruction following, and health decision helpfulness. We then reported the fraction of responses across sources rated perfectly across all axes, across 20,000 total axis ratings. GPT-5.6 Sol appeared strongest, although all GPT-5.6 models performed significantly better than physicians.
热门回复 4
基于症状的诊断应该由AI来做。百分之百都是这样。如果原因已经确定,医生就可以介入进行操作程序。
Symptom based diagnosis should be AI . 100% of them.
If the cause is decided , doctors come in for operational procedures.
#keep4o #OpenSource4o #GPT4o
#keep4o #OpenSource4o #GPT4o

