Op

gpt-5.1 Benchmark & Insights

OpenAI OpenAI API
Updated Jul 31, 2026 All models
Sample size
170 runs
in window
Accuracy
87.7%
consensus match · 211d
Confidence
92%
over 213 runs
Window end
Jul 31, 2026
most recent run
Input price
$1.25/MTok
prompt tokens
Output price
$10.00/MTok
completion tokens
Model insights
  • 01 The only strongly optimistic model in the GPT line, and its weakness lands where it costs most: five days it rated a "high" consensus day as "medium" and called it "safe", four of those in July.
  • 02 An early-March run of "low" for "medium" is the second cluster.
Recent forecasts