Op
gpt-5.6-luna Benchmark & Insights
OpenAI OpenAI API
Updated Jul 18, 2026 All models
Sample size
157 runs
in window
Accuracy
90.0%
consensus match · 160d
Confidence
92%
over 162 runs
Window end
Jul 18, 2026
most recent run
Input price
$1.00/MTok
prompt tokens
Output price
$6.00/MTok
completion tokens
Model insights
- 01 The best of the three gpt-5.6 variants and unusually two-sided: 11 overratings vs 9 underratings, so it swings rather than leans.
- 02 Its bad habit is jumping to "high" and declaring "unsafe" on safe days — five false "unsafe" calls, more than any model above it.
Recent forecasts