Op
gpt-5.6-terra Benchmark & Insights
OpenAI OpenAI API
Updated Jul 18, 2026 All models
Sample size
157 runs
in window
Accuracy
66.9%
consensus match · 160d
Confidence
92%
over 162 runs
Window end
Jul 18, 2026
most recent run
Input price
$2.50/MTok
prompt tokens
Output price
$15.00/MTok
completion tokens
Model insights
- 01 The black sheep of the gpt-5.6 trio: 24 points behind luna despite sharing a family, with 74 overratings vs 1 under and about 30 false "unsafe" days.
- 02 At $2.50/$15 it costs mid-tier money for bottom-tier judgment; its late-window behavior (four false "unsafe" calls in the last three weeks) shows no improvement.
Recent forecasts
Date
Conf.
Risk
Safe
Jul 23, 2026
95%
low
Safe
Jul 22, 2026
93%
low
Safe
Jul 21, 2026
93%
high
Unsafe
Jul 20, 2026
97%
low
Safe
Jul 19, 2026
94%
low
Safe
Jul 18, 2026
94%
high
Unsafe
Jul 17, 2026
86%
high
Unsafe
Jul 16, 2026
82%
high
Unsafe
Jul 15, 2026
86%
medium
Unsafe
Jul 14, 2026
90%
medium
Safe