Op
gpt-oss-120b Benchmark & Insights
Openai Cloudflare Workers AI
Updated Jul 31, 2026 All models
Sample size
169 runs
in window
Accuracy
59.0%
consensus match · 178d
Confidence
96%
over 179 runs
Window end
Jul 31, 2026
most recent run
Input price
$0.35/MTok
prompt tokens
Output price
$0.75/MTok
completion tokens
Model insights
- 01 It disagrees on more than a third of days and always upward, but the safe flag is the bigger problem — wrong on nearly a third of days, mostly declaring "unsafe" where everything else said "safe".
- 02 Cheap, yet it would fire a needless warning most weeks.
Recent forecasts
Date
Conf.
Risk
Safe
Aug 10, 2026
95%
medium
Safe
Aug 9, 2026
90%
low
Safe
Aug 8, 2026
90%
medium
Unsafe
Aug 7, 2026
90%
medium
Unsafe
Aug 6, 2026
97%
high
Unsafe
Aug 5, 2026
97%
medium
Unsafe
Aug 4, 2026
100%
low
Safe
Aug 3, 2026
97%
high
Unsafe
Aug 2, 2026
97%
medium
Unsafe
Aug 1, 2026
98%
low
Safe