An
claude-opus-4-6 Benchmark & Insights
Anthropic Claude API
Updated Jul 31, 2026 All models
Sample size
169 runs
in window
Accuracy
90.3%
consensus match · 206d
Confidence
89%
over 207 runs
Window end
Jul 31, 2026
most recent run
Input price
$5.00/MTok
prompt tokens
Output price
$25.00/MTok
completion tokens
Model insights
- 01 Same score as Opus 4.7 but a different shape: it overrated through late February, then flipped to underrating for the rest of the run and never went back.
- 02 Alone among the Claude models it never mislabeled an "unsafe" day as "safe".