An
claude-sonnet-4-6 Benchmark & Insights
Anthropic Claude API
Updated Jul 31, 2026 All models
Sample size
170 runs
in window
Accuracy
81.6%
consensus match · 206d
Confidence
87%
over 208 runs
Window end
Jul 31, 2026
most recent run
Input price
$3.00/MTok
prompt tokens
Output price
$15.00/MTok
completion tokens
Model insights
- 01 The worst Claude by a wide margin and the family lone pessimist: forty upward misses and none downward.
- 02 The damage is front-loaded, with February and March almost solid divergence before it settles; Sonnet 5 fixes the bias outright and costs less.
Recent forecasts