Go
gemma-4-31b-it Benchmark & Insights
Google Gemini API
Updated Jul 18, 2026 All models
Sample size
157 runs
in window
Accuracy
88.8%
consensus match · 160d
Confidence
100%
over 162 runs
Window end
Jul 18, 2026
most recent run
Input price
$0.13/MTok
prompt tokens
Output price
$0.38/MTok
completion tokens
Model insights
- 01 An outlier among open models: decent risk accuracy but the worst "safe" score in the top half, with ten false "unsafe" alarms on calm days.
- 02 It is firmly "pessimistic" (20 over vs 2 under), yet at $0.13/$0.38 it is one of the cheapest ways to get near-90% agreement.
Notes
Open-weight model; also available free via OpenRouter
Recent forecasts
Date
Conf.
Risk
Safe
Jul 23, 2026
100%
low
Safe
Jul 22, 2026
100%
low
Safe
Jul 21, 2026
100%
high
Unsafe
Jul 20, 2026
100%
low
Safe
Jul 19, 2026
100%
low
Safe
Jul 18, 2026
90%
high
Unsafe
Jul 17, 2026
90%
medium
Unsafe
Jul 16, 2026
100%
medium
Safe
Jul 15, 2026
100%
medium
Safe
Jul 14, 2026
100%
medium
Safe