Zh
glm-4.7-flash Benchmark & Insights
Zhipu AI Cloudflare Workers AI
Updated Jul 18, 2026 All models
Sample size
156 runs
in window
Accuracy
68.8%
consensus match · 160d
Confidence
97%
over 161 runs
Window end
Jul 18, 2026
most recent run
Input price
$0.06/MTok
prompt tokens
Output price
$0.40/MTok
completion tokens
Model insights
- 01 Cries wolf on roughly one day in five: 34 false "unsafe" calls and 28 "high" ratings against a consensus that saw only 9 "high" days.
- 02 It also managed the worst-case inversion, saying "safe" on the consensus-"unsafe" 2026-07-05.
- 03 Cheap, but the noise never quiets down — July was as bad as February.
Recent forecasts
Date
Conf.
Risk
Safe
Jul 23, 2026
100%
low
Safe
Jul 22, 2026
90%
low
Safe
Jul 21, 2026
95%
high
Unsafe
Jul 20, 2026
100%
low
Safe
Jul 19, 2026
95%
low
Safe
Jul 18, 2026
90%
high
Unsafe
Jul 17, 2026
90%
medium
Unsafe
Jul 16, 2026
90%
medium
Safe
Jul 15, 2026
100%
medium
Safe
Jul 14, 2026
100%
high
Unsafe