Sales coaching benchmark
Compare models
Compare quality, response time, and coach cost across the same 50 calls.
Selected model comparison
6 models selected
Green marks stronger results; red marks weaker results. Lower cost and response time are better.
| Model | Score | Avg | Recall | Inst. | Floor | Cost | Time |
|---|---|---|---|---|---|---|---|
| GPT-5.6 Solmax | 90.0 | 89.7 | 88.8 | 91.5 | 80.3 | $0.37 | - |
| GPT-5.6 Terralow | 89.9 | 89.5 | 89.0 | 90.9 | 84.5 | $0.05 | - |
| GPT-5.6 Lunaxhigh | 89.9 | 89.7 | 89.5 | 90.8 | 81.7 | $0.04 | - |
| GPT-5.4xhigh | 89.3 | 89.0 | 88.1 | 90.6 | 79.9 | $0.28 | - |
| Claude Fable 5high | 87.7 | 87.5 | 86.9 | 90.4 | 77.0 | $0.46 | - |
| Claude Opus 4.8medium | 85.6 | 85.8 | 84.7 | 88.1 | 72.1 | $0.14 | - |
Model
GPT-5.6 Solmax$0.37/call
GPT-5.6 Terralow$0.05/call
GPT-5.6 Lunaxhigh$0.04/call
GPT-5.4xhigh$0.28/call
Claude Fable 5high$0.46/call
Claude Opus 4.8medium$0.14/call
Benchmark
90.0
89.9
89.9
89.3
87.7
85.6
Raw average
89.7
89.5
89.7
89.0
87.5
85.8
Answer-key recall
88.8
89.0
89.5
88.1
86.9
84.7
Sales instinct
91.5
90.9
90.8
90.6
90.4
88.1
Downside floor
80.3
84.5
81.7
79.9
77.0
72.1
Coach cost
$0.37
$0.05
$0.04
$0.28
$0.46
$0.14
Coach response
Unavailable
Unavailable
Unavailable
Unavailable
Unavailable
Unavailable
All eight scoring dimensionsCompare all eight judging dimensions across 62 models.
| Model | Overall | Answer-key recall | Evidence grounding | False-positive control | Prioritization | Actionability | Sales instinct | Technical accuracy |
|---|---|---|---|---|---|---|---|---|
gpt-5.6 sol max GPT-5.6 Sol · max | 89.7 | 88.8 | 94.3 | 89.9 | 89.5 | 94.3 | 91.5 | 91.3 |
gpt-5.6 terra max GPT-5.6 Terra · max | 89.6 | 87.9 | 94.0 | 91.1 | 89.9 | 94.2 | 91.1 | 91.6 |
gpt-5.6 luna max GPT-5.6 Luna · max | 89.7 | 89.6 | 93.9 | 89.2 | 88.7 | 93.8 | 90.8 | 91.4 |
gpt-5.6 luna xhigh GPT-5.6 Luna · xhigh | 89.7 | 89.5 | 93.7 | 89.7 | 88.4 | 93.7 | 90.8 | 91.5 |
gpt-5.6 terra low GPT-5.6 Terra · low | 89.5 | 89.0 | 93.9 | 89.4 | 89.2 | 93.4 | 90.9 | 91.7 |
gpt-5.6 sol low GPT-5.6 Sol · low | 89.5 | 89.1 | 93.9 | 89.3 | 88.4 | 93.8 | 91.2 | 91.3 |
gpt-5.6 sol xhigh GPT-5.6 Sol · xhigh | 89.5 | 88.8 | 94.2 | 89.5 | 88.4 | 94.1 | 91.3 | 91.4 |
gpt-5.6 terra xhigh GPT-5.6 Terra · xhigh | 89.2 | 87.5 | 94.2 | 91.0 | 89.3 | 93.9 | 91.1 | 91.7 |
gpt-5.6 terra high GPT-5.6 Terra · high | 89.3 | 88.3 | 93.7 | 89.6 | 89.2 | 93.6 | 91.0 | 91.6 |
gpt-5.6 sol none GPT-5.6 Sol · none | 89.3 | 88.5 | 93.8 | 89.1 | 88.9 | 93.8 | 90.8 | 91.2 |
gpt-5.6 sol high GPT-5.6 Sol · high | 89.2 | 88.9 | 93.8 | 89.0 | 88.1 | 93.8 | 90.7 | 90.9 |
gpt-5.4 xhigh GPT-5.4 · xhigh | 89.0 | 88.1 | 93.9 | 89.6 | 88.2 | 92.9 | 90.6 | 91.5 |
gpt-5.6 sol medium GPT-5.6 Sol · medium | 89.0 | 88.1 | 93.4 | 89.0 | 88.2 | 93.9 | 90.8 | 91.3 |
gpt-5.4 high GPT-5.4 · high | 89.0 | 87.6 | 93.4 | 89.9 | 88.0 | 92.8 | 90.6 | 91.3 |
gpt-5.5 medium GPT-5.5 · medium | 88.8 | 88.5 | 93.7 | 88.0 | 87.8 | 93.3 | 90.6 | 91.2 |
gpt-5.5 xhigh GPT-5.5 · xhigh | 89.0 | 88.5 | 93.9 | 89.2 | 87.6 | 93.3 | 89.9 | 91.3 |
gpt-5.6 terra none GPT-5.6 Terra · none | 88.8 | 87.2 | 93.7 | 88.6 | 88.8 | 93.7 | 90.6 | 91.3 |
gpt-5.6 terra medium GPT-5.6 Terra · medium | 88.9 | 87.3 | 93.7 | 88.6 | 88.6 | 93.2 | 90.4 | 91.3 |
gpt-5.5 high GPT-5.5 · high | 88.6 | 88.4 | 93.6 | 89.0 | 86.7 | 93.4 | 90.2 | 91.1 |
gpt-5.6 luna medium GPT-5.6 Luna · medium | 88.6 | 88.4 | 92.7 | 87.1 | 87.4 | 93.1 | 89.8 | 91.2 |
gpt-5.6 luna high GPT-5.6 Luna · high | 88.6 | 87.7 | 92.9 | 87.7 | 87.4 | 93.4 | 89.9 | 91.3 |
gpt-5.6 luna low GPT-5.6 Luna · low | 88.5 | 87.5 | 93.1 | 88.1 | 87.5 | 93.0 | 89.8 | 90.8 |
gpt-5.4 medium GPT-5.4 · medium | 88.3 | 86.5 | 93.3 | 88.8 | 87.2 | 92.3 | 90.0 | 90.9 |
gpt-5.5 none GPT-5.5 · none | 88.1 | 86.8 | 93.5 | 87.9 | 86.9 | 92.9 | 90.0 | 91.1 |
gpt-5.6 luna none GPT-5.6 Luna · none | 87.7 | 87.3 | 92.3 | 86.7 | 86.8 | 92.8 | 89.2 | 90.6 |
gpt-5.5 low GPT-5.5 · low | 87.7 | 87.0 | 92.9 | 87.1 | 86.3 | 92.4 | 89.2 | 90.4 |
fable 5 high Claude Fable 5 · high | 87.5 | 86.9 | 90.1 | 83.5 | 86.7 | 93.0 | 90.4 | 89.6 |
gpt-5.4 low GPT-5.4 · low | 87.4 | 86.0 | 92.1 | 86.8 | 86.3 | 91.8 | 89.0 | 90.4 |
gpt-5.4 none GPT-5.4 · none | 87.4 | 85.8 | 92.7 | 86.7 | 86.7 | 91.4 | 88.7 | 90.5 |
opus 4.7 max Claude Opus 4.7 · max | 87.3 | 86.9 | 90.6 | 83.5 | 85.9 | 92.7 | 89.1 | 89.3 |
kimi k3 max Kimi K3 · max | 86.6 | 85.3 | 90.6 | 84.0 | 86.0 | 93.1 | 89.9 | 89.1 |
opus 5 max Claude Opus 5 · max | 86.6 | 88.3 | 89.4 | 80.3 | 84.3 | 93.4 | 89.6 | 88.6 |
opus 5 xhigh Claude Opus 5 · xhigh | 86.6 | 88.3 | 88.6 | 80.2 | 83.9 | 93.7 | 89.6 | 88.8 |
opus 4.7 high Claude Opus 4.7 · high | 86.8 | 86.1 | 89.0 | 82.4 | 85.3 | 92.1 | 88.8 | 88.8 |
muse spark 1.1 high Muse Spark 1.1 · high | 86.4 | 86.1 | 89.4 | 83.1 | 85.8 | 90.4 | 88.4 | 88.9 |
muse spark 1.1 medium Muse Spark 1.1 · medium | 86.2 | 85.0 | 89.2 | 83.0 | 86.5 | 90.2 | 88.9 | 87.8 |
opus 5 medium Claude Opus 5 · medium | 86.3 | 86.5 | 88.8 | 80.5 | 84.7 | 93.2 | 89.4 | 88.3 |
opus 5 low Claude Opus 5 · low | 85.9 | 85.5 | 89.4 | 81.3 | 84.2 | 92.6 | 89.2 | 88.1 |
muse spark 1.1 minimal Muse Spark 1.1 · minimal | 85.7 | 83.8 | 88.4 | 82.9 | 85.8 | 88.9 | 87.9 | 87.9 |
opus 5 high Claude Opus 5 · high | 85.5 | 86.0 | 88.5 | 80.3 | 83.6 | 92.9 | 88.8 | 88.0 |
muse spark 1.1 low Muse Spark 1.1 · low | 85.4 | 83.2 | 89.7 | 84.0 | 85.2 | 90.1 | 88.1 | 88.6 |
opus 4.8 medium Claude Opus 4.8 · medium | 85.8 | 84.7 | 89.7 | 81.8 | 83.8 | 90.5 | 88.1 | 88.2 |
opus 4.7 medium Claude Opus 4.7 · medium | 85.6 | 84.2 | 89.1 | 82.6 | 83.9 | 91.0 | 88.2 | 87.8 |
opus 4.7 xhigh Claude Opus 4.7 · xhigh | 85.6 | 84.6 | 89.0 | 82.3 | 83.9 | 91.6 | 87.8 | 88.0 |
opus 4.7 low Claude Opus 4.7 · low | 85.6 | 84.1 | 89.6 | 82.7 | 84.1 | 90.9 | 87.6 | 88.5 |
opus 4.8 max Claude Opus 4.8 · max | 85.4 | 85.4 | 88.6 | 80.9 | 83.8 | 91.4 | 87.6 | 88.2 |
opus 4.8 xhigh Claude Opus 4.8 · xhigh | 85.2 | 84.8 | 88.7 | 81.3 | 83.8 | 90.4 | 87.6 | 88.3 |
opus 4.8 high Claude Opus 4.8 · high | 84.9 | 83.6 | 89.2 | 81.1 | 83.7 | 90.5 | 87.1 | 88.6 |
sonnet 4.6 Claude Sonnet 4.6 · default | 84.6 | 83.8 | 87.0 | 79.6 | 83.2 | 91.2 | 87.6 | 86.6 |
sonnet 5 Claude Sonnet 5 · default | 84.6 | 83.8 | 88.5 | 81.0 | 82.3 | 89.3 | 85.9 | 87.8 |
opus 4.8 low Claude Opus 4.8 · low | 84.0 | 82.8 | 88.5 | 80.4 | 81.9 | 89.3 | 85.5 | 87.5 |
glm 5.2 GLM 5.2 · default | 84.0 | 82.2 | 88.4 | 80.8 | 81.6 | 89.4 | 85.7 | 87.3 |
deepseek v4 pro DeepSeek V4 Pro · default | 83.5 | 81.9 | 86.9 | 79.5 | 81.9 | 88.5 | 84.9 | 86.6 |
gemini 3.6 flash minimal Gemini 3.6 Flash · minimal | 81.6 | 79.1 | 85.6 | 77.7 | 79.3 | 84.9 | 82.8 | 85.2 |
gemini 3.6 flash medium Gemini 3.6 Flash · medium | 79.8 | 76.6 | 86.6 | 78.5 | 77.6 | 84.0 | 81.5 | 85.4 |
gemini 3.6 flash high Gemini 3.6 Flash · high | 79.3 | 75.9 | 85.0 | 75.7 | 76.7 | 83.6 | 80.8 | 84.4 |
gemini 3.1 pro preview Gemini 3.1 Pro Preview · default | 78.9 | 74.5 | 86.2 | 78.0 | 76.9 | 84.1 | 81.4 | 84.1 |
gemini 3.6 flash low Gemini 3.6 Flash · low | 78.5 | 75.0 | 84.8 | 75.2 | 75.6 | 81.2 | 79.7 | 85.1 |
gemini 3.5 flash lite high Gemini 3.5 Flash-Lite · high | 77.4 | 73.1 | 84.0 | 75.5 | 74.1 | 79.2 | 78.9 | 83.9 |
gemini 3.5 flash lite minimal Gemini 3.5 Flash-Lite · minimal | 74.8 | 71.2 | 81.9 | 72.2 | 70.0 | 75.8 | 75.1 | 83.0 |
gemini 3.5 flash lite medium Gemini 3.5 Flash-Lite · medium | 74.6 | 70.3 | 82.0 | 73.3 | 69.7 | 75.1 | 74.9 | 82.2 |
gemini 3.5 flash lite low Gemini 3.5 Flash-Lite · low | 72.7 | 68.2 | 80.5 | 70.2 | 68.0 | 72.9 | 72.4 | 81.7 |
| Mean | 85.9 | 84.7 | 90.3 | 83.9 | 84.5 | 90.5 | 87.7 | 88.9 |
Coach cost and response timeWhat each coach costs and how long it takes to answer. Call generation and judging are excluded.
- Estimated 50-call total
- $359.82
- Median / call
- $0.10
- Models
- 62
Model
Score
Cost / call
Response time
50-call total
Input / response
Reasoning
Input / output rate
gemini 3.5 flash lite low
Gemini 3.5 Flash-Lite · low
Raw provider cost · 50 calls
Score
72.7
Cost / call
$0.0047
Response time
4.9s
50-call total
$0.23
Input / response
4,918 / 1,276
Reasoning
0
Input / output rate
$0.30 / $2.50
deepseek v4 pro
DeepSeek V4 Pro · default
Estimated from saved response
Score
83.5
Cost / call
$0.0047
Response time
—
50-call total
$0.24
Input / response
4,559 / 3,180
Reasoning
—
Input / output rate
$0.43 / $0.87
gemini 3.5 flash lite minimal
Gemini 3.5 Flash-Lite · minimal
Raw provider cost · 50 calls
Score
74.8
Cost / call
$0.0053
Response time
5.6s
50-call total
$0.26
Input / response
4,918 / 1,513
Reasoning
0
Input / output rate
$0.30 / $2.50
gemini 3.5 flash lite medium
Gemini 3.5 Flash-Lite · medium
Raw provider cost · 50 calls
Score
74.6
Cost / call
$0.0053
Response time
5.7s
50-call total
$0.27
Input / response
4,918 / 1,332
Reasoning
208
Input / output rate
$0.30 / $2.50
gemini 3.5 flash lite high
Gemini 3.5 Flash-Lite · high
Raw provider cost · 50 calls
Score
77.4
Cost / call
$0.01
Response time
11s
50-call total
$0.53
Input / response
4,918 / 1,421
Reasoning
2,219
Input / output rate
$0.30 / $2.50
gemini 3.6 flash low
Gemini 3.6 Flash · low
Raw provider cost · 50 calls
Score
78.5
Cost / call
$0.02
Response time
8.8s
50-call total
$0.99
Input / response
4,918 / 1,533
Reasoning
129
Input / output rate
$1.50 / $7.50
gpt-5.6 luna low
GPT-5.6 Luna · low
Measured on 1 call
Score
88.5
Cost / call
$0.02
Response time
—
50-call total
$1.00
Input / response
2,879 / 2,843
Reasoning
21
Input / output rate
$1.00 / $6.00
muse spark 1.1 minimal
Muse Spark 1.1 · minimal
Measured on 1 call
Score
85.7
Cost / call
$0.02
Response time
—
50-call total
$1.03
Input / response
2,554 / 2,557
Reasoning
1,533
Input / output rate
$1.25 / $4.25
gpt-5.6 luna none
GPT-5.6 Luna · none
Measured on 1 call
Score
87.7
Cost / call
$0.02
Response time
—
50-call total
$1.03
Input / response
2,879 / 2,970
Reasoning
0
Input / output rate
$1.00 / $6.00
gpt-5.6 luna medium
GPT-5.6 Luna · medium
Measured on 1 call
Score
88.6
Cost / call
$0.02
Response time
—
50-call total
$1.04
Input / response
2,879 / 2,936
Reasoning
67
Input / output rate
$1.00 / $6.00
gemini 3.6 flash minimal
Gemini 3.6 Flash · minimal
Raw provider cost · 50 calls
Score
81.6
Cost / call
$0.02
Response time
9.7s
50-call total
$1.07
Input / response
4,918 / 1,871
Reasoning
0
Input / output rate
$1.50 / $7.50
muse spark 1.1 low
Muse Spark 1.1 · low
Measured on 1 call
Score
85.4
Cost / call
$0.02
Response time
—
50-call total
$1.10
Input / response
2,554 / 2,854
Reasoning
1,587
Input / output rate
$1.25 / $4.25
muse spark 1.1 medium
Muse Spark 1.1 · medium
Measured on 1 call
Score
86.2
Cost / call
$0.02
Response time
—
50-call total
$1.25
Input / response
2,554 / 2,680
Reasoning
2,451
Input / output rate
$1.25 / $4.25
muse spark 1.1 high
Muse Spark 1.1 · high
Measured on 1 call
Score
86.4
Cost / call
$0.03
Response time
—
50-call total
$1.28
Input / response
2,554 / 2,539
Reasoning
2,728
Input / output rate
$1.25 / $4.25
glm 5.2
GLM 5.2 · default
Estimated from saved response
Score
84.0
Cost / call
$0.03
Response time
—
50-call total
$1.30
Input / response
4,559 / 4,455
Reasoning
—
Input / output rate
$1.40 / $4.40
gpt-5.6 luna high
GPT-5.6 Luna · high
Measured on 1 call
Score
88.6
Cost / call
$0.03
Response time
—
50-call total
$1.42
Input / response
2,879 / 3,560
Reasoning
678
Input / output rate
$1.00 / $6.00
gemini 3.6 flash medium
Gemini 3.6 Flash · medium
Raw provider cost · 50 calls
Score
79.8
Cost / call
$0.03
Response time
15s
50-call total
$1.48
Input / response
4,918 / 1,699
Reasoning
1,264
Input / output rate
$1.50 / $7.50
gemini 3.1 pro preview
Gemini 3.1 Pro Preview · default
Estimated from saved response
Score
78.9
Cost / call
$0.03
Response time
—
50-call total
$1.60
Input / response
4,559 / 1,900
Reasoning
—
Input / output rate
$2.00 / $12.00
gpt-5.6 luna xhigh
GPT-5.6 Luna · xhigh
Measured on 1 call
Score
89.7
Cost / call
$0.04
Response time
—
50-call total
$2.09
Input / response
2,879 / 3,644
Reasoning
2,837
Input / output rate
$1.00 / $6.00
gemini 3.6 flash high
Gemini 3.6 Flash · high
Raw provider cost · 50 calls
Score
79.3
Cost / call
$0.04
Response time
22s
50-call total
$2.11
Input / response
4,918 / 1,849
Reasoning
2,803
Input / output rate
$1.50 / $7.50
gpt-5.4 none
GPT-5.4 · none
Measured on 1 call
Score
87.4
Cost / call
$0.05
Response time
—
50-call total
$2.40
Input / response
2,879 / 2,726
Reasoning
0
Input / output rate
$2.50 / $15.00
gpt-5.4 low
GPT-5.4 · low
Measured on 1 call
Score
87.4
Cost / call
$0.05
Response time
—
50-call total
$2.41
Input / response
2,879 / 2,728
Reasoning
10
Input / output rate
$2.50 / $15.00
gpt-5.6 terra low
GPT-5.6 Terra · low
Measured on 1 call
Score
89.5
Cost / call
$0.05
Response time
—
50-call total
$2.49
Input / response
2,879 / 2,787
Reasoning
57
Input / output rate
$2.50 / $15.00
sonnet 5
Claude Sonnet 5 · default
Estimated from saved response
Score
84.6
Cost / call
$0.05
Response time
—
50-call total
$2.50
Input / response
4,559 / 4,079
Reasoning
—
Input / output rate
$2.00 / $10.00
gpt-5.6 terra medium
GPT-5.6 Terra · medium
Measured on 1 call
Score
88.9
Cost / call
$0.05
Response time
—
50-call total
$2.65
Input / response
2,879 / 3,014
Reasoning
44
Input / output rate
$2.50 / $15.00
gpt-5.6 terra none
GPT-5.6 Terra · none
Measured on 1 call
Score
88.8
Cost / call
$0.05
Response time
—
50-call total
$2.66
Input / response
2,879 / 3,067
Reasoning
0
Input / output rate
$2.50 / $15.00
gpt-5.4 medium
GPT-5.4 · medium
Measured on 1 call
Score
88.3
Cost / call
$0.06
Response time
—
50-call total
$2.78
Input / response
2,879 / 2,799
Reasoning
430
Input / output rate
$2.50 / $15.00
gpt-5.6 terra high
GPT-5.6 Terra · high
Measured on 1 call
Score
89.3
Cost / call
$0.06
Response time
—
50-call total
$2.90
Input / response
2,879 / 3,234
Reasoning
159
Input / output rate
$2.50 / $15.00
gpt-5.6 luna max
GPT-5.6 Luna · max
Measured on 1 call
Score
89.7
Cost / call
$0.06
Response time
—
50-call total
$3.19
Input / response
2,879 / 3,557
Reasoning
6,588
Input / output rate
$1.00 / $6.00
gpt-5.6 terra xhigh
GPT-5.6 Terra · xhigh
Measured on 1 call
Score
89.2
Cost / call
$0.08
Response time
—
50-call total
$3.80
Input / response
2,879 / 3,041
Reasoning
1,552
Input / output rate
$2.50 / $15.00
gpt-5.6 terra max
GPT-5.6 Terra · max
Measured on 1 call
Score
89.6
Cost / call
$0.10
Response time
—
50-call total
$5.02
Input / response
2,879 / 3,620
Reasoning
2,588
Input / output rate
$2.50 / $15.00
gpt-5.4 high
GPT-5.4 · high
Measured on 1 call
Score
89.0
Cost / call
$0.10
Response time
—
50-call total
$5.19
Input / response
2,879 / 3,045
Reasoning
3,389
Input / output rate
$2.50 / $15.00
sonnet 4.6
Claude Sonnet 4.6 · default
Estimated from saved response
Score
84.6
Cost / call
$0.10
Response time
—
50-call total
$5.21
Input / response
4,559 / 6,034
Reasoning
—
Input / output rate
$3.00 / $15.00
gpt-5.6 sol low
GPT-5.6 Sol · low
Measured on 1 call
Score
89.5
Cost / call
$0.12
Response time
—
50-call total
$5.96
Input / response
2,879 / 3,377
Reasoning
116
Input / output rate
$5.00 / $30.00
gpt-5.5 medium
GPT-5.5 · medium
Measured on 1 call
Score
88.8
Cost / call
$0.12
Response time
—
50-call total
$6.03
Input / response
2,879 / 3,474
Reasoning
64
Input / output rate
$5.00 / $30.00
opus 4.8 low
Claude Opus 4.8 · low
Measured on 1 call
Score
84.0
Cost / call
$0.12
Response time
—
50-call total
$6.16
Input / response
5,993 / 3,729
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.6 sol high
GPT-5.6 Sol · high
Measured on 1 call
Score
89.2
Cost / call
$0.12
Response time
—
50-call total
$6.17
Input / response
2,879 / 3,451
Reasoning
180
Input / output rate
$5.00 / $30.00
gpt-5.6 sol none
GPT-5.6 Sol · none
Measured on 1 call
Score
89.3
Cost / call
$0.12
Response time
—
50-call total
$6.21
Input / response
2,879 / 3,660
Reasoning
0
Input / output rate
$5.00 / $30.00
gpt-5.5 none
GPT-5.5 · none
Measured on 1 call
Score
88.1
Cost / call
$0.13
Response time
—
50-call total
$6.26
Input / response
2,879 / 3,694
Reasoning
0
Input / output rate
$5.00 / $30.00
gpt-5.6 sol medium
GPT-5.6 Sol · medium
Measured on 1 call
Score
89.0
Cost / call
$0.13
Response time
—
50-call total
$6.31
Input / response
2,879 / 3,659
Reasoning
66
Input / output rate
$5.00 / $30.00
opus 4.7 low
Claude Opus 4.7 · low
Measured on 1 call
Score
85.6
Cost / call
$0.13
Response time
—
50-call total
$6.67
Input / response
5,998 / 4,140
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.5 low
GPT-5.5 · low
Measured on 1 call
Score
87.7
Cost / call
$0.14
Response time
—
50-call total
$6.80
Input / response
2,879 / 4,051
Reasoning
0
Input / output rate
$5.00 / $30.00
gpt-5.5 high
GPT-5.5 · high
Measured on 1 call
Score
88.6
Cost / call
$0.14
Response time
—
50-call total
$7.01
Input / response
2,879 / 3,679
Reasoning
516
Input / output rate
$5.00 / $30.00
opus 4.8 medium
Claude Opus 4.8 · medium
Measured on 1 call
Score
85.8
Cost / call
$0.14
Response time
—
50-call total
$7.18
Input / response
5,993 / 4,548
Reasoning
0
Input / output rate
$5.00 / $25.00
kimi k3 max
Kimi K3 · max
Raw provider cost · 50 calls
Score
86.6
Cost / call
$0.15
Response time
—
50-call total
$7.32
Input / response
4,485 / 5,322
Reasoning
3,801
Input / output rate
$3.00 / $15.00
opus 4.8 high
Claude Opus 4.8 · high
Measured on 1 call
Score
84.9
Cost / call
$0.15
Response time
—
50-call total
$7.69
Input / response
5,993 / 4,956
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.8 xhigh
Claude Opus 4.8 · xhigh
Measured on 1 call
Score
85.2
Cost / call
$0.16
Response time
—
50-call total
$8.17
Input / response
5,993 / 5,341
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.6 sol xhigh
GPT-5.6 Sol · xhigh
Measured on 1 call
Score
89.5
Cost / call
$0.17
Response time
—
50-call total
$8.35
Input / response
2,879 / 3,681
Reasoning
1,407
Input / output rate
$5.00 / $30.00
opus 4.7 high
Claude Opus 4.7 · high
Measured on 1 call
Score
86.8
Cost / call
$0.17
Response time
—
50-call total
$8.60
Input / response
6,244 / 5,630
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.7 xhigh
Claude Opus 4.7 · xhigh
Measured on 1 call
Score
85.6
Cost / call
$0.17
Response time
—
50-call total
$8.67
Input / response
5,998 / 5,737
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.8 max
Claude Opus 4.8 · max
Measured on 1 call
Score
85.4
Cost / call
$0.18
Response time
—
50-call total
$9.15
Input / response
5,992 / 6,118
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 5 low
Claude Opus 5 · low
Raw provider cost · 50 calls
Score
85.9
Cost / call
$0.20
Response time
1m 30s
50-call total
$10.01
Input / response
7,939 / 6,421
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.5 xhigh
GPT-5.5 · xhigh
Measured on 1 call
Score
89.0
Cost / call
$0.20
Response time
—
50-call total
$10.16
Input / response
2,879 / 3,704
Reasoning
2,588
Input / output rate
$5.00 / $30.00
opus 4.7 max
Claude Opus 4.7 · max
Measured on 1 call
Score
87.3
Cost / call
$0.21
Response time
—
50-call total
$10.34
Input / response
5,997 / 7,071
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 5 medium
Claude Opus 5 · medium
Raw provider cost · 50 calls
Score
86.3
Cost / call
$0.25
Response time
1m 59s
50-call total
$12.53
Input / response
7,939 / 8,439
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.4 xhigh
GPT-5.4 · xhigh
Measured on 1 call
Score
89.0
Cost / call
$0.28
Response time
—
50-call total
$13.90
Input / response
2,879 / 3,690
Reasoning
14,369
Input / output rate
$2.50 / $15.00
opus 5 high
Claude Opus 5 · high
Raw provider cost · 50 calls
Score
85.5
Cost / call
$0.30
Response time
2m 27s
50-call total
$14.86
Input / response
7,939 / 10,300
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 4.7 medium
Claude Opus 4.7 · medium
Measured on 1 call
Score
85.6
Cost / call
$0.33
Response time
—
50-call total
$16.64
Input / response
12,353 / 10,843
Reasoning
0
Input / output rate
$5.00 / $25.00
opus 5 xhigh
Claude Opus 5 · xhigh
Raw provider cost · 50 calls
Score
86.6
Cost / call
$0.35
Response time
2m 53s
50-call total
$17.36
Input / response
7,939 / 12,298
Reasoning
0
Input / output rate
$5.00 / $25.00
gpt-5.6 sol max
GPT-5.6 Sol · max
Measured on 1 call
Score
89.7
Cost / call
$0.37
Response time
—
50-call total
$18.72
Input / response
2,879 / 4,041
Reasoning
7,959
Input / output rate
$5.00 / $30.00
opus 5 max
Claude Opus 5 · max
Raw provider cost · 50 calls
Score
86.6
Cost / call
$0.38
Response time
3m 13s
50-call total
$19.20
Input / response
7,939 / 13,772
Reasoning
0
Input / output rate
$5.00 / $25.00
fable 5 high
Claude Fable 5 · high
Measured on 1 call
Score
87.5
Cost / call
$0.46
Response time
—
50-call total
$22.85
Input / response
5,993 / 7,943
Reasoning
0
Input / output rate
$10.00 / $50.00
Costs for 57 of 62 models use recorded provider inference usage or a controlled same-call measurement, priced at uncached list rates. Gateway markup and reporting charges are excluded. Response time is shown where it was recorded with the coach run. 5 models use an estimate from the saved coaching response because exact token usage was not recorded.
Performance by transcript sourceCompare performance on the 25 GPT-generated calls and 25 Sonnet-generated calls.
| Model | Overall | GPT-generated n=25 | Sonnet-generated n=25 |
|---|---|---|---|
gpt-5.6 sol max GPT-5.6 Sol | 89.7 | 93.0+3.3 | 86.5-3.3 |
gpt-5.6 terra max GPT-5.6 Terra | 89.6 | 93.4+3.8 | 85.8-3.8 |
gpt-5.6 luna max GPT-5.6 Luna | 89.7 | 92.4+2.8 | 86.9-2.8 |
gpt-5.6 luna xhigh GPT-5.6 Luna | 89.7 | 92.7+3.1 | 86.6-3.1 |
gpt-5.6 terra low GPT-5.6 Terra | 89.5 | 92.6+3.1 | 86.4-3.1 |
gpt-5.6 sol low GPT-5.6 Sol | 89.5 | 92.9+3.4 | 86.1-3.4 |
gpt-5.6 sol xhigh GPT-5.6 Sol | 89.5 | 93.1+3.6 | 86.0-3.6 |
gpt-5.6 terra xhigh GPT-5.6 Terra | 89.2 | 92.6+3.4 | 85.8-3.4 |
gpt-5.6 terra high GPT-5.6 Terra | 89.3 | 92.6+3.3 | 86.1-3.3 |
gpt-5.6 sol none GPT-5.6 Sol | 89.3 | 92.6+3.2 | 86.1-3.2 |
gpt-5.6 sol high GPT-5.6 Sol | 89.2 | 92.7+3.5 | 85.7-3.5 |
gpt-5.4 xhigh GPT-5.4 | 89.0 | 92.0+3.0 | 86.0-3.0 |
gpt-5.6 sol medium GPT-5.6 Sol | 89.0 | 92.4+3.4 | 85.6-3.4 |
gpt-5.4 high GPT-5.4 | 89.0 | 92.0+3.1 | 85.9-3.1 |
gpt-5.5 medium GPT-5.5 | 88.8 | 91.7+2.8 | 86.0-2.8 |
gpt-5.5 xhigh GPT-5.5 | 89.0 | 92.0+3.0 | 86.0-3.0 |
gpt-5.6 terra none GPT-5.6 Terra | 88.8 | 92.1+3.3 | 85.5-3.3 |
gpt-5.6 terra medium GPT-5.6 Terra | 88.9 | 92.4+3.5 | 85.5-3.5 |
gpt-5.5 high GPT-5.5 | 88.6 | 91.7+3.2 | 85.4-3.2 |
gpt-5.6 luna medium GPT-5.6 Luna | 88.6 | 91.2+2.6 | 86.0-2.6 |
gpt-5.6 luna high GPT-5.6 Luna | 88.6 | 91.4+2.9 | 85.7-2.9 |
gpt-5.6 luna low GPT-5.6 Luna | 88.5 | 91.1+2.6 | 86.0-2.6 |
gpt-5.4 medium GPT-5.4 | 88.3 | 90.9+2.6 | 85.6-2.6 |
gpt-5.5 none GPT-5.5 | 88.1 | 92.0+3.9 | 84.3-3.9 |
gpt-5.6 luna none GPT-5.6 Luna | 87.7 | 91.2+3.4 | 84.3-3.4 |
gpt-5.5 low GPT-5.5 | 87.7 | 90.8+3.1 | 84.6-3.1 |
fable 5 high Claude Fable 5 | 87.5 | 91.0+3.4 | 84.1-3.4 |
gpt-5.4 low GPT-5.4 | 87.4 | 90.3+3.0 | 84.4-3.0 |
gpt-5.4 none GPT-5.4 | 87.4 | 90.8+3.5 | 83.9-3.5 |
opus 4.7 max Claude Opus 4.7 | 87.3 | 90.2+3.0 | 84.3-3.0 |
kimi k3 max Kimi K3 | 86.6 | 90.6+4.0 | 82.6-4.0 |
opus 5 max Claude Opus 5 | 86.6 | 89.2+2.6 | 84.1-2.6 |
opus 5 xhigh Claude Opus 5 | 86.6 | 88.9+2.3 | 84.2-2.3 |
opus 4.7 high Claude Opus 4.7 | 86.8 | 89.6+2.7 | 84.1-2.7 |
muse spark 1.1 high Muse Spark 1.1 | 86.4 | 88.8+2.4 | 84.0-2.4 |
muse spark 1.1 medium Muse Spark 1.1 | 86.2 | 88.1+1.9 | 84.3-1.9 |
opus 5 medium Claude Opus 5 | 86.3 | 89.8+3.5 | 82.8-3.5 |
opus 5 low Claude Opus 5 | 85.9 | 89.2+3.3 | 82.6-3.3 |
muse spark 1.1 minimal Muse Spark 1.1 | 85.7 | 87.3+1.7 | 84.0-1.7 |
opus 5 high Claude Opus 5 | 85.5 | 89.3+3.8 | 81.7-3.8 |
muse spark 1.1 low Muse Spark 1.1 | 85.4 | 88.6+3.3 | 82.1-3.3 |
opus 4.8 medium Claude Opus 4.8 | 85.8 | 89.2+3.4 | 82.4-3.4 |
opus 4.7 medium Claude Opus 4.7 | 85.6 | 89.0+3.3 | 82.3-3.3 |
opus 4.7 xhigh Claude Opus 4.7 | 85.6 | 89.4+3.8 | 81.8-3.8 |
opus 4.7 low Claude Opus 4.7 | 85.6 | 87.6+2.0 | 83.6-2.0 |
opus 4.8 max Claude Opus 4.8 | 85.4 | 88.6+3.1 | 82.3-3.1 |
opus 4.8 xhigh Claude Opus 4.8 | 85.2 | 88.1+2.9 | 82.4-2.9 |
opus 4.8 high Claude Opus 4.8 | 84.9 | 89.4+4.5 | 80.4-4.5 |
sonnet 4.6 Claude Sonnet 4.6 | 84.6 | 88.8+4.3 | 80.3-4.3 |
sonnet 5 Claude Sonnet 5 | 84.6 | 87.4+2.8 | 81.8-2.8 |
opus 4.8 low Claude Opus 4.8 | 84.0 | 88.0+4.0 | 80.0-4.0 |
glm 5.2 GLM 5.2 | 84.0 | 86.1+2.1 | 81.8-2.1 |
deepseek v4 pro DeepSeek V4 Pro | 83.5 | 85.8+2.3 | 81.2-2.3 |
gemini 3.6 flash minimal Gemini 3.6 Flash | 81.6 | 84.2+2.6 | 79.0-2.6 |
gemini 3.6 flash medium Gemini 3.6 Flash | 79.8 | 83.7+3.9 | 76.0-3.9 |
gemini 3.6 flash high Gemini 3.6 Flash | 79.3 | 83.2+4.0 | 75.3-4.0 |
gemini 3.1 pro preview Gemini 3.1 Pro Preview | 78.9 | 82.4+3.5 | 75.4-3.5 |
gemini 3.6 flash low Gemini 3.6 Flash | 78.5 | 82.4+3.8 | 74.7-3.8 |
gemini 3.5 flash lite high Gemini 3.5 Flash-Lite | 77.4 | 80.2+2.8 | 74.6-2.8 |
gemini 3.5 flash lite minimal Gemini 3.5 Flash-Lite | 74.8 | 78.5+3.7 | 71.1-3.7 |
gemini 3.5 flash lite medium Gemini 3.5 Flash-Lite | 74.6 | 79.3+4.6 | 70.0-4.6 |
gemini 3.5 flash lite low Gemini 3.5 Flash-Lite | 72.7 | 76.4+3.8 | 68.9-3.8 |
Performance by call qualityCompare how models coach excellent, mixed, and flawed calls.
| Model | Overall | Excellent n=18 | Mixed n=14 | Flawed n=18 |
|---|---|---|---|---|
gpt-5.6 sol max GPT-5.6 Sol | 89.7 | 88.7-1.1 | 88.7-1.0 | 91.6+1.9 |
gpt-5.6 terra max GPT-5.6 Terra | 89.6 | 88.3-1.3 | 90.2+0.6 | 90.5+0.9 |
gpt-5.6 luna max GPT-5.6 Luna | 89.7 | 88.1-1.6 | 89.9+0.2 | 91.2+1.5 |
gpt-5.6 luna xhigh GPT-5.6 Luna | 89.7 | 88.1-1.5 | 89.0-0.7 | 91.7+2.1 |
gpt-5.6 terra low GPT-5.6 Terra | 89.5 | 88.9-0.6 | 88.4-1.1 | 90.9+1.4 |
gpt-5.6 sol low GPT-5.6 Sol | 89.5 | 89.2-0.3 | 87.6-1.8 | 91.2+1.7 |
gpt-5.6 sol xhigh GPT-5.6 Sol | 89.5 | 89.0-0.5 | 88.2-1.3 | 91.1+1.6 |
gpt-5.6 terra xhigh GPT-5.6 Terra | 89.2 | 88.8-0.5 | 87.4-1.8 | 91.1+1.9 |
gpt-5.6 terra high GPT-5.6 Terra | 89.3 | 89.6+0.2 | 88.1-1.3 | 90.1+0.8 |
gpt-5.6 sol none GPT-5.6 Sol | 89.3 | 90.0+0.7 | 87.9-1.4 | 89.8+0.4 |
gpt-5.6 sol high GPT-5.6 Sol | 89.2 | 88.8-0.4 | 87.8-1.4 | 90.7+1.5 |
gpt-5.4 xhigh GPT-5.4 | 89.0 | 88.9-0.1 | 87.6-1.3 | 90.1+1.1 |
gpt-5.6 sol medium GPT-5.6 Sol | 89.0 | 89.4+0.4 | 86.1-3.0 | 90.9+1.9 |
gpt-5.4 high GPT-5.4 | 89.0 | 88.2-0.7 | 87.7-1.2 | 90.7+1.7 |
gpt-5.5 medium GPT-5.5 | 88.8 | 90.1+1.2 | 86.7-2.1 | 89.3+0.4 |
gpt-5.5 xhigh GPT-5.5 | 89.0 | 89.8+0.8 | 87.7-1.3 | 89.2+0.2 |
gpt-5.6 terra none GPT-5.6 Terra | 88.8 | 90.0+1.2 | 87.6-1.1 | 88.4-0.3 |
gpt-5.6 terra medium GPT-5.6 Terra | 88.9 | 89.0+0.1 | 88.0-0.9 | 89.6+0.7 |
gpt-5.5 high GPT-5.5 | 88.6 | 90.1+1.6 | 84.5-4.1 | 90.2+1.6 |
gpt-5.6 luna medium GPT-5.6 Luna | 88.6 | 88.4-0.1 | 86.4-2.1 | 90.3+1.8 |
gpt-5.6 luna high GPT-5.6 Luna | 88.6 | 88.3-0.2 | 87.1-1.5 | 89.9+1.4 |
gpt-5.6 luna low GPT-5.6 Luna | 88.5 | 88.2-0.3 | 87.1-1.4 | 89.9+1.4 |
gpt-5.4 medium GPT-5.4 | 88.3 | 88.1-0.2 | 85.6-2.7 | 90.6+2.3 |
gpt-5.5 none GPT-5.5 | 88.1 | 89.6+1.5 | 85.0-3.1 | 89.1+1.0 |
gpt-5.6 luna none GPT-5.6 Luna | 87.7 | 87.80.0 | 86.7-1.0 | 88.5+0.8 |
gpt-5.5 low GPT-5.5 | 87.7 | 89.4+1.7 | 85.6-2.1 | 87.70.0 |
fable 5 high Claude Fable 5 | 87.5 | 86.9-0.6 | 85.9-1.6 | 89.3+1.8 |
gpt-5.4 low GPT-5.4 | 87.4 | 87.9+0.6 | 84.8-2.6 | 88.8+1.4 |
gpt-5.4 none GPT-5.4 | 87.4 | 87.9+0.5 | 84.5-2.9 | 89.1+1.7 |
opus 4.7 max Claude Opus 4.7 | 87.3 | 88.8+1.6 | 83.6-3.6 | 88.5+1.2 |
kimi k3 max Kimi K3 | 86.6 | 86.4-0.2 | 83.7-2.9 | 89.1+2.4 |
opus 5 max Claude Opus 5 | 86.6 | 81.5-5.1 | 86.9+0.2 | 91.6+5.0 |
opus 5 xhigh Claude Opus 5 | 86.6 | 82.0-4.6 | 86.7+0.2 | 91.0+4.4 |
opus 4.7 high Claude Opus 4.7 | 86.8 | 87.1+0.3 | 83.4-3.5 | 89.2+2.4 |
muse spark 1.1 high Muse Spark 1.1 | 86.4 | 87.1+0.6 | 83.6-2.8 | 87.9+1.5 |
muse spark 1.1 medium Muse Spark 1.1 | 86.2 | 86.7+0.5 | 84.2-2.0 | 87.2+1.0 |
opus 5 medium Claude Opus 5 | 86.3 | 82.3-4.0 | 86.30.0 | 90.3+4.0 |
opus 5 low Claude Opus 5 | 85.9 | 84.4-1.5 | 83.1-2.8 | 89.6+3.7 |
muse spark 1.1 minimal Muse Spark 1.1 | 85.7 | 85.9+0.2 | 81.3-4.4 | 88.8+3.2 |
opus 5 high Claude Opus 5 | 85.5 | 82.4-3.1 | 83.0-2.5 | 90.6+5.1 |
muse spark 1.1 low Muse Spark 1.1 | 85.4 | 85.8+0.4 | 83.3-2.1 | 86.6+1.2 |
opus 4.8 medium Claude Opus 4.8 | 85.8 | 86.6+0.8 | 81.1-4.7 | 88.6+2.9 |
opus 4.7 medium Claude Opus 4.7 | 85.6 | 87.2+1.6 | 80.0-5.6 | 88.4+2.8 |
opus 4.7 xhigh Claude Opus 4.7 | 85.6 | 86.8+1.2 | 81.1-4.5 | 87.9+2.3 |
opus 4.7 low Claude Opus 4.7 | 85.6 | 86.6+1.0 | 80.8-4.8 | 88.4+2.8 |
opus 4.8 max Claude Opus 4.8 | 85.4 | 85.9+0.5 | 80.4-5.0 | 88.8+3.4 |
opus 4.8 xhigh Claude Opus 4.8 | 85.2 | 88.1+2.8 | 79.4-5.9 | 87.0+1.8 |
opus 4.8 high Claude Opus 4.8 | 84.9 | 87.3+2.4 | 77.4-7.5 | 88.4+3.5 |
sonnet 4.6 Claude Sonnet 4.6 | 84.6 | 84.7+0.1 | 80.1-4.4 | 87.9+3.3 |
sonnet 5 Claude Sonnet 5 | 84.6 | 84.0-0.6 | 83.4-1.2 | 86.2+1.6 |
opus 4.8 low Claude Opus 4.8 | 84.0 | 86.7+2.7 | 77.0-7.0 | 86.7+2.7 |
glm 5.2 GLM 5.2 | 84.0 | 85.8+1.8 | 79.2-4.8 | 85.9+1.9 |
deepseek v4 pro DeepSeek V4 Pro | 83.5 | 86.2+2.7 | 76.6-6.9 | 86.2+2.7 |
gemini 3.6 flash minimal Gemini 3.6 Flash | 81.6 | 85.9+4.3 | 71.7-9.9 | 84.9+3.4 |
gemini 3.6 flash medium Gemini 3.6 Flash | 79.8 | 84.8+4.9 | 70.1-9.7 | 82.4+2.6 |
gemini 3.6 flash high Gemini 3.6 Flash | 79.3 | 84.1+4.8 | 68.6-10.6 | 82.8+3.5 |
gemini 3.1 pro preview Gemini 3.1 Pro Preview | 78.9 | 81.6+2.7 | 69.8-9.1 | 83.3+4.4 |
gemini 3.6 flash low Gemini 3.6 Flash | 78.5 | 84.4+5.9 | 66.3-12.2 | 82.1+3.6 |
gemini 3.5 flash lite high Gemini 3.5 Flash-Lite | 77.4 | 84.7+7.3 | 67.0-10.4 | 78.2+0.8 |
gemini 3.5 flash lite minimal Gemini 3.5 Flash-Lite | 74.8 | 83.6+8.8 | 64.9-10.0 | 73.8-1.0 |
gemini 3.5 flash lite medium Gemini 3.5 Flash-Lite | 74.6 | 82.9+8.2 | 65.6-9.1 | 73.4-1.2 |
gemini 3.5 flash lite low Gemini 3.5 Flash-Lite | 72.7 | 82.2+9.5 | 63.3-9.4 | 70.5-2.2 |
Performance by call typeCompare model performance across discovery, demos, renewals, QBRs, and competitive calls.
| Model | Overall | Discovery n=16 | Product demo n=20 | Renewal save n=4 | QBR n=4 | Competitive displacement n=6 |
|---|---|---|---|---|---|---|
gpt-5.6 sol max GPT-5.6 Sol | 89.7 | 91.0+1.3 | 88.5-1.3 | 88.5-1.2 | 88.5-1.2 | 92.3+2.6 |
gpt-5.6 terra max GPT-5.6 Terra | 89.6 | 90.3+0.7 | 88.8-0.8 | 90.5+0.9 | 87.8-1.9 | 91.0+1.4 |
gpt-5.6 luna max GPT-5.6 Luna | 89.7 | 90.9+1.3 | 88.6-1.1 | 88.3-1.4 | 89.5-0.2 | 91.0+1.3 |
gpt-5.6 luna xhigh GPT-5.6 Luna | 89.7 | 91.0+1.3 | 88.0-1.7 | 90.3+0.6 | 88.3-1.4 | 92.3+2.7 |
gpt-5.6 terra low GPT-5.6 Terra | 89.5 | 91.8+2.3 | 87.7-1.8 | 89.8+0.3 | 88.5-1.0 | 90.0+0.5 |
gpt-5.6 sol low GPT-5.6 Sol | 89.5 | 91.3+1.8 | 88.2-1.3 | 86.3-3.2 | 89.50.0 | 91.2+1.7 |
gpt-5.6 sol xhigh GPT-5.6 Sol | 89.5 | 91.4+1.8 | 87.8-1.7 | 90.0+0.5 | 88.0-1.5 | 91.0+1.5 |
gpt-5.6 terra xhigh GPT-5.6 Terra | 89.2 | 91.0+1.8 | 87.3-1.9 | 87.0-2.2 | 90.0+0.8 | 91.8+2.6 |
gpt-5.6 terra high GPT-5.6 Terra | 89.3 | 90.4+1.0 | 88.3-1.0 | 89.8+0.4 | 88.0-1.3 | 90.5+1.2 |
gpt-5.6 sol none GPT-5.6 Sol | 89.3 | 90.7+1.3 | 88.2-1.2 | 88.5-0.8 | 89.8+0.4 | 90.0+0.7 |
gpt-5.6 sol high GPT-5.6 Sol | 89.2 | 91.7+2.5 | 86.8-2.5 | 89.0-0.2 | 89.5+0.3 | 90.7+1.5 |
gpt-5.4 xhigh GPT-5.4 | 89.0 | 89.7+0.7 | 88.5-0.5 | 86.1-2.9 | 88.0-1.0 | 91.5+2.5 |
gpt-5.6 sol medium GPT-5.6 Sol | 89.0 | 91.6+2.6 | 88.4-0.6 | 83.3-5.8 | 87.3-1.8 | 89.2+0.1 |
gpt-5.4 high GPT-5.4 | 89.0 | 90.4+1.5 | 87.6-1.4 | 89.8+0.8 | 87.5-1.5 | 90.0+1.0 |
gpt-5.5 medium GPT-5.5 | 88.8 | 89.9+1.1 | 87.9-0.9 | 87.0-1.8 | 90.8+1.9 | 89.0+0.2 |
gpt-5.5 xhigh GPT-5.5 | 89.0 | 89.3+0.2 | 88.4-0.6 | 89.8+0.7 | 88.8-0.3 | 90.2+1.1 |
gpt-5.6 terra none GPT-5.6 Terra | 88.8 | 89.6+0.8 | 88.3-0.5 | 87.0-1.8 | 89.5+0.7 | 88.8+0.1 |
gpt-5.6 terra medium GPT-5.6 Terra | 88.9 | 90.1+1.2 | 88.0-0.9 | 89.8+0.8 | 87.8-1.2 | 89.2+0.2 |
gpt-5.5 high GPT-5.5 | 88.6 | 90.4+1.8 | 87.8-0.8 | 80.8-7.8 | 87.8-0.8 | 92.0+3.4 |
gpt-5.6 luna medium GPT-5.6 Luna | 88.6 | 90.6+2.1 | 87.4-1.2 | 87.0-1.6 | 85.5-3.1 | 90.0+1.4 |
gpt-5.6 luna high GPT-5.6 Luna | 88.6 | 90.4+1.9 | 87.2-1.4 | 86.8-1.8 | 86.5-2.1 | 90.8+2.3 |
gpt-5.6 luna low GPT-5.6 Luna | 88.5 | 90.5+2.0 | 86.8-1.8 | 87.8-0.8 | 88.3-0.3 | 89.8+1.3 |
gpt-5.4 medium GPT-5.4 | 88.3 | 90.1+1.8 | 87.3-0.9 | 84.3-4.0 | 86.8-1.5 | 90.2+1.9 |
gpt-5.5 none GPT-5.5 | 88.1 | 90.1+2.0 | 86.5-1.7 | 86.0-2.1 | 88.0-0.1 | 90.0+1.9 |
gpt-5.6 luna none GPT-5.6 Luna | 87.7 | 89.4+1.7 | 86.6-1.1 | 86.5-1.2 | 86.8-1.0 | 88.5+0.8 |
gpt-5.5 low GPT-5.5 | 87.7 | 89.0+1.3 | 86.5-1.2 | 87.0-0.7 | 86.5-1.2 | 89.3+1.6 |
fable 5 high Claude Fable 5 | 87.5 | 89.7+2.2 | 84.8-2.7 | 87.3-0.3 | 88.5+1.0 | 90.2+2.6 |
gpt-5.4 low GPT-5.4 | 87.4 | 89.3+1.9 | 86.5-0.9 | 86.3-1.1 | 87.0-0.4 | 86.2-1.2 |
gpt-5.4 none GPT-5.4 | 87.4 | 88.9+1.6 | 86.5-0.9 | 86.3-1.1 | 86.5-0.9 | 87.5+0.1 |
opus 4.7 max Claude Opus 4.7 | 87.3 | 91.3+4.1 | 84.5-2.8 | 89.3+2.0 | 85.8-1.5 | 85.5-1.8 |
kimi k3 max Kimi K3 | 86.6 | 88.6+2.0 | 84.6-2.0 | 81.8-4.9 | 88.3+1.6 | 90.2+3.5 |
opus 5 max Claude Opus 5 | 86.6 | 89.3+2.7 | 83.7-2.9 | 87.8+1.1 | 83.3-3.4 | 90.8+4.2 |
opus 5 xhigh Claude Opus 5 | 86.6 | 88.8+2.2 | 84.6-2.0 | 88.5+1.9 | 84.0-2.6 | 87.7+1.1 |
opus 4.7 high Claude Opus 4.7 | 86.8 | 90.1+3.2 | 83.7-3.2 | 86.8-0.1 | 86.3-0.6 | 89.2+2.3 |
muse spark 1.1 high Muse Spark 1.1 | 86.4 | 88.4+2.0 | 83.9-2.5 | 86.3-0.2 | 87.3+0.8 | 89.0+2.6 |
muse spark 1.1 medium Muse Spark 1.1 | 86.2 | 88.1+1.9 | 85.0-1.3 | 83.0-3.2 | 86.30.0 | 87.5+1.3 |
opus 5 medium Claude Opus 5 | 86.3 | 88.1+1.8 | 83.7-2.6 | 88.8+2.5 | 84.8-1.5 | 89.8+3.5 |
opus 5 low Claude Opus 5 | 85.9 | 89.4+3.5 | 83.6-2.3 | 79.5-6.4 | 84.3-1.7 | 89.7+3.8 |
muse spark 1.1 minimal Muse Spark 1.1 | 85.7 | 88.3+2.7 | 83.0-2.7 | 83.3-2.4 | 86.0+0.3 | 89.0+3.3 |
opus 5 high Claude Opus 5 | 85.5 | 89.3+3.8 | 84.0-1.5 | 77.8-7.8 | 82.5-3.0 | 87.5+2.0 |
muse spark 1.1 low Muse Spark 1.1 | 85.4 | 86.4+1.0 | 83.4-2.0 | 83.0-2.4 | 85.8+0.4 | 90.7+5.3 |
opus 4.8 medium Claude Opus 4.8 | 85.8 | 89.4+3.6 | 81.5-4.2 | 87.3+1.5 | 86.5+0.7 | 88.7+2.9 |
opus 4.7 medium Claude Opus 4.7 | 85.6 | 89.8+4.2 | 81.6-4.0 | 86.3+0.6 | 84.8-0.9 | 88.0+2.4 |
opus 4.7 xhigh Claude Opus 4.7 | 85.6 | 87.9+2.3 | 83.0-2.5 | 86.3+0.7 | 84.0-1.6 | 88.5+2.9 |
opus 4.7 low Claude Opus 4.7 | 85.6 | 90.2+4.6 | 82.0-3.6 | 86.8+1.1 | 84.5-1.1 | 85.5-0.1 |
opus 4.8 max Claude Opus 4.8 | 85.4 | 88.8+3.4 | 81.3-4.1 | 84.8-0.7 | 87.8+2.3 | 88.8+3.4 |
opus 4.8 xhigh Claude Opus 4.8 | 85.2 | 88.5+3.3 | 81.5-3.8 | 82.8-2.5 | 86.8+1.5 | 89.8+4.6 |
opus 4.8 high Claude Opus 4.8 | 84.9 | 89.9+5.0 | 81.4-3.5 | 78.5-6.4 | 86.0+1.1 | 86.8+1.9 |
sonnet 4.6 Claude Sonnet 4.6 | 84.6 | 86.9+2.3 | 83.3-1.2 | 82.0-2.6 | 82.0-2.6 | 85.8+1.3 |
sonnet 5 Claude Sonnet 5 | 84.6 | 86.6+2.0 | 82.0-2.6 | 88.0+3.4 | 82.8-1.8 | 87.2+2.6 |
opus 4.8 low Claude Opus 4.8 | 84.0 | 88.1+4.1 | 80.5-3.5 | 85.8+1.8 | 81.8-2.2 | 85.0+1.0 |
glm 5.2 GLM 5.2 | 84.0 | 86.3+2.3 | 83.0-1.0 | 85.5+1.5 | 79.5-4.5 | 83.2-0.8 |
deepseek v4 pro DeepSeek V4 Pro | 83.5 | 87.4+3.9 | 81.5-2.0 | 82.0-1.5 | 82.3-1.3 | 81.8-1.7 |
gemini 3.6 flash minimal Gemini 3.6 Flash | 81.6 | 86.4+4.9 | 78.5-3.0 | 79.0-2.6 | 78.5-3.1 | 82.5+0.9 |
gemini 3.6 flash medium Gemini 3.6 Flash | 79.8 | 85.2+5.3 | 77.0-2.9 | 74.8-5.1 | 74.8-5.1 | 82.0+2.2 |
gemini 3.6 flash high Gemini 3.6 Flash | 79.3 | 84.7+5.4 | 74.8-4.4 | 75.3-4.0 | 77.5-1.8 | 83.5+4.2 |
gemini 3.1 pro preview Gemini 3.1 Pro Preview | 78.9 | 84.4+5.5 | 74.7-4.2 | 74.3-4.7 | 77.5-1.4 | 82.3+3.4 |
gemini 3.6 flash low Gemini 3.6 Flash | 78.5 | 83.2+4.7 | 77.0-1.6 | 69.5-9.0 | 75.5-3.0 | 79.3+0.8 |
gemini 3.5 flash lite high Gemini 3.5 Flash-Lite | 77.4 | 81.4+4.0 | 76.0-1.4 | 71.5-5.9 | 75.0-2.4 | 77.0-0.4 |
gemini 3.5 flash lite minimal Gemini 3.5 Flash-Lite | 74.8 | 79.3+4.5 | 72.5-2.3 | 71.3-3.6 | 74.5-0.3 | 73.0-1.8 |
gemini 3.5 flash lite medium Gemini 3.5 Flash-Lite | 74.6 | 78.6+3.9 | 74.8+0.2 | 65.0-9.6 | 71.8-2.9 | 72.0-2.6 |
gemini 3.5 flash lite low Gemini 3.5 Flash-Lite | 72.7 | 78.5+5.8 | 71.3-1.4 | 66.0-6.7 | 70.5-2.2 | 67.8-4.8 |