Skip to results

Sales coaching benchmark / Jul 25, 2026

Which models know sales?

We gave 62 models the same 50 synthetic sales calls. Hidden answer keys measured whether each coach found the real strengths, flaws, and next moves.

All models

Full leaderboard

62 models / 50 calls each

How scoring works
Model
1GPT-5.6 Solmax90.0
P10 80.3, P90 96.2, benchmark 90.0
$0.37
2GPT-5.6 Terramax90.0
P10 80.6, P90 95.6, benchmark 90.0
$0.10
3GPT-5.6 Lunamax89.9
P10 83.7, P90 95.7, benchmark 89.9
$0.06
4GPT-5.6 Lunaxhigh89.9
P10 81.7, P90 96.1, benchmark 89.9
$0.04
5GPT-5.6 Terralow89.9
P10 84.5, P90 95.1, benchmark 89.9
$0.05
6GPT-5.6 Sollow89.8
P10 81.3, P90 96.2, benchmark 89.8
$0.12
7GPT-5.6 Solxhigh89.7
P10 83.1, P90 95.8, benchmark 89.7
$0.17
8GPT-5.6 Terraxhigh89.7
P10 78.6, P90 96.2, benchmark 89.7
$0.08
9GPT-5.6 Terrahigh89.7
P10 82.0, P90 95.7, benchmark 89.7
$0.06
10GPT-5.6 Solnone89.6
P10 81.4, P90 96.0, benchmark 89.6
$0.12
11GPT-5.6 Solhigh89.4
P10 79.8, P90 95.6, benchmark 89.4
$0.12
12GPT-5.4xhigh89.3
P10 79.9, P90 94.7, benchmark 89.3
$0.28
13GPT-5.6 Solmedium89.3
P10 79.7, P90 95.5, benchmark 89.3
$0.13
14GPT-5.4high89.2
P10 81.9, P90 95.2, benchmark 89.2
$0.10
15GPT-5.5medium89.1
P10 82.0, P90 95.5, benchmark 89.1
$0.12
16GPT-5.5xhigh89.1
P10 78.5, P90 95.5, benchmark 89.1
$0.20
17GPT-5.6 Terranone89.1
P10 79.7, P90 95.6, benchmark 89.1
$0.05
18GPT-5.6 Terramedium89.1
P10 80.1, P90 95.7, benchmark 89.1
$0.05
19GPT-5.5high88.9
P10 79.4, P90 95.2, benchmark 88.9
$0.14
20GPT-5.6 Lunamedium88.7
P10 82.0, P90 95.8, benchmark 88.7
$0.02
21GPT-5.6 Lunahigh88.7
P10 78.9, P90 95.4, benchmark 88.7
$0.03
22GPT-5.6 Lunalow88.6
P10 79.4, P90 96.2, benchmark 88.6
$0.02
23GPT-5.4medium88.4
P10 80.1, P90 94.6, benchmark 88.4
$0.06
24GPT-5.5none88.3
P10 76.6, P90 95.8, benchmark 88.3
$0.13
25GPT-5.6 Lunanone88.0
P10 79.7, P90 95.6, benchmark 88.0
$0.02
26GPT-5.5low87.8
P10 78.3, P90 94.6, benchmark 87.8
$0.14
27Claude Fable 5high87.7
P10 77.0, P90 95.5, benchmark 87.7
$0.46
28GPT-5.4low87.5
P10 78.3, P90 94.7, benchmark 87.5
$0.05
29GPT-5.4none87.5
P10 78.2, P90 94.8, benchmark 87.5
$0.05
30Claude Opus 4.7max87.2
P10 75.0, P90 94.8, benchmark 87.2
$0.21
31Kimi K3max86.9
P10 72.0, P90 94.5, benchmark 86.9
$0.15
32Claude Opus 5max86.8
P10 78.2, P90 95.0, benchmark 86.8
$0.383m 13s
33Claude Opus 5xhigh86.7
P10 77.8, P90 94.5, benchmark 86.7
$0.352m 53s
34Claude Opus 4.7high86.6
P10 77.8, P90 94.8, benchmark 86.6
$0.17
35Muse Spark 1.1high86.6
P10 77.0, P90 95.0, benchmark 86.6
$0.03
36Muse Spark 1.1medium86.4
P10 78.4, P90 94.9, benchmark 86.4
$0.02
37Claude Opus 5medium86.3
P10 75.9, P90 94.4, benchmark 86.3
$0.251m 59s
38Claude Opus 5low86.0
P10 74.3, P90 94.2, benchmark 86.0
$0.201m 30s
39Muse Spark 1.1minimal85.7
P10 75.8, P90 93.9, benchmark 85.7
$0.02
40Claude Opus 5high85.7
P10 72.8, P90 94.7, benchmark 85.7
$0.302m 27s
41Muse Spark 1.1low85.6
P10 75.1, P90 94.6, benchmark 85.6
$0.02
42Claude Opus 4.8medium85.6
P10 72.1, P90 94.8, benchmark 85.6
$0.14
43Claude Opus 4.7medium85.5
P10 74.0, P90 93.7, benchmark 85.5
$0.33
44Claude Opus 4.7xhigh85.5
P10 73.1, P90 95.1, benchmark 85.5
$0.17
45Claude Opus 4.7low85.5
P10 75.6, P90 93.7, benchmark 85.5
$0.13
46Claude Opus 4.8max85.4
P10 72.9, P90 95.0, benchmark 85.4
$0.18
47Claude Opus 4.8xhigh85.3
P10 73.0, P90 95.0, benchmark 85.3
$0.16
48Claude Opus 4.8high84.9
P10 68.8, P90 95.4, benchmark 84.9
$0.15
49Claude Sonnet 4.6default84.5
P10 71.9, P90 94.1, benchmark 84.5
$0.10
50Claude Sonnet 5default84.3
P10 71.6, P90 92.2, benchmark 84.3
$0.05
51Claude Opus 4.8low83.7
P10 67.2, P90 94.9, benchmark 83.7
$0.12
52GLM 5.2default83.6
P10 71.3, P90 94.3, benchmark 83.6
$0.03
53DeepSeek V4 Prodefault83.1
P10 69.1, P90 93.0, benchmark 83.1
$0.0047
54Gemini 3.6 Flashminimal81.0
P10 60.6, P90 92.8, benchmark 81.0
$0.029.7s
55Gemini 3.6 Flashmedium79.6
P10 63.7, P90 92.1, benchmark 79.6
$0.0315s
56Gemini 3.6 Flashhigh78.7
P10 60.5, P90 91.7, benchmark 78.7
$0.0422s
57Gemini 3.1 Pro Previewdefault78.7
P10 66.6, P90 91.0, benchmark 78.7
$0.03
58Gemini 3.6 Flashlow77.9
P10 60.4, P90 92.8, benchmark 77.9
$0.028.8s
59Gemini 3.5 Flash-Litehigh76.8
P10 60.5, P90 89.5, benchmark 76.8
$0.0111s
60Gemini 3.5 Flash-Liteminimal74.0
P10 53.7, P90 89.2, benchmark 74.0
$0.00535.6s
61Gemini 3.5 Flash-Litemedium73.7
P10 57.0, P90 89.0, benchmark 73.7
$0.00535.7s
62Gemini 3.5 Flash-Litelow71.7
P10 50.2, P90 88.5, benchmark 71.7
$0.00474.9s

How to read this benchmark

Synthetic cases, hidden answer keys

50 calls, 62 models, and 3,100 judged coaching notes.

Full methodology
01 / Dataset

50 researched sales calls

The calls were generated with GPT and Claude Sonnet 4.6 and span varied call types and seller performance.

02 / Protocol

Every model sees the same evidence

Models receive the setup, research, participants, and transcript. The answer key stays out of the prompt.

03 / Scoring

An AI judge scores each coaching note

The judge scores 8 dimensions, including sales instinct, evidence, prioritization, accuracy, and restraint.

Benchmark calls

Calls that separated the models

A mix of call types, seller performance, and transcript sources.

Browse all calls