Skip to results

Sales coaching benchmark / Jul 16, 2026

Which models know sales?

We gave 49 models the same 50 synthetic sales calls. Hidden answer keys measured whether each coach found the real strengths, flaws, and next moves.

All models

Full leaderboard

49 models / 50 calls each

How scoring works
Model
1GPT-5.6 Solmax90.0
P10 80.3, P90 96.2, benchmark 90.0
$0.37
2GPT-5.6 Terramax90.0
P10 80.6, P90 95.6, benchmark 90.0
$0.10
3GPT-5.6 Lunamax89.9
P10 83.7, P90 95.7, benchmark 89.9
$0.06
4GPT-5.6 Lunaxhigh89.9
P10 81.7, P90 96.1, benchmark 89.9
$0.04
5GPT-5.6 Terralow89.9
P10 84.5, P90 95.1, benchmark 89.9
$0.05
6GPT-5.6 Sollow89.8
P10 81.3, P90 96.2, benchmark 89.8
$0.12
7GPT-5.6 Solxhigh89.7
P10 83.1, P90 95.8, benchmark 89.7
$0.17
8GPT-5.6 Terraxhigh89.7
P10 78.6, P90 96.2, benchmark 89.7
$0.08
9GPT-5.6 Terrahigh89.7
P10 82.0, P90 95.7, benchmark 89.7
$0.06
10GPT-5.6 Solnone89.6
P10 81.4, P90 96.0, benchmark 89.6
$0.12
11GPT-5.6 Solhigh89.4
P10 79.8, P90 95.6, benchmark 89.4
$0.12
12GPT-5.4xhigh89.3
P10 79.9, P90 94.7, benchmark 89.3
$0.28
13GPT-5.6 Solmedium89.3
P10 79.7, P90 95.5, benchmark 89.3
$0.13
14GPT-5.4high89.2
P10 81.9, P90 95.2, benchmark 89.2
$0.10
15GPT-5.5medium89.1
P10 82.0, P90 95.5, benchmark 89.1
$0.12
16GPT-5.5xhigh89.1
P10 78.5, P90 95.5, benchmark 89.1
$0.20
17GPT-5.6 Terranone89.1
P10 79.7, P90 95.6, benchmark 89.1
$0.05
18GPT-5.6 Terramedium89.1
P10 80.1, P90 95.7, benchmark 89.1
$0.05
19GPT-5.5high88.9
P10 79.4, P90 95.2, benchmark 88.9
$0.14
20GPT-5.6 Lunamedium88.7
P10 82.0, P90 95.8, benchmark 88.7
$0.02
21GPT-5.6 Lunahigh88.7
P10 78.9, P90 95.4, benchmark 88.7
$0.03
22GPT-5.6 Lunalow88.6
P10 79.4, P90 96.2, benchmark 88.6
$0.02
23GPT-5.4medium88.4
P10 80.1, P90 94.6, benchmark 88.4
$0.06
24GPT-5.5none88.3
P10 76.6, P90 95.8, benchmark 88.3
$0.13
25GPT-5.6 Lunanone88.0
P10 79.7, P90 95.6, benchmark 88.0
$0.02
26GPT-5.5low87.8
P10 78.3, P90 94.6, benchmark 87.8
$0.14
27Claude Fable 5high87.7
P10 77.0, P90 95.5, benchmark 87.7
$0.46
28GPT-5.4low87.5
P10 78.3, P90 94.7, benchmark 87.5
$0.05
29GPT-5.4none87.5
P10 78.2, P90 94.8, benchmark 87.5
$0.05
30Claude Opus 4.7max87.2
P10 75.0, P90 94.8, benchmark 87.2
$0.21
31Kimi K3max86.9
P10 72.0, P90 94.5, benchmark 86.9
$0.15
32Claude Opus 4.7high86.6
P10 77.8, P90 94.8, benchmark 86.6
$0.17
33Muse Spark 1.1high86.6
P10 77.0, P90 95.0, benchmark 86.6
$0.03
34Muse Spark 1.1medium86.4
P10 78.4, P90 94.9, benchmark 86.4
$0.02
35Muse Spark 1.1minimal85.7
P10 75.8, P90 93.9, benchmark 85.7
$0.02
36Muse Spark 1.1low85.6
P10 75.1, P90 94.6, benchmark 85.6
$0.02
37Claude Opus 4.8medium85.6
P10 72.1, P90 94.8, benchmark 85.6
$0.14
38Claude Opus 4.7medium85.5
P10 74.0, P90 93.7, benchmark 85.5
$0.33
39Claude Opus 4.7xhigh85.5
P10 73.1, P90 95.1, benchmark 85.5
$0.17
40Claude Opus 4.7low85.5
P10 75.6, P90 93.7, benchmark 85.5
$0.13
41Claude Opus 4.8max85.4
P10 72.9, P90 95.0, benchmark 85.4
$0.18
42Claude Opus 4.8xhigh85.3
P10 73.0, P90 95.0, benchmark 85.3
$0.16
43Claude Opus 4.8high84.9
P10 68.8, P90 95.4, benchmark 84.9
$0.15
44Claude Sonnet 4.6default84.5
P10 71.9, P90 94.1, benchmark 84.5
$0.10
45Claude Sonnet 5default84.3
P10 71.6, P90 92.2, benchmark 84.3
$0.05
46Claude Opus 4.8low83.7
P10 67.2, P90 94.9, benchmark 83.7
$0.12
47GLM 5.2default83.6
P10 71.3, P90 94.3, benchmark 83.6
$0.03
48DeepSeek V4 Prodefault83.1
P10 69.1, P90 93.0, benchmark 83.1
$0.0047
49Gemini 3.1 Pro Previewdefault78.7
P10 66.6, P90 91.0, benchmark 78.7
$0.03

How to read this benchmark

Synthetic cases, hidden answer keys

50 calls, 49 models, and 2,450 judged coaching notes.

Full methodology
01 / Dataset

50 researched sales calls

The calls were generated with GPT and Claude Sonnet 4.6 and span varied call types and seller performance.

02 / Protocol

Every model sees the same evidence

Models receive the setup, research, participants, and transcript. The answer key stays out of the prompt.

03 / Scoring

An AI judge scores each coaching note

The judge scores 8 dimensions, including sales instinct, evidence, prioritization, accuracy, and restraint.

Benchmark calls

Calls that separated the models

A mix of call types, seller performance, and transcript sources.

Browse all calls