01 / Dataset
50 researched sales calls
The calls were generated with GPT and Claude Sonnet 4.6 and span varied call types and seller performance.
Sales coaching benchmark / Jul 16, 2026
We gave 49 models the same 50 synthetic sales calls. Hidden answer keys measured whether each coach found the real strengths, flaws, and next moves.
All models
49 models / 50 calls each
How scoring works| Model | |||
|---|---|---|---|
| 1GPT-5.6 Solmax | 90.0 | P10 80.3, P90 96.2, benchmark 90.0 | $0.37 |
| 2GPT-5.6 Terramax | 90.0 | P10 80.6, P90 95.6, benchmark 90.0 | $0.10 |
| 3GPT-5.6 Lunamax | 89.9 | P10 83.7, P90 95.7, benchmark 89.9 | $0.06 |
| 4GPT-5.6 Lunaxhigh | 89.9 | P10 81.7, P90 96.1, benchmark 89.9 | $0.04 |
| 5GPT-5.6 Terralow | 89.9 | P10 84.5, P90 95.1, benchmark 89.9 | $0.05 |
| 6GPT-5.6 Sollow | 89.8 | P10 81.3, P90 96.2, benchmark 89.8 | $0.12 |
| 7GPT-5.6 Solxhigh | 89.7 | P10 83.1, P90 95.8, benchmark 89.7 | $0.17 |
| 8GPT-5.6 Terraxhigh | 89.7 | P10 78.6, P90 96.2, benchmark 89.7 | $0.08 |
| 9GPT-5.6 Terrahigh | 89.7 | P10 82.0, P90 95.7, benchmark 89.7 | $0.06 |
| 10GPT-5.6 Solnone | 89.6 | P10 81.4, P90 96.0, benchmark 89.6 | $0.12 |
| 11GPT-5.6 Solhigh | 89.4 | P10 79.8, P90 95.6, benchmark 89.4 | $0.12 |
| 12GPT-5.4xhigh | 89.3 | P10 79.9, P90 94.7, benchmark 89.3 | $0.28 |
| 13GPT-5.6 Solmedium | 89.3 | P10 79.7, P90 95.5, benchmark 89.3 | $0.13 |
| 14GPT-5.4high | 89.2 | P10 81.9, P90 95.2, benchmark 89.2 | $0.10 |
| 15GPT-5.5medium | 89.1 | P10 82.0, P90 95.5, benchmark 89.1 | $0.12 |
| 16GPT-5.5xhigh | 89.1 | P10 78.5, P90 95.5, benchmark 89.1 | $0.20 |
| 17GPT-5.6 Terranone | 89.1 | P10 79.7, P90 95.6, benchmark 89.1 | $0.05 |
| 18GPT-5.6 Terramedium | 89.1 | P10 80.1, P90 95.7, benchmark 89.1 | $0.05 |
| 19GPT-5.5high | 88.9 | P10 79.4, P90 95.2, benchmark 88.9 | $0.14 |
| 20GPT-5.6 Lunamedium | 88.7 | P10 82.0, P90 95.8, benchmark 88.7 | $0.02 |
| 21GPT-5.6 Lunahigh | 88.7 | P10 78.9, P90 95.4, benchmark 88.7 | $0.03 |
| 22GPT-5.6 Lunalow | 88.6 | P10 79.4, P90 96.2, benchmark 88.6 | $0.02 |
| 23GPT-5.4medium | 88.4 | P10 80.1, P90 94.6, benchmark 88.4 | $0.06 |
| 24GPT-5.5none | 88.3 | P10 76.6, P90 95.8, benchmark 88.3 | $0.13 |
| 25GPT-5.6 Lunanone | 88.0 | P10 79.7, P90 95.6, benchmark 88.0 | $0.02 |
| 26GPT-5.5low | 87.8 | P10 78.3, P90 94.6, benchmark 87.8 | $0.14 |
| 27Claude Fable 5high | 87.7 | P10 77.0, P90 95.5, benchmark 87.7 | $0.46 |
| 28GPT-5.4low | 87.5 | P10 78.3, P90 94.7, benchmark 87.5 | $0.05 |
| 29GPT-5.4none | 87.5 | P10 78.2, P90 94.8, benchmark 87.5 | $0.05 |
| 30Claude Opus 4.7max | 87.2 | P10 75.0, P90 94.8, benchmark 87.2 | $0.21 |
| 31Kimi K3max | 86.9 | P10 72.0, P90 94.5, benchmark 86.9 | $0.15 |
| 32Claude Opus 4.7high | 86.6 | P10 77.8, P90 94.8, benchmark 86.6 | $0.17 |
| 33Muse Spark 1.1high | 86.6 | P10 77.0, P90 95.0, benchmark 86.6 | $0.03 |
| 34Muse Spark 1.1medium | 86.4 | P10 78.4, P90 94.9, benchmark 86.4 | $0.02 |
| 35Muse Spark 1.1minimal | 85.7 | P10 75.8, P90 93.9, benchmark 85.7 | $0.02 |
| 36Muse Spark 1.1low | 85.6 | P10 75.1, P90 94.6, benchmark 85.6 | $0.02 |
| 37Claude Opus 4.8medium | 85.6 | P10 72.1, P90 94.8, benchmark 85.6 | $0.14 |
| 38Claude Opus 4.7medium | 85.5 | P10 74.0, P90 93.7, benchmark 85.5 | $0.33 |
| 39Claude Opus 4.7xhigh | 85.5 | P10 73.1, P90 95.1, benchmark 85.5 | $0.17 |
| 40Claude Opus 4.7low | 85.5 | P10 75.6, P90 93.7, benchmark 85.5 | $0.13 |
| 41Claude Opus 4.8max | 85.4 | P10 72.9, P90 95.0, benchmark 85.4 | $0.18 |
| 42Claude Opus 4.8xhigh | 85.3 | P10 73.0, P90 95.0, benchmark 85.3 | $0.16 |
| 43Claude Opus 4.8high | 84.9 | P10 68.8, P90 95.4, benchmark 84.9 | $0.15 |
| 44Claude Sonnet 4.6default | 84.5 | P10 71.9, P90 94.1, benchmark 84.5 | $0.10 |
| 45Claude Sonnet 5default | 84.3 | P10 71.6, P90 92.2, benchmark 84.3 | $0.05 |
| 46Claude Opus 4.8low | 83.7 | P10 67.2, P90 94.9, benchmark 83.7 | $0.12 |
| 47GLM 5.2default | 83.6 | P10 71.3, P90 94.3, benchmark 83.6 | $0.03 |
| 48DeepSeek V4 Prodefault | 83.1 | P10 69.1, P90 93.0, benchmark 83.1 | $0.0047 |
| 49Gemini 3.1 Pro Previewdefault | 78.7 | P10 66.6, P90 91.0, benchmark 78.7 | $0.03 |
How to read this benchmark
50 calls, 49 models, and 2,450 judged coaching notes.
The calls were generated with GPT and Claude Sonnet 4.6 and span varied call types and seller performance.
Models receive the setup, research, participants, and transcript. The answer key stays out of the prompt.
The judge scores 8 dimensions, including sales instinct, evidence, prioritization, accuracy, and restraint.
Benchmark calls
A mix of call types, seller performance, and transcript sources.