01 / Dataset
50 researched sales calls
The calls were generated with GPT and Claude Sonnet 4.6 and span varied call types and seller performance.
Sales coaching benchmark / Jul 25, 2026
We gave 62 models the same 50 synthetic sales calls. Hidden answer keys measured whether each coach found the real strengths, flaws, and next moves.
All models
62 models / 50 calls each
How scoring works| Model | |||
|---|---|---|---|
| 1GPT-5.6 Solmax | 90.0 | P10 80.3, P90 96.2, benchmark 90.0 | $0.37 |
| 2GPT-5.6 Terramax | 90.0 | P10 80.6, P90 95.6, benchmark 90.0 | $0.10 |
| 3GPT-5.6 Lunamax | 89.9 | P10 83.7, P90 95.7, benchmark 89.9 | $0.06 |
| 4GPT-5.6 Lunaxhigh | 89.9 | P10 81.7, P90 96.1, benchmark 89.9 | $0.04 |
| 5GPT-5.6 Terralow | 89.9 | P10 84.5, P90 95.1, benchmark 89.9 | $0.05 |
| 6GPT-5.6 Sollow | 89.8 | P10 81.3, P90 96.2, benchmark 89.8 | $0.12 |
| 7GPT-5.6 Solxhigh | 89.7 | P10 83.1, P90 95.8, benchmark 89.7 | $0.17 |
| 8GPT-5.6 Terraxhigh | 89.7 | P10 78.6, P90 96.2, benchmark 89.7 | $0.08 |
| 9GPT-5.6 Terrahigh | 89.7 | P10 82.0, P90 95.7, benchmark 89.7 | $0.06 |
| 10GPT-5.6 Solnone | 89.6 | P10 81.4, P90 96.0, benchmark 89.6 | $0.12 |
| 11GPT-5.6 Solhigh | 89.4 | P10 79.8, P90 95.6, benchmark 89.4 | $0.12 |
| 12GPT-5.4xhigh | 89.3 | P10 79.9, P90 94.7, benchmark 89.3 | $0.28 |
| 13GPT-5.6 Solmedium | 89.3 | P10 79.7, P90 95.5, benchmark 89.3 | $0.13 |
| 14GPT-5.4high | 89.2 | P10 81.9, P90 95.2, benchmark 89.2 | $0.10 |
| 15GPT-5.5medium | 89.1 | P10 82.0, P90 95.5, benchmark 89.1 | $0.12 |
| 16GPT-5.5xhigh | 89.1 | P10 78.5, P90 95.5, benchmark 89.1 | $0.20 |
| 17GPT-5.6 Terranone | 89.1 | P10 79.7, P90 95.6, benchmark 89.1 | $0.05 |
| 18GPT-5.6 Terramedium | 89.1 | P10 80.1, P90 95.7, benchmark 89.1 | $0.05 |
| 19GPT-5.5high | 88.9 | P10 79.4, P90 95.2, benchmark 88.9 | $0.14 |
| 20GPT-5.6 Lunamedium | 88.7 | P10 82.0, P90 95.8, benchmark 88.7 | $0.02 |
| 21GPT-5.6 Lunahigh | 88.7 | P10 78.9, P90 95.4, benchmark 88.7 | $0.03 |
| 22GPT-5.6 Lunalow | 88.6 | P10 79.4, P90 96.2, benchmark 88.6 | $0.02 |
| 23GPT-5.4medium | 88.4 | P10 80.1, P90 94.6, benchmark 88.4 | $0.06 |
| 24GPT-5.5none | 88.3 | P10 76.6, P90 95.8, benchmark 88.3 | $0.13 |
| 25GPT-5.6 Lunanone | 88.0 | P10 79.7, P90 95.6, benchmark 88.0 | $0.02 |
| 26GPT-5.5low | 87.8 | P10 78.3, P90 94.6, benchmark 87.8 | $0.14 |
| 27Claude Fable 5high | 87.7 | P10 77.0, P90 95.5, benchmark 87.7 | $0.46 |
| 28GPT-5.4low | 87.5 | P10 78.3, P90 94.7, benchmark 87.5 | $0.05 |
| 29GPT-5.4none | 87.5 | P10 78.2, P90 94.8, benchmark 87.5 | $0.05 |
| 30Claude Opus 4.7max | 87.2 | P10 75.0, P90 94.8, benchmark 87.2 | $0.21 |
| 31Kimi K3max | 86.9 | P10 72.0, P90 94.5, benchmark 86.9 | $0.15 |
| 32Claude Opus 5max | 86.8 | P10 78.2, P90 95.0, benchmark 86.8 | $0.383m 13s |
| 33Claude Opus 5xhigh | 86.7 | P10 77.8, P90 94.5, benchmark 86.7 | $0.352m 53s |
| 34Claude Opus 4.7high | 86.6 | P10 77.8, P90 94.8, benchmark 86.6 | $0.17 |
| 35Muse Spark 1.1high | 86.6 | P10 77.0, P90 95.0, benchmark 86.6 | $0.03 |
| 36Muse Spark 1.1medium | 86.4 | P10 78.4, P90 94.9, benchmark 86.4 | $0.02 |
| 37Claude Opus 5medium | 86.3 | P10 75.9, P90 94.4, benchmark 86.3 | $0.251m 59s |
| 38Claude Opus 5low | 86.0 | P10 74.3, P90 94.2, benchmark 86.0 | $0.201m 30s |
| 39Muse Spark 1.1minimal | 85.7 | P10 75.8, P90 93.9, benchmark 85.7 | $0.02 |
| 40Claude Opus 5high | 85.7 | P10 72.8, P90 94.7, benchmark 85.7 | $0.302m 27s |
| 41Muse Spark 1.1low | 85.6 | P10 75.1, P90 94.6, benchmark 85.6 | $0.02 |
| 42Claude Opus 4.8medium | 85.6 | P10 72.1, P90 94.8, benchmark 85.6 | $0.14 |
| 43Claude Opus 4.7medium | 85.5 | P10 74.0, P90 93.7, benchmark 85.5 | $0.33 |
| 44Claude Opus 4.7xhigh | 85.5 | P10 73.1, P90 95.1, benchmark 85.5 | $0.17 |
| 45Claude Opus 4.7low | 85.5 | P10 75.6, P90 93.7, benchmark 85.5 | $0.13 |
| 46Claude Opus 4.8max | 85.4 | P10 72.9, P90 95.0, benchmark 85.4 | $0.18 |
| 47Claude Opus 4.8xhigh | 85.3 | P10 73.0, P90 95.0, benchmark 85.3 | $0.16 |
| 48Claude Opus 4.8high | 84.9 | P10 68.8, P90 95.4, benchmark 84.9 | $0.15 |
| 49Claude Sonnet 4.6default | 84.5 | P10 71.9, P90 94.1, benchmark 84.5 | $0.10 |
| 50Claude Sonnet 5default | 84.3 | P10 71.6, P90 92.2, benchmark 84.3 | $0.05 |
| 51Claude Opus 4.8low | 83.7 | P10 67.2, P90 94.9, benchmark 83.7 | $0.12 |
| 52GLM 5.2default | 83.6 | P10 71.3, P90 94.3, benchmark 83.6 | $0.03 |
| 53DeepSeek V4 Prodefault | 83.1 | P10 69.1, P90 93.0, benchmark 83.1 | $0.0047 |
| 54Gemini 3.6 Flashminimal | 81.0 | P10 60.6, P90 92.8, benchmark 81.0 | $0.029.7s |
| 55Gemini 3.6 Flashmedium | 79.6 | P10 63.7, P90 92.1, benchmark 79.6 | $0.0315s |
| 56Gemini 3.6 Flashhigh | 78.7 | P10 60.5, P90 91.7, benchmark 78.7 | $0.0422s |
| 57Gemini 3.1 Pro Previewdefault | 78.7 | P10 66.6, P90 91.0, benchmark 78.7 | $0.03 |
| 58Gemini 3.6 Flashlow | 77.9 | P10 60.4, P90 92.8, benchmark 77.9 | $0.028.8s |
| 59Gemini 3.5 Flash-Litehigh | 76.8 | P10 60.5, P90 89.5, benchmark 76.8 | $0.0111s |
| 60Gemini 3.5 Flash-Liteminimal | 74.0 | P10 53.7, P90 89.2, benchmark 74.0 | $0.00535.6s |
| 61Gemini 3.5 Flash-Litemedium | 73.7 | P10 57.0, P90 89.0, benchmark 73.7 | $0.00535.7s |
| 62Gemini 3.5 Flash-Litelow | 71.7 | P10 50.2, P90 88.5, benchmark 71.7 | $0.00474.9s |
How to read this benchmark
50 calls, 62 models, and 3,100 judged coaching notes.
The calls were generated with GPT and Claude Sonnet 4.6 and span varied call types and seller performance.
Models receive the setup, research, participants, and transcript. The answer key stays out of the prompt.
The judge scores 8 dimensions, including sales instinct, evidence, prioritization, accuracy, and restraint.
Benchmark calls
A mix of call types, seller performance, and transcript sources.