Frontier model leaderboard
A blind expert evaluation of six frontier models over 598 executive-reasoning tasks, ranked by Reward, the share of tasks a model solves. The rubrics, win rate, and rank distributions that follow are the diagnostic detail behind that ranking.
Reward: the share of tasks solved
A task counts as solved only when the experts give the answer a mean rubric score of at least 3.0/5 and it covers at least 60% of the checklist, so Reward credits substance over presentation. Fable 5 leads at 58%, and the field falls to 19%.
Full leaderboard
i Reward % share of tasks solved (mean rubric ≥ 3/5 AND coverage ≥ 60%), the headline score.
Overall mean of the five rubrics, as a % of the 5-point max.
Dom / Strat / Act / Exec / Local the five rubrics, each a % of the 5-point max.
Cov. mean checklist coverage.
Avg rank mean blind rank (1 = best).
Win % share of blind judgments ranking the model first.
Top-3 % share ranking it top three.
Small grey numbers are 95% bootstrap confidence intervals.
Rows are sorted by Reward %: the share of tasks a model solves (mean rubric ≥ 3.0/5 AND checklist coverage ≥ 60%, experts averaged per task). Small grey numbers are 95% bootstrap CIs (1,000 resamples). Judgments are not independent (two correlated judgments per task), so every CI is clustered over the 598 tasks. Overlapping CIs mean the separation is not significant.
| # | Model | Reward % | Overall | Dom. | Strat. | Act. | Exec. | Local Fid. | Cov. | Avg rank | Win % | Top-3 % |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 58% 54–62 | 83% 82–84 | 84% | 86% | 83% | 85% | 77% | 63% | 2.133 | 50% 47–53 | 82% | |
| 2 | 52% 48–56 | 78% 77–79 | 82% | 81% | 81% | 75% | 70% | 59% | 2.849 | 22% 20–25 | 69% | |
| 3 | 36% 33–40 | 76% 75–77 | 76% | 80% | 74% | 82% | 66% | 53% | 3.242 | 9% 8–11 | 61% | |
| 4 | 27% 23–30 | 72% 71–73 | 71% | 73% | 74% | 79% | 62% | 48% | 3.778 | 7% 6–9 | 38% | |
| 5 | 23% 20–26 | 68% 67–69 | 66% | 69% | 72% | 77% | 59% | 44% | 4.156 | 6% 5–7 | 30% | |
| 6 | 19% 16–22 | 61% 60–62 | 58% | 60% | 68% | 69% | 52% | 40% | 4.842 | 6% 4–7 | 20% |
Blind preference distribution
Where each model lands across all blind rankings, from #1 (best) to #6.
Head-to-head
Each cell is the row model's win rate over the column model, computed only on the blind rankings where an expert placed both. Violet = the row wins the majority of those matchups; red = it loses. Beats counts how many opponents each model beats outright.
The ranking is perfectly transitive: the “Beats” column runs 5, 4, …, 0 with no ties, so every model beats exactly the ones below it and loses to those above, with no rock-paper-scissors cycles. That is a strong internal-consistency signal for the ordering, independent of the rubric means.
| Model | Fable | GPT-5.5 | Claude | Gemini | GLM-5.2 | Mistral | Beats |
|---|---|---|---|---|---|---|---|
| — | 66% | 76% | 80% | 82% | 82% | 5/5 | |
| 34% | — | 56% | 68% | 74% | 83% | 4/5 | |
| 24% | 44% | — | 62% | 69% | 77% | 3/5 | |
| 20% | 32% | 38% | — | 58% | 74% | 2/5 | |
| 18% | 26% | 31% | 42% | — | 68% | 1/5 | |
| 18% | 17% | 23% | 26% | 32% | — | 0/5 |
Rubrics (diagnostic detail)
Secondary to Reward %: the per-rubric means (% of max) that explain why a model lands where it does. Domain Accuracy and Strategic Reasoning separate the models most (spread 26%), and Local / Regulatory Fidelity is every model's weakest rubric — exactly the gap EuroExec exists to measure.
Domain Accuracy (% of max · spread 26%)
Strategic Reasoning (% of max · spread 26%)
Actionability & Specificity (% of max · spread 15%)
Executive Communication (% of max · spread 16%)
Local / Regulatory Fidelity (% of max · spread 25%)
Rubric scores by domain
Mean rubric score (% of the 5-point max) by executive domain. Darker = stronger. Per-domain n: Business 162 · Finance 114 · Marketing 134 · Product & Tech 188.
| Model | Business | Finance | Marketing | Product & Tech |
|---|---|---|---|---|
| 80% | 92% | 84% | 79% | |
| 78% | 80% | 80% | 75% | |
| 74% | 81% | 76% | 74% | |
| 70% | 78% | 72% | 69% | |
| 68% | 69% | 70% | 67% | |
| 64% | 59% | 61% | 60% |
Explore one graded task end to end in the sample task, see where models fail on the failure modes page, or see how the set is built on the methodology page.