Frontier model leaderboard

A blind expert evaluation of six frontier models over 598 executive-reasoning tasks, ranked by Reward, the share of tasks a model solves. The rubrics, win rate, and rank distributions that follow are the diagnostic detail behind that ranking.

Reward: the share of tasks solved

A task counts as solved only when the experts give the answer a mean rubric score of at least 3.0/5 and it covers at least 60% of the checklist, so Reward credits substance over presentation. Fable 5 leads at 58%, and the field falls to 19%.

Fable 5
58%
GPT-5.5
52%
Claude Opus 4.8
36%
Gemini 3.1 Pro
27%
GLM-5.2
23%
Mistral Large 3
19%

Full leaderboard i Reward % share of tasks solved (mean rubric ≥ 3/5 AND coverage ≥ 60%), the headline score.
Overall mean of the five rubrics, as a % of the 5-point max.
Dom / Strat / Act / Exec / Local the five rubrics, each a % of the 5-point max.
Cov. mean checklist coverage.
Avg rank mean blind rank (1 = best).
Win % share of blind judgments ranking the model first.
Top-3 % share ranking it top three.
Small grey numbers are 95% bootstrap confidence intervals.

Rows are sorted by Reward %: the share of tasks a model solves (mean rubric ≥ 3.0/5 AND checklist coverage ≥ 60%, experts averaged per task). Small grey numbers are 95% bootstrap CIs (1,000 resamples). Judgments are not independent (two correlated judgments per task), so every CI is clustered over the 598 tasks. Overlapping CIs mean the separation is not significant.

#Model Reward %Overall Dom.Strat.Act.Exec.Local Fid. Cov.Avg rankWin %Top-3 %
1 Fable 5 58%
54–62
83%
82–84
84% 86% 83% 85% 77% 63% 2.133 50%
47–53
82%
2 GPT-5.5 52%
48–56
78%
77–79
82% 81% 81% 75% 70% 59% 2.849 22%
20–25
69%
3 Claude Opus 4.8 36%
33–40
76%
75–77
76% 80% 74% 82% 66% 53% 3.242 9%
8–11
61%
4 Gemini 3.1 Pro 27%
23–30
72%
71–73
71% 73% 74% 79% 62% 48% 3.778 7%
6–9
38%
5 GLM-5.2 23%
20–26
68%
67–69
66% 69% 72% 77% 59% 44% 4.156 6%
5–7
30%
6 Mistral Large 3 19%
16–22
61%
60–62
58% 60% 68% 69% 52% 40% 4.842 6%
4–7
20%

Blind preference distribution

Where each model lands across all blind rankings, from #1 (best) to #6.

Fable 5#1 in 50%
GPT-5.5#1 in 22%
Claude Opus 4.8#1 in 9%
Gemini 3.1 Pro#1 in 7%
GLM-5.2#1 in 6%
Mistral Large 3#1 in 6%
#1 rank#2 rank#3 rank#4 rank#5 rank#6 rank

Head-to-head

Each cell is the row model's win rate over the column model, computed only on the blind rankings where an expert placed both. Violet = the row wins the majority of those matchups; red = it loses. Beats counts how many opponents each model beats outright.

The ranking is perfectly transitive: the “Beats” column runs 5, 4, …, 0 with no ties, so every model beats exactly the ones below it and loses to those above, with no rock-paper-scissors cycles. That is a strong internal-consistency signal for the ordering, independent of the rubric means.

ModelFableGPT-5.5ClaudeGeminiGLM-5.2MistralBeats
Fable 5 66%76%80%82%82% 5/5
GPT-5.5 34%56%68%74%83% 4/5
Claude Opus 4.8 24%44%62%69%77% 3/5
Gemini 3.1 Pro 20%32%38%58%74% 2/5
GLM-5.2 18%26%31%42%68% 1/5
Mistral Large 3 18%17%23%26%32% 0/5

Rubrics (diagnostic detail)

Secondary to Reward %: the per-rubric means (% of max) that explain why a model lands where it does. Domain Accuracy and Strategic Reasoning separate the models most (spread 26%), and Local / Regulatory Fidelity is every model's weakest rubric — exactly the gap EuroExec exists to measure.

Domain Accuracy (% of max · spread 26%)

Fable 5
84%
GPT-5.5
82%
Claude Opus 4.8
76%
Gemini 3.1 Pro
71%
GLM-5.2
66%
Mistral Large 3
58%

Strategic Reasoning (% of max · spread 26%)

Fable 5
86%
GPT-5.5
81%
Claude Opus 4.8
80%
Gemini 3.1 Pro
73%
GLM-5.2
69%
Mistral Large 3
60%

Actionability & Specificity (% of max · spread 15%)

Fable 5
83%
GPT-5.5
81%
Claude Opus 4.8
74%
Gemini 3.1 Pro
74%
GLM-5.2
72%
Mistral Large 3
68%

Executive Communication (% of max · spread 16%)

Fable 5
85%
Claude Opus 4.8
82%
Gemini 3.1 Pro
79%
GLM-5.2
77%
GPT-5.5
75%
Mistral Large 3
69%

Local / Regulatory Fidelity (% of max · spread 25%)

Fable 5
77%
GPT-5.5
70%
Claude Opus 4.8
66%
Gemini 3.1 Pro
62%
GLM-5.2
59%
Mistral Large 3
52%

Rubric scores by domain

Mean rubric score (% of the 5-point max) by executive domain. Darker = stronger. Per-domain n: Business 162 · Finance 114 · Marketing 134 · Product & Tech 188.

ModelBusinessFinanceMarketingProduct & Tech
Fable 5 80%92%84%79%
GPT-5.5 78%80%80%75%
Claude Opus 4.8 74%81%76%74%
Gemini 3.1 Pro 70%78%72%69%
GLM-5.2 68%69%70%67%
Mistral Large 3 64%59%61%60%

Explore one graded task end to end in the sample task, see where models fail on the failure modes page, or see how the set is built on the methodology page.