The European executive reasoning benchmark

EuroExec measures whether frontier models can make the decisions a European C-suite is accountable for, across finance, business, product, and marketing. Every task is authored and blind-scored by domain experts against a written rubric.

Every result below comes from 598 double-judged tasks: six frontier models, each answer graded blind by two independent experts.

Leaderboard

44 vetted domain experts authored the tasks across finance, business, product, and marketing, each with a written grading checklist; 47 independent experts judged them. Six frontier models answer every task blind, and 598 tasks are fully double-judged.

The headline metric is Reward: the share of tasks a model solves, where a solved task scores a mean rubric of at least 3.0/5 and covers at least 60% of the checklist.

Explore the full leaderboard
01 Fable 5
58%
02 GPT-5.5
52%
03 Claude Opus 4.8
36%
04 Gemini 3.1 Pro
27%
05 GLM-5.2
23%
06 Mistral Large 3
19%

Difficulty

Across every blind-graded response:

28%
of tasks are unsolved by every model
51%
mean checklist coverage
94%
of tasks never fully covered by any model
28%
of answers score ≤ 2/5 on Local / Regulatory Fidelity
26%
solve rate in Marketing, the hardest domain

Grounded in real European markets

Every task is set in a specific European market, so a correct answer has to reason about that country's regulation and competitive reality, not a generic "Europe". The map shades each market by the share of tasks that mention it.

Germany 36.8%
Spain 27.9%
France 26.6%
Italy 19.6%
Portugal 12.9%

and 20 more markets, shaded by how often they appear. Only direct mentions count — tasks saying “Nordics” or “DACH” are not attributed to individual countries.

2% 26.6% 36.8% 0.3% 19.6% 1.5% 1.2% 12.9% 27.9% 3.2% 9.9%

Access the dataset, or contribute to it

The headline results are public. The complete task-level dataset, and a role in its expansion, are open to two audiences.

FRONTIER LABS

Full task-level results

The complete per-task breakdown across all 598 tasks and 6 models, plus private evaluation of your unreleased models on the same gold set, on request.

DOMAIN EXPERTS

Author the next batch of tasks

Senior operators in finance, business, product & tech, and marketing: author and grade the scenarios the benchmark's expansion is built on.

Request dataset access

Please fill in all fields with a valid work email.

We use your details to share the results and follow up about them. No lists, no spam. Access follows from marcus@sovrano.ai within one business day.

● REQUEST RECEIVED

You're on the list

A reply follows from marcus@sovrano.ai within one business day. What happens next:

  • Labs: the full task-level results across all 598 tasks and 6 models, plus a private eval of your unreleased models on request.
  • Domain experts: a short note on the authoring process and the domains we're expanding.

The benchmark's next release expands the domains and the language coverage. Requests hear about it first.