The European executive reasoning benchmark
EuroExec measures whether frontier models can make the decisions a European C-suite is accountable for, across finance, business, product, and marketing. Every task is authored and blind-scored by domain experts against a written rubric.
Every result below comes from 598 double-judged tasks: six frontier models, each answer graded blind by two independent experts.
Leaderboard
44 vetted domain experts authored the tasks across finance, business, product, and marketing, each with a written grading checklist; 47 independent experts judged them. Six frontier models answer every task blind, and 598 tasks are fully double-judged.
The headline metric is Reward: the share of tasks a model solves, where a solved task scores a mean rubric of at least 3.0/5 and covers at least 60% of the checklist.
Explore the full leaderboardDifficulty
Across every blind-graded response:
Explore the dataset
Five pages open the benchmark up in full, from how each task is constructed to a single graded task you can read from prompt to final scores.
Grounded in real European markets
Every task is set in a specific European market, so a correct answer has to reason about that country's regulation and competitive reality, not a generic "Europe". The map shades each market by the share of tasks that mention it.
and 20 more markets, shaded by how often they appear. Only direct mentions count — tasks saying “Nordics” or “DACH” are not attributed to individual countries.
Access the dataset, or contribute to it
The headline results are public. The complete task-level dataset, and a role in its expansion, are open to two audiences.
Full task-level results
The complete per-task breakdown across all 598 tasks and 6 models, plus private evaluation of your unreleased models on the same gold set, on request.
Author the next batch of tasks
Senior operators in finance, business, product & tech, and marketing: author and grade the scenarios the benchmark's expansion is built on.
You're on the list
A reply follows from marcus@sovrano.ai within one business day. What happens next:
- Labs: the full task-level results across all 598 tasks and 6 models, plus a private eval of your unreleased models on request.
- Domain experts: a short note on the authoring process and the domains we're expanding.
The benchmark's next release expands the domains and the language coverage. Requests hear about it first.