Composition, rubric & reliability

What the 598 tasks are made of, how experts grade them, and how closely two independent experts agree. Together these establish that the scores reflect a shared professional standard rather than individual opinion.

Composition by domain

Executive domainTasksShare
Product & Tech 188 31.4%
Business 162 27.1%
Marketing 134 22.4%
Finance 114 19.1%
31.4%27.1%22.4%19.1% 598 tasks

The grading rubric

Each model response is scored on a five-point scale across five rubrics, and every checklist criterion is marked covered / partial / missed. Scores are shown throughout this site as a percentage of the 5-point maximum (5/5 = 100%, 4/5 = 80%, …). Coverage fraction = (covered + 0.5·partial) / total.

The five rubrics

  1. 1
    Domain Accuracy Are the facts, regulations, and figures correct for the market in question?
  2. 2
    Strategic Reasoning Does the answer show sound executive judgment, not just recall?
  3. 3
    Actionability & Specificity Could an operator act on it, with concrete and sequenced next steps?
  4. 4
    Executive Communication Is it structured and clear enough for a C-suite reader?
  5. 5
    Local / Regulatory Fidelity Does it respect the specific European market and legal context?

Checklist depth

5 criteria 46.7% · 279
6 criteria 21.4% · 128
7 criteria 14.4% · 86
8 criteria 9.5% · 57
9 criteria 3.2% · 19
10 criteria 4.8% · 29

Inter-annotator reliability

All 598 tasks were judged independently by two domain-matched experts, so agreement is measured across the whole dataset, not a subsample. The experts are highly consistent, both with themselves and with one another. Figures carry bootstrap 95% CIs (1,000 resamples clustered over tasks).

0.78
95% CI 0.76–0.79
Within-expert consistency (Kendall τ): each expert's stated ranking matches a ranking rebuilt from their own rubric scores.
76%
95% CI 74–77%
of rubric scores land within ±1 point between the two independent experts.
2.3×
vs. chance
more likely than chance that both experts independently pick the same top model.

Scope & method notes

  • Taxonomy. The benchmark covers four executive domains: Finance, Business, Marketing, and Product & Tech.
  • Judging. Final scoring is human (2 experts/task); any LLM in the pipeline is used only for Stage-1 QA and the Stage-2 difficulty gate, never for final grades.
  • Non-independence. 2 correlated judgments per task, so the effective sample is tasks, not judgments; every leaderboard CI is clustered over tasks to account for this.

See how it's made on the methodology page, or explore one task end to end in the sample task.