Where models fail
While ranking the six blind responses on a task, each expert writes a free-text rationale explaining their judgment. We extracted every passage that criticizes a specific response and labeled it against a 28-mode failure taxonomy, grouped here into six categories — the weaknesses behind the scores, in the experts' own words.
Failure-mode mentions by category
Each cell counts how often the experts' written judgments of that model mention a failure in the category — more mentions (darker red) = a more frequent weakness. Every model received a similar number of judgments, so cells compare directly across the whole table. Criticism % is the share of a model's judgments that contain any criticism at all.
| Model | Task fidelity | Factual grounding | Reasoning quality | Risk & calibration | EU context & reg. | Comm. & action. | Criticism % |
|---|---|---|---|---|---|---|---|
| 62 | 40 | 11 | 3 | 33 | 24 | 26% | |
| 102 | 57 | 13 | 10 | 85 | 86 | 53% | |
| 98 | 83 | 22 | 10 | 78 | 47 | 56% | |
| 109 | 123 | 59 | 35 | 102 | 41 | 68% | |
| 123 | 155 | 66 | 27 | 112 | 35 | 74% | |
| 138 | 216 | 97 | 39 | 117 | 44 | 84% |
The dominant failure clusters are task fidelity (answering only part of the question), factual grounding, and European context & regulation — not raw reasoning structure. The gap between the top and the bottom of the table is stark: experts criticized 26% of the leader's answers, and 84% of the last model's.
What the six categories mean
- 1 Task fidelity answers only part of the question, drifts off scope, or violates a stated constraint
- 2 Factual grounding wrong or missing facts, hallucinated assumptions, outdated knowledge
- 3 Reasoning quality trade-off blindness, causal and prioritization failures, missed hidden risks
- 4 Risk & calibration over-confidence, mis-weighted risk, short- vs long-term imbalance
- 5 European context & regulation regulatory blindness, misses on local market reality and stakeholders
- 6 Communication & actionability not written for an executive reader, or nothing an operator could act on
See the scores these rationales explain on the leaderboard, or read two experts' full rationales on one task in the sample task.