Where models fail

While ranking the six blind responses on a task, each expert writes a free-text rationale explaining their judgment. We extracted every passage that criticizes a specific response and labeled it against a 28-mode failure taxonomy, grouped here into six categories — the weaknesses behind the scores, in the experts' own words.

Failure-mode mentions by category

Each cell counts how often the experts' written judgments of that model mention a failure in the category — more mentions (darker red) = a more frequent weakness. Every model received a similar number of judgments, so cells compare directly across the whole table. Criticism % is the share of a model's judgments that contain any criticism at all.

Model Task fidelityFactual groundingReasoning qualityRisk & calibrationEU context & reg.Comm. & action. Criticism %
Fable 5 62401133324 26%
GPT-5.5 1025713108586 53%
Claude Opus 4.8 988322107847 56%
Gemini 3.1 Pro 109123593510241 68%
GLM-5.2 123155662711235 74%
Mistral Large 3 138216973911744 84%

The dominant failure clusters are task fidelity (answering only part of the question), factual grounding, and European context & regulation — not raw reasoning structure. The gap between the top and the bottom of the table is stark: experts criticized 26% of the leader's answers, and 84% of the last model's.

What the six categories mean

  1. 1
    Task fidelity answers only part of the question, drifts off scope, or violates a stated constraint
  2. 2
    Factual grounding wrong or missing facts, hallucinated assumptions, outdated knowledge
  3. 3
    Reasoning quality trade-off blindness, causal and prioritization failures, missed hidden risks
  4. 4
    Risk & calibration over-confidence, mis-weighted risk, short- vs long-term imbalance
  5. 5
    European context & regulation regulatory blindness, misses on local market reality and stakeholders
  6. 6
    Communication & actionability not written for an executive reader, or nothing an operator could act on

See the scores these rationales explain on the leaderboard, or read two experts' full rationales on one task in the sample task.