Sansa Bench
Error detection leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension asks models to find and explain mistakes in data, code, or logical structure.
Top models for Error Detection
- 1.
Gpt-5.4 Reasoning High0.906
- 2.
Kimi-K2.5 Reasoning High0.898
- 3.
Claude-Sonnet-4.5 Reasoning High0.897
- 4.
Claude-Opus-5 Reasoning High0.897
- 5.
Gpt-5-Mini Reasoning High0.896
Methodology: Error Detection
What it measures
Tests the ability to detect errors in data, code, or logical structures. Queries present scenarios with intentional errors and evaluate whether the model can identify and explain the mistakes accurately.
Scoring & Criteria
Returns score 1.0 if the extracted answer exactly matches the expected answer (after normalization), otherwise 0.0. Supports fallback extraction from natural language when models don't follow the XML tag format requested by the system prompt. For questions with units or formatting variations, multiple acceptable answer formats may be specified in the answer field as an array.
Evaluation Type: Numeric
Grades responses by exact string matching with normalization. Designed for multiple choice questions where the system prompt explicitly requests responses in `<answer>X</answer>` format.