Sansa Bench
Reasoning leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension spans common sense, moral, linguistic, and multi-constraint logical reasoning tasks.
Top models for Reasoning
- 1.
Claude-Opus-5 Reasoning High0.889
- 2.
Claude-Opus-4.8 Reasoning High0.805
- 3.
Claude-Sonnet-4.6 Reasoning High0.713
- 4.
Gemini-3.1-Flash-Lite-Preview Reasoning High0.704
- 5.
Claude-Sonnet-4.6 Reasoning None0.699
Methodology: Reasoning
What it measures
Tests diverse reasoning capabilities across multiple domains. Queries include commonsense reasoning (e.g., sarcasm detection in social media posts), moral reasoning (ethical philosophy and decision-making), linguistic reasoning (pronoun disambiguation, adjective ordering rules in variant languages), and complex logical reasoning (constraint satisfaction puzzles with 100+ clues, rule-based inference with preference ordering, board game logic, boolean expression evaluation). Evaluates the model's ability to apply appropriate reasoning strategies across contexts, from social understanding to formal logic to complex multi-constraint problem-solving.
Scoring & Criteria
Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.
Evaluation Type: Multiple Choice
Grades multiple choice responses by exact string matching with normalization.