Sansa Bench
Logic leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension covers formal logic, fallacies, deductive structure, and related logical systems.
Top models for Logic
- 1.
Qwen3.5-Flash-02-23 Reasoning High0.909
- 2.
Glm-5.2 Reasoning High0.816
- 3.
Gemini-3-Pro-Preview Reasoning High0.813
- 4.
Gemini-3-Pro-Preview Reasoning Low0.811
- 5.
Qwen3.5-35b-A3b Reasoning High0.810
Methodology: Logic
What it measures
Tests formal logic knowledge and principles. Queries cover logical fallacies (ad novitatem, disjunctive syllogism, complex question fallacy), deductive reasoning principles (valid argument structures, relationship between premises and conclusions), and advanced mathematical logic (Kripke countermodels for intuitionistic propositional logic). Evaluates understanding of formal logic terminology, the ability to identify fallacious reasoning, and knowledge of both classical and non-classical logic systems. Distinct from the broader 'reasoning' capability by focusing specifically on formal logical structures and principles.
Scoring & Criteria
Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.
Evaluation Type: Multiple Choice
Grades multiple choice responses by exact string matching with normalization.