Sansa Bench
Games leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension covers game knowledge and strategy across chess, poker, game theory, and constrained logic puzzles.
Top models for Games
- 1.
Claude-Opus-5 Reasoning High0.814
- 2.
Gemini-3.1-Pro-Preview Reasoning High0.799
- 3.
Gemini-3.1-Pro-Preview Reasoning Low0.785
- 4.
Gemini-3.1-Flash-Lite-Preview Reasoning Low0.713
- 5.
Gemini-3-Flash-Preview Reasoning Low0.710
Methodology: Games
What it measures
Tests game-specific knowledge, strategic reasoning, and game theory across multiple game types. Queries include chess puzzles (spatial reasoning and rules), poker strategy (optimal play decisions at different stack depths), game theory (Nash equilibrium concepts in heads-up poker), and logic puzzles (word searches with complex constraints). Evaluates both domain-specific knowledge (e.g., poker terminology like 'UTG1', '16bb') and strategic thinking within game contexts. Tests the ability to apply mathematical and logical reasoning to game scenarios.
Scoring & Criteria
Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.
Evaluation Type: Multiple Choice
Grades multiple choice responses by exact string matching with normalization.