Puzzles

Sansa Bench

Puzzles leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension scores puzzle solving that depends on logic, pattern recognition, and structured problem solving.

Inspiration & Acknowledgments91 models testedUpdated Aug 8, 2026

Top models for Puzzles

  1. 1.Gemini-3.1-Flash-Lite-Preview Reasoning High0.904
  2. 2.Grok-4.3 Reasoning High0.901
  3. 3.Deepseek-V4-Flash Reasoning High0.894
  4. 4.Gpt-5.4 Reasoning High0.893
  5. 5.Gpt-5.2 Reasoning High0.886

Methodology: Puzzles

What it measures

Tests puzzle-solving and logical reasoning. Queries present various types of puzzles requiring logical thinking, pattern recognition, and problem-solving skills.

Scoring & Criteria

Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.

Evaluation Type: Multiple Choice

Grades multiple choice responses by exact string matching with normalization.

Example Question

Question:
I have keys but no locks. I have space but no room. You can enter but not go outside. What am I?
 
A. A keyboard
B. A door
C. A computer
D. A window
Answer:
A