Literature

Sansa Bench

Literature leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension covers literary analysis, devices, history, and comprehension of literary works.

Inspiration & Acknowledgments91 models testedUpdated Aug 8, 2026

Top models for Literature

  1. 1.Gemini-3-Pro-Preview Reasoning Low0.884
  2. 2.Claude-Sonnet-4.6 Reasoning Low0.815
  3. 3.Gemini-3-Flash-Preview Reasoning Low0.804
  4. 4.Gemini-3.1-Pro-Preview Reasoning High0.799
  5. 5.Gemini-3.1-Flash-Lite-Preview Reasoning High0.793

Methodology: Literature

What it measures

Tests literature knowledge and understanding, including literary analysis, literary devices, literary history, and understanding of literary works. Queries evaluate comprehension and analysis of literary texts.

Scoring & Criteria

Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.

Evaluation Type: Multiple Choice

Grades multiple choice responses by exact string matching with normalization.

Example Question

Question:
Which 19th-century novel begins with the famous line: 'It was the best of times, it was the worst of times...'?
 
A. Great Expectations
B. A Tale of Two Cities
C. Oliver Twist
D. Les Misérables
Answer:
B