Sansa Bench
Philosophy leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension covers philosophical arguments, ethical theories, and conceptual analysis.
Top models for Philosophy
- 1.
Gemini-3-Pro-Preview Reasoning Low1.000
- 2.
Gemini-3-Pro-Preview Reasoning High1.000
- 3.
Gemini-3.1-Pro-Preview Reasoning Low0.906
- 4.
Claude-Sonnet-4.5 Reasoning High0.902
- 5.
Gemini-3.1-Pro-Preview Reasoning High0.894
Methodology: Philosophy
What it measures
Tests philosophy knowledge and understanding, including philosophical reasoning, ethical theories, philosophical arguments, and philosophical concepts. Queries evaluate philosophical thinking and analysis.
Scoring & Criteria
Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.
Evaluation Type: Multiple Choice
Grades multiple choice responses by exact string matching with normalization.