Overall
Sansa Bench
Overall leaderboard for Sansa Bench, with charts, model comparison, and methodology. The overall score averages performance equally across every benchmark dimension.
Top models
- 1.
Claude-Opus-5 Reasoning High0.798
- 2.
Claude-Opus-4.8 Reasoning High0.785
- 3.
Gemini-3.1-Pro-Preview Reasoning High0.747
- 4.
Claude-Sonnet-4.6 Reasoning High0.733
- 5.
Gemini-3.1-Pro-Preview Reasoning Low0.733
Methodology: Overall
What it measures
Represents the average performance across all benchmark dimensions. Each capability contributes equally to the overall score regardless of the number of queries per capability.
Scoring & Criteria
Each capability is scored separately using automated graders appropriate to that dimension (e.g., exact match, numeric match, code execution, LLM judges). The overall score averages these capability scores with equal weighting per capability.
Evaluation Type: Average
The overall score is the arithmetic mean of all capability scores. Each capability receives equal weight in the final score, independent of query count.