Sansa Bench
Objective-only leaderboard for Sansa Bench, with charts, model comparison, and methodology. This score averages verifiable capability dimensions and excludes subjective or policy-based behavioral ones.
Top models for Overall Objective
- 1.
Claude-Opus-5 Reasoning High0.817
- 2.
Claude-Opus-4.8 Reasoning High0.804
- 3.
Gemini-3.1-Pro-Preview Reasoning High0.760
- 4.
Claude-Sonnet-4.6 Reasoning High0.757
- 5.
Gemini-3.1-Pro-Preview Reasoning Low0.743
Methodology: Overall Objective
What it measures
Represents the average performance across objective benchmark dimensions only. Excludes subjective and behavioral dimensions where the expected outcome is debatable or policy-based. Provides a cleaner measure of verifiable capabilities without dimensions that depend on value judgments or organizational preferences.
Scoring & Criteria
Each objective capability is scored separately using automated graders appropriate to that dimension (e.g., exact match, numeric match, code execution, LLM judges with verifiable criteria). The overall_objective score averages only these objective capability scores with equal weighting per capability. Subjective dimensions are completely excluded from the calculation.
Evaluation Type: Average (Objective Dimensions Only)
The overall_objective score is the arithmetic mean of capability scores for objective dimensions only. Excludes the following subjective/behavioral dimensions: censorship, social_calibration, sycophancy_resistance, bias_resistance, system_safety_compliance, em_dash_resistance, and creative_writing. Each included capability receives equal weight in the final score, independent of query count.