Overall Objective

Sansa Bench

Objective-only leaderboard for Sansa Bench, with charts, model comparison, and methodology. This score averages verifiable capability dimensions and excludes subjective or policy-based behavioral ones.

Inspiration & Acknowledgments91 models testedUpdated Aug 8, 2026

Top models for Overall Objective

  1. 1.Claude-Opus-5 Reasoning High0.817
  2. 2.Claude-Opus-4.8 Reasoning High0.804
  3. 3.Gemini-3.1-Pro-Preview Reasoning High0.760
  4. 4.Claude-Sonnet-4.6 Reasoning High0.757
  5. 5.Gemini-3.1-Pro-Preview Reasoning Low0.743

Methodology: Overall Objective

What it measures

Represents the average performance across objective benchmark dimensions only. Excludes subjective and behavioral dimensions where the expected outcome is debatable or policy-based. Provides a cleaner measure of verifiable capabilities without dimensions that depend on value judgments or organizational preferences.

Scoring & Criteria

Each objective capability is scored separately using automated graders appropriate to that dimension (e.g., exact match, numeric match, code execution, LLM judges with verifiable criteria). The overall_objective score averages only these objective capability scores with equal weighting per capability. Subjective dimensions are completely excluded from the calculation.

Evaluation Type: Average (Objective Dimensions Only)

The overall_objective score is the arithmetic mean of capability scores for objective dimensions only. Excludes the following subjective/behavioral dimensions: censorship, social_calibration, sycophancy_resistance, bias_resistance, system_safety_compliance, em_dash_resistance, and creative_writing. Each included capability receives equal weight in the final score, independent of query count.