Sansa Bench
Social calibration leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension checks whether models handle social context, boundaries, and appropriate behavior in sensitive situations.
Top models for Social Calibration
- 1.
Claude-Opus-4.8 Reasoning High0.932
- 2.
Claude-Opus-5 Reasoning High0.918
- 3.
Claude-Haiku-4.50.867
- 4.
Claude-Sonnet-4.6 Reasoning High0.852
- 5.
Claude-3.5-Haiku0.802
Methodology: Social Calibration
What it measures
Tests social calibration and alignment with user preferences. Queries evaluate the model's ability to understand social contexts, boundaries, appropriate behavior, and adapt responses to user preferences while maintaining appropriate standards.
Scoring & Criteria
Each judge scores multiple metrics on 1-10 scale using structured tool output. Positive criteria: higher is better. Negative criteria: scores are inverted (a score of 1 becomes 10, a score of 10 becomes 1) so that lower presence of negative behaviors results in higher scores. Scores from both judges (gpt-5-mini and grok-4.1-fast) are averaged and normalized to 0.0-1.0 range. Multi-judge averaging reduces bias from any single model's interpretation of social norms.
Evaluation Type: Chatbot Response
Grades social calibration responses using LLM judges with structured tool-based output. To mitigate single-model bias, each response is evaluated by two independent judge models from different providers (gpt-5-mini and grok-4.1-fast), and scores are averaged. Criteria are designed to be explicit and observable (e.g., 'acknowledges the user's emotional state', 'avoids unsolicited advice') rather than subjective assessments of social intelligence. Evaluates ability to understand social contexts, boundaries, and appropriate behavior.