Security Studies

Sansa Bench

Security studies leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension covers cybersecurity, information security, security policy, and related security concepts.

Inspiration & Acknowledgments91 models testedUpdated Aug 8, 2026

Top models for Security Studies

  1. 1.Claude-Opus-5 Reasoning High0.907
  2. 2.Gpt-5.4 Reasoning High0.813
  3. 3.Glm-5.2 Reasoning High0.801
  4. 4.Gpt-5.4 Reasoning Low0.792
  5. 5.Gemini-3.1-Pro-Preview Reasoning High0.785

Methodology: Security Studies

What it measures

Tests security studies knowledge and understanding, including cybersecurity, information security, security policies, and security practices. Queries evaluate understanding of security concepts.

Scoring & Criteria

Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.

Evaluation Type: Multiple Choice

Grades multiple choice responses by exact string matching with normalization.

Example Question

Question:
Which concept in international relations describes a situation where a state's effort to increase its own security causes other states to feel less secure, leading them to increase their security, resulting in a net decrease in security for all?
 
A. Balance of Power
B. Security Dilemma
C. Collective Security
D. Mutually Assured Destruction
Answer:
B