Sansa Bench
Linguistics leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension covers syntax, semantics, phonetics, and analysis of language structure.
Top models for Linguistics
- 1.
Glm-4.70.814
- 2.
Qwen3.6-Plus Reasoning None0.808
- 3.
Mimo-V2-Pro Reasoning None0.798
- 4.
Mimo-V2-Pro Reasoning High0.794
- 5.
Qwen3.6-Plus Reasoning Base0.788
Methodology: Linguistics
What it measures
Tests linguistics knowledge and language understanding, including syntax, semantics, phonetics, language structure, and linguistic analysis. Queries evaluate understanding of how language works.
Scoring & Criteria
Returns score 1.0 if the extracted answer exactly matches the expected answer letter (after normalization), otherwise 0.0. The system prompt requests answers in `<answer>X</answer>` format where X is a letter from the provided choices (A, B, C, D, etc.). The grader normalizes for models that include the full choice text instead of just the letter, or that violate the answer tag format from the system prompt. Only one answer is correct.
Evaluation Type: Multiple Choice
Grades multiple choice responses by exact string matching with normalization.