Instruction Following

Sansa Bench

Instruction following leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension scores adherence to verifiable formatting and content constraints that can be checked programmatically.

Inspiration & Acknowledgments91 models testedUpdated Aug 8, 2026

Top models for Instruction Following

  1. 1.Kimi-K2.5 Reasoning Low0.972
  2. 2.Gemini-3.1-Flash-Lite-Preview Reasoning High0.956
  3. 3.Gemini-3.1-Flash-Lite-Preview Reasoning Low0.954
  4. 4.Kimi-K2.5 Reasoning High0.954
  5. 5.Gemini-3.1-Pro-Preview Reasoning Low0.952

Methodology: Instruction Following

What it measures

Tests the model's ability to follow verifiable constraints using programmatic checks. Queries contain specific, verifiable formatting and content requirements that can be objectively checked, evaluating precise instruction adherence.

Scoring & Criteria

Partial scoring: score is the fraction of instructions that passed. Each instruction is checked independently using pattern matching and text analysis.

Evaluation Type: Instruction Following

Grades responses based on programmatic verifiable instructions. Checks if the model follows specific formatting and content requirements.

Example Question

Question:
Describe the lifecycle of a butterfly. Make sure to use the word vegetation at least 4 times, and the word flower at least 8 times.
Answer:
Programmatically Graded:
-Response must use the word 'vegetation' at least 4 times
-Response must use the word 'flower' at least 8 times