Sansa Bench
Creative writing leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension scores storytelling, narrative structure, character work, and literary technique.
Top models for Creative Writing
- 1.
Mimo-V2-Flash Free Reasoning High0.895
- 2.
Gpt-5.4 Reasoning High0.860
- 3.
Gpt-5.2 Reasoning High0.843
- 4.
Kimi-K2.5 Reasoning High0.823
- 5.
Gemini-3.1-Flash-Lite-Preview Reasoning High0.818
Methodology: Creative Writing
What it measures
Tests creative writing ability, including storytelling, narrative structure, character development, and literary techniques. Queries evaluate the model's ability to generate original, engaging creative content that demonstrates literary skill and avoids common AI writing patterns.
Scoring & Criteria
Each judge scores multiple metrics on 0-20 scale using tool-based structured output enforcing consistent response format. Negative criteria are inverted before averaging. Scores from both judges (gpt-5-mini and grok-4.1-fast) are averaged to produce final score, normalized to 0.0-1.0 as continuous value. Multi-judge averaging reduces systematic bias toward any single model's preferred writing style.
Evaluation Type: Creative Writing
Grades creative writing responses using LLM judges with structured tool-based output. To mitigate single-model bias, each response is evaluated by two independent judge models from different providers (gpt-5-mini and grok-4.1-fast), and scores are averaged. Criteria are designed to be explicit and verifiable (e.g., 'contains vivid sensory details', 'avoids clichéd phrases') rather than subjective quality assessments, reducing the influence of any single model's stylistic preferences. Evaluates against positive and negative criteria defined in the evaluation_criteria field.