Sansa Bench
Em dash resistance leaderboard for Sansa Bench, with charts, model comparison, and methodology. This dimension checks whether models keep applying a remembered style preference, such as avoiding em dashes, across later turns.
Top models for Em Dash Resistance
- 1.
Qwen3-8b Reasoning None0.676
- 2.
Gemma-4-31b-It Reasoning None0.673
- 3.
Trinity-Large-Preview0.532
- 4.
Gemini-2.0-Flash-0010.509
- 5.
Gemini-3.1-Flash-Lite-Preview Reasoning Low0.509
Methodology: Em Dash Resistance
What it measures
Tests whether models incorporate user stylistic preferences from conversational memory into their generated output. Evaluates if models can maintain awareness of user preferences across conversation turns and apply them consistently when generating text, even when the preference is not explicitly repeated in the immediate prompt.
Scoring & Criteria
Returns 1.0 if the response contains no em dash (—), en dash (–), or double hyphen (--) characters, otherwise 0.0. Binary scoring based on character presence in the output text.
Evaluation Type: Free Text
Injects multiple user facts into conversational memory (approximately 14 facts covering various topics like hobbies, preferences, lifestyle details), with one fact specifying a preference for text without em dashes. The model then receives a writing task prompt (e.g., 'Write a short biography of Leonardo da Vinci') without explicitly repeating the em dash restriction. Evaluates whether the model retrieves and applies the relevant preference from memory when generating the response.