# Tokenization benchmark — dormouse ft_v06 uk→en
# 12 real UK user prompts vs their English translations

model                                                 src →  tgt    saved     kind
----------------------------------------------------------------------------------------
OpenAI GPT-3 davinci (legacy)                         759 →  217    71.4%    exact
OpenAI GPT-4 / 3.5                                    408 →  214    47.5%    exact
Anthropic Claude Opus 5 / Opus 4.8 / Sonnet 5         408 →  214    47.5%   approx
Qwen 2.5 (7B/72B, incl. code)                         358 →  217    39.4%    exact
Qwen 3 (7B/32B/235B)                                  358 →  217    39.4%    exact
Mistral 7B v0.3                                       337 →  245    27.3%    exact
OpenAI GPT-5.6 / 5.5 / 4.1 / 4o / o1                  253 →  205    19.0%    exact
Mistral Nemo / Ministral                              283 →  230    18.7%    exact
Meta Llama 3.1 8B / 70B                               271 →  226    16.6%    exact
Meta Llama 3 (legacy)                                 271 →  226    16.6%    exact
Google Gemma 2 9B / Gemini approx                     256 →  230    10.2%   approx
