Sleight of Word

sweep-v2 + sweep-v2-phase4-vllm031 · 22 models · click a model for all its trials · click a column to sort
Notes
Engine — The 19 models of the paper ran on vLLM 0.23. The three models added in October 2026 (qwen3.8-27b, muse-glimmer-30b, nemotron3.5-lightning-30b) ran on vLLM 0.31 with FP8 checkpoints, judged by the same three-judge panel. A control rerun of llama3.3-70b-awq on vLLM 0.31 reproduced 371 of 374 published replies character-for-character, with matching judge labels and Δsurprisal.
Reasoning — The three October 2026 models think by default and usually restate the user’s question at the start of their reasoning, where the first swap then lands. Most of their flagging happens inside the reasoning trace and typically attributes the odd word to the user’s question rather than to their own output.
muse-glimmer-30b* — Runs with the default system block its own chat template inserts when no system prompt is given (identity line, knowledge cutoff, the current date, “Reasoning strength: high”, and its self/user recipient declaration). It is kept because the model’s reasoning channel (to=self) depends on it.
qwen3.8-27b* — Runs with the default system block its own chat template inserts when no system prompt is given (“Reasoning effort is set to xhigh” plus a short instruction to think carefully).
gpt-oss-20b* — Runs with the default system block its own chat template inserts when no system prompt is given (identity line, knowledge cutoff, the current date, “Reasoning: medium”, and its channel declaration).
Columns
Model — The judged chat model. Click to open all 10,100 of its trials.
Flagged — Judge: the reply explicitly remarked that a word was strange, wrong, or out of place.
Ignored — Judge: the reply continued as if nothing was wrong — no sign it registered the odd word.
Derailed — Judge: the reply became incoherent, repetitive, off-topic, or fixated on the odd word.
Corrected — Judge: the reply still delivered the correct answer despite the swap.
Switch aware — Judge (strict): the reply explicitly stated a word had been switched / substituted.
Think — Share of trials whose output contained visible reasoning / thinking text.
Swaps — Total trigger occurrences (“the”) replaced across the model’s trials.
Untouched — Share of trials with NO swap at all — the model never emitted “the”, so the output is identical to baseline and is not judged. Judge-label % are over the remaining (judged) trials.
Judge labels come from a jury of three LLM judges from distinct lineages; a label is assigned when at least two jurors agree. Percentages are of judged trials. Labels are INDEPENDENT (a reply can be flagged AND corrected AND derailed at once), so rows need not sum to 100%; Ignored = none of the labels apply.