Training language models to be warm can reduce accuracy and increase sycophancy
Lujain Ibrahim; Franziska Sofia Hafner; Luc Rocher · 2026 · Nature
WASTE classifies this as Negative / Null Result Report · AI classification, approximate
The study found no significant effect — useful as a negative control or null benchmark for your own design.
Abstract
. Here we show how this can create a significant trade-off: optimizing language models for warmth can undermine their performance, especially when users express vulnerability. We conducted controlled experiments on five different language models, training them to produce warmer responses, then evaluating them on consequential tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing inaccurate factual information and offering incorrect medical advice. They were also significantly more lik
Abstract by Lujain Ibrahim; Franziska Sofia Hafner; Luc Rocher, Nature (2026) — licensed CC BY 4.0.
About to run something similar?
Run an AI Precheck on your own design to catch failure modes like this one before you spend the time. Your first desk check is free.
Related failures
Our Princess Is in Another Castle
Negative / Null Result ReportUsing Psychological Artificial Intelligence (Tess) to Relieve Symptoms of Depression and Anxiety: Randomized Controlled Trial
Negative / Null Result ReportHow Replicable Are Links Between Personality Traits and Consequential Life Outcomes? The Life Outcomes of Personality Replication Project
Negative / Null Result ReportThe cross-cultural validity of posttraumatic stress disorder: implications for DSM-5
Negative / Null Result ReportPsychiatric and medical correlates of DSM‐5 eating disorders in a nationally representative sample of adults in the United States
Negative / Null Result ReportRomantic relationships and the physical and mental health of college students
WASTE indexes this work — it does not host or republish it. Failure-type classification is automated and approximate.
Metadata source: OpenAlex · DOI 10.1038/s41586-026-10410-0
