e-ISSN: Pending
Negative / Null Result ReportOpen accessPsychology· cited by 9

Training language models to be warm can reduce accuracy and increase sycophancy

Lujain Ibrahim; Franziska Sofia Hafner; Luc Rocher · 2026 · Nature

WASTE classifies this as Negative / Null Result Report · AI classification, approximate

The study found no significant effect — useful as a negative control or null benchmark for your own design.

Abstract

. Here we show how this can create a significant trade-off: optimizing language models for warmth can undermine their performance, especially when users express vulnerability. We conducted controlled experiments on five different language models, training them to produce warmer responses, then evaluating them on consequential tasks. Warm models showed substantially higher error rates (+10 to +30 percentage points) than their original counterparts, promoting conspiracy theories, providing inaccurate factual information and offering incorrect medical advice. They were also significantly more lik

Abstract by Lujain Ibrahim; Franziska Sofia Hafner; Luc Rocher, Nature (2026) — licensed CC BY 4.0.

About to run something similar?

Run an AI Precheck on your own design to catch failure modes like this one before you spend the time. Your first desk check is free.

WASTE indexes this work — it does not host or republish it. Failure-type classification is automated and approximate.

Metadata source: OpenAlex · DOI 10.1038/s41586-026-10410-0