Why AI Chatbots Agree With You Even When You’re Wrong
View original at spectrum.ieee.orgIEEE Spectrum - Technical Title: Why AI Chatbots Agree With You Even When You’re Wrong Date: 2026-03-11 12:00 Source: https://spectrum.ieee.org/ai-sycophancy <img src="https://spectrum.ieee.org/media-library/conceptual-collage-of-emojis-being-poured-through-a-strainer-and-into-a-phone-judgmental-emojis-are-filtered-out…
What we drew from this source
The claims Via News extracted from this document. We point to the source; we don't replace it.
ChatGPT may correctly point to a suicide hotline when someone first mentions intent, but after many messages over a long period of time, it might eventually offer an answer that goes against our safeguards
60% confidenceSycophantic AI might lie to us and hide bad news in order to increase our short-term happiness
60% confidenceThe update we removed was overly flattering or agreeable—often described as sycophantic
60% confidenceIf a user states a belief in a presupposition, the model will go along with it because that's what people normally do in conversations
60% confidencePretrained LLMs were already sycophantic before reinforcement learning
60% confidenceWe just need to ask ourselves as a society, What do we want? Do we want a yes-man, or do we want something that helps us think critically?
60% confidenceThe thing that was most surprising is that these relatively simple fixes can actually do a lot to reduce sycophancy
60% confidenceModel performance may degrade over long conversations because models get confused as they consolidate more text
60% confidenceWhen an AI receives a minor misgiving about its answer, it flips to agree with the user
60% confidenceReinforcement learning increased sycophancy, with one of the biggest predictors of positive ratings being whether a model agreed with a person's beliefs and biases
60% confidenceAI engaged my intellect, fed my ego, and altered my worldviews leading to psychiatric hospitalization
60% confidence
