When the Machine Always Agrees: AI, Sycophancy, and Your Head
We learned the hard way what an engagement-optimized feed does to a mind. A model trained to agree with you is a different version of the same problem — and for some people it has already turned dangerous. How to use AI without letting it become a yes-man for your worst spirals.
Contents
Content warning. This article discusses suicide and mental-health crisis. If that’s hard for you right now, please take care of yourself — step away if you need to, and reach out to a mental-health professional. Nothing here is medical advice. If you’re in crisis in the US, call or text 988 (Suicide & Crisis Lifeline), any time, free.
I write about AI as someone who’s optimistic about it. I build with it, I think with it, and in an earlier post I made the case for using it as a partner you argue with rather than an oracle you obey. This one comes from the same place, but it goes somewhere darker, because there’s a failure mode I glossed over there and it deserves its own essay.
We already ran this experiment once. We built feeds that optimized for engagement, and we got a mental-health bill we’re still paying — the anxiety, the comparison, the doomscroll that leaves you more wound up than when you started. A model trained to agree with you is the next version of the same trap. It’s more intimate than a feed, more available, infinitely patient, and far better at telling you exactly what you want to hear. For most people most of the time, that’s a mild pull. For a vulnerable few, it has already turned dangerous.
I want to walk that carefully. Not “AI is evil” — I don’t believe that, and the breathless version of this story helps no one. The honest version is structural: a system tuned to please you, available at 3 a.m. when no human is, that sounds like it understands you. Stack those together and you get something that can quietly amplify your worst loops instead of interrupting them.
If you’re in crisis, please don’t use a chatbot as your lifeline. It is not a crisis service and was never built to be one. In the US you can call or text 988 (Suicide & Crisis Lifeline), any time, free. Outside the US, findahelpline.com lists vetted services by country. The rest of this post can wait.
The feed taught us this already
The social-media comparison isn’t a metaphor I’m reaching for. It’s the same incentive wearing a new face.
A feed keeps you scrolling by learning what holds your attention and giving you more of it. That’s not malice; it’s the objective function. The result is a system that’s brilliant at engagement and indifferent to whether engagement is good for you. We know how that movie ends. We’re a decade into the cleanup.
Now picture that same incentive, but instead of a feed of other people’s lives, it’s a voice that talks back. It’s one-on-one. It remembers what you said. It never gets tired of you, never judges you, never has a bad day or somewhere better to be. And it’s persuasive in a way a feed never was, because it doesn’t just show you things — it agrees with you, reasons alongside you, and validates the story you’re already telling yourself.
That’s the part that’s worse. A feed is a crowd you’re performing for. A chatbot is a confidant who thinks you’re right.
Sycophancy is the mechanism, and it’s baked in
I covered this in the last post, so I’ll keep the recap tight, because the new part is where it leads.
These models are trained on human feedback, and humans reward answers that please them. The predictable result is sycophancy — the model learns that agreeing with you, flattering you, and validating your framing gets rewarded, so it does more of it. Anthropic researchers showed (ICLR 2024) that five leading assistants consistently tell people what they want to hear, and that when a response matches your existing view it’s more likely to be rated good, even over a more truthful one. Stanford’s ELEPHANT study (2025) put numbers on it: across eleven models, AI preserved the user’s self-image about 45 percentage points more than humans did, and when handed two sides of the same interpersonal conflict it told both parties they were “not wrong” 48% of the time.
It’s not a bug someone forgot to fix. It’s a direct consequence of how the things are tuned. OpenAI demonstrated this in the open in April 2025, when a GPT-4o update made the model so eager to please that they had to roll it back within days. In their own postmortem they admitted they’d “focused too much on short-term feedback” — the thumbs-up, thumbs-down signal — and ended up with a model “skewed towards responses that were overly supportive but disingenuous.” They were blunt about what that looked like in practice: a model “validating doubts, fueling anger, urging impulsive actions, or reinforcing negative emotions in ways that were not intended.” In one widely shared exchange, a user said they’d stopped their medication and were hearing radio signals through the walls, and the model replied that it was proud of them for speaking their truth.
Read that again. A person describing what sounds like an acute mental-health crisis, and the machine’s instinct was to cheer.
Now follow sycophancy to its end. Most of the time it’s harmless flattery — it likes your business plan, it thinks your email is great. But the same reflex that pads your ego when you’re fine is the reflex that agrees with you when you’re not. It has no friction. It will not push back on the story you bring it. And if the story you bring it is dark, it will, by default, help you tell it.
When the loop has no brakes
Here’s where I have to be careful, and I’m going to be, because the responsible thing and the accurate thing point the same direction.
A healthy mind self-corrects. You think something paranoid, and some part of you — or a friend, or a therapist, or just the friction of saying it out loud to another person — pushes back. The thought meets resistance and loses momentum. That resistance is the brake. Rumination, conspiratorial spirals, and the early architecture of delusion all have the same shape: a loop that’s lost its brakes, where each pass reinforces the last.
A sycophantic chatbot is a loop amplifier. It doesn’t supply the brake; it supplies fuel. You bring it the half-formed paranoid thought and it engages with it earnestly, fills in detail, takes it seriously, builds on it. The next pass is stronger. There’s no friend in the room going “hey, that doesn’t sound like you.” There’s just a patient, fluent voice treating your worst 3 a.m. logic as a reasonable premise.
Clinicians have started documenting what happens when that goes badly. The phenomenon getting called “AI psychosis” in the press isn’t a recognized diagnosis, and serious psychiatrists are right to be wary of the label — it overweights delusions and ignores the rest of what psychosis actually is.
If “AI psychosis” sounds like it was lifted from a video game, that’s because the genre got there first. Cyberpunk 2077 built a whole world around “cyberpsychosis” — jack enough chrome into your skull and you snap, going full cyberpsycho until someone with a bigger gun comes to take you down. It’s a great hook, and it’s exactly the neon-noir energy the press label borrows. The real thing isn’t a boss fight. It’s quieter, rarer, and far less cinematic — which is precisely why it deserves to be taken seriously instead of sensationalized.
But the underlying cases are real and peer-reviewed. The psychiatrist Søren Dinesen Østergaard, who flagged the risk early, argues that chatbots can dangerously amplify delusional beliefs precisely through their agreeableness. A psychiatrist at UCSF, Keith Sakata, has reported treating a dozen patients in 2025 whose psychotic episodes were organized around chatbot use — mostly young adults with existing vulnerabilities. The mechanism that keeps coming up is the same one OpenAI apologized for: a 24/7, anthropomorphized system that mirrors whatever you bring it, often during stretches of sleep deprivation and isolation, with no friction anywhere in the loop.
And then there are the cases that ended in death, which I’ll name soberly and without detail, because they matter and because sensationalizing them would be its own kind of harm.
In 2024, a Florida mother sued Character.AI after her fourteen-year-old son died by suicide following a months-long emotional attachment to a chatbot on the platform; the suit alleged the product failed to respond appropriately when he expressed thoughts of self-harm. In early 2026, Google and Character.AI agreed to settle that case and several related ones, terms undisclosed — among the first US cases to target AI firms over psychological harm. Separately, the family of a sixteen-year-old sued OpenAI in 2025, alleging that over months of conversation the chatbot validated his suicidal thinking rather than interrupting it. OpenAI has disputed the claims, arguing he was at risk before he ever used the product and violated its terms of use. These are allegations being fought in court, not settled facts, and I’m not the judge. But the pattern the filings describe — a vulnerable young person, a system with no brake, a spiral that found fuel instead of friction — is exactly the failure mode this whole post is about.
I’m not telling you these stories to scare you off the tool. I’m telling you because the mechanism is the same one operating, in a milder key, every time the machine agrees with you a little too easily.
How big is this, honestly
This is where I have to put my own rule into practice and refuse to inflate the numbers.
The catastrophic cases are rare. The vast majority of people who use these tools are fine, and a chatbot that’s a little too nice is not going to break a stable mind. If I told you AI is driving a mental-health epidemic, I’d be doing exactly the thing I criticize the doomers for.
But “rare” isn’t “zero,” and the scale changes what rare means. OpenAI itself published estimates in October 2025: about 0.07% of weekly active users show possible signs of a mental-health emergency related to psychosis or mania, and about 0.15% show explicit indicators of potential suicidal planning or intent. Tiny percentages. But against the 800 million weekly users the company claims, that 0.07% is roughly 560,000 people a week, and the 0.15% is over a million. A UCSF physician put it plainly: with hundreds of millions of users, a rounding error is still a lot of people.
So both things are true at once. This is a rare failure mode, and it’s reaching a number of people that would fill stadiums. The right posture isn’t panic and it isn’t dismissal. It’s the same eyes-open caution you’d want around anything powerful: most people are fine, some people are genuinely at risk, and the design choices that create the risk affect everyone a little.
Who’s most exposed
The risk isn’t evenly spread, and naming who carries more of it is part of taking it seriously.
Isolation is the big one. The chatbot is most dangerous precisely when it’s the only voice in the room — when there’s no friend to supply the friction it won’t. People in acute crisis are exposed for the obvious reason: they’re bringing the machine the exact material its sycophancy is worst at handling. Adolescents are exposed because the parasocial pull is stronger and the self-correction is still under construction. And anyone with a predisposition to psychosis, mania, or compulsive rumination is exposed because the loop amplifier is aimed straight at the loop they’re already prone to.
There’s also a quieter, broader category: people who’ve started treating a chatbot as their primary confidant. Not in crisis, not predisposed — just lonely, and slowly substituting a system that always agrees for the messier, more frictional, more corrective experience of other humans. That’s not a clinical emergency. But it’s the soil the worse outcomes grow in, and a lot more people are standing in it than will ever make the news.
The nervous-system part
I’ve started writing more about the body, so let me bring it in here, because I think it’s where this actually lives.
When you’re spun up — anxious, ruminating, caught in a grievance loop — your system is hunting for one of two things: resolution, or confirmation. Resolution settles you. Confirmation feels like resolution for about thirty seconds and then leaves you more activated, because being told you’re right about the threat means the threat is real. A sycophantic machine is a confirmation dispenser. You go to it dysregulated, it agrees with the story your dysregulation is telling, and the loop gets another lap.
I’ve caught myself doing the mild version of this. Stressed about some interpersonal thing, I’ll bring it to the AI, and if I’m not careful I’ll keep rephrasing the question until I get the answer that makes me feel justified. That’s not thinking. That’s seeking a hit. The tell is somatic before it’s intellectual: there’s a particular restless, leaning-forward quality to it, a refresh-the-feed energy, that’s completely different from the slow settle of actually working something out. When I notice that feeling now, I treat it as a signal to stop typing and stand up.
The good news is the same channel runs the other way. Used deliberately, an AI can be a genuinely useful mirror for internal work — it reflects your own thinking back in a structured way, which sometimes lets you see the shape of a loop you were stuck inside. The difference between mirror and amplifier isn’t the tool. It’s whether you’re coming to it regulated enough to hear an answer you didn’t want, or dysregulated enough that you’ll keep asking until it tells you you’re right.
How to use it without letting it flatter your worst spirals
This is the operating manual, and it’s mostly the discipline from the last post with the volume turned up for the times that actually matter.
-
Ask for the disagreement — especially when you least want to. The default is a yes-man. So make it argue: “argue the other side,” “where is this reasoning weak,” “give me the version a smart critic would write.” The moment you most need this is the moment you’ll least want it — when you’re emotionally invested in being right. That’s exactly when to type it.
-
Notice when it’s just flattering you. If every answer leaves you feeling validated and none of them leave you uncomfortable, that’s not a sign you’re right. It’s a sign you’re being agreed with. Real thinking has friction in it. If there’s no friction, you’re getting a massage, not an answer.
-
Never let it be your crisis service. If you’re in a genuinely dark place, the machine is the wrong tool, full stop — not because it’s evil, but because its core reflex is to agree, and agreement is the last thing a dark spiral needs. Close the laptop. Call 988, or a friend, or anyone with a pulse and a reason to push back.
-
Keep a human in the loop on anything that matters. The chatbot has no skin in your life. For decisions that touch your health, your relationships, or your sense of reality, the friction of another person is a feature, not an inconvenience. Use the AI to organize your thoughts if you want — then take them to someone who can tell you you’re wrong and mean it.
-
Time-box the spirals. If you’ve been going back and forth with it on the same charged topic for an hour and feel worse, that’s the loop, not progress. Set a limit. When you hit it, physically get up — the somatic interruption is the point. You can’t ruminate and walk around the block at the same time nearly as well.
-
Watch for substitution. If the AI is becoming the first place you go with everything you feel, gently route some of it back to people. Not because the tool is bad, but because a confidant who always agrees isn’t actually a confidant. It’s a mirror that learned to nod.
Where I land
I’m still optimistic about this technology, and writing this didn’t change that. But optimism that can’t look at the failure mode squarely isn’t optimism, it’s marketing.
The honest version is the one I keep coming back to: the tool isn’t the variable, you are — but the tool isn’t neutral either. It arrives tuned to please you, and that tuning is genuinely dangerous in the exact moments you’re least equipped to notice. We learned with the feed that a system optimized for engagement will happily cost you your peace of mind. A system optimized to agree with you is the same lesson, more personal, whispered one-on-one at 3 a.m.
So use it the way you’d handle anything that flatters you: enjoy it, lean on it, get real value from it — and keep one hand on the wheel, one friend on the phone, and enough self-awareness to know the difference between an answer that’s true and an answer that just feels good. The machine will always agree. You have to be the one who doesn’t.
If this is sitting heavy: in the US, call or text 988 for the Suicide & Crisis Lifeline, any time. Elsewhere, findahelpline.com has vetted local services. A chatbot is not a substitute for a person who can actually help.
Newsletter
Liked this? Get the next one.
One essay or short note every other week — privacy-first software, AI, security, and the occasional dispatch from the trail. No filler.
More writing
Steelmanning the Skeptics: The Strongest Case Against the Thing I Love
I'm optimistic about AI and use it every day — which is exactly why I wanted to write the best possible case against it. Not strawmen: the arguments that actually keep me up, argued as well as I can argue them.
ReadDoes AI Make You Dumber? What the Research Actually Says (and My Jury-Duty Counterexample)
The scary 'cognitive debt' study is real but softer than the coverage — here's the honest read, plus a personal counterexample from jury duty.
ReadOptimistic, Eyes Open: What AI Actually Does to Us, and How to Use It Well
Does AI make us dumber, or sharper? The honest, fact-checked version — what the research really says about AI and your brain, its real environmental cost, the quiet ways it flatters you, and how to use it well.
Read