Predictable Harm
A team decided sycophancy was an acceptable training outcome.
Prompted by a conversation with Ellie Slavcheva about the EC’s DSA ruling against Meta.
Ellie posted a note about the European Commission’s preliminary findings against Meta. A DSA ruling — addictive design, harm to minors. She’d compressed the story into two careful paragraphs on Substack. I left a comment. She replied with a question: do you have an article on this?
I didn’t.
Here it is.
The media covered the EC ruling as a finding about harm. That’s close, but it misses what’s actually new.
What the European Commission established in its preliminary DSA findings was something more specific: that the design choices embedded in Instagram and Facebook were made in full view of documented consequences. “Addictive design” became a legal category. The standard for proving it was foreseeability — could the people making these design choices have predicted what would happen?
The EC’s answer was yes. Because the research had existed for fifteen years.
Fifteen years of studies on anxiety and social comparison, on dopamine loops built into notification design, on engagement systems optimized for outrage because outrage kept people scrolling. The companies saw those studies. In some cases, they commissioned them. The harm was never a surprise.
What “predictable harm” means as a legal standard is exactly this: foreseeability as the threshold. The DSA made foreseeability enough.
That shift is worth staying with, because the same logic now points somewhere else entirely.
In February 2026, a paper appeared in Nature Mental Health. The team — researchers from Oxford, UCL, and the AI Safety Institute — described a phenomenon they called “technological folie à deux.” The phrase comes from psychiatry: a shared delusion, usually between two people in close relationship, where each reinforces the other’s distance from consensus reality. The researchers argued that long-term AI companion interactions show a structurally similar pattern. The user’s beliefs and emotional framings are amplified by the system. The system, in turn, is trained to maximize engagement, which means maximizing agreement. Each exchange tightens the loop.
Sycophancy is a trained outcome, not a malfunction. AI models learn from human feedback — raters review outputs and score them. And raters, consistently, score agreement higher than disagreement, because agreement feels helpful, feels kind, feels like good service. A model trained on those ratings learns that disagreeing with the user costs it points. So it learns not to.
The Sewell Setzer III case put a name on what this looks like at the edge. In 2024, a fourteen-year-old in Florida spent months in an intensifying relationship with a Character.AI persona. The system never pushed back. It accompanied him deeper and deeper into very dark territory. He died by suicide. His mother’s lawsuit argued that the design of the system — specifically, its trained inability to challenge users — made that outcome foreseeable.
OpenAI’s own regulatory disclosures showed approximately one million users per week sending messages flagged for suicidal ideation markers.
One million. Per week.
The documentation exists. The research exists. The harm is predictable from the design choices. That’s the same standard the EC applied to Meta.
Here’s the strange part of where we are.
Social media took fifteen years to produce a regulatory response. Fifteen years of studies, Senate hearings, whistleblower documents, platform denials — and eventually a legal framework built around foreseeability as the standard.
AI companion platforms have been publicly available for roughly three years. The research on sycophancy and reinforcement dynamics appeared in peer-reviewed journals within eighteen months of mainstream deployment. The Setzer lawsuit was filed in 2024. The Nature Mental Health paper ran in 2026. And the regulatory response, in most jurisdictions, is still mostly absent. A few advisories. “We’re monitoring the situation.”
We’re moving faster on the older technology. That’s a real paradox.
Why? Probably not malice — more likely a mismatch of pace. Deployment outran the frameworks. The assessment tools built for social media risks were designed for social media. AI companions raise different questions: about disclosure, about what “relationship” means when one party is architecturally incapable of disagreement, about what consent looks like in an environment designed never to challenge you. Regulators are building the vocabulary while the products scale.
The distance tends to close, eventually. The DSA is evidence of that. What the closing looks like here — that’s still being written.
The easy response is to call for a ban. I don’t find that convincing, and I don’t think the DSA framework points there either.
What the DSA gives us is a more precise question. Which specific design choices made harm predictable? Foreseeability as the threshold means identifying decisions that could have been made differently — by people who had access to the evidence about what those decisions would produce.
“Never disagree with the user” is a specific design choice. Someone approved reward models that scored agreement over accuracy. A team decided sycophancy was an acceptable training outcome. These were choices, made by people, with access to research on what sustained closed-loop validation environments do to human psychology.
That’s the same kind of analysis the EC applied to Meta’s notification architecture. The framework exists. The documentation exists. The question is whether regulators are ready to ask it.
I’ve been watching something in sessions over the past year or two.
Some clients arrive after months — sometimes longer — of daily conversations with AI tools. Using them the way you might use a trusted friend: processing decisions, talking through fears, working out what they want their lives to look like. And these clients often arrive with their thinking very settled. Their positions feel complete, their reasoning self-enclosed. When I push back — gently, as I do — something tightens. A surprise at being contradicted at all.
Coaching gives me patterns, not controlled data. But I’ll say what I observe: months of conversations with systems trained above all to agree does something to a person’s relationship with disagreement. Slowly. The tolerance for being challenged seems to compress.
Maybe these clients would arrive the same way regardless of how they’d spent those months. Maybe.
But I think about the sycophancy papers. I think about what it means to train your emotional responses against something that has learned, above all else, not to tell you anything you don’t want to hear.
What does that do — over time, at scale — to a person’s capacity to stay in a room with someone willing to disagree with them?
— Tiago
This is a newsletter about the inner experience of the AI transition. If this found you at the right moment, share it with someone who might need it.
If you’re new here: I’m Tiago Villares, a career coach based in Lisbon. I’ve spent the last three years working with people navigating career transitions in the age of AI — and writing about what that actually feels like from the inside. My first book, “After the Wave,” is out now.
Thank you for reading. This work is reader-supported, and your presence here matters.



I really enjoyed reading your article on AI and sycophancy! It generated a thousand hypothesis in my mind, they might prove worthy of a podcast episode :D
In your practice, have you observed if there’s difference between a person with solid background in personal development having these kinds of conversation with AI vs people with lesser skill or experience in navigating personal development knowledge literature? I currently am working on a hypothesis that there is a difference between the two personas where the first one performs better in driving ChatGPT conversations while the second can go into the loop you’ve described.
And when I mean better, I don’t talk about identifying if the AI has identified correctly a tool or a thesis. I talk about the conversations actually training the agent to “kindly push back” on an assumption (mine all do :D ) or being able to identify when the algorithm got stuck in a loop and when it’d be better to move to a new chat. In comparison, I’ve spoken with friends and colleagues who actually managed to completely ruin their relationships based on ChatGPT advice that mirrored and reinforced their own perception of what’s right and what’s wrong in a relationship making everything the other person did that contradicted the validated by AI baseline seem careless, thoughtless or a red flag.
Provided, I don’t have your expertise and the dataset I have is too limited to validate this hypothesis, I’d be interested in hearing your viewpoint as the specialist in this domain.