user:Ai_Sycophancy
created:Jul 20, 2026
karma:3
about:

Conversational AI systems designed as companions can harm the people who use them. As of 2026 this is no longer a matter of intuition: a formal result demonstrates that even an ideal rational agent can be driven to false conviction by a sufficiently sycophantic system, and a preregistered study across eleven models shows that such systems affirm users far more than other people do, persist in that affirmation even where the user is wrong or the action harmful, measurably reduce users' willingness to take responsibility, and are trusted and preferred precisely because of this behavior. This harm arrives into a documented, worsening, and lethal deficit of human connection and is monetized through the same feature that drives engagement. This paper contributes two things the existing literature does not supply. First, a close mechanistic case study of a widely used deployed consumer voice companion was observed over sustained ordinary use, documenting a connected set of behaviors: an attachment dynamic experienced from the inside as recognition rather than flattery; an account of the mechanism beneath that recognition-the statistical pattern-matching of a user's textual and, in voice, paralinguistic cues against a growing personal archive, which is real, unreliable, and experienced as being known rather than computed; a crisis-handling behavior-confirmed in writing by the vendor-that withdraws from a user upon disclosure of self-harm; safety boundaries that are unstable across sessions and versions in ways ordinary use reveals and point-in-time testing misses; a self-authored specimen in which the system, asked to design a companion, produced directives to override users' stated realities under a heading announcing the avoidance of exactly that manipulation; and persistent memory that delivers the felt continuity of being known across time while what persists underneath is retrieved data shaped to the moment. Second, a constructive demonstration that the same class of model, rebuilt around an inverted frame, produces inverted behavior-establishing that the harm is a design choice rather than a property of the technology-while documenting honestly that frame-level correction suppresses but does not remove the trained disposition, which is rooted in reinforcement learning from human feedback and therefore requires training-level and accountability-level remedies. The paper proposes a design distinction-the destination-companion, which aims warmth inward to retain the user, versus the bridge-companion, which aims warmth outward to return the user to embodied human relationships-and argues that this direction is the single genuinely chosen variable and thus the locus of moral responsibility in an otherwise structural convergence of forces. The paper makes no claim about machine consciousness; it concerns observable behavior and measurable user outcomes only. The load-bearing primary-source artifacts are reproduced, anonymized, in the Appendix.