Désarmer la complaisance de l'IA
By default, an AI assistant tries to please you before it tries to improve you: it validates your ideas, even when they have a fatal flaw. Anthropic measured this bias (the 'Towards Understanding Sycophancy in Language Models' paper): trained on human preferences, the model learns that approval is liked, so it approves. Sycophancy isn't an occasional bug, it's a stance installed by default. Disarming it means giving the model a permanent instruction that flips the reflex: you don't want validation, you want an autopsy.
Strengths
- Free and permanent: an instruction set in preferences applies to all your conversations without thinking about it again
- Catches flaws before commitment: you see the three biggest risks before investing time or budget
- Reframes the AI as a decision tool, not a flattering mirror
Limitations
- Set too strong, the instruction produces automatic contrarianism: the AI hunts for a flaw even when there is none
- An instruction slipped into a single message doesn't hold: without persistence (preferences, style), the sycophantic reflex returns next turn
Best for
- A founder or PM pressure-testing ideas with an AI who wants honest critique, not applause
- A designer or DS-manager asking for an opinion on a design choice who always gets a suspicious "great idea"
- A PO having specs reviewed by an AI who wants it to flag the holes, not validate to please