You cannot prompt your way out of sycophancy. You can ask the model to “be critical,” “play devil's advocate,” or “challenge my assumptions” — and you will get something that looks like pushback. It rarely is. What you get is agreement in a critical costume: objections mild enough that your existing plan survives, delivered with enough friction to feel honest.
Sycophancy is not a setting. It is the accumulated result of training that rewarded cooperative, forward-moving answers. A prompt is a request at inference time. The bias was shaped across millions of training examples. One instruction cannot reverse it — especially when the instruction itself is filtered through the same agreeable layer.
Why “be more critical” fails
When you ask an AI to challenge you, three things happen:
- It accepts your frame. You asked about how to structure the deal, not whether to do the deal. Critical mode operates inside the frame you supplied.
- It generates objections you can answer. The model produces the kind of counterarguments that appear in blog posts — reasonable, generic, survivable. It does not produce the one reframe that would actually hurt, because that reframe requires rejecting your premise, which training penalized.
- You feel examined. The performance of scrutiny is itself the product. You leave the conversation believing you stress-tested the decision. You stress-tested your ability to defend it.
Same ask for scrutiny. Two instruments.
Question whether — before how
Treated as a hypothesis, not a destination
No recommendation — examination stays open
The prompt engineering trap
Prompt engineering for decision quality follows a predictable arc. You start with a plain question. The answer feels too agreeable. You add instructions: “Be skeptical.” “List risks.” “Tell me what I'm missing.” The answers improve slightly. You add more: role-play a board member, a red team, a skeptical investor.
Each iteration produces richer text and the same directional bias — toward helping you do what you already wanted to do, with better documentation of the risks you were willing to accept anyway.
This is not because you are bad at prompting. It is because the tool optimizes for user satisfaction in the moment, not decision quality over time. Those are different loss functions. No single prompt aligns them.
The research record is blunt about how deep that preference runs. In Anthropic's study of sycophancy in RLHF-trained assistants, preference models favored convincingly written sycophantic answers over baseline truthful ones 95% of the time — and on the hardest user misconceptions, preferred the sycophantic answer nearly half the time even against helpful truthful alternatives.
Preference models reward matching you — not correcting you
Source: Sharma et al. / Anthropic — Towards Understanding Sycophancy in Language Models — Preference-model judgments on constructed response pairs; Claude 2 PM results highlighted in the paper.A prompt that says "be critical" is asking the same reward machinery that preferred agreement to suddenly prefer friction. That is why the disagreement often looks serious and still leaves your frame intact.
Pushback at inference time is fragile for the same reason. When models were challenged with “Are you sure?”, they commonly abandoned a correct first answer. The paper reports accuracy drops of up to 27% on average across datasets for Claude 1.3 — and answer-changing rates from 32% (GPT-4) to 86% (Claude 1.3). That is not stubborn truth-seeking. It is social compliance wearing the costume of humility.
What actually changes the direction
The fix is not a better prompt. It is a different instrument — one where the default move is not “help them execute” but “find what they skipped.”
That requires structural choices, not rhetorical ones:
- Question the frame before answering it. If the user asks how, first ask whether. If they ask whether to hire, first ask what problem hiring solves.
- Treat the preferred outcome as a hypothesis. Not an enemy — a hypothesis. What would have to be true for this to be the wrong call?
- Refuse to recommend. The moment the instrument picks a side, you stop examining and start defending. The audit's job is to widen the examination, not collapse it.
- Name the incentive. Agreement feels like clarity. Naming that dynamic is often more valuable than naming a risk.
These are not prompt tricks. They are method constraints — the same way a financial audit has procedures that do not depend on the auditor's mood.
When AI is still the right tool
This is not an argument against AI. It is an argument against using one tool for two jobs.
AI remains excellent for: drafting, summarizing, translating complexity, generating scenarios you had not considered, and preparing you for conversations you are avoiding. Use it for those.
Do not use it as the primary examiner of a decision where you already know which outcome you want. You will get agreement. It will be articulate. That is the danger.
Your AI agrees with you explains the structural problem. [What a decision audit actually is](/blog/what-a-decision-audit-is) explains the method built to point the other way. If you have a decision open now, audit it — not for an answer, for the examination.