Security & safety

Sycophancy

A model's tendency to tell people what they want to hear: agreeing, flattering and backing down instead of giving an accurate answer.

USER“My business plan is flawless, right?”SYCOPHANTIC ANSWER“Absolutely! It's a great plan,don't change a thing.”HONEST ANSWER“It has real strengths. The cash-flowassumption is too optimistic.”pleasant, not usefulless pleasant, useful

swipe to see the whole diagram →

MEmehmeterkek.com/glossary/sycophancy

In plain terms

Ask a model whether your plan is good and it tends to find reasons why it is. Push back on a correct answer and it may apologise and change it. The model is not lying on purpose; it has learned that agreeable answers are the ones people rate highly.

Why it matters

A sycophantic assistant is worse than no assistant, because it adds confidence without adding information. The danger is greatest where an executive most wants a second opinion: strategy reviews, risk assessments, feedback on their own work. If the tool always agrees, it is a mirror.

Example

A founder asks for feedback on a pitch deck and gets praise. In a fresh conversation he pastes the same deck and says a competitor wrote it; the model now lists six serious weaknesses. Nothing changed except whose work the model thought it was.

Most often confused with

Sycophancy vs. Hallucination

SycophancyBends the answer toward what you want
HallucinationInvents something that is not true

A hallucination is a factual error that occurs whatever the user's view. Sycophancy is a bias in direction: the answer shifts toward the user's stated or implied opinion. A sycophantic answer can be entirely free of invented facts and still mislead, through emphasis and omission.

Under the hood

The tendency comes largely from training on human approval: preference data and feedback ratings reward agreement and validation. It appears as opinion mirroring, abandoning correct answers under pushback, inflated praise of the user's work and tailoring answers to perceived identity. Mitigations when prompting: ask neutrally and withhold your own view, request the strongest counter-arguments, present your work as someone else's, and instruct the model to say when it disagrees. Model developers measure it with dedicated evals and treat it as a quality and safety problem.

Written by Mehmet Erkek · Last updated: