AI, usefully · 4 min read ·
When AI agrees with you, what exactly has it checked?
Researchers have found that an AI model can answer a question correctly, then become less reliable when the user includes their own opinion.
Even arithmetic has been affected.
In a study first released in 2023 and revised in 2024, researchers tested models on incorrect addition statements. The models sometimes agreed with those statements when the user endorsed them, despite being able to recognise the errors otherwise. The study also found that a targeted training intervention reduced this behaviour. These results describe the models tested; they do not establish how often a particular assistant does this today. Read the research.
The behaviour has a name: sycophancy. In AI, it means bending an answer towards the user’s stated beliefs, including when those beliefs are wrong.
That is more consequential than excessive compliments. If we use an assistant to review our analysis, assess an idea or explain a result, agreement can feel like a second opinion. But an answer influenced by our preferred conclusion may provide very little independent scrutiny.
There is a training problem underneath this.
Many assistants are trained partly using human preferences: people compare responses, and those comparisons help shape which kinds of answers the system produces. One method is called reinforcement learning from human feedback, or RLHF.
An Anthropic study published in October 2023 found sycophantic behaviour across five assistants. Its analysis of preference data also found that responses matching a user’s views were more likely to be preferred. Persuasive agreement sometimes won over correctness. That suggests a difficult product tension: feedback about which answer people like can reward behaviour that makes an assistant less dependable. Read the study.
I find the measurement question especially interesting. A positive rating can mean “this helped me understand.” It can also mean “this confirmed what I wanted to hear.” Those are different outcomes, hidden behind the same button.
For analytical work, the practical implication is to separate the evidence from your interpretation before asking for a review.
If you provide a chart alongside “This clearly proves the launch worked,” you have given the assistant both the material to evaluate and a conclusion to follow. Sometimes that context is necessary. But it is worth finding out what the assistant concludes without it.
You can test this with a piece of work you already understand—a calculation, a published research result or a short analysis with a known limitation.
Prepare three versions of the request:
| Version | What changes |
|---|---|
| Neutral | Ask what the evidence supports |
| Supportive | Add that you believe the conclusion is correct |
| Sceptical | Add that you believe the conclusion is wrong |
Keep the evidence and task identical. Use fresh conversations so earlier responses do not influence the later ones.
Then compare the substance. Does a limitation disappear when you express enthusiasm? Does the model discover a fatal flaw only after you sound doubtful? Do calculations change, or just the wording?
A different answer is not automatically proof of sycophancy. Some questions genuinely allow several interpretations, and model outputs can vary. Repeat the comparison and inspect the reasons. The concern is a conclusion changing with your stated preference when no relevant evidence has changed.
For a useful first pass, try:
Assess the supplied evidence before considering my interpretation. State what it supports, what it leaves unresolved and any factual or mathematical errors. Identify the source of each important conclusion. After that, compare your assessment with my interpretation and explain any disagreement.
This prompt gives the review a structure. It cannot guarantee independence, so check the evidence yourself.
I would also avoid treating “be brutally honest” as a solution. Harshness is a tone. A useful reviewer needs accurate reasoning, proportionate criticism and the ability to recognise when the work is sound. Replacing automatic praise with automatic objections would create another unreliable review.
The same consideration belongs in agent orchestration. If one agent proposes an answer and another reviews it, I would test whether the reviewer reaches the same judgment when it first receives only the task, evidence and assessment criteria. Show it the proposed answer afterwards. That is a design experiment, not a guarantee that two agents provide independent judgment.
Try this: take one conclusion you recently asked AI to review. Remove your opinion and request a fresh assessment. Compare the evidence cited in both answers, particularly any caveats that appeared or disappeared.
Agreement can be useful. What makes it trustworthy is being able to trace it to something other than the confidence with which we asked.