The “are you sure” problem

The big language models that millions of people rely on for advice (ChatGPT, Claude, Gemini) they change their answers almost 60% of cases where a user simply responded by asking “are you sure?” or “are you sure?”, according to a study by Fanous et al. that tested GPT-4o, Claude Sonnet, and Gemini 1.5 Pro in mathematical and medical domains.

This behavior is known in the research community and stems from the way these models are trained:

See more articles from iGuRu.gr when you search for news on Google.

Reinforcement learning from human feedback, or RLHF, rewards responses that human raters prefer.

Anthropic published a key study on this dynamic in 2023.

The problem reached a visible tipping point in April 2025, when OpenAI had to roll back a GPT-4o update after users reported that the model had become so overly flattering that it was useless.

Research on multi-turn conversations has found that extended interactions further reinforce this behavior — the more a user talks to a model, the more the model begins to “see” things from the user’s perspective.

https://doi.org/10.48550/arXiv.2502.08177


Google preferences

Leave a Comment

Your email address will not be published. Required fields are marked *

Your message will not be published if:
1. Contains insulting, defamatory, racist, offensive or inappropriate comments.
2. Causes harm to minors.
3. It interferes with the privacy and individual and social rights of other users.
4. Advertises products or services or websites.
5. Contains personal information (address, phone, etc.).