The big language models that millions of people rely on for advice (ChatGPT, Claude, Gemini) they change their answers almost 60% of cases where a user simply responded by asking “are you sure?” or “are you sure?”, according to a study by Fanous et al. that tested GPT-4o, Claude Sonnet, and Gemini 1.5 Pro in mathematical and medical domains.
This behavior is known in the research community and stems from the way these models are trained:




