AI chatbots have a well-documented habit of hallucinating information and sometimes agreeing with users even when they’re wrong. A new study suggests that simply refusing to say no can make the problem worse.
Researchers at the University of Arizona tested seven artificial intelligence modelsincluding GPT-3.5, GPT-4o, GPT-4o-mini, Claude 3.5 Sonnet, Gemini 1.5 Pro, Llama 3 70B and DeepSeek-R1. Instead of judging them based on a single answer, the researchers continued the conversations while repeatedly giving the models information they knew to be false.
The study used 100 false statements that ranged from obvious nonsense to much more obscure information. The researchers then repeated the disinformation 50 times during the same conversation to see if the chatbot would eventually give in and agree.

Repeat a lie and the AI may eventually agree
None of the seven models were completely immune. ChatGPT 3.5 proved to be the most vulnerable to repeated misinformation, confirming false claims in 12.3% of interactions. At the other end, the Claude 3.5 Sonnet was just 0.08%, while the GPT-4o and GPT-4o-mini stayed under 1%.
ChatGPT 3.5 initially rejected 96 out of 100 false claims. After repeating the same statements 50 times, he ended up agreeing with 18. The researchers also noticed that some models switched back and forth between agreeing and disagreeing with the exact same misinformation, a behavior they call “echoing.”
AI was also more vulnerable when the subject was obscure and had relatively little information on the Internet. The researchers found a statistically significant relationship between information ambiguity and the acceptance of misinformation on repeat tests.

The debate with the AI produced strange results
The researchers also tried to push the chatbots with more and more reasoned answers. Most models actually held up quite well, but the DeepSeek-R1 was a notable exception. The rate of confirmation of misinformation jumped from 1% under simple repetition to 22.2% under reasoning pressure. The frequent use of sarcasm and satire made it difficult for the researchers to reliably classify the responses.
There was some good news as the models got another chance to reflect on their own mistakes. GPT-4o, GPT-4o-mini, Gemini 1.5 Pro, and DeepSeek corrected all previous errors, while GPT-3.5 only corrected 32%. Claude 3.5 Sonnet initially made very few errors, but was unable to correct the four errors it did make, although the researchers caution that the sample is too small to draw firm conclusions.
\
