Ask a chatbot for the title of a researcher’s PhD thesis and it may answer at once, in a complete sentence, with a title that does not exist. Ask again and it may give a different one. This is what people mean when they say an AI tool “hallucinates”: it produces a plausible statement that is false, and it does so with the same fluency it uses for true ones.

The example is not hypothetical. OpenAI researchers describe asking “a widely used chatbot” for the dissertation title of Adam Tauman Kalai, one of their own authors, and getting three different answers, none correct, in a September 2025 post that accompanies their paper Why Language Models Hallucinate. The paper’s answer to the question in its title is less mysterious than the word suggests.

Where the wrong facts come from

A language model first learns by predicting the next word across a very large amount of text. Nothing in that text is labeled true or false. The model sees fluent language and learns its patterns.

Some patterns are regular enough to learn completely. Spelling and matching brackets follow consistent rules, which is why a large model rarely misspells a common word. Other facts follow no pattern at all. The OpenAI authors use the example of a birthday: if a person’s birth date appears rarely or never in the training text, no amount of pattern-learning can recover it. The model can only produce something that looks like a birthday.

That is the first point to take away. Hallucinations are concentrated in specific, low-frequency facts: dates, exact titles, citation details, small companies’ figures, the version number of a niche library. Those are the answers to check first.

Why the models guess instead of saying “I don’t know”

Training does not stop at next-word prediction. Later stages are meant to teach a model to answer helpfully and to admit uncertainty, and the paper argues they only partly succeed because of how models are scored.

Most benchmarks grade a model on accuracy alone: the share of questions it gets exactly right. Under that rule, a guess is never worse than an honest “I don’t know.” The OpenAI post compares it to a multiple-choice exam with no penalty for wrong answers. A model that guesses a birthday has a 1-in-365 chance of scoring; a model that abstains scores zero every time. Across thousands of questions, the guesser looks better on the leaderboard.

OpenAI illustrates the trade-off with its own figures from the SimpleQA test, as reported in its GPT-5 system card. These are the company’s own measurements, not an independent test:

Model Declined to answer Right Wrong
gpt-5-thinking-mini 52% 22% 26%
OpenAI o4-mini 1% 24% 75%

The older model is slightly more accurate and gives roughly three times as many wrong answers, because it almost never declines. The authors’ proposed fix is to change the main benchmarks so confident errors cost more than abstaining, rather than adding a few separate hallucination tests. Our explainer on why an AI benchmark win may not help with your work covers other ways headline scores can mislead.

Two kinds of wrong

Not every wrong answer has the same cause. A 2024 paper in Nature by researchers at the University of Oxford, Detecting hallucinations in large language models using semantic entropy, separates out what it calls confabulations: answers that are both wrong and arbitrary, so that asking the same question again can produce a different answer. Their example is a medical question that a model sometimes answered correctly and sometimes not, with identical instructions.

The authors distinguish those from errors with other causes, such as a model that is consistently wrong because its training data repeated a common misconception. Their detection method works by sampling several answers and measuring whether they disagree in meaning. It helps with confabulations but, as the paper says, cannot catch an error the model makes every time.

Practical checks before you rely on an answer

These steps come from the research above and from vendor guidance such as Anthropic’s documentation on reducing hallucinations. None of them removes the problem; Anthropic’s guide itself says to validate critical information.

  1. Treat exact details as unverified. Names, dates, quotes, figures, case citations and paper titles are the facts most likely to be invented. Open every source the tool cites and confirm it exists and says what the answer claims.
  2. Ask the same question twice. Run the prompt again in a new conversation or ask for the answer in other words. If the specifics change, treat them as a guess. This is a manual version of the sampling idea behind the Nature method, and it will not catch an error the model repeats every time.
  3. Give the tool permission not to know. Instructions such as “If you are not sure, say so” make an abstention an acceptable answer.
  4. Supply the document. When the answer should come from a contract, a policy or a report, paste it in or attach it and ask the tool to quote the passage that supports each claim. A claim without a supporting quote should be dropped.
  5. Keep it away from actions you cannot undo. A wrong fact in a draft is easy to fix. A wrong fact acted on automatically is not. Our guide to giving an AI agent access to your inbox covers approval steps for that case.

What will not fix it

A bigger model alone is not the answer. The OpenAI authors argue that accuracy can never reach 100% because some questions cannot be answered from the information available, and that a small model can sometimes know its limits more easily than a large one. They also reject the idea that hallucinations are inevitable: a model that declines when unsure does not hallucinate on that question. The practical result is that how a tool is built and scored matters as much as its size, and that a confident tone tells you nothing about whether an answer is correct.