AI Explained

AI Hallucinations: Why Chatbots Make Things Up

Why AI chatbots hallucinate, making up facts, quotes and sources with confidence, which tasks are riskiest, and how to reduce errors and check answers.

A hand holding a magnifying glass over a blank notebook on a desk beside a laptop and a stack of books
Illustration: AIEmulate / AI-generated.

Key takeaways

  • Hallucinations are confident, fluent answers that are false, from wrong dates to invented quotes and sources.
  • They happen because models generate plausible text rather than retrieving verified facts.
  • Give the model sources, ask it to quote them, allow it to say it doesn’t know, and verify anything that matters.
On this page

An AI hallucination is when a chatbot gives an answer that sounds confident and fluent but is wrong: a made-up statistic, a quote nobody said, a book that doesn’t exist, a legal case that never happened. It’s not a glitch that will vanish with the next update; it follows from how language models work. Newer models hallucinate less on many tasks, but the risk never reaches zero, so the practical skill is knowing when it’s likely and how to catch it.

What hallucinations look like

  • Invented facts: a wrong date, figure or name stated plainly.
  • Fabricated sources: real-looking citations, links or DOIs for papers that don’t exist, or real papers that don’t say what’s claimed.
  • False quotes: plausible words attributed to a real person.
  • Made-up details: features of a product, clauses in a contract or steps in a procedure that aren’t there.
  • Wrong summaries: a summary that adds conclusions the document doesn’t contain.

The NIST generative AI risk framework calls this confabulation, which captures it well: the model fills gaps with something that fits.

Why models make things up

A large language model predicts the next piece of text based on patterns from training; our explainer on how large language models work covers this in detail. That design has consequences:

  • No built-in fact check. The model generates what is plausible, and plausible and true usually overlap, but not always.
  • Gaps in knowledge. Rare topics, recent events and very specific details are thinly covered in training data, so the model improvises.
  • Pressure to answer. Models are trained to be helpful, and a confident answer can look more helpful than “I don’t know”.
  • Leading questions. Ask “Why did X happen?” and a model may explain an event that never occurred.
  • Long or messy inputs. Details in the middle of very long documents are easier to miss or blend together.

Where the risk is highest

TaskRiskWhy
Brainstorming, drafting, rewordingLowThere’s no single right answer
Summarising a document you provideMediumCan omit or add details
General knowledge questionsMediumUsually right on common topics, shaky on specifics
Citations, quotes, statisticsHighEasy to invent convincingly
Legal, medical, financial specificsHighDetails matter and errors are costly
Niche or recent topicsHighThin or missing training data

The cost of a mistake matters as much as its likelihood. Lawyers have been sanctioned by courts for filing briefs containing case citations invented by a chatbot, a reminder that “the AI said so” is no defence.

How to reduce hallucinations

  1. Give it the source. Paste or upload the document, data or web page, and ask the model to answer only from it. This is the idea behind retrieval-augmented generation.
  2. Ask for quotes. “Quote the sentence that supports each point” makes invented claims easier to spot.
  3. Allow uncertainty. Add: “If you’re not sure, or the source doesn’t say, tell me.”
  4. Use search-enabled modes for current facts, and open the links it cites.
  5. Break big tasks down. Several focused questions beat one sprawling request.
  6. Avoid leading questions. Ask “Did X happen?” before “Why did X happen?”
  7. Use reasoning modes for logic and maths, and still check the working.

Our prompt engineering guide has templates that build these habits in.

How to check an answer

  • Verify citations exist. Search for the exact title, check the author and publication, and read the relevant section.
  • Check numbers at the source, such as the official statistics office, company report or study.
  • Read laterally. Look for two or three independent, reputable sources that confirm the claim.
  • Ask again differently. If rewording the question produces a different answer, treat both with suspicion.
  • Look for vagueness. Hedged, generic answers on a specific question often hide a gap.

Beyond text: images and code

The same problem shows up in other kinds of AI output. Image generators draw extra fingers, impossible reflections and signs full of nonsense letters, and coding assistants call functions or packages that don’t exist. The fix is the same: treat the output as a draft and test it. Our guide to AI coding assistants covers the checks that catch invented code before it causes trouble.

When not to rely on AI answers

Don’t use a chatbot as the final word on medical diagnoses, medication doses, legal advice, tax or safety-critical instructions. Use it to understand the topic and prepare questions, then consult a qualified professional or official source. The same caution applies to AI-generated images and audio presented as evidence; see how to spot deepfakes.

Frequently asked questions

Why does ChatGPT make up sources?

Because it generates text that looks like a citation based on patterns, not by looking up a database of papers. Search-enabled modes reduce this, but always open and check the source.

Will AI hallucinations ever be solved?

They’re becoming less frequent, and grounding models in trusted documents helps a lot, but a system that generates plausible text can always generate a plausible mistake.

Which AI chatbot hallucinates least?

It varies by task and changes with each model update. Any model can be wrong, so the checking habits above matter more than the brand.

Sources

  1. NIST — AI Risk Management Framework: Generative AI Profile
  2. Ji et al. — Survey of Hallucination in Natural Language Generation

Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.

Keep reading