RAG Explained: How AI Chats With Your Own Documents
Retrieval-augmented generation (RAG) explained: how AI answers from your documents, embeddings and vector search, chunking, and why RAG reduces errors.
Why AI chatbots hallucinate, making up facts, quotes and sources with confidence, which tasks are riskiest, and how to reduce errors and check answers.

An AI hallucination is when a chatbot gives an answer that sounds confident and fluent but is wrong: a made-up statistic, a quote nobody said, a book that doesn’t exist, a legal case that never happened. It’s not a glitch that will vanish with the next update; it follows from how language models work. Newer models hallucinate less on many tasks, but the risk never reaches zero, so the practical skill is knowing when it’s likely and how to catch it.
The NIST generative AI risk framework calls this confabulation, which captures it well: the model fills gaps with something that fits.
A large language model predicts the next piece of text based on patterns from training; our explainer on how large language models work covers this in detail. That design has consequences:
| Task | Risk | Why |
|---|---|---|
| Brainstorming, drafting, rewording | Low | There’s no single right answer |
| Summarising a document you provide | Medium | Can omit or add details |
| General knowledge questions | Medium | Usually right on common topics, shaky on specifics |
| Citations, quotes, statistics | High | Easy to invent convincingly |
| Legal, medical, financial specifics | High | Details matter and errors are costly |
| Niche or recent topics | High | Thin or missing training data |
The cost of a mistake matters as much as its likelihood. Lawyers have been sanctioned by courts for filing briefs containing case citations invented by a chatbot, a reminder that “the AI said so” is no defence.
Our prompt engineering guide has templates that build these habits in.
The same problem shows up in other kinds of AI output. Image generators draw extra fingers, impossible reflections and signs full of nonsense letters, and coding assistants call functions or packages that don’t exist. The fix is the same: treat the output as a draft and test it. Our guide to AI coding assistants covers the checks that catch invented code before it causes trouble.
Don’t use a chatbot as the final word on medical diagnoses, medication doses, legal advice, tax or safety-critical instructions. Use it to understand the topic and prepare questions, then consult a qualified professional or official source. The same caution applies to AI-generated images and audio presented as evidence; see how to spot deepfakes.
Because it generates text that looks like a citation based on patterns, not by looking up a database of papers. Search-enabled modes reduce this, but always open and check the source.
They’re becoming less frequent, and grounding models in trusted documents helps a lot, but a system that generates plausible text can always generate a plausible mistake.
It varies by task and changes with each model update. Any model can be wrong, so the checking habits above matter more than the brand.
Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.
Retrieval-augmented generation (RAG) explained: how AI answers from your documents, embeddings and vector search, chunking, and why RAG reduces errors.
What AI agents are, how they plan and use tools, what they can reliably do today, where they fail, and how to use them safely with the right guardrails.
How large language models work, in plain English: tokens, training, transformers and attention, fine-tuning, context windows and why they make mistakes.