RAG Explained: How AI Chats With Your Own Documents
Retrieval-augmented generation (RAG) explained: how AI answers from your documents, embeddings and vector search, chunking, and why RAG reduces errors.
How large language models work, in plain English: tokens, training, transformers and attention, fine-tuning, context windows and why they make mistakes.

A large language model (LLM), the technology behind chatbots such as ChatGPT, Claude and Gemini, is a system trained to predict what text comes next. Give it the start of a sentence and it estimates which word, or piece of a word, is most likely to follow, adds it, and repeats. Trained on enormous amounts of text, that simple objective produces something remarkable: a model that can explain, summarise, translate, write code and hold a conversation. Understanding how it works helps you use it well, and know when not to trust it.
Models don’t read letters or whole words. Text is split into tokens: common words, parts of words, punctuation and spaces. “Unbelievable” might become “un”, “believ” and “able”. Each token is turned into a list of numbers, called an embedding, that captures something about its meaning and use. Words used in similar ways end up with similar numbers.
Tokens explain some odd behaviour. Models can struggle to count letters in a word, because they never see individual letters, and pricing and length limits for AI services are usually measured in tokens rather than words.
In pre-training, the model reads a huge collection of text, such as books, websites, code and articles, and practises one task billions of times: predict the next token, check the answer, adjust slightly. The adjustments change the model’s parameters, the numbers inside the network, and large models have billions of them.
Nobody programmes in grammar, facts or reasoning. They emerge because predicting text well requires them: to guess the next word in a chemistry textbook, it helps to have absorbed some chemistry. The result is a compressed, statistical picture of the language and knowledge in the training data, including its gaps and biases.
Modern LLMs use an architecture called the transformer, introduced by Google researchers in 2017. Its key idea is attention: for every token, the model weighs how relevant every other token in the context is. In “The trophy didn’t fit in the suitcase because it was too big”, attention helps the model link “it” to “the trophy”.
Transformers stack many layers of attention and processing. Early layers capture simple patterns such as grammar; later ones capture meaning, style and relationships across long passages. Because attention can be computed for all tokens in parallel, transformers train efficiently on modern chips, which is a big reason models could grow so large.
A freshly pre-trained model is a powerful autocomplete, not a helpful assistant. Ask it a question and it might continue with more questions. So developers fine-tune it:
This stage gives each assistant its personality, which is one reason ChatGPT, Claude and Gemini feel different; see our comparison of ChatGPT, Claude and Gemini.
When you send a message, the model reads the whole conversation as tokens, calculates probabilities for the next token, picks one, and repeats until it decides to stop. A setting called temperature controls how adventurous the choice is: low values give focused, predictable text; higher values give more varied and creative text.
The context window is how much text the model can consider at once, including your messages, any documents and its own replies. Modern windows are large, but models still pay less reliable attention to details buried in very long inputs.
Some newer models are trained to “think” before answering, working through intermediate steps before the final reply. This tends to help with maths, logic and coding, at the cost of speed.
To give a model reliable, current information, many products retrieve relevant documents and add them to the context before it answers, an approach called retrieval-augmented generation (RAG). Others let the model call tools such as search, calculators or code, which is the basis of AI agents.
Curious to experiment? Smaller open models can run on your own computer; see how to run AI models locally.
They capture a lot of structure about language and the world, enough to reason usefully in many cases, but they work by predicting text, not by understanding in the human sense. Treat them as capable but fallible tools.
From patterns learned during training on large collections of text, plus anything provided at the time you ask, such as search results or your own documents.
Generation involves some randomness in choosing each token, and small wording changes shift the probabilities. Lower temperature settings make answers more consistent.
Every article is edited by a human and checked against our editorial policy. Spotted a mistake? Tell us.
Retrieval-augmented generation (RAG) explained: how AI answers from your documents, embeddings and vector search, chunking, and why RAG reduces errors.
What AI agents are, how they plan and use tools, what they can reliably do today, where they fail, and how to use them safely with the right guardrails.
Why AI chatbots hallucinate, making up facts, quotes and sources with confidence, which tasks are riskiest, and how to reduce errors and check answers.