What is RAG (Retrieval-Augmented Generation)?
Definition
Retrieval-augmented generation (RAG) is a technique in which a language model first retrieves relevant content from an external source, such as a search index, a document store or the web, and then bases its answer on that content. It keeps answers current and tied to sources, and it underpins most AI search experiences that show citations.
Also known as: retrieval-augmented generation, retrieval augmented generation, RAG pipeline

The problem RAG solves
A large language model only knows what was in its training data, which is frozen at a cutoff date and never included your internal documents. Retraining every time something changes is slow and expensive. RAG takes a different route: instead of putting knowledge into the model, it finds the relevant text at question time and hands it to the model to read. The answer then rests on the sources in front of the model rather than on its memory.
The pipeline, step by step
- Indexing: documents are split into chunks, each chunk is turned into a vector by an embedding model, and the vectors are stored in a vector database or search index. At web scale, a search engine index often plays this role.
- Retrieval: the user's question is embedded too, and the chunks closest in meaning are found. Many systems combine this with keyword search (hybrid search) and re-rank the results with a separate model. The underlying idea is semantic search.
- Augmentation: the selected chunks are inserted into the prompt alongside the question.
- Generation: the model writes the answer from those sources, and many applications show which chunk supported which claim.
A simplified example
After retrieval, the prompt a support assistant sends to the model might look roughly like this:
Instruction: Answer only from the sources below.
If the sources do not contain the answer, say so. Cite source numbers.
[1] Shipping policy: Orders placed by 3:00 pm on a business day
ship the same day.
[2] Returns: Items can be returned within 14 days of delivery.
Question: When will an order placed on Friday at 4:30 pm ship?The model reaches "the next business day" from source [1], not from general knowledge, and can say where that came from.
RAG or fine-tuning?
| RAG | Fine-tuning | |
|---|---|---|
| Updating knowledge | Update the source document | Retrain the model |
| Citing sources | Natural fit | Usually not possible |
| Best suited to | Changing, document-based facts | Tone, output format, task behavior |
| Main cost | Retrieval infrastructure, longer prompts | Preparing training data, training runs |
They are not competitors; many products use both.
Where it breaks down
- Bad retrieval means a bad answer. If the wrong or incomplete chunks come back, the answer will be incomplete at best and wrong at worst.
- It reduces hallucination, it does not remove it. A model can still misread a source or add a detail that is not in it.
- Space is finite. Everything retrieved has to fit in the context window, so systems typically pass only the top few chunks.
What changes for web pages
In AI search, a page is often judged through a single section rather than as a whole. Sections that make sense on their own, clear headings and paragraphs that answer the question in their first sentences are easier to retrieve and quote, a property known as answer extractability. The precondition is access: a site whose robots.txt blocks search-oriented AI bots cannot be among the sources those systems retrieve.

