What is Retrieval?
Definition
Retrieval is the step that takes a query and returns a ranked list of the most relevant documents or passages from a large collection. The component that does it is called a retriever. It is the first stage of a search engine and, in RAG systems, the part that decides which sources a language model gets to read. Retrieval can be sparse (based on word matching), dense (based on embeddings) or a combination of both.
Also known as: information retrieval, retriever, document retrieval, first-stage retrieval

A retriever returns candidates, not answers
A retriever does not answer anything. It takes a query, searches a collection, and hands back the top k documents or passages most likely to be relevant, in ranked order. The collection might be a help center, a product catalog or the web itself. The first cut a search engine makes before building a results page relies on this step, and so does a RAG system choosing what its language model will read.
Retrieval is under time pressure: it has to narrow millions of documents to a few hundred candidates in milliseconds. That is why retrievers use relatively simple but fast similarity computations over indexes built in advance, and leave expensive, fine-grained judgments to later stages.
Sparse versus dense
| Sparse | Dense | |
|---|---|---|
| Representation | Terms and term weights (e.g. BM25) | An embedding vector produced by a model |
| Index | Inverted index | Approximate nearest neighbor index |
| Good at | Product codes, names, error messages, exact phrases | Synonyms, the same question asked in different words |
| Weak at | Text that says the same thing in other words | Rare terms and exact identifiers |
The mechanics of dense retrieval are covered under vector search, and combining both result lists is the subject of hybrid search. Morphology matters too: in an agglutinative language like Turkish, one noun can appear in dozens of suffixed forms, so sparse retrieval depends heavily on a good language analyzer, which is one reason dense or hybrid retrieval is attractive for Turkish content.
Recall, precision and the two-stage pipeline
Retrieval quality is discussed in two measures. Recall is the share of all relevant items in the collection that made it into the result list. Precision is the share of the result list that is actually relevant. Say a collection holds 10 passages relevant to a question and the retriever returns 20 results, 6 of them relevant: recall@20 is 60 percent and precision@20 is 30 percent.
The two pull against each other. Returning more results raises recall but lets in more noise. Modern systems therefore usually work in two stages: a fast retriever casts a wide net to keep recall high, then a slower, more accurate model re-ranks the candidates to raise precision. A passage missed in the first stage can never be recovered in the second, which is why recall gets so much attention.
How retrieval failures surface in answers
- The relevant passage never arrives: the model either says it does not know or falls back on what it learned in training, which raises the risk of hallucination.
- The wrong passage arrives: the answer can be confident and wrong, because models trust the text placed in front of them.
- Documents were split badly: the right information is spread across two fragments and neither looks similar enough on its own, which is why the chunking strategy has a direct effect on retrieval.
When evaluating a RAG system, measure retrieval separately from generation: build a test set from real user questions and check whether the correct passage appears in the top k. Only then can you tell whether a bad answer is the model's fault or the retriever's.
What it means for web content
AI search systems often judge a page section by section rather than as a whole, a pattern known as passage retrieval. It helps to write for both kinds of retriever: spell out product names, model numbers and technical terms for sparse matching, and keep each section focused on one question with the answer up front for dense matching.

