What is Chunking?
Definition
Chunking is the practice of splitting long documents into smaller, self-contained pieces of text before they are embedded and indexed. RAG and AI search systems retrieve these chunks rather than whole documents. Chunk size, the overlap between neighboring chunks and where the boundaries fall all have a direct effect on whether the right information can be found.
Also known as: document chunking, text chunking, semantic chunking, document chunk, text splitting

Why split documents at all?
When a RAG system answers a question, it does not hand whole documents to the language model, only the handful of pieces most relevant to the question. There are three reasons. Embedding models accept only so much text. Squeezing a twenty-page document into a single vector averages out the meaning of dozens of topics, and the one paragraph that answers a specific question disappears into that average. And the model's context window is finite while every token costs money, so what it reads has to be selective.
So documents are cut into chunks before indexing. Each chunk gets its own embedding and competes as a separate candidate at retrieval time.
The main strategies
| Strategy | How it splits | Good fit for |
|---|---|---|
| Fixed-size | Cuts every N tokens, usually with some overlap between chunks | Getting started quickly; plain text with no clear structure |
| Recursive, structure-aware | Splits on headings first, then paragraphs, then sentences if needed | Web pages and documentation with a heading hierarchy |
| Semantic chunking | Starts a new chunk where embedding similarity between consecutive sentences drops sharply | Long prose where topic shifts are not marked by headings |
| Format-specific | Splits code by function, tables by row groups, FAQs by question-answer pair | Code repositories, product catalogs, help centers |
Semantic chunking sounds like the clever option, but it needs an embedding for every sentence and, without a well-tuned threshold, produces chunks of wildly different sizes. On a document with sensible headings, structure-aware splitting often gets a similar result for far less effort.
Size, overlap and metadata
Chunk size is a balancing act. Small chunks of, say, 150 to 200 tokens match questions precisely but lose their surroundings: "This setting applies only to the enterprise plan" is useless once it is separated from the paragraph saying which setting is meant. Large chunks of 1,000 tokens or more keep context but carry several topics at once, which blurs their embeddings. These numbers are illustrations only; the right size comes from testing against a set of real questions.
Two refinements are common. Overlap repeats the last few sentences of one chunk at the start of the next, so information that straddles a boundary survives. Metadata attaches the document title, heading path and URL to every chunk:
{
"text": "On the enterprise plan, backups are kept for 90 days...",
"doc_title": "Backup and Restore",
"section_path": "Plans > Enterprise > Backups",
"url": "https://example.com/help/backups#enterprise"
}The chunk then carries its context into retrieval, and the final answer can link to the exact source.
Symptoms of bad chunking
- A table's header row lands in one chunk and its data in another, so the model cannot tell what the numbers mean.
- The answer to a question is split across two chunks, and neither looks similar enough on its own to be retrieved.
- Chunks open with "as mentioned above" or "this method" and say nothing when read in isolation.
- Navigation, footer and cookie-banner text leaks into every chunk and distorts similarity scores.
What it means for people writing web content
You cannot control how an AI search system splits your page, but you can write content that survives whatever splitter it uses. A meaningful heading structure, sections that each address one question, and paragraphs that make sense without leaning on earlier text all produce usable chunks wherever the cuts fall. That property is closely tied to answer extractability, and how chunks are then selected for a query is the subject of passage retrieval.

