What is Context Window?
Definition
A context window is the maximum number of tokens a language model can take into account in a single request. System instructions, conversation history, attached documents, retrieved sources and the model's own response all share this budget. Anything that does not fit in the window simply does not exist for the model during that request.
Also known as: context length, token limit, context size, model context

A working space measured in tokens
A language model processes text as tokens, and its context window is measured in the same unit. Tokens do not map one-to-one to words: a word can be a single token or split into several. Languages with long, suffix-heavy words, such as Turkish, often take more tokens than English for the same content. The exact count depends on the model, so use the provider's own token counter when precision matters.
Window sizes range from a few thousand tokens to a million or more, depending on the model. Providers change these limits between versions, so check the current documentation for the model you use rather than relying on a remembered figure.
What shares the window
- System instructions and tool definitions
- The conversation so far
- Files the user attached or text they pasted
- Source chunks retrieved through search or RAG
- The model's response (most models also cap output length separately)
When a long chat seems to "forget" its opening messages, it is often because the application trimmed or dropped older content to stay within the window.
Bigger is not automatically better
A larger window lets you process more material at once, but it comes with trade-offs:
- Cost and latency: every token processed is billed and adds time.
- Diluted attention: research has shown that with very long inputs, models can use information in some positions, especially the middle, less reliably. Being in the window does not guarantee being used correctly.
- Noise: adding irrelevant material can weaken the effect of the relevant material and does nothing to prevent hallucination.
That is why many systems do not fill the window with everything available. A retrieval step picks the few most relevant chunks, usually with embedding-based similarity search.
Implications for web content
In AI search, the model usually sees selected sections of your page, not the whole thing, and several sources compete for the same window. That rewards a few habits:
- Answer the question in a section's first sentences
- Write each section so it makes sense without what came before
- Do not push the key information down with long intros or repetition
- Use tables and lists where they actually carry data
Writing so that a passage can be selected and quoted directly is called answer extractability, and the finite context window is one of the reasons it matters.

