What is Context Engineering?
Definition
Context engineering is the practice of deciding what information enters a language model's context window on each call, in what order and at what token cost, and of managing that information as a task unfolds. It covers system instructions, tool definitions, retrieved documents, conversation history and memory, treating prompt writing as one part of a wider problem alongside retrieval and memory management.
Also known as: LLM context engineering, context management, context assembly, context design

Beyond the prompt
Early LLM applications revolved around one question: how should the instruction be worded? That is prompt engineering. As models became agents that call tools, read documents and work through long tasks, the question shifted to what exactly the model should be looking at on each step. In a September 2025 engineering post, Anthropic defines context engineering as the set of strategies for curating and maintaining the optimal set of tokens during inference, including everything that lands in the context besides the prompt. Seen that way, prompt engineering is one part of context engineering.
What gets assembled on each call
For a customer support assistant, the context for a single call might be assembled like this:
[system prompt] role, tone, constraints, output format
[tool definitions] lookup_order(), start_return()
[persistent notes] customer prefers email contact
[retrieved docs] returns policy, section 2 (top 3 chunks)
[history] last 6 messages + summary of earlier ones
[user message] "My parcel arrived damaged. What now?"All of these share one context window. Context engineering decides which parts stay fixed across calls, which are fetched on demand, when long history gets summarised and how irrelevant material is kept out.
Why more context is not better context
Anthropic describes "context rot": as the context grows, the model's ability to recall information from it accurately declines. Attention behaves like a finite budget, and every extra token draws on it. Irrelevant documents, verbose tool output and repeated instructions add cost and latency, and they also make the useful information harder to find. The goal is the smallest set of high-signal tokens that gets the job done.
Techniques in common use
- Retrieval: selecting relevant chunks instead of whole documents. RAG is the best-known form, but context engineering is the broader frame: it also covers how retrieved content is balanced against instructions, tools and history.
- Just-in-time loading: the agent keeps lightweight references such as file paths, links or queries, and pulls data in with tools only when it needs it.
- Compaction: summarising the conversation as the window fills, keeping critical details and dropping the rest.
- Structured note-taking: the agent writes progress and decisions to a file outside the window and reads them back later.
- Sub-agents: focused tasks run in separate agents with clean contexts, returning only a condensed summary to the coordinator.
Protocols such as the Model Context Protocol standardise how tools and data sources connect to a model, but deciding which tools are exposed on which call remains a context engineering decision.
Why it matters beyond app builders
AI search systems practise their own form of context engineering: they pass selected passages from your page to the model rather than the whole page, and several sources compete for the same window. Sections that stand on their own, answer early and avoid padding survive that selection and trimming best. If you are building an internal AI assistant, context design is as decisive for the result as the choice of model.

