Contact

What is Inference?

Definition

Inference is the process of running a trained AI model on new input to produce an output. For a large language model, it means reading the prompt and generating the answer token by token. Training changes a model's weights; inference uses those weights unchanged and happens again for every request, which is why latency, cost per request and serving capacity are decided at this stage.

Also known as: AI inference, model inference, LLM inference, inference time

Inference diagram: an app request runs a forward pass on the trained model and the answer streams back

Training versus inference

A model lives through two very different phases. During training it is run over a huge dataset again and again, and its billions of weights are nudged at every step. During inference the weights are frozen: the model just processes the input in front of it and produces an output. Training is a large, occasional project. Inference is an operating cost that recurs with every request from the day a model goes live.

TrainingInference
What changesThe model's weightsNothing; the weights stay fixed
How oftenOnce per model version, occasionally more via fine-tuningOn every request
What matters mostTotal compute and data qualityLatency, cost per request, concurrent capacity
Source of knowledgeThe training dataWhat was learned in training plus the current context

Every message someone types into ChatGPT, Claude or an internal company assistant turns into an inference request behind the scenes.

Inside a request: prefill and decode

When a request reaches an LLM, the text is first split into tokens. Inference then runs in two phases:

  1. Prefill: every prompt token is processed in parallel, and the intermediate results the attention mechanism will reuse later (the KV cache) are built. Longer prompts mean a longer prefill.
  2. Decode: the model writes its answer one token at a time. Each new token is chosen in light of everything generated so far, and sampling settings such as temperature control how much randomness goes into that choice.

That split explains the two delays users actually notice. Time to first token depends mostly on prefill; how fast the rest of the text streams depends on decode. A request that summarizes a 3,000-token document has to read the entire document before it can show a single word, while a short question with a long answer starts quickly and then keeps streaming. Streaming does not reduce the total time, but it does reduce the time someone spends staring at an empty screen.

What drives cost and speed

  • Token counts. APIs usually price input and output tokens separately. Bloated system instructions and documents resent on every call land directly on the bill.
  • Model size. Larger models do more computation per token. Simple classification or extraction jobs often run perfectly well on a smaller one.
  • Context length. The fuller the context window, the more memory each request needs and the longer prefill takes.
  • Batching and caching. Providers process concurrent requests together to keep hardware busy, and some APIs cache repeated prompt prefixes to bring costs down.

Teams that host their own models have one more lever. Quantization, storing weights at lower numerical precision, cuts memory use and cost, sometimes at the price of quality. That trade-off should be measured on your own tasks rather than assumed.

What the model knows at inference time

During inference a model draws on two things: the parametric knowledge baked into its weights during training, and whatever context arrives with the request. It can only know about events after its training data ends if someone places that information in the context. RAG systems and AI assistants with web search do exactly that: they find documents relevant to the question and add them to the inference request.

For site owners this has a concrete consequence. Being cited in an AI search answer usually depends on your page being fetched and read at inference time, not on whether it ended up in a training set. It is the reasoning behind blocking training crawlers while still allowing search and user-request bots.

Not the statistician's inference

In statistics, inference means drawing conclusions about a population from a sample. In AI engineering the word is far more concrete: it means running the model. Nor is it a synonym for a model's ability to reason. Whenever people talk about inference cost, inference servers or inference latency, they mean this operational sense.

Related terms

← Back to the glossary