Contact

What is Knowledge Cutoff?

Definition

A knowledge cutoff is the latest date covered by a language model's training data. The model carries no built-in knowledge of events, price changes or new products after that date, and can only learn about them at answer time through web search, document retrieval or context supplied by the user or application. Cutoff dates differ between models and between versions of the same model.

Also known as: knowledge cutoff date, training data cutoff, training cutoff, cutoff date, model cutoff

Timeline showing a model's training data ending at a cutoff date, leaving later events unknown to it

A memory that stops on a date

A language model is trained on a large text corpus that is frozen at some point. What the model learned from that corpus lives in its weights as parametric knowledge. An election result, a new regulation, a company rebrand or an updated price list from after the cutoff simply is not part of that internal knowledge.

That is a natural consequence of training, not a defect. The trouble is that a model does not always notice the boundary. It may present outdated information as current, or fill the gap with something plausible and wrong, which is one of the most common routes to hallucination.

Often more than one date

Some providers publish two dates. Anthropic's model documentation, for example, lists a "training data cutoff" and a "reliable knowledge cutoff" separately: the first is the latest date reached by the data used in training, the second marks the end of the period for which the model's knowledge is most extensive and reliable. The gap makes sense. Writing about any event accumulates on the web over time, so the months just before a cutoff are thinly represented in the data. In practice, a model's knowledge of the period right before its cutoff can be patchy too.

How applications get current information

ApproachHow it worksLimitation
Web search and groundingCurrent pages are retrieved at answer time and the response is based on themOnly as current as the sources retrieved
RAGAn organisation's own documents are searched and added to the contextThe document store must be kept up to date
Information in the contextThe user or app adds today's date and the facts needed to the requestBounded by the context window and by the accuracy of what is supplied
Retraining or fine-tuningThe model itself is updated with newer dataExpensive, and it simply sets a new cutoff

One simple and often forgotten step when building an AI application is putting the current date in the system prompt. Otherwise the model may assume "today" is close to its cutoff and get phrases like "last year" or "this month" wrong.

What it means for brands and web content

  • Changes you make after a model's cutoff, such as a new name, address, price or service, reach users only when the system retrieves current sources from the web. That makes it decisive that the current facts are on your site, crawlable and stated plainly.
  • Old information that keeps circulating online makes it easier for a model to repeat the old version. Updating outdated profiles, directory listings and press materials where you can is worth the effort.
  • Showing publication and update dates on pages helps both readers and answer systems judge how recent a piece of information is.

Models and their cutoff dates change frequently, so check the current documentation for the model you use rather than relying on a remembered date.

Related terms

← Back to the glossary