Contact

What is Foundation Model?

Definition

A foundation model is an AI model trained on broad, diverse data at scale, usually with self-supervised methods, that can be adapted to a wide range of tasks through fine-tuning, instructions or added components. The term was popularized by Stanford researchers in 2021. Large language models are the best-known examples, but image, speech and multimodal models can be foundation models too.

Also known as: foundational model, general-purpose AI model, base model, pretrained model

Tree diagram of a foundation model pre-trained on web-scale data and adapted to chat, code, search, vision and agents

Where the term comes from

"Foundation model" took hold after a long report published in August 2021 by Stanford's Center for Research on Foundation Models. It defines them as models "trained on broad data at scale" and "adaptable to a wide range of downstream tasks", naming BERT, GPT-3 and DALL-E as examples. The authors chose "foundation" deliberately: these models are the shared layer underneath many applications, yet unfinished on their own. A flaw in the foundation can be inherited by everything built on it.

The EU's Artificial Intelligence Act (Regulation (EU) 2024/1689) uses the term "general-purpose AI model" for similar systems, describing models that display significant generality, perform a wide range of distinct tasks competently and can be integrated into many downstream systems. The legal scope and obligations are their own topic; what the two ideas share is generality and adaptability.

Is a foundation model the same as an LLM?

No. A large language model is trained on text and produces text, and LLMs are the most visible foundation models today. "Foundation model" is the broader umbrella and is not tied to one kind of data:

Model familyInput and outputTypical uses
Language modelsText to textChat assistants, summarization, code generation
Image generation modelsText to imageDesign work, illustration
Speech modelsAudio to text and backTranscription, voice interfaces
Multimodal modelsText, images and audio togetherAnswering questions about images, reading documents
Embedding modelsText or images to vectorsSemantic search, recommendations

A spam classifier trained on a small dataset for one job is not a foundation model, however well it works; generality and adaptability are the point. Models that handle several data types at once are covered under multimodal model.

How foundation models are adapted

  • Prompting: the model stays as it is; the task is described in the prompt, often with a few examples.
  • Fine-tuning: the weights are trained further on a smaller, task-specific dataset. Turning a raw base model into a chat assistant is itself the result of this kind of additional training. See fine-tuning.
  • Retrieval: the model is unchanged, but relevant documents are fetched at question time and added to its context. That pattern is RAG, and for anything that depends on current information it is often more practical than fine-tuning.

Why it matters for websites and brands

Many of today's chat assistants, search features and business tools are built on a fairly small number of foundation models. Two things follow. First, what a model learned during training, its parametric knowledge, is frozen at a point in time; current facts about your company only reach an answer through outside sources such as web search. Second, an error or gap in a foundation model can surface in many products at once. Clear, consistent and verifiable information about your brand across the web is the practical defense.

Foundation models are capable but not infallible. They can carry biases from their training data and produce fluent statements that are simply untrue, a problem known as hallucination. Adaptation can reduce it; it doesn't eliminate it.

Related terms

← Back to the glossary