What is LLM (Large Language Model)?
Definition
A large language model (LLM) is an AI model trained on very large collections of text to learn patterns in language, which then generates text by predicting it one token at a time. LLMs power assistants such as ChatGPT, Claude and Gemini and are used to answer questions, summarize, translate and write code.
Also known as: large language model, language model, LLMs

How an LLM produces text
An LLM does not read or write whole words. It works with tokens: pieces of text that may be a word, part of a word or a punctuation mark. Given a prompt, the model assigns a probability to every possible next token, one is selected and appended, and the loop repeats until the answer is complete. Nearly all widely used models today are built on the Transformer architecture introduced in 2017, whose attention mechanism lets every part of the input be weighed against every other part, which is how the model keeps track of context.
There is a hard limit on how much text the model can take into account at once, called its context window. Systems that work with long documents or many sources have to be designed around it.
How models are trained
- Pre-training: the model learns to predict the next token across a broad corpus such as web pages, books and code. Grammar, facts and reasoning patterns end up encoded in its parameters.
- Fine-tuning: the pre-trained model is trained further on curated examples so it follows instructions and holds a conversation.
- Alignment: many providers use methods based on human feedback (such as RLHF) to steer the model toward more helpful and safer responses.
Providers rarely disclose their training data in detail, so from the outside you cannot reliably say whether a given model "knows" a given website.
Training data versus the live web
What a model holds in its parameters is frozen at its knowledge cutoff. To mention today's price, a newly launched product or yesterday's news, it has to receive that information at answer time. Most AI search products do this with RAG: they first retrieve relevant pages from an index or the web, then write the answer based on those sources, often with citations.
For site owners this has a practical consequence. Having content used to train a model and having it fetched and cited in an AI search answer are separate processes, usually run by different AI crawlers. Blocking one does not automatically block the other.
Limitations worth knowing
- Fabrication: the objective is a plausible continuation, not a verified fact, so fluent but wrong statements happen. This is known as hallucination.
- Freshness: without external sources, anything after the cutoff is unknown.
- Consistency: the same question can produce different answers on different runs.
- Traceability: knowledge drawn from parameters usually cannot be traced back to a specific document.
Common misunderstandings
- "The model searches the internet." Search happens in the product layer around the model (search tools, retrieval systems), not inside the model. The same model can answer very differently with search switched on or off.
- "An LLM stores facts like a database." Knowledge is not kept as records but spread across billions of parameters as statistical associations, which is why "deleting" or "correcting" a specific fact in a model is not a simple operation.
What it means for publishers
LLM-based assistants describe a brand or topic using whatever sources they can access. Pages that define things plainly, state facts consistently and make clear who you are and what you offer are easier to represent accurately and to cite, although no technique guarantees inclusion in a particular answer. The broader practice of working on this is called generative engine optimization (GEO).

