What is Transformer Architecture?
Definition
The Transformer is a neural network architecture built on the attention mechanism, introduced in the 2017 paper “Attention Is All You Need”. Instead of reading text one word at a time, it processes all tokens in a sequence together and calculates how strongly each one relates to every other. Today's large language models, and many image and multimodal models, are built on this architecture.
Also known as: Transformer, Transformer model, attention mechanism, self-attention, multi-head attention

The problem it replaced
Before 2017, most language systems used recurrent networks such as RNNs and LSTMs. They read a sentence left to right, one step at a time. That caused two problems: information from early in a long passage faded by the time the network reached the end, and because each step depended on the previous one, training was hard to parallelise on modern hardware.
The Google paper “Attention Is All You Need” proposed an architecture based solely on attention, dropping recurrence and convolutions altogether. Its first test case was machine translation. Within a few years the design had spread to almost every language task and then to images and audio.
Attention in plain terms
In “The trophy didn't fit in the suitcase because it was too big”, a reader knows that “it” is the trophy. Attention is the mechanism that lets a model build that kind of link numerically. Conceptually, every token produces three vectors:
- Query: what this token is looking for.
- Key: how each token advertises what it contains.
- Value: the information a token passes on if it is attended to.
Each token's query is scored against every key in the sequence. The scores become weights, and the token's new representation is a weighted blend of the values. So the representation of “it” ends up drawing heavily on “trophy”. Models do this with many attention heads in parallel, each free to specialise: one may track pronouns, another subject–verb agreement. Attention on its own ignores word order, so position information is added to each token separately.
Encoders, decoders and the model families they produced
| Layout | Example families | Typical use |
|---|---|---|
| Encoder-only | BERT and its descendants | Classification, entity recognition, producing embeddings |
| Decoder-only | GPT-style models | Generating text and code, chat |
| Encoder–decoder | The original Transformer, T5 | Translation, summarisation |
Most large language models behind today's chat assistants are decoder-only and generate text one token at a time. Encoder models remain a common source of the embeddings that power semantic search.
The price of attention
In standard attention, every token is compared with every other token, so doubling the input length roughly quadruples the number of comparisons. That is a big part of why models have a finite context window and why long inputs cost more. Many efficient attention variants exist to soften the curve, but the trade-off between compute, memory and context length is still a design decision in every model. It is also why usage is metered in tokens.
Why a non-researcher should care
A working picture of the Transformer explains behaviour you will meet in practice: why a model can base its answer on a document placed in its context, why details buried in the middle of a very long input sometimes get missed, and why the same word gets a different representation in different sentences. Representing words through their surroundings rather than as fixed dictionary entries is exactly where modern natural language processing parted ways with older keyword-matching methods.

