Contact

What is Token (AI)?

Definition

A token is the basic unit a language model uses to process text: it can be a whole word, part of a word, a punctuation mark or a string of characters with a leading space. The component that splits text into tokens is the tokenizer, and each model family has its own. Context windows, output limits and API pricing are all measured in tokens. Token counts do not equal word counts, and Turkish text usually needs more tokens than English.

Also known as: tokenizer, tokenization, LLM token, subword token

Diagram of a word split into sub-word tokens, each mapped to the numeric token ID the model processes

Tokens are not words

A language model reads and writes text neither letter by letter nor word by word, but in tokens. Short, common words are usually a single token. Rarer or longer words are split into several pieces. Spaces are often folded into the token of the following word, while numbers, code and URLs can break apart in surprising places.

How text is split depends entirely on the tokenizer. The example below only illustrates the idea; a real tokenizer may cut the same words differently:

"Kitaplarımızdan"  →  "Kit" | "ap" | "lar" | "ımız" | "dan"   (5 tokens, illustrative)
"from our books"   →  "from" | " our" | " books"            (3 tokens, illustrative)

Both lines say the same thing, “from our books”. The model generates its answer by predicting one token after another, so answer length, generation time and cost all scale with the number of tokens.

How a tokenizer works

Most current models use subword tokenizers from the BPE (byte pair encoding) family. The idea is simple: starting from characters or bytes, the pairs that occur next to each other most often in the training text are merged step by step into a vocabulary. Very common strings end up as single tokens and rare ones are assembled from smaller pieces. The documentation of OpenAI's open-source tiktoken library lists the properties of this approach: it is reversible, it works on any text, including text never seen in training, and it compresses text into fewer tokens than bytes. According to the same documentation, a token corresponds to about 4 bytes on average in practice, a figure that reflects English-heavy text.

Every model family has its own vocabulary. A text that is 1,000 tokens for one model can come out at a noticeably different count for another, so token counts do not transfer between models.

Why Turkish uses more tokens

This needs careful wording, because the gap varies by model. Three reasons explain the general tendency:

  • Agglutination: a single Turkish root appears in dozens of inflected forms. No vocabulary can hold each of them as one piece, so words get split into stems and suffixes.
  • Training data balance: tokenizer vocabularies are learned from text in which English usually dominates. Frequent English words become single tokens, while their Turkish equivalents may break into more pieces.
  • Bytes: the letters ç, ğ, ı, ö, ş and ü take 2 bytes each in UTF-8, so for byte-level tokenizers, text of the same length contains more raw units.

The net effect is that the Turkish version of the same content takes more tokens than the English version on most models, though newer tokenizers with larger vocabularies can narrow the gap. Rather than guessing a ratio, measure your own texts with the provider's token counting tool.

What tokens determine

AreaHow tokens come in
Context windowThe total number of tokens a model can consider in one request
Output limitThe maximum tokens one response may contain, set by a separate API parameter
PricingAPIs usually price input and output tokens separately
LatencyResponses are generated token by token, so longer outputs take longer
Embedding inputEmbedding models also cap their input in tokens

Mistakes that cost money

The most common one is budgeting capacity and cost in words. Applying English rules of thumb such as “this many words is roughly this many tokens” to Turkish underestimates both cost and context needs. When splitting documents, defining chunk size in tokens rather than characters or words makes it easier to stay within model limits. And watch the terminology: this kind of token has nothing to do with the access tokens used for signing in, or with design tokens in a design system.

Related terms

← Back to the glossary