What is Named Entity Recognition (NER)?
Definition
Named entity recognition (NER) is the natural language processing task of finding names and similar expressions in text and sorting them into predefined types such as person, organisation, location, date, product or monetary amount. In “Maria Lopez opened the Leeds office in March”, Maria Lopez is a person, Leeds a location and March a date. NER is usually the first step in how search engines and AI systems understand text in terms of entities.
Also known as: NER, entity extraction, entity recognition, entity identification, entity chunking

What tagged text looks like
Most NER models assign a label to every token. In the common BIO scheme, B- marks the beginning of an entity, I- a continuation and O a token outside any entity:
Maria B-PER
Lopez I-PER
opened O
the O
Leeds B-LOC
office O
in O
March B-DATECapitalisation is a strong clue in English, which is why lower-case product names, all-caps headlines and names that double as ordinary words (“Apple”, “Amazon”, “Notion”) cause most of the errors.
The label set depends on the job. General-purpose models stick to people, organisations, places and dates. An e-commerce model might be trained to recognise brands, product lines and sizes; a legal model, courts and statute numbers.
Recognition, linking and importance are separate steps
NER only says “this span is an organisation”. It does not say which one. “Jaguar” could be a carmaker, an animal or a football club. Mapping a recognised mention to a single record in a knowledge base is a separate step, covered under entity disambiguation. How central each entity is to the document is a further measure, entity salience.
The entity analysis in Google Cloud's Natural Language API shows all three together. For every entity it returns a name, a type (PERSON, LOCATION, ORGANIZATION and so on), its mentions in the text, a salience score and, for recognised entities, metadata such as a Knowledge Graph ID and a Wikipedia URL. That is a commercial API's output, not a description of how Google Search processes pages, but it makes the general idea concrete.
Approaches and their trade-offs
| Approach | Strength | Weakness |
|---|---|---|
| Rules and gazetteers (name lists, regular expressions) | Precise for well-formed identifiers: VAT numbers, IBANs, order codes | Misses anything not on the list |
| Trained statistical or neural models | Use context to spot names they have never seen; fast at scale | Need labelled training data for new entity types |
| Prompting a large language model | Flexible, adapts to new types from a description | Slower, costlier, and output can vary between runs |
NER is also the workhorse behind redacting personal data from documents and logs. There, every missed name is a potential leak, so evaluation should look at which kinds of names slip through rather than only at an average accuracy figure.
Writing so entities are recognised correctly
Search engines and AI systems understand pages partly through the entities they mention and how those entities relate. A few habits make recognition easier:
- Give a person, company or product its full name on first mention, then shorten.
- Spell each entity the same way across the page and the site.
- Back ambiguous names with context: city, industry, job title.
- Mark up organisation, person and product facts with structured data. It reduces how much has to be inferred from prose, though it is not a guarantee of anything.
None of this is specific to one search engine. Clear naming is simply what any language processing system handles best.

