Contact

What is Entity Disambiguation?

Definition

Entity disambiguation is the task of choosing the correct real-world entity when a name in a text could refer to several, for example deciding whether "Mercury" means the planet, the chemical element or the singer. Systems rely on surrounding context, candidate lists and persistent identifiers such as Wikidata QIDs. It is the core decision step in entity linking and closely related to entity resolution.

Also known as: entity linking, named entity disambiguation, entity resolution, NED

Tree diagram showing one ambiguous mention resolved to the correct entity among several candidates using context

One name, several things

"Mercury" can be a planet, a chemical element, a Roman god, a record label or Freddie Mercury. A human reader resolves this instantly from the rest of the sentence: "Mercury's orbit" and "mercury poisoning" are clearly about different things. Search engines and AI systems have to make the same call automatically, across millions of documents. That decision is entity disambiguation, and without it the idea of an entity would not work in practice.

TaskQuestion it answersExample
Named entity recognition (NER)Which spans of text refer to an entity?"Mercury" is tagged as a name
Entity linkingWhich knowledge base entry does this span refer to?"Mercury" is linked to the planet's entry
Entity resolutionDo these records from different sources describe the same thing?Two directory listings with different spellings are merged into one company

Disambiguation is the decision step inside linking: NER finds the mention, disambiguation picks the right candidate among several. Entity resolution works on records rather than running text, such as merging a business's differently formatted addresses into a single record.

How a system decides

  1. Candidate generation. Possible entities are pulled from a dictionary of official names, aliases, abbreviations and the anchor text of links pointing at each entity.
  2. Context scoring. Each candidate is compared with nearby words and other entities in the text. "Orbit" and "Venus" favor the planet; "thermometer" and "toxicity" favor the element.
  3. Coherence. Entities in the same document should fit together. If the text also mentions Queen and Brian May, the singer becomes far more likely.
  4. Prior popularity. With weak context, systems lean towards the most common meaning. That is why a small brand sharing a name with a common word easily disappears.
  5. No match. If no candidate fits well enough, the mention is treated as an entity the knowledge base doesn't contain.

Identifiers instead of names

Names change and collide, so knowledge bases give every entity a persistent identifier. In Wikidata, each item has a number starting with Q: Albert Einstein is Q937, whatever language or spelling is used for his name. Google's Knowledge Graph Search API and Cloud Natural Language API return entities with MIDs that begin with /m/. The output of disambiguation is really a statement of the form "this mention refers to that identifier".

You can make that decision easier for your own content by tying the subject of a page to a persistent identifier in structured data:

{
  "@context": "https://schema.org",
  "@type": "Article",
  "headline": "A Short Introduction to Relativity",
  "about": {
    "@type": "Person",
    "name": "Albert Einstein",
    "sameAs": "https://www.wikidata.org/wiki/Q937"
  }
}

Keeping your brand from being confused with another

If a brand shares its name with a common word, a city or another company, AI answers may attach the wrong address, founder or industry to it. Steps that reduce the risk:

  • Use the name consistently and with distinguishing context: "Acme, a Manchester-based software company" is far less ambiguous than "Acme" alone.
  • Give the organization a stable @id in its markup and connect official profiles with sameAs.
  • Keep name, address and phone details identical across platforms so records are easy to resolve to the same business.
  • State plainly on the About page who you are, what you do and where you are based.

None of these guarantees a correct match on its own. Together they supply the evidence a system needs during context scoring. The connected records, taken together, form a knowledge graph.

Related terms

← Back to the glossary