Contact

What is Embedding?

Definition

An embedding is a numerical vector that represents the meaning of a piece of text, an image or other data, produced by an embedding model. Content with similar meaning ends up with vectors that sit close together, so related items can be found even when they share no words. Embeddings power semantic search, retrieval-augmented generation (RAG) and recommendation systems.

Also known as: vector embedding, text embedding, embeddings, vector representation

Diagram of text turned into a numeric vector by an embedding model, with similar meanings close together

Meaning as coordinates

A computer cannot tell from the letters alone that "shipping" and "delivery" are related. An embedding model bridges that gap by turning a piece of text into a list of hundreds or thousands of numbers. No single number means anything on its own; what matters is where vectors sit relative to each other. During training the model learns to place expressions that occur in similar contexts near one another.

Measuring similarity

Closeness between two vectors is most often measured with cosine similarity: vectors pointing in the same direction score close to 1, and the score drops as they diverge. The example below only illustrates the idea; real vectors are far longer and these values are made up for illustration:

"when will my parcel arrive"   → [0.81, 0.12, -0.33, ...]
"delivery time for my order"   → [0.78, 0.15, -0.29, ...]   similarity: high
"how do I reset my password"   → [-0.10, 0.64, 0.42, ...]   similarity: low

The first two sentences share almost no words, yet their vectors are close because their meaning is. That relationship, which plain keyword matching misses, is the foundation of semantic search.

Where embeddings are used

  • Search: finding the documents closest in meaning to a query.
  • RAG: choosing which document chunks a language model should read before answering.
  • Recommendations: surfacing articles or products similar to what someone just viewed.
  • Clustering and classification: grouping support tickets, reviews or pages by topic.
  • Near-duplicate detection: catching content that says the same thing in different words.

Mistakes to avoid

  • Mixing models: every embedding model defines its own space. A document vector from one model cannot be meaningfully compared with a query vector from another, and switching models means re-embedding everything.
  • Treating dimensions as quality: more dimensions do not automatically mean better results, but they do increase storage and query cost. Test on your own data before choosing.
  • Oversized chunks: squeezing a long text covering many topics into one vector blurs its meaning, which is why documents are usually split by heading or paragraph.
  • Confusing embeddings with the language model: an embedding model produces representations, not text. The answer is written by a separate LLM.

Good practice when building with them

  • Use the same model and preprocessing for queries and documents.
  • Store metadata such as source URL, title and date next to each vector for filtering and citations.
  • Consider pairing semantic search with keyword search for exact-match needs like product codes or proper names.
  • Measure retrieval quality regularly against a test set built from real user questions.

What it means for content

Meaning-based matching reduces the need to sprinkle every synonym of a term across a page. What helps more is a page that treats its topic clearly and stays focused, with sections that each answer one question well. When a section wanders across topics, its vector ends up close to no question in particular.

Related terms

← Back to the glossary