Contact

What is Vector Database?

Definition

A vector database is a data management system designed to store high-dimensional vectors, such as embeddings, together with their source text and metadata, and to run fast similarity queries over them. On top of approximate nearest neighbor indexes it provides persistence, updates, filtering, access control and scaling. It can be a standalone product or an extension of an existing database, such as pgvector for PostgreSQL.

Also known as: vector store, vector DB, vector index database

Vector database table storing text chunks as embeddings, where the rows closest in meaning to a query vector are found by similarity score

More than a vector index

The difference between a vector search library and a vector database is a bit like the difference between a sorting algorithm and a relational database. The library puts vectors in an in-memory index and finds the nearest ones. The database wraps that in everything a production system needs: data persisted to disk, inserts, updates and deletes, metadata filters, backups, replication, access control and scaling as the data grows.

“Vector database” is therefore the name of a product category, not an algorithm. The index inside is usually one of the well-known methods such as HNSW or IVF.

What a record holds

In a typical RAG project a record is much more than a vector:

{
  "id": "help-returns-03",
  "embedding": [0.021, -0.113, 0.087, ...],
  "text": "Returns can be requested within 14 days of delivery ...",
  "url": "https://example.com/help/returns/",
  "title": "Return policy",
  "lang": "en",
  "updated": "2026-09-12",
  "tenant_id": "tenant-42",
  "model": "embedding-model-v2"
}

The text is what the model will read; the URL and title make citations possible; language and date are for filtering; the tenant ID keeps customers apart in a multi-tenant system; and the model version tells you which records need recomputing when you switch models. How documents are cut into these records is a chunking decision, and it shapes retrieval quality directly.

Extension or dedicated product?

OptionUpsideWatch out for
Extension to an existing database (e.g. PostgreSQL with pgvector)One system; SQL, joins and transactions alongside your main dataMemory and tuning needs grow with very large vector volumes
Vector features of a search engine (e.g. Elasticsearch, OpenSearch)Lives next to keyword search, so hybrid queries are easyRequires experience running a search cluster
Dedicated vector database (e.g. Qdrant, Weaviate, Milvus, Pinecone)Built for large-scale vector workloads; most offer a managed serviceA second system that has to stay in sync with the source of truth

The choice usually comes down to data volume, the systems the team already runs and whether hybrid search is needed. For an internal knowledge base of a few hundred thousand passages, an existing PostgreSQL setup is often enough.

Operational traps

  • Changing models: vectors from different models cannot be compared. Switching models means re-embedding the entire collection, so plan for it from the start.
  • Synchronization: if a source document is deleted or edited but its vector copy stays behind, the system keeps answering with stale information. Deletes and updates must flow through to the vector side.
  • Access control: in a multi-tenant system every query must carry the tenant filter, or retrieval can pull one customer's document into another customer's answer. That is a security flaw, not a quality issue.
  • Memory cost: HNSW indexes perform best when held in memory, so hardware cost climbs quickly with vector count and dimension. Quantization and smaller embedding models are the usual levers.

When you do not need one

For a collection of a few thousand documents, a separate system is usually overkill; exact search in memory is fine. If the real problem is exact matching on a product code or an ID rather than similarity of meaning, an ordinary database index is the right tool. Vector databases earn their place in large, meaning-based workloads such as RAG and recommendation systems.

Related terms

← Back to the glossary