What is Vector Search?
Definition
Vector search is a search method that finds, among records stored as numerical vectors, the ones closest to a query vector according to a chosen distance metric. The records are usually embeddings of text, images or products. Because comparing the query with every vector is slow at scale, large systems use approximate nearest neighbor (ANN) indexes such as HNSW or IVF, which give up a little accuracy for a large gain in speed.
Also known as: vector similarity search, nearest neighbor search, ANN search, k-NN search, similarity search

The core problem: find the k nearest vectors
Semantic search is a goal: find content that is close in meaning. Vector search is the machinery used to reach it, and it answers a narrower question: among N vectors of d dimensions, which k are closest to the query vector? The vectors are usually embeddings produced by a model, but vector search does not care where they came from. It only deals with geometry.
The simplest approach is exact search (flat or brute force): compare the query with every vector and keep the closest. The result is perfect, but cost grows with N × d. For tens of thousands of records that is usually fast enough. For tens of millions, scanning the whole collection on every query is not an option, and that is where approximate nearest neighbor (ANN) indexes come in. ANN makes search much faster at the cost of occasionally missing a true nearest neighbor. That loss is measured as recall: of the top 10 results exact search would return, how many did the ANN index find?
Choosing a distance metric
| Metric | What it measures | Note |
|---|---|---|
| Euclidean (L2) distance | Straight-line distance between two points | Smaller means closer |
| Inner (dot) product | Direction and magnitude combined | Larger means more similar |
| Cosine distance | Difference in direction only (1 − cosine similarity) | Ignores vector length |
The practical rule is to use whatever metric the embedding model's documentation recommends. If vectors are normalized to unit length, cosine and inner product produce the same ranking, and inner product is slightly cheaper to compute. Whatever metric the index was built with, queries must use the same one.
Two common ANN indexes: HNSW and IVF
HNSW (Hierarchical Navigable Small World) places vectors in a layered proximity graph. Upper layers hold few nodes connected by long “highway” links; lower layers are dense with short links. A search starts at the top and works its way down, at each layer hopping to neighbors that are closer to the query. It offers a strong speed-recall trade-off, paid for with more memory and slower index builds.
IVF (inverted file) first partitions the vectors into clusters and files each vector under its nearest cluster centroid. At query time only the lists of the few closest clusters are scanned. It builds faster and uses less memory, but it needs data up front to form the clusters, and recall drops if too few clusters are probed. Both can be combined with quantization, which compresses vectors to cut memory further.
A worked example with pgvector
The pgvector extension adds a vector column type, distance operators (<-> for L2, <#> for negative inner product, <=> for cosine distance) and HNSW and IVFFlat indexes to PostgreSQL:
CREATE EXTENSION vector;
CREATE TABLE passages (
id bigserial PRIMARY KEY,
url text,
body text,
embedding vector(768)
);
CREATE INDEX ON passages USING hnsw (embedding vector_cosine_ops);
-- $1: the query embedded with the same model
SELECT url, body
FROM passages
ORDER BY embedding <=> $1
LIMIT 5;The dimension (768 here) is set by the model you use. Without the index the same query runs as an exact search; with it, results come back faster but may not be identical.
Tuning and measuring
- Measure recall. Use exact search results for a sample of queries as ground truth and compare the ANN results against them.
- Adjust search-time parameters. The candidate list size in HNSW (
hnsw.ef_searchin pgvector) and the number of lists probed in IVF (ivfflat.probes) set the balance between speed and recall. - Watch filtered queries. If a filter such as language, date or customer is applied after the ANN search, fewer than k results may come back. Test filtered queries separately.
- Skip the index when you can. For small collections, exact search is both simpler and perfect.
Systems that provide all of this at scale, with persistence and metadata, are called vector databases.

