Contact

What is NoSQL?

Definition

NoSQL is an umbrella term for databases that store data in models other than relational tables, such as documents, key-value pairs, wide columns or graphs. They are usually designed around the queries an application will run and often spread data across many servers. In return they accept different trade-offs from relational systems, typically around joins, enforced schemas and how quickly every copy of the data becomes consistent.

Also known as: Not only SQL, non-relational database, document database, key-value store

Diagram of the four main NoSQL data models with flexible schemas: document, key-value, wide-column and graph databases

A name defined by what it is not

NoSQL names an exclusion rather than a technology: “not the relational table model”. The label spread in the late 2000s alongside systems that large web companies built for data volumes no single server could hold, and today it is usually read as “Not only SQL”. The systems under this umbrella differ more from one another than from relational databases; a document store and a graph database share little beyond not using tables. Several NoSQL products even offer query languages that look a lot like SQL, so “NoSQL means no SQL” is not quite right either.

Four models, four design questions

ModelUnit of dataHow you query itHard design decision
DocumentA self-contained record, e.g. an order with its line itemsFilters on fields, secondary indexesEmbed related data or reference another document?
Key-valueKey → valueBy key onlyNo searching inside values, so key design is everything
Wide-columnSorted rows grouped under a partition keyBy supplying the partition keyQueries must be known up front; data is often written to several tables, one per query
GraphNodes and the edges between themBy traversing relationshipsVery large graphs are hard to split across servers

Redis, the best-known key-value store, stretches that table a little, because its values can be lists, sets or sorted sets rather than opaque blobs. Vector databases, which store embeddings and run similarity search, are often counted as part of the same family.

Model the queries first

Relational design normalises first, storing each fact once, and assembles answers later with joins. NoSQL design usually runs the other way: list which screen reads which data by which key, then shape the data to serve those reads. An e-commerce order in a document store might look like this:

{
  "_id": "ord_8412",
  "customer": { "id": "cus_77", "name": "Jane Doe" },
  "shippingAddress": { "city": "Leeds", "postcode": "LS1 4AP" },
  "lines": [
    { "sku": "TSH-BLK-M", "title": "Black T-shirt", "qty": 2, "unitPrice": 18.50 }
  ],
  "status": "shipped"
}

The order page loads in one read, and the address and product title are frozen as they were at purchase time, which is exactly what an invoice needs. The price is duplication: if the customer's name changes and old orders must reflect it, the update fans out across hundreds of documents. This deliberate duplication is called denormalisation, and it sits at the heart of NoSQL modelling.

Consistency: when will you see a write?

Most NoSQL systems replicate data across several servers. When the network splits, a system can either refuse some requests to stay consistent or keep answering and let copies diverge for a while. That trade-off, known as the CAP theorem, shows up in daily work as eventual consistency: a user changes their display name, refreshes, and a lagging replica still shows the old one.

Many systems let you tune this per request. In a cluster that keeps three copies, requiring two acknowledgements for writes (W=2) and two for reads (R=2) means R+W exceeds the number of copies, so every read overlaps at least one up-to-date replica. Lower settings cut latency but raise the chance of stale reads. Operations on a single document are atomic in most document databases; transactions spanning several documents depend on the product and configuration and usually cost extra.

Where it fits and where it fights you

  • Good fit: stable, well-understood access patterns; write volumes beyond what one server handles; records that are naturally self-contained (event logs, product catalogues, sessions); deep relationship traversal.
  • Poor fit: ad-hoc reporting, rules that must hold across many entities at once (stock, balances, invoices) and query needs that keep changing. Missing joins and constraints then move into application code.

It is rarely all or nothing. Core business data can live in a relational database while caches and queues sit in a key-value store, and before adding a separate system for flexible attributes it is worth looking at middle grounds such as JSONB columns in PostgreSQL. If horizontal scalability is not a real requirement, the modelling overhead of NoSQL often does not pay for itself.

Related terms

← Back to the glossary