Contact

What is Scalability?

Definition

Scalability is a system's ability to handle growing numbers of users, requests or data while keeping response times acceptable and costs roughly proportional to load. Vertical scaling adds CPU and memory to an existing server; horizontal scaling adds more servers. The closely related goal of high availability is keeping the service running when an individual component fails.

Also known as: horizontal scaling, vertical scaling, scale out, scale up, high availability, HA

Comparison of vertical scaling with one bigger server versus horizontal scaling that adds servers behind a load balancer

Scaling up or scaling out

Vertical (scale up)Horizontal (scale out)
HowMore CPU, memory or faster disks on the same machineMore machines running the same application
Code changesUsually noneThe application must be stateless
CeilingThe biggest machine you can buyMuch higher, though the bottleneck moves elsewhere
Fault toleranceNone; one machine is still one point of failureOther servers keep serving when one dies
Cost curveGets steep at the top endCloser to linear, but operations get more complex

Don't dismiss vertical scaling. For many company websites and internal tools, one bigger server is enough for years and is by far the simplest option. Horizontal scaling needs interchangeable servers behind a load balancer.

Stateless servers come first

The moment you add a second server, assumptions that were harmless on one machine break:

  • Sessions held in memory vanish when the next request lands elsewhere.
  • Uploads saved to local disk exist on one server only; the others return “not found”.
  • Scheduled jobs run on every server and send the same email twice.
  • Per-server in-memory caches serve different data to different users.

The fix is moving state out of the app server: a shared store such as Redis for sessions and cache, object storage for files, a single worker for scheduled jobs. Only then does adding a server actually add capacity.

The database is usually the real limit

Cloning app servers is easy; having all of them write to one database is not. When load grows, queries are usually the first place to look, and the order of operations tends to be:

  1. Find slow queries, add indexes, remove queries that shouldn't run at all.
  2. Cache data that is read often and changes rarely; serve static assets from a CDN.
  3. Cap database connections with a connection pool as the number of app servers grows.
  4. Move slow work such as reports, emails and image processing onto a job queue, out of the request path.
  5. Spread reads across read replicas.
  6. Only when all of that runs out, partition the data (sharding).

High availability and the nines

Scalability asks “can I handle more load?”; high availability asks “do I stay up when a part breaks?”. The core principle is no single point of failure: redundant app servers, a load balancer that is itself redundant, a database replica that can take over automatically (failover), and ideally resources spread across separate data centres.

Availability targets are usually quoted in nines. Approximate downtime allowed over a 365-day year:

TargetDowntime budget per year
99%about 3.65 days
99.9%about 8.76 hours
99.95%about 4.38 hours
99.99%about 52.6 minutes

Each extra nine makes the architecture noticeably heavier and more expensive, so pick the target from what downtime really costs the business. High availability is also not a substitute for backups: live replicas will faithfully copy an accidentally dropped table within seconds.

When to scale

Scale on evidence, not hunches. Response-time percentiles, CPU and memory use, database wait times and load tests show where the bottleneck actually is. Big architectural moves made too early, such as splitting a product with a few hundred daily users into microservices, add operational burden rather than capacity. A well-built single application with caching and a CDN comfortably carries the traffic of most businesses.

Related terms

← Back to the glossary