What is Serverless?
Definition
Serverless is a cloud execution model in which the provider runs application code without the developer provisioning, patching or sizing any servers. Code is typically written as short-lived functions triggered by events such as HTTP requests, queue messages or schedules. The platform scales instances automatically with demand, and billing is based on requests and execution time rather than on servers kept running.
Also known as: serverless computing, serverless architecture, FaaS, Function as a Service

There are still servers
The name is misleading: your code still runs on servers. What changes is who looks after them. Operating-system patches, capacity planning and adding or removing machines as traffic moves are the provider's problem; you supply a function and the event that should trigger it. The best-known form is FaaS (Function as a Service), with offerings such as AWS Lambda, Google Cloud Run functions and Azure Functions. The term BaaS (Backend as a Service) usually covers the other half: managed authentication, databases and file storage you consume instead of building.
Serverless systems are event-driven by nature. An HTTP request, a file landing in storage, a message on a queue or a nightly schedule wakes a function up, it does its work and it goes back to sleep.
Lifecycle and cold starts
The first time a function is invoked, the platform has to prepare an execution environment: download the code, start the runtime and run any initialisation outside the handler. The extra latency this adds is the cold start. Afterwards the environment is frozen and kept around for a while; if another request arrives for the same function, it is thawed and reused, which is a warm start.
AWS's own Lambda documentation says cold starts typically affect under 1% of invocations and range from under 100 ms to over a second. Rarely called functions hit them far more often than busy ones. Keeping deployment packages lean, lazy-loading heavy libraries and, for latency-sensitive paths, paying for pre-initialised capacity (provisioned concurrency on Lambda) all reduce the impact. Objects created outside the handler, such as an SDK client, survive between warm invocations, which is why initialisation code belongs there.
The limits shape the design
Functions are built for short, stateless work. On AWS Lambda the main defaults are:
| Limit | AWS Lambda |
|---|---|
| Maximum duration per invocation | 900 seconds (15 minutes) |
| Memory | 128 MB to 10,240 MB; CPU scales with memory |
| Synchronous request and response payload | 6 MB each |
Ephemeral /tmp storage | 512 MB to 10,240 MB |
In practice that means long video transcodes or bulk exports have to be chunked or moved elsewhere, nothing kept in memory or on disk can be relied on between calls, and databases need care. A burst of traffic can start hundreds of instances that each open their own connections, so a shared connection pool or an HTTP-based data API becomes essential.
How the bill is calculated
FaaS pricing has two parts: the number of requests and the compute time. Compute is measured in GB-seconds, meaning duration multiplied by allocated memory, rounded up to the nearest millisecond. A function with 512 MB that runs for 200 ms uses 0.1 GB-seconds per call, so a million calls a month come to 100,000 GB-seconds. Lambda's free tier currently includes one million requests and 400,000 GB-seconds per month.
That model is excellent for spiky or sparse traffic, since an idle function costs close to nothing. For a service under steady, heavy load all day, the same work can be cheaper on a fixed-price VPS or a container platform. Do the arithmetic with realistic call volumes and durations before committing.
Good fits and poor fits
- Good fits: webhook receivers, form handling, image resizing and other event-triggered jobs; scheduled tasks; APIs with unpredictable or bursty traffic; prototypes.
- Poor fits: long-lived connections such as WebSockets; services under constant high load; endpoints that need consistently low tail latency; projects that cannot accept deep coupling to one provider's proprietary services.

