KnowSys

Compute Abstractions

Follow one small photo-thumbnail handler up the ladder from a server to a VM, a container, a function and a Kubernetes pod: what you rent at each rung, what a cold start is made of, and how a scheduler decides which machine runs your code.

⏱ 49 min read◆ BeginnerAssumes: a terminal, Python and Docker; processes, containers and basic virtualization help
Start reading

You've written a small program, thumbnail.py. When a photo is uploaded, it shrinks the photo and saves the thumbnail to S3, Amazon's file-storage service. It works on your laptop. Now it has to run somewhere else, every time a photo arrives, and you find that "run this somewhere" has half a dozen answers: rent a whole server, rent a virtual machine, put the program in a container, upload it as a function, or hand it to Kubernetes as a pod.

Choosing is a bit like finding a place to live. Buy a house and you have everything to yourself, and it takes months. Rent a flat and you share the building, and it takes days. Book a hotel room and you're in tonight, with less control. Take a taxi and you own nothing, arrive in minutes, and pay only for the ride. Each step gives up some control for speed of getting started. Compute has the same ladder, and the price of climbing it is paid in how long the first request has to wait (its latency) and in how much of the machine you can no longer see.

This chapter follows thumbnail up the ladder and asks three questions the whole way. What am I renting at each rung? How long does the first photo wait before its thumbnail is made? And who decides which machine runs my code? We start by timing how long it takes to start something on your own computer, then see what each rung hides, take apart a function's cold start (the extra wait when a request arrives and nothing is ready to serve it), and finish by watching Kubernetes pick a machine for your code.

01Timing three ways to start something

1.1A process, a process in a container, a whole container

Before comparing rungs, you can feel the cost of climbing one. We'll start the smallest possible program three ways and time each. The program is /usr/bin/true, which does nothing and exits at once, so whatever time we measure is the cost of starting it and nothing else.

The three ways are these. First, run it directly as a process, a running program, on your own computer. Second, run it inside a container that's already running. A container is a program that has been given its own private view of the machine's files, processes and network, so it behaves as though it has the machine to itself; Docker is the usual tool for making one, and section 2 explains how it works. The command for this is docker exec warm true. Third, create a brand-new container just to run true, with docker run --rm alpine true. Here alpine is the name of a tiny packaged Linux, called an image, that the container is made from, and --rm deletes the container when the command exits.

The script runs each command 20 times and prints the median, the middle value, which a few unusually slow runs can't distort. Save it as startup.py. The second code block starts the container named warm (with -d, so it keeps running in the background, sleeping for 600 seconds) and then runs the script, so run those two lines first.

Time starting /usr/bin/true on the host, inside a running container, and as a brand-new container
python
Python
import statistics, subprocess, time
 
def median_ms(cmd, n=20):
    times = []
    for _ in range(n):
        t = time.perf_counter()
        subprocess.run(cmd, stdout=subprocess.DEVNULL, check=True)
        times.append((time.perf_counter() - t) * 1000)
    return statistics.median(times)
 
print(f"start a process on the host        (/usr/bin/true)              : {median_ms(['/usr/bin/true']):8.1f} ms")
print(f"start a process in a running container (docker exec)            : {median_ms(['docker', 'exec', 'warm', 'true']):8.1f} ms")
print(f"start a whole new container        (docker run --rm alpine true): {median_ms(['docker', 'run', '--rm', 'alpine', 'true']):8.1f} ms")
Shell
docker run -d --name warm alpine sleep 600     # the running container used by the second line
python3 startup.py
output
C++
start a process on the host        (/usr/bin/true)              :      1.4 ms
start a process in a running container (docker exec)            :     56.3 ms
start a whole new container        (docker run --rm alpine true): 178.0 ms

Starting a process on the host took 1.4 ms. Starting a process in a container that was already running took 56.3 ms, and creating a whole new container took 178.0 ms. Running the script again gives numbers like 53.1 and 176.0 ms for the last two, so expect the first row to be stable and the other two to wobble by a few milliseconds.

1.2Reading the numbers with care

This isn't a clean comparison. On a Mac or Windows machine, Docker runs containers inside a Linux virtual machine (a simulated computer with its own operating system, which section 2 explains), so the container timings include a call from the Docker command to a background service and into that virtual machine, and the plain process doesn't. (On Linux there's no virtual machine in between.) The gap between the two container rows is the more trustworthy part: building a new container from scratch costs about 120 ms more than starting one more process in a container that already exists.

That's the pattern the ladder predicts. Each abstraction wraps another, and starting a higher-level thing means setting up more layers around the same tiny program. To know what we're paying for, we need to name the layers and say what each one gives us.

02What you rent at each rung

The three timings differ because each one builds more machinery around true. So let's build up the layers the way a cloud provider would, starting from the bottom and adding one each time the previous one runs into a problem.

2.1From a whole server to a function

The lowest rung is bare metal: you rent a whole physical server and manage everything from its firmware (the software built into the hardware that starts it up) upward. Nothing is hidden from you and nothing is shared, and getting one takes minutes to hours.

Everything running on a server leans on one program, the kernel. The kernel is the core of the operating system: it controls the hardware, decides which program runs when, and hands out memory, files and network connections when programs ask. Two programs on the same machine share its kernel. That's fine for your own programs, but cloud providers make money by cutting one big server into many small pieces and renting them to different customers, and a bug in one customer's program must not reach another's.

One fix is to give every customer their own kernel. A hypervisor is software, helped by features in the CPU, that makes one physical machine look like several separate computers. Each of these is a virtual machine, or VM, and runs its own kernel and operating system, called the guest. The wall between two customers is the hypervisor, whose job is narrow and well guarded. Amazon's EC2 and Google's Compute Engine (GCE) rent these out. The price is that every VM has to boot an entire operating system before it can do anything, which takes tens of seconds.

The opposite fix is to share one kernel and make it hide things. A container is an ordinary process (or group of them) that the kernel restricts in two ways. Namespaces give it a private view: its own list of processes, its own files, its own network. Cgroups (control groups) cap how much CPU and memory it can use. A container starts in tens of milliseconds because the kernel is already running and only has to set up the restrictions. The price is that every container on the machine shares one kernel, so the wall between two customers is the kernel's own interface, and a bug in the kernel can let code escape its container. Chapter 11 builds a container by hand.

Three stacks side by side. Traditional: apps on one operating system on hardware. Virtualized: two virtual machines, each with apps, libraries and its own operating system, on a hypervisor, an operating system and hardware. Container: three containers, each with an app and libraries only, on a container runtime, one operating system and hardware
The two fixes side by side, with plain sharing on the left. In the middle, every virtual machine carries an operating system of its own, which is what has to boot before a VM can do anything. On the right, the containers hold only an app and its libraries, and all of them sit on the one operating system underneath, sharing its kernel.Image: The Kubernetes Authors, CC BY 4.0, from kubernetes.io

Between the two sits the microVM: a VM with nearly everything trimmed away, so that it still has its own kernel but boots in a fraction of a second. Firecracker, built by AWS (Amazon Web Services, Amazon's cloud), is the best-known one, and section 5 shows what it leaves out.

At the top, you stop thinking about machines at all. A function is the platform's offer to take only your handler, the piece of code that answers one request, and decide for you when to start somewhere to run it. You are billed for the milliseconds it runs. AWS Lambda, the most-used function platform, runs each function inside a microVM. Here are the rungs side by side:

RungIsolation boundaryYou manageTime to a new oneBilled by
Bare metalA physical machineEverything from firmware upMinutes to hoursThe month or hour
VM (EC2, GCE)A hypervisor; your own kernelThe guest OS and upTens of secondsThe second, with a 60-second minimum on EC2 Linux (2017)
microVM (Firecracker)A minimal hypervisor; your own kernelUsually nothing: a platform runs it for youUnder 125 ms until the guest's first program starts (NSDI 2020)Whatever the platform on top charges
ContainerNamespaces and cgroups; a shared kernelThe image and upTens of milliseconds (runc run took 48 ms in chapter 11), plus the image pullThe node it runs on
Function (Lambda)A microVM per execution environmentYour handlerAs little as 50 ms (Brooker et al. 2023); seconds for heavy runtimesThe millisecond, since December 2020

Each lower rung's mechanics have their own chapter. Chapter 47 covers how a hypervisor traps a guest and what each exit costs, and chapter 11 covers how namespaces, cgroups and overlayfs (the layered filesystem containers use) make a container. This chapter treats the rungs as products: what they promise and what they charge.

?Why not always pick the top rung?

Because each rung up trades control for convenience, and the price of that trade is paid in latency you can't see coming. A function costs nothing while it's idle, but the platform decides when to throw away the process you've already started, and the next request pays to start a new one. A container starts fast, but it shares a kernel with its neighbours.

A long-running service with steady traffic gets little from a function's scale-to-zero (paying nothing while nobody calls it) and pays for every cold start. A burst of short, unrelated jobs gets a lot.

2.2Two questions that pick the rung

Most choices come down to two questions.

  1. Whose kernel is it? If you don't trust your neighbours, or they don't trust you, you want a separate kernel per tenant. That's a VM or microVM. A container's boundary is the shared kernel's syscall surface, meaning the set of requests a program can make to the kernel.
  2. Who keeps it warm? If you keep processes running, you pay for idle time but never wait for a start. If the platform keeps them, you pay only for work, and sometimes wait.

The top rung hides the most, which means it's also the one where you can least see what happens between a photo arriving and your handler running. That gap is where cold starts live, so we'll look inside it.

03What happens when a function is called

Lambda is the most-used function platform, and AWS documents its lifecycle in unusual detail, so it's the clearest place to watch what "serverless" does between a request and your handler. Say you've uploaded thumbnail to Lambda as a function, and a photo arrives.

3.1Init, Invoke and Shutdown

Lambda doesn't keep thumbnail running while it waits for photos. When one arrives, it runs your code in an execution environment, one microVM set up for your function. Inside it are three things. The runtime is the program that runs your language, such as CPython, the JVM or Node. Next to it is your code. And there may be extensions, optional helper programs for things like monitoring. An internal extension runs inside the runtime's own process, and an external one runs as a separate process beside it. Each environment handles one request at a time. (The newer Lambda Managed Instances, which run on EC2 virtual machines in your own account and take several requests at once, have a different lifecycle.)

An environment's life has three phases, per the lifecycle docs. Init builds the environment and runs your static code, meaning everything in your file outside the handler. For thumbnail that's importing boto3 (the library for talking to AWS), creating the S3 client, and making an ID prefix that it stamps on the name of every thumbnail. Invoke calls your handler, the function that answers one request. Shutdown ends the environment. Between invocations, Lambda freezes the environment so it uses no CPU while it waits.

Here's thumbnail handling three photos. Watch the environment move through the phases, and notice which requests have to wait.

One execution environment, from the first photo to shutdown
WaitingrequestsInitup to 10 sInvokeyour handlerFrozenno CPUphoto 1requestenvironmentmicroVM bootingphoto 2requestphoto 3request
Step 1. A photo is uploaded and request 1 arrives for thumbnail. Nothing is running for this function yet, so the request has to wait while Lambda builds somewhere to run it.
1 / 7

A few details from the same page change how you write the code:

  • Frozen means frozen. Background threads and callbacks that didn't finish before the handler returned resume on the next thaw, possibly minutes later.
  • /tmp survives. It's 512 MB to 10,240 MB, and it keeps its contents across invocations in one environment. It isn't wiped even when a failed invoke resets the environment.
  • No environment lives forever. Lambda "terminates execution environments every few hours" for maintenance, even under continuous traffic.
  • A crash costs you an Init. If the handler crashes or times out, Lambda resets the environment and runs Init again on the next request, a suppressed init that doesn't get its own log line.
  • Extensions get a little time to finish at Shutdown. An external extension, a separate process, gets 2,000 ms to send off whatever it has buffered, and an internal one gets 500 ms. Anything still running after that is killed.

Everything in the picture happened somewhere, and something decided to build the environment and chose where. That decision is the part of the cold start you never see.

3.2One cold invoke, end to end

The Lambda paper by Brooker, Danilov, Greenwood and Piwonka (USENIX ATC 2023, arXiv) describes the invoke path. Here's what happens to photo 1 when no warm environment exists, as messages between the machines involved. (The paper calls an execution environment a sandbox, so "start sandbox" below means "create an environment".)

A cold Lambda invoke
CallerFrontendWorker ManagerWorkermicroVMInvokecapacity?start sandboxbootInitpayloadresponse
Step 1. A frontend, one of many identical servers that share the incoming traffic and keep nothing between requests, receives the request, loads the function's metadata, and checks who is calling and whether they are allowed to.
1 / 7

On a warm invoke the Worker Manager already knows a free environment, and everything from "start sandbox" to "Init" is skipped. Notice the phrase "loaded lazily" on the microVM's second disk. It hides a real problem, because the code that disk holds can be enormous.

3.3Loading a 10 GiB image in milliseconds

When Lambda launched, functions were zip files of up to 250 MB, downloaded and unpacked in full before the microVM could do anything. Then Lambda began accepting container images, a function's code and libraries packaged as a complete filesystem, of up to 10 GiB. Downloading 10 GiB before every cold start would make cold starts take minutes, so the paper's system loads images a block at a time, on demand.

?How can a 10 GiB image start in tens of milliseconds?

Because almost none of it is read at startup. Brooker et al. cite Harter et al.'s finding that "on average only 6.4% of container data is needed at startup." So Lambda flattens each image into one filesystem, cuts it into fixed 512 KiB chunks, and fetches a chunk only when the guest first reads from it.

The guest sees an ordinary virtual disk. Behind it, a per-function agent on the worker answers each read from the nearest place that has the chunk. Here are two reads of thumbnail's image, one from the shared cache and one that has to go to the origin:

Two chunk reads during Init: one cache hit, one trip to S3
Your microVMsees a virtual diskWorker cacheon this serverAZ cachemedian 550 µsS3, the originmedian 36 mschunk 7handler.pychunk 41boto3 fileschunk 902never readchunk 903never readchunk 41cachedchunk 41keptchunk 41boto3 fileschunk 7cachedchunk 7keptchunk 7handler.py
Step 1. The thumbnail image is 10 GiB, mostly files it never touches. Lambda has cut it into 512 KiB chunks held in S3, the origin (the original copy), and a shared cache for the whole availability zone (AZ, one group of data centres in a region) already holds chunk 41. Nothing has been downloaded to this worker.
1 / 6

The three tiers, in the order a read tries them:

TierShare of chunks servedLatency from the worker
Per-worker local cache67% (median, one week, one large region)Local
AZ-level distributed cache32%Median 550 µs, p99.9 3.7 ms
S3, the origin0.06%Median 36 ms, p99.9 175 ms

Two design choices make those hit rates possible:

  • Deduplication without shared keys. Each chunk is encrypted with a key derived from its own SHA-256 hash (convergent encryption), so identical chunks from different customers become identical ciphertext and are stored once. About 80% of newly uploaded functions had zero unique chunks, mostly re-uploads from automated build systems; of the rest, the median upload was 2.5% unique.
  • Erasure coding in the cache. Chunks are stored in the AZ cache as a 4-of-5 code, meaning each chunk is split into five stripes of which any four rebuild it. A reader asks for five stripes and reconstructs from the first four, which costs 25% more storage and requests but cuts tail latency (the slowest few reads) and hides a failed cache node without retries.

With the code arriving lazily and from nearby, the platform's share of the cold start is small. What's left is mostly Init, which is your code, so the next question is how long that takes.

04What a cold start is made of

"Cold start" gets used for any slow first request, and it's more useful to split it into the parts we've just met, because each part has a different owner and a different fix.

4.1The pieces, in order

A cold start for thumbnail has six pieces, in this order. The platform finds a server with room (placement), gets the code onto it (fetch), and either boots a microVM or brings back a saved copy of one, called a snapshot (section 5). Then the runtime starts, your static code runs, and finally the first request meets whatever your code was too lazy to set up earlier: the first connection to open, the first cache to fill.

Placementfind a host with room01Fetch codeimage or zip, lazily02Boot or restoremicroVM kernel, or snapshot03Runtime startJVM, CPython, Node04Your static initimports, clients, config05First requestlazy paths, empty caches06

The platform owns the first three. You own the last three, and the Lambda docs say plainly which one dominates: "The largest contributor of latency before function execution comes from initialization code."

4.2Measuring the pieces you own

To see the size of the pieces you own, here's the cost of starting processes that do progressively more before they're ready. The row "imports boto3 and creates an S3 client" is exactly what thumbnail's Init does. The first two columns come from a small 4-CPU Linux virtual machine running on a Mac, and the third from the Mac itself. The two Linux columns differ only in where the Python packages live. In the first, the virtual environment (a folder of installed Python packages) sits on the Linux VM's own disk. In the second, it sits in a folder the Mac shares into the VM. The two java rows compare the JVM's default, which loads its standard classes from a prebuilt archive (class data sharing), with the same program when the archive is switched off.

Each figure is the median of 21 runs, and the VM shared its CPUs with other busy work, so the spread between runs was wide. The milliseconds will be different on your machine; the ratios between rows are what carry over.

Start a new process that…Linux VM, packages on its own diskLinux VM, packages on a share from the MacmacOS
exec /bin/true0.3 ms0.3 ms2.5 ms
runs python3 -S -c pass5.8 ms6.0 ms18.6 ms
runs python3 -c pass7.8 ms9.8 ms25.8 ms
imports json12.4 ms17.4 ms28.7 ms
imports boto3126 ms252 ms187 ms
imports boto3 and creates an S3 client182 ms1,040 ms277 ms
java -version (class data sharing on)18.2 ms——
java -version -Xshare:off35.0 ms——
is forked from a parent that already did all of the above0.58 ms0.58 ms1.50 ms

Read down the first column. An empty program starts in 0.3 ms, the Python interpreter alone adds a few milliseconds, and importing boto3 and building one S3 client brings thumbnail to 182 ms. Nothing about the photo has been touched yet, and almost all of the time goes to getting ready.

?Why did the same code take almost six times longer on one disk?

The middle column ran the same virtual environment from /scratch, a directory shared in from the Mac host, where every metadata operation (listing a directory, checking a file) crosses the VM boundary and is slow. Creating one S3 client there spends roughly 65% of its time in listdir(), the call that lists a directory. botocore, the library underneath boto3, ships a description of every AWS service as files in its own folders, and to find S3's it calls listdir() 895 times and stat() (which reads one file's details) 2,777 times.

That's a cold start in miniature. The interpreter's own start barely moved between the two columns. What changed was the cost of your code's file access pattern, thousands of small lookups, on a disk where each lookup is slow. A Lambda function reading its code from a lazily loaded disk (section 3.3) is in the same position, and that's the effect the chunk caches, and the prefetching in section 5.3, exist to hide.

There's a second lesson in the last row. fork() is the system call that makes a copy of a running process, memory and all. A copy made from a parent that has already done the imports is ready in well under a millisecond, roughly 300 times faster than repeating the init. If a platform could do that for a whole machine, Init would disappear from the cold start, and that is what a snapshot does.

05Starting from a snapshot instead of booting

A forked process is a copy of something that's already initialised. For a function, the thing to copy is a whole machine, which takes two ingredients: a machine that's cheap to create, and a way to save one in the middle of its life and bring copies back.

5.1What a microVM leaves out

A VMM (virtual machine monitor) is the program that creates and runs a VM. A general-purpose one such as QEMU emulates a whole PC, down to its old hardware: a BIOS, PCI buses, USB, a floppy controller. Firecracker emulates almost nothing. It offers a virtual network card and a virtual disk (both built on virtio, a standard for simple virtual devices), a serial console and a keyboard controller used only to reset the guest, and it boots a Linux kernel directly.

A host with several Firecracker microVMs side by side in host user space, each containing a guest, each connected to a file-backed block device and to a shared network bridge, and each running through KVM in the host kernel
Firecracker on one host. Each microVM is an ordinary process in the host's user space with one guest inside, run with the help of KVM, the Linux kernel's built-in support for virtual machines (red lines). Its devices are a disk backed by a file on the host (yellow) and a network link to a shared bridge (blue). A separate orchestrator program starts and stops the microVMs.Image: The Firecracker authors, Apache 2.0, from the Firecracker design docs

With so little to set up, a Firecracker microVM gets from "start" to running the guest's first program (Linux calls it init) in under 125 ms, and it costs about 3 MB of memory on top of what the guest itself uses. Chapter 47, section 8.4 has the full table from the NSDI paper, including the I/O throughput it gives up.

?If a microVM boots in 125 ms, why are cold starts ever slow?

Because the kernel is probably the smallest part of what has to start. After the guest kernel comes the language runtime, then your code, then everything your code does before it can serve a request. For a Java app that loads a framework, Init can run for seconds (SnapStart's launch example, in section 5.4, took over 6), so a 125 ms guest boot is a small slice of the wait.

So the platforms stopped booting. They snapshot a machine that has already booted and initialised, then restore copies of it.

5.2A snapshot is a paused machine in a file

A Firecracker snapshot has three parts, per its snapshot docs: the guest memory file, the microVM state file (the CPU registers and device state of the paused machine), and any disk files, which you manage yourself.

What's interesting is how memory comes back. You'd expect the restore to read the whole memory file into RAM. Firecracker maps it instead, using mmap, a system call that makes a file appear as a region of the program's memory without reading it:

Rust
/// Creates a GuestMemoryMmap given a `file` containing the data
/// and a `state` containing mapping information.
pub fn snapshot_file(
    file: File,
    regions: impl Iterator<Item = (GuestAddress, usize)>,
    track_dirty_pages: bool,
    huge_pages: HugePageConfig,
) -> Result<Vec<GuestRegionMmap>, MemoryError> {
    /* ... check the regions fit inside the file ... */
    create(
        regions.into_iter(),
        libc::MAP_PRIVATE,
        Some(file),
        track_dirty_pages,
        huge_pages.madvise_flags(),
    )
}

MAP_PRIVATE on a file means copy-on-write: every restored VM shares the same file pages until it writes to one, and at that moment the writer gets its own private copy of that page. The docs call the result "runtime on-demand loading of memory pages". Restore returns quickly because it has barely loaded anything yet.

That raises a question, because the memory has to come from somewhere eventually. When the guest touches a page that isn't in RAM yet, the CPU raises a page fault, a signal that tells the host "this memory isn't loaded", and the host loads that one page and lets the guest carry on. Here's what that looks like for two clones of thumbnail restored from one snapshot:

Two clones restored from one snapshot, page by page
Snapshot fileon disk, saved after InitHost memoryRAMClone Arestored microVMClone Brestored microVMboto3 codepage 8S3 clientpage 9ID prefixpage 10 · 1a8fa67bboto3 codemapped, not readS3 clientmapped, not readID prefixmapped, not readboto3 codepage 8boto3 codeshared pageID prefix1a8fa67bID prefix1a8fa67b
Step 1. After thumbnail finishes Init, the platform pauses the machine and saves its memory to a file. Three of its pages matter here: boto3's code, the S3 client, and the ID prefix 1a8fa67b that Init generated.
1 / 6

5.3Where the restore time goes

Look back at frame 4. Every page the guest touches for the first time faults, and the host reads it from the snapshot file one page at a time.

Ustiugov et al. measured this in Benchmarking, Analysis, and Optimization of Serverless Function Snapshots (ASPLOS 2021). A function started from a Firecracker snapshot ran 95% longer, on average, than the same function already in memory, and the time went to those faults. Functions touched the same pages on every invocation, so their REAP prototype recorded that working set once and prefetched it in bulk. That cut cold-start delay by 3.7× on average.

Firecracker's own answer to the same problem is a userfaultfd backend. userfaultfd is a Linux feature that lets a separate process you supply answer the guest's page faults, and that process can fetch pages from wherever and in whatever order you like.

Faults are what a snapshot costs in time. The last frame of the scene showed what it can cost in correctness, and that is the trap SnapStart walks into.

5.4SnapStart, and state that must stay unique

SnapStart (November 2022, Java first; Python 3.12+ and .NET 8+ now) is Lambda's use of this idea. It runs Init once when you publish a version, takes a Firecracker snapshot of memory and disk, and restores new environments from it. AWS's launch example went from an Init of over 6 seconds to a start under 200 ms.

Restoring one snapshot into many environments means every one of them starts with the same memory. Anything your Init made that was supposed to be unique is now shared. Firecracker's docs are blunt about it: without a mechanism that keeps unique things unique, they "consider resuming execution from the same state more than once insecure."

Here's the problem on one Linux box, using fork() as a stand-in for a snapshot restore. The parent does the "Init": it creates a random generator and an ID prefix, just like thumbnail's. Three children then "invoke". Each prints four values: a number from the private generator the parent made, the ID prefix, a number from Python's global random generator, and a number from random.SystemRandom, which asks the kernel for fresh randomness every time.

Three clones of one initialised process
python
Python
import os, random, uuid
 
# "Init": runs once, before the snapshot (here, before fork)
rng = random.Random()                      # a private generator, seeded now
request_id_prefix = uuid.uuid4().hex[:8]
 
for child in range(3):
    if os.fork() == 0:                     # "restore" a clone
        print(f"clone {child}: Random()={rng.random():.6f} "
              f"prefix={request_id_prefix} "
              f"global random={random.random():.6f} "
              f"SystemRandom={random.SystemRandom().random():.6f}", flush=True)
        os._exit(0)
    os.wait()
output
Output
clone 0: Random()=0.695238 prefix=1a8fa67b global random=0.785913 SystemRandom=0.897001
clone 1: Random()=0.695238 prefix=1a8fa67b global random=0.406180 SystemRandom=0.599397
clone 2: Random()=0.695238 prefix=1a8fa67b global random=0.310840 SystemRandom=0.585724

Both the private generator and the ID prefix are identical in every clone. The global random differs, and so does SystemRandom, which reads the kernel's entropy each time.

The private generator and the prefix were made before the fork, so every clone carries the same copy, exactly as both clones in the scene held 1a8fa67b. The global generator and SystemRandom differ. SystemRandom differs because it asks the kernel every time. The global generator comes out different because CPython reseeds it in every forked child:

Python
## ------------------ fork support  ---------------------
 
if hasattr(_os, "fork"):
    _os.register_at_fork(after_in_child=_inst.seed)

A VM snapshot restore has no such hook. No fork() happens inside the guest, so even that global generator would come back identical in every clone. What Lambda does provide, per its uniqueness docs, is a reseed of the kernel's /dev/random and /dev/urandom (the files programs read to get kernel randomness) on restore. Code that asks the kernel for randomness every time stays safe. Code that made its random state during Init and kept it in the program's own memory, like rng and the prefix above, does not.

A function whose Init has been snapshotted starts fast, but a burst of traffic needs many environments at once, and each environment serves one request at a time. How many does a burst need, and what are the other ways to avoid cold starts?

06Concurrency and the cold-start levers

So far we've followed one request. A real service gets bursts, and every photo that arrives while all existing environments are busy needs a new one, which means a cold start. So we need to know how many environments a given amount of traffic keeps busy.

6.1Concurrency is the unit you're sold

Lambda calls the number of environments busy at the same moment your function's concurrency. Concurrency is what your account's limits are written in, and since each environment serves one request at a time, it's also how many cold starts a burst can cause. Lambda's concurrency docs give the way to estimate it:

Output
concurrency = average requests per second × average duration in seconds

That's Little's law: the average number of things in a system equals the rate at which they arrive times how long each stays. For thumbnail, suppose photos arrive at 100 a second and each takes 500 ms to process:

Requests per secondarrival rate100
Average duration500 ms ÷ 1,0000.5 s
Environments busy at once100 × 0.550
environments needed on average50

That's an average, and averages hide the queueing, as chapter 16 shows. Lambda throttles requests that go over a limit, meaning it refuses them instead of running them. Each account has a concurrency quota, a cap on how many environments all its functions together may have busy at once in one Region (one of AWS's geographic areas). You can also reserve part of the quota for one function, so that the others can't use it up. Here are the limits that matter, from the same page:

LimitDefaultWhat it means
Account concurrency, per Region1,000Shared by every function without a reservation
Scaling rate, per function1,000 new environments every 10 secondsHow fast a burst can be absorbed
Requests per second10 × your concurrency quotaBites for very short functions
Reservable concurrencyQuota minus 100100 always stays unreserved
Predict before you read on

A function averages 20 ms, and at peak it gets 30,000 requests per second. The account has the default quota of 1,000. What happens?

Fifty environments in the first estimate means fifty cold starts the first time a burst arrives. So what can you do about cold starts?

6.2The levers, and what each one costs

Each part of the cold start has a lever, and each lever has a price.

LeverWhat it removesWhat it costs you
Import less, lazilyStatic init time: import only the clients you use, create rarely used ones on first useCode discipline; the first request down a lazy path pays instead
More memoryCPU-bound init time: Lambda allocates CPU in proportion to memory, one virtual CPU at 1,769 MB (docs)Higher price per millisecond
Provisioned concurrency (section 6.3)The whole cold start, for up to N concurrent requests, by initialising N environments in advanceYou pay for the environments while they sit idle
SnapStartRuntime start and static init, replaced by a snapshot restoreUniqueness bugs (section 5.4); can't be combined with provisioned concurrency
A smaller image or packageFetch time, for cold chunksMostly already hidden by lazy loading; helps less than it used to
A keep-warm pingNothing reliableIt keeps one environment warm, not the N a burst needs

6.3Provisioned concurrency

Provisioned concurrency runs Init ahead of time for a fixed number of environments and keeps them initialised. From the concurrency docs:

  • It takes "a minute or two" to come online, and none of the environments are usable until the whole allocation is ready.
  • Past N concurrent requests, traffic spills into ordinary on-demand environments, which cold-start as usual.
  • Lambda still recycles provisioned environments in the background, so you can still see an occasional cold start after a reset.
  • Init for provisioned (and SnapStart) functions may run for up to 15 minutes, not 10 seconds.

?Why does provisioned concurrency sometimes still log a slow first request?

Because Init finished long before the request arrived, and some work is lazy. The docs warn of "variable latency on the first invocation on an initialized execution environment". JIT compilation (the JVM compiling hot code as it runs), connection pools that open on first use and lazily loaded classes all land on that first request. Warm them up in your static init if they matter.

6.4Init is billed now

Until August 2025, Init time wasn't billed for functions uploaded as zip files and run on the runtimes AWS provides (its managed runtimes). From August 1, 2025, it is, the same as it already was for container images, custom runtimes (ones you supply yourself) and provisioned concurrency (AWS blog).

AWS expects the effect to be small, since Init runs for a small fraction of invocations; the docs say cold starts "typically occur in under 1% of invocations". But a slow Init is now a cost as well as a latency, and so is a crash loop that keeps re-running it.

With a function, Lambda decided where each environment ran, when to start it and when to throw it away, and you never saw those decisions. One rung down you make some of them yourself, and the machine that makes the rest is a scheduler.

07How a pod gets placed

Suppose you run thumbnail yourself, as a container on a cluster, instead of as a function. Kubernetes is the system most teams use for that. It manages a cluster, a group of machines called nodes, and the unit it places on a node is a pod: one or more containers that run together on the same node. You write down what a pod needs, and a program called the scheduler (kube-scheduler) decides which node runs it. On this rung you can see your process, but you still don't pick its machine. The scheduler picks it, and it decides from numbers you wrote down in advance, without looking at how busy each node is.

A Kubernetes cluster: a control plane containing the API server, etcd, the scheduler, a controller manager and a cloud controller manager, and two nodes each running a kubelet and kube-proxy, with every component connected to the API server
The parts of a Kubernetes cluster. The control plane holds the API server (api), which every other part talks to, etcd, where the cluster's state is stored, the scheduler (sched) that this section follows, and the controller managers. Each node runs a kubelet, which starts the pods assigned to that node, and kube-proxy. Notice that the scheduler talks only to the API server: it never reaches into a node itself.Image: The Kubernetes Authors, CC BY 4.0, from kubernetes.io

7.1Requests, limits and QoS

Each container can declare a request and a limit for CPU and memory. The request is what the container says it needs, and the limit is the most it may use. They do different jobs, per the resource management docs, and the table below uses three more terms. The kubelet is the agent on each node that starts the pods assigned to it and sets up their cgroups. The CFS quota is how Linux's CPU scheduler (CFS, the Completely Fair Scheduler) caps a cgroup at so many milliseconds of CPU in every 100 ms period; chapter 11 measures it. And a container that is OOM-killed has been ended by the kernel's out-of-memory killer for using more memory than it was allowed.

RequestLimit
Used byThe scheduler, to pick a nodeThe kernel, at run time, via the kubelet's cgroup settings
CPUA share of CPU under contentionA CFS quota: the container is throttled when it uses its share of each 100 ms period
MemorySpace reserved on the nodeA hard ceiling: exceed it and the container is OOM-killed
If unsetCopied from the limit, if there is oneNone

?Why can a pod stay Pending on a cluster that's mostly idle?

Because the scheduler never looks at usage. A node is full when the sum of its pods' requests reaches its allocatable capacity (what's left of the machine after the system reserves its own share), however little CPU they burn. A team that requests 4 CPUs "to be safe" and uses 0.3 has reserved the rest of that node from everyone else, which seems harmless until the cluster fills up. A pod that hasn't been given a node is Pending.

Requests and limits together also set a pod's QoS class (quality of service). Here's the rule, per resource:

pkg/apis/core/v1/helper/qos/qos.go
kubernetes/kubernetes @ v1.37.0 ↗
Go
// resourceQOS determines the QOS "shape" of the given resource in the requirements:
// - BestEffort: Request and Limit are both zero
// - Burstable: Request != Limit
// - Guaranteed: Request and Limit are equal and non-zero
func resourceQOS(resources *v1.ResourceRequirements, res v1.ResourceName) v1.PodQOSClass {
	req := resources.Requests[res]
	lim := resources.Limits[res]
 
	if !req.Equal(lim) {
		return v1.PodQOSBurstable
	} else if req.IsZero() {
		return v1.PodQOSBestEffort
	} else {
		return v1.PodQOSGuaranteed
	}
}

A pod is Guaranteed only if every container has equal, non-zero requests and limits for both CPU and memory. Any mismatch makes it Burstable. The class decides who dies first when the node runs out of memory, through the oom_score_adj the kubelet writes for each container (pkg/kubelet/qos/policy.go). That's a number the kernel's OOM killer reads for each process: the higher it is, the sooner the process is picked.

QoS classoom_score_adjMeaning
Guaranteed−997Almost never picked by the kernel's OOM killer
Burstable1000 − 1000 × request ÷ node memory, clamped to 3–999The bigger your request, the safer you are
BestEffort1000Killed first

Requests are numbers the scheduler reads. Next we'll see what it does with them when a pod arrives.

7.2The scheduling cycle

kube-scheduler watches for pods with no node assigned and gives each one a node. It does so as a pipeline of stages, each a fixed extension point where small pieces of logic called plugins run. Roughly: the pod waits in a queue, Filter plugins throw out every node that can't run it, Score plugins rate the nodes that are left, the scheduler reserves room on the winner and binds the pod to it (writes the choice to the cluster's central record, the API server), and the kubelet on that node starts the containers. The scheduling framework docs name every hook. Here's one pod's trip through it:

One pod, from Pending to running
●
≡
Queue
ordered by priority
⊘
Filter
PreFilter · Filter
★
Score
PreScore · Score · Normalize
◧
Reserve · Permit
in the scheduler's cache
⇄
Bind
PreBind · Bind
◉
Kubelet
admit · sandbox · start
Step 1. The pod enters the scheduling queue. PreEnqueue plugins (scheduling gates) can hold it back; QueueSort orders the queue, by priority by default.
1 / 7

?Why is the scheduling cycle serial when binding isn't?

The docs put it in one line: "Scheduling cycles are run serially, while binding cycles may run concurrently." Each decision has to see the reservations of the one before it, or two pods could be placed into the same free space. Binding involves API writes and maybe volume attachment, which are slow and don't change the decision, so they run off the critical path.

Every default plugin is enabled at every extension point it implements, with weights for the scoring ones:

pkg/scheduler/apis/config/v1/default_plugins.go
kubernetes/kubernetes @ v1.37.0 ↗
Go
// getDefaultPlugins returns the default set of plugins.
func getDefaultPlugins() *v1.Plugins {
	plugins := &v1.Plugins{
		MultiPoint: v1.PluginSet{
			Enabled: []v1.Plugin{
				{Name: names.SchedulingGates},
				{Name: names.PrioritySort},
				{Name: names.NodeName},
				{Name: names.NodeUnschedulable},
				{Name: names.TaintToleration, Weight: ptr.To[int32](3)},
				{Name: names.NodeAffinity, Weight: ptr.To[int32](2)},
				{Name: names.NodePorts},
				{Name: names.NodeResourcesFit, Weight: ptr.To[int32](1)},
				{Name: names.VolumeRestrictions},
				{Name: names.NodeVolumeLimits},
				{Name: names.VolumeBinding},
				{Name: names.VolumeZone},
				{Name: names.PodTopologySpread, Weight: ptr.To[int32](2)},
				{Name: names.InterPodAffinity, Weight: ptr.To[int32](2)},
				{Name: names.DefaultPreemption},
				{Name: names.NodeResourcesBalancedAllocation, Weight: ptr.To[int32](1)},
				{Name: names.ImageLocality, Weight: ptr.To[int32](1)},
				{Name: names.DefaultBinder},
			},
		},
	}
	/* ... */

When a pod can't be placed, the Filter plugins' reasons are counted and put in the pod's FailedScheduling event, in the format 0/N nodes are available: <reasons>. That message is the first thing to read.

Filtering every node for every pod gets expensive when a cluster has thousands of them, and the scheduler has a shortcut for that.

7.3Not every node gets looked at

On a large cluster, running every Filter against every node for every pod is too slow. So the scheduler stops once it has found "enough" feasible nodes:

pkg/scheduler/schedule_one.go
kubernetes/kubernetes @ v1.37.0 ↗
Go
func (sched *Scheduler) numFeasibleNodesToFind(percentageOfNodesToScore *int32, numAllNodes int32) (numNodes int32) {
	if numAllNodes < minFeasibleNodesToFind {       // 100
		return numAllNodes
	}
	/* ... use the profile's percentageOfNodesToScore if set ... */
	if percentage == 0 {
		percentage = int32(50) - numAllNodes/125
		if percentage < minFeasibleNodesPercentageToFind {   // 5
			percentage = minFeasibleNodesPercentageToFind
		}
	}
 
	numNodes = numAllNodes * percentage / 100
	if numNodes < minFeasibleNodesToFind {
		return minFeasibleNodesToFind
	}
	return numNodes
}

Under 100 nodes, it checks all of them. At 1,000 nodes it stops after 420 feasible ones (42%); at 5,000, after 500 (10%). It walks nodes round-robin across zones and resumes where the last pod left off, per the performance tuning docs, so the sample moves around the cluster.

?Doesn't that mean the pod might not get the best node?

Yes, on purpose. The best of 500 feasible nodes, taken from a sample that moves around the cluster, is usually close to the best of all 5,000, and it takes about a tenth of the work to find. Borg, Google's cluster manager and Kubernetes' ancestor, made the same trade and called it relaxed randomization (section 8).

Among the nodes it did look at, the scheduler still has to pick one.

7.4Spread or pack

For resources, the main scoring plugin is NodeResourcesFit, and its default strategy is LeastAllocated: prefer the node with the most room left.

pkg/scheduler/framework/plugins/noderesources/least_allocated.go
kubernetes/kubernetes @ v1.37.0 ↗
Go
// The unused capacity is calculated on a scale of 0-MaxNodeScore
// 0 being the lowest priority and `MaxNodeScore` being the highest.
// The more unused resources the higher the score is.
func leastRequestedScore(requested, capacity int64) int64 {
	if capacity == 0 {
		return 0
	}
	if requested > capacity {
		return 0
	}
 
	return ((capacity - requested) * fwk.MaxNodeScore) / capacity
}

An empty 4-CPU node scores 100, and one with 1 CPU already requested scores (4 − 1) × 100 ÷ 4 = 75. Let's run it on a small cluster: three nodes with 4 CPUs each, and thumbnail running as six pods that each request 1 CPU.

LeastAllocated places six 1-CPU pods on three 4-CPU nodes
Pendingwaiting for a nodenode-a4 CPUsnode-b4 CPUsnode-c4 CPUsthumb-11 CPUthumb-21 CPUthumb-31 CPUthumb-41 CPUthumb-51 CPUthumb-61 CPU
Step 1. thumbnail runs as six pods, each requesting 1 CPU, and the cluster has three empty nodes. The scheduler takes the pods one at a time: it filters the nodes (all three have room) and scores them.
1 / 5

Spreading leaves headroom on every node for a load spike and limits the blast radius of one node dying. It has a cost, though:

Predict before you read on

Three nodes with 4 allocatable CPUs each hold the six 1-CPU pods placed by LeastAllocated, two per node. Now a pod requesting 3 CPUs arrives. Where does it go?

Borg's paper describes the same tension. Its original scorer spread load ("we sometimes call this 'worst fit'") and fragmented machines for large tasks. Pure best fit packed tightly but "penalizes any mis-estimations in resource requirements". The hybrid it settled on tries to minimise stranded resources, capacity left unusable because another resource on the machine is full, and gave "about 3–5% better packing efficiency" than best fit.

Resources decide whether a pod can fit. Operators usually also want to say where a pod must not go, and where it should.

7.5Taints, tolerations and affinity

Resource fit decides whether a pod can go on a node. A handful of other mechanisms decide whether it should. A taint is a mark on a node that repels pods. A toleration is a pod's permission to ignore a taint. Affinity is a rule on a pod that pulls it towards nodes, or towards other pods, with certain labels (key=value tags attached to nodes and pods).

MechanismSet onHard or softTypical use
Taint NoSchedule / PreferNoScheduleThe nodeHard / softKeep general pods off GPU or dedicated nodes
Taint NoExecuteThe nodeHard, and evicts running podsNode problems: not-ready, unreachable
TolerationThe pod—Let a pod onto a tainted node; it doesn't attract it there
Node affinityThe podrequired… or preferred… (weight 1–100)Pin to a zone, an instance type, an architecture
Pod (anti-)affinityThe podEitherCo-locate with a cache; keep replicas apart
Topology spreadThe podmaxSkew (how uneven the counts may get), hard or softEven replicas across zones

A few details from the taints and affinity docs catch people:

  • A toleration isn't an attraction. It permits a pod onto a tainted node. To keep a pod on dedicated nodes, you need a taint and node affinity.
  • IgnoredDuringExecution means what it says. Change a node's labels after scheduling and its pods keep running.
  • Every pod tolerates dead nodes for five minutes. The DefaultTolerationSeconds admission controller (a plugin in the API server that edits pods as they're created) adds 300-second tolerations for not-ready and unreachable, so pods on a lost node aren't evicted for five minutes.
  • Inter-pod affinity is expensive. The docs say it "require[s] substantial amounts of processing" and "we do not recommend using them in clusters larger than several hundred nodes." Topology spread constraints cover most of the same needs.

When even these rules can't find a node, the scheduler has one more move.

7.6Preemption

When Filter finds no node, the default PostFilter plugin looks for lower-priority pods it could evict. Priority comes from a PriorityClass; user classes go up to 1,000,000,000, and the two system classes sit above that, per the priority docs. A PodDisruptionBudget (PDB) is a limit an application sets on how many of its pods may be taken down at once, and preemption tries to respect it.

A high-priority pod preempts its way onto a node
High-priority podSchedulerAPI serverVictim podsschedule mefind victimsnominatedNodeNamedeleteretry
Step 1. Filter rejects every node. PostFilter runs the default preemption plugin.
1 / 5

?Why doesn't preemption just bind the pod straight away?

Because the space isn't free until the victims have exited, which can take their full grace period. Recording a nomination and retrying lets the scheduler keep working, and lets the pod take a better spot if one opens up.

PodDisruptionBudgets are respected only on a best-effort basis: if no other option exists, preemption will violate one. A class with preemptionPolicy: Never gets its place in the queue by priority but never evicts anything.

The scheduler has now chosen a node and written its name down. Whether the pod runs there is up to someone else.

7.7The kubelet has the last word

Binding only writes a node name. The kubelet on that node then runs its own admission check against its own view of the node, and it can refuse. The reasons are fixed strings in pkg/kubelet/lifecycle/predicate.go: OutOfcpu, OutOfmemory, OutOfephemeral-storage, OutOfpods. A pod rejected here goes to Failed, and its controller (the Kubernetes component that keeps the right number of copies running) creates a replacement.

After admission the kubelet does the pod's cold start. It creates the pod's sandbox and its network through two standard interfaces (CRI for the container runtime, CNI for networking), pulls images, runs init containers (containers that must finish before the app starts), starts the app containers, and waits for readiness probes (health checks that must pass before traffic arrives). Image pulls are serialised by default: serializeImagePulls defaults to true unless you set maxParallelImagePulls to 2 or more (defaults.go). One large image can hold up every other pod starting on the node.

Later, under memory or disk pressure, the kubelet evicts pods itself, with no PDBs and no grace period for hard thresholds (memory.available under 100 Mi by default). It goes first for pods using more than they requested, then by priority, then by how far over their request they are, per the node-pressure eviction docs.

A scheduler with filters, scores, sampling and a second check at the node can look over-engineered. Each part has a history.

08Borg, Omega, and why the scheduler looks like this

Kubernetes' scheduler descends from two Google systems with public papers, a lineage its authors describe in Borg, Omega, and Kubernetes (ACM Queue, 2016). The papers explain design choices that look arbitrary on their own.

8.1Borg: feasibility, scoring, and shortcuts

Large-scale cluster management at Google with Borg (Verma et al., EuroSys 2015) describes a scheduler in two parts, "feasibility checking, to find machines on which the task could run, and scoring, which picks one of the feasible machines". Kubernetes' Filter and Score are the same split.

The paper's median cell (Borg's name for a cluster) was about 10,000 machines, and three shortcuts made that tractable:

Borg techniqueWhat it doesKubernetes' version
Score cachingKeep a machine's score until the machine or task changesThe scheduler's cache and per-cycle state
Equivalence classesScore one task per group of identical tasksAn equivalence cache, removed in 2018; now OpportunisticBatching (KEP-5598, a Kubernetes Enhancement Proposal; beta and on by default since 1.35) reuses one pod's ranked nodes for the next pod with the same signature
Relaxed randomizationExamine machines in random order until "enough" are feasiblepercentageOfNodesToScore (section 7.3)

Without them, scheduling a cell's whole workload from scratch "did not finish after more than 3 days"; with them it took "a few hundred seconds".

?What did Borg find was the slow part of starting a task?

The scheduler wasn't it. Median task startup was "about 25 s", and package installation took "about 80% of the total". So Borg's scorer preferred machines that already had the task's packages. That's the ImageLocality score plugin in Kubernetes, and it's the same lesson as section 4: fetching code and running init dominate.

Borg also over-committed on purpose. It reserved what a task was predicted to use, not what it asked for, and "about 20% of the workload" ran in the reclaimed gap. Kubernetes' requests and limits are a simpler form of that idea: the request is what you're promised, and the gap up to the limit is borrowed.

8.2Omega: shared state and optimistic placement

Omega (Schwarzkopf et al., EuroSys 2013) compared three architectures: a monolithic scheduler, a two-level design like Mesos (another cluster manager) where a central allocator offers resources to framework schedulers, and its own shared state, where several schedulers see the whole cluster and commit placements with "lock-free optimistic concurrency control". A conflicting commit is retried.

Kubernetes is mostly monolithic, with an Omega flavour. One scheduler makes decisions serially, but the cluster state lives in the API server and every write is a compare-and-swap on resourceVersion (an update that succeeds only if nobody has changed the object since you read it). You can run a second scheduler beside the default one, and both commit through the API server. But nothing in the API server checks a node's capacity when a pod is bound, so two schedulers can place pods into the same free space. The kubelet's admission check (section 7.7) catches the loser.

Now that every rung has been taken apart, the remaining question is how to see these costs and decisions on a system you run.

09Operating it

9.1What each start costs

Here are the numbers from this chapter side by side. They come from different systems, so they show the scale of each cost and aren't a benchmark of one against another.

1.4 ms
Start a process on the host
section 1, /usr/bin/true
178 ms
Start a whole new container
docker run --rm alpine true, includes the call into Docker's background service
48 ms
Start a container with runc
chapter 11, without the Docker layer
< 125 ms
A Firecracker microVM reaches guest init
NSDI 2020, with about 3 MB of memory overhead
182 ms
thumbnail's Init: import boto3, create an S3 client
section 4.2, venv on the VM's local disk
1,040 ms
The same Init on a slow shared disk
section 4.2, venv on a host share
0.58 ms
The same Init, forked from an initialised parent
section 4.2, the idea snapshots reuse

9.2Where to look

Each question the chapter raised has a tool that answers it on a running system.

Shell
# Kubernetes: why is this pod Pending? (sections 7.1 and 7.2)
kubectl describe pod thumbnail-7d9f -n prod | sed -n '/Events/,$p'   # "0/N nodes are available: ..."
kubectl get events -A --field-selector reason=FailedScheduling
kubectl describe node ip-10-0-3-7 | sed -n '/Allocated resources/,/Events/p'  # requests vs allocatable
 
# Is a container throttled? (section 7.1)
cat /sys/fs/cgroup/cpu.stat          # from inside it: nr_periods, nr_throttled, throttled_usec
 
# Lambda: which invocations were cold, and how long was Init? (sections 3 and 4)
#   REPORT lines carry "Init Duration: … ms" only on cold starts.
#   INIT_REPORT lines appear on Init timeouts and failures, and always with SnapStart or provisioned concurrency.
aws logs filter-log-events --log-group-name /aws/lambda/my-fn --filter-pattern '"Init Duration"'

In Python, python -X importtime -c 'import yourmodule' prints the time each import took, which is usually where a function's static init goes (section 4.2).

9.3Rules that hold up

  1. Keep static init small and lazy. Import only the clients you use, and create rarely used ones on first use (section 4).
  2. Generate anything that must be unique inside the handler or with a CSPRNG that reads the kernel, never in static init, as soon as snapshots are involved (section 5.4).
  3. Size concurrency with rate × duration, then check the requests-per-second limit too, because very fast functions hit it first (section 6.1).
  4. Use provisioned concurrency for a latency floor. A keep-warm ping keeps one environment alive and no more (section 6.2).
  5. Set requests from real usage, and make memory request equal to limit for anything important, so the pod is Guaranteed or at least well protected (section 7.1).
  6. Think twice before a tight CPU limit on a latency-sensitive service, and watch nr_throttled (section 7.1).
  7. Plan for the empty cache. If your cold path depends on a cache, decide what happens on the day it's empty (section 3.3).

9.4What you give up

You getYou payWhen the bill arrives
A kernel of your own (VM, microVM)Boot or restore time, and memory for a second kernelEvery scale-out, unless something is kept warm
Density on a shared kernel (containers)A boundary that's the kernel's syscall surfaceThe next kernel escape CVE (a numbered public record of a security bug)
Scale to zero (functions)Cold starts, and a platform that recycles your processThe first request after a quiet period, or a burst
Snapshots instead of initClones that share every byte of stateThe day two requests get the same ID
A scheduler that places for youDecisions made from requests, not usageThe first time a Pending pod meets an idle dashboard
Spreading by defaultFragmented free spaceWhen a large pod arrives

9.5Symptom, cause, fix

Each row starts from something you'd see on a running system and names the mechanism from this chapter behind it. One new term appears: a descheduler evicts running pods so that the scheduler can place them again, this time packed more tightly.

SymptomLikely causeFix
Pod Pending, cluster CPU graphs show plenty freeRequests, not usage, fill nodesRight-size requests from real usage; look at Allocated resources on nodes
Pod Pending with enough total free, but no single node fitsFragmentation from spreadingMostAllocated scoring, bigger nodes, or a descheduler; let the autoscaler add a node
0/N nodes are available: N node(s) had untolerated taintMissing toleration, or a node still tainted not-readyAdd the toleration, or fix the node
Latency spikes, CPU usage well under the limitCFS throttling within 100 ms periodsRaise or remove the CPU limit; check nr_throttled; kernel 5.4 or later
BestEffort pods die first under pressureoom_score_adj 1000 and eviction orderSet requests; memory request = limit for anything important
New pods on one node start slowlySerialised image pulls behind one large imageSmaller images, pre-pulled images, maxParallelImagePulls
Lambda p99 bad after quiet periods, fine under steady loadCold starts on scale-out and after recyclingTrim static init, then provisioned concurrency or SnapStart
Lambda throttles well under the concurrency quotaThe 10× requests-per-second limit, or the scaling rateRaise the quota; smooth the burst with a queue
Duplicate IDs or identical random values after enabling SnapStartUniqueness generated during InitGenerate in the handler, use a kernel CSPRNG, or an after-restore hook
A function "occasionally" doubles its duration after an errorSuppressed init after a crash resetFix the crash; watch for Init time inside the REPORT duration

10Summary

  1. Each rung hides one more layer. VMs hide hardware, containers hide the OS install, functions hide the process and whether it's running at all.
  2. Serverless platforms keep the separate kernel. Lambda and Fargate give each tenant a Firecracker microVM; density comes from making VMs small, not from sharing a kernel.
  3. A function's life is Init, Invoke and Shutdown, frozen between requests. A cold start is the wait for Init, and a warm start thaws an environment whose globals are still set.
  4. Lambda loads code lazily and from caches. Images are cut into 512 KiB chunks, deduplicated by convergent encryption, and almost every read is served by the worker or AZ cache; only about 0.06% go to S3.
  5. Booting isn't the slow part; initialising is. A microVM reaches init in under 125 ms. Runtime start and your static code usually cost more, and a process forked after its imports starts about 300 times faster.
  6. Snapshots replace init with page faults. Memory is mapped copy-on-write and loaded as touched; prefetching the working set is what makes restores fast.
  7. Clones share everything, including secrets. With SnapStart, anything unique created during Init is duplicated in every environment.
  8. Concurrency is Little's law. Requests per second × duration, with a separate 10× requests-per-second limit that fast functions hit first.
  9. The scheduler reads requests, never usage. Pods go Pending on idle clusters, and requests versus limits also set the QoS class that decides who gets OOM-killed.
  10. Filter, then score, on a sample. Large clusters score only a slice of nodes, and the default scorer spreads, which fragments free space.
  11. Placement isn't finished at bind. The kubelet can still reject the pod, pull images one at a time, and later evict it under pressure.

11Build this

A cold-start budget for one service.

  • Pick a real function or service and break its start into the six pieces of section 4.1. Measure your part with python -X importtime, -Xlog:startuptime or --cpu-prof, whichever fits your runtime.
  • Build a "zygote": a parent that does all static init, then forks a child per request. Compare its time-to-first-byte with a fresh start, as in section 4.2.
  • Add the uniqueness test from section 5.4 to your code base: fork after init and assert that every ID, seed and token differs between the children.
  • On a kind or minikube cluster (a small Kubernetes cluster that runs on one machine) with three small nodes, reproduce the Predict in section 7.4, then switch NodeResourcesFit to MostAllocated in a scheduler profile and watch the large pod fit.

12Interview questions

beginnerWhat's the difference between a request and a limit in Kubernetes?›

The scheduler uses requests to place pods: a node is full when the sum of requests reaches its allocatable capacity. Limits are enforced at run time: a CPU limit becomes a CFS quota and throttles, and a memory limit is a ceiling past which the container is OOM-killed. If you set only a limit, the request defaults to it. Together they set the QoS class.

beginnerWhat is a Lambda cold start?›

A request that arrives when no initialised execution environment is free. Lambda has to place a new microVM, load the code, start the runtime and run the function's static initialisation before the handler runs. The static init is usually the biggest part, and the part you control. Warm requests reuse a frozen environment and skip all of it.

intermediateA pod is Pending, but the cluster's CPU usage is 30%. Why?›

Probably requests. The scheduler compares requests, not usage, against allocatable capacity, so nodes can be "full" of reservations while idle. Or the free capacity is fragmented: plenty in total, but no single node has enough, which the default spreading scorer makes likely. Read the FailedScheduling event, then compare Allocated resources with actual usage on the nodes.

intermediateProvisioned concurrency or SnapStart: how do you choose?›

Provisioned concurrency keeps N environments fully initialised, so the first N concurrent requests never see a cold start, at a cost for idle time. SnapStart restores new environments from a post-init snapshot, which cuts the cold start instead of removing it, at no extra charge, but it duplicates any state that Init made unique. You can't use both on one version. Pick provisioned concurrency for a predictable floor, SnapStart for a heavy init with bursty traffic.

intermediateHow does Lambda start a 10 GiB container image quickly?›

It doesn't download it. The image is flattened and split into 512 KiB chunks, and the microVM sees a block device whose reads are served lazily, chunk by chunk, from a worker-local cache, then an AZ-level cache, then S3. Containers read only a small fraction of their data at startup, and convergent encryption lets identical chunks across customers be stored and cached once.

deepWhat goes wrong when you restore many VMs from one snapshot?›

Every clone starts with identical memory, so every generator state, cached random bytes, UUID and nonce created before the snapshot is shared. Unlike fork(), a VM restore runs no at-fork hooks in the guest's processes. Lambda reseeds the kernel's /dev/random and /dev/urandom on restore, so code reading the kernel every time is safe; user-space RNGs and anything cached at init need regenerating in the handler or an after-restore hook.

deepWhy does kube-scheduler not score every node in a 5,000-node cluster?›

Filtering and scoring every node for every pod doesn't scale. By default it stops after finding a percentage of feasible nodes, 50 − nodes/125, floored at 5% and at 100 nodes: 500 nodes at 5,000. It walks nodes round-robin across zones and resumes where it left off. The best of a large sample is usually close to the global best. Borg did the same and called it relaxed randomization.

deepWhy can a pod that the scheduler bound still fail to start on the node?›

Binding only records a node name. The kubelet runs its own admission against its own view of the node's resources and can reject with OutOfcpu, OutOfmemory or OutOfpods, for example when a second scheduler or a static pod (one the kubelet starts itself from a file, without going through the scheduler) used the same space. The shared state is optimistic, in Omega's sense, and the kubelet is the final check.

13Go deeper

check yourself
A container requests 500m CPU and 1 Gi memory, and has limits of 1 CPU and 1 Gi. What's its QoS class?›

Burstable. Memory's request equals its limit, but CPU's doesn't, and any mismatch in any resource makes the pod Burstable.

Your Lambda's static init takes 12 seconds. What happens on a cold start?›

Init is cut off at 10 seconds, then retried at the first invocation under the function's own timeout. Provisioned concurrency and SnapStart functions can run Init for much longer.

How many nodes does the default scheduler look for in a 1,000-node cluster?›

420: the percentage is 50 − 1000/125 = 42%.

What does a Firecracker snapshot restore read from the memory file up front?›

Almost nothing. It maps the file MAP_PRIVATE, so pages are loaded on first touch and copied on first write.

Brooker et al., On-demand Container Loading in AWS Lambda (USENIX ATC 2023)

Best paper. Lazy block loading, convergent encryption, tiered caches, erasure coding in the cache, and how a very good cache can turn its own outage into an overload of the origin. arXiv.

Agache et al., Firecracker: Lightweight Virtualization for Serverless Applications (NSDI 2020)

Why Lambda moved from containers in per-customer VMs to microVMs, and what the minimal device model buys. USENIX.

Ustiugov et al., Benchmarking, Analysis, and Optimization of Serverless Function Snapshots (ASPLOS 2021)

Where snapshot restore time goes (page faults) and how recording a working set fixes it. arXiv.

Verma et al., Large-scale cluster management at Google with Borg (EuroSys 2015)

Feasibility and scoring, relaxed randomization, resource reclamation, and the 25-second task start. Google Research.

Schwarzkopf et al., Omega (EuroSys 2013)

Monolithic, two-level and shared-state schedulers compared on Google traces. Google Research.

Shahrad et al., Serverless in the Wild (USENIX ATC 2020)

Azure Functions' production workload: 81% of applications invoked at most once a minute on average, and a histogram keep-alive policy that beats a fixed timeout. arXiv.

pkg/scheduler/schedule_one.go

The scheduling cycle itself: schedulePod, findNodesThatFitPod, prioritizeNodes. v1.37.0.

Lambda execution environment lifecycle

The primary source for section 3.1, including suppressed inits and the shutdown limits. AWS docs.

Virtualization, Hypervisors & microVMs

The mechanism under the VM and microVM rungs: exits, nested paging, virtio, and Firecracker's numbers. Chapter 47.

Containers from Scratch

Namespaces, cgroups and overlayfs, and a measured CPU limit. Chapter 11.

Contention, Queueing & Tail Latency

Little's law and why the average hides the cold-start tail. Chapter 16.

Processes, Threads & Scheduling

The kernel's scheduler, one level below kube-scheduler: CFS, quotas and run queues. Chapter 06.