KnowSys

Designing Zomato

It's New Year's Eve, Aditi orders four biryanis for a party in Gurugram, and the app promises them in 38 minutes. We'll design the system that keeps that promise while a kitchen, a rider and a customer each do their part on their own phones: the order as a durable state machine, which restaurants can deliver to you right now, predicting how long the kitchen will take, sending the rider so they arrive as the food is ready, batching two orders on one scooter, live tracking over patchy mobile networks, refunds, and the busiest night of the year.

⏱ 55 min read◆ IntermediateAssumes: the Uber case study (chapter 50), the UPI case study (chapter 70) helps, chapter 40 (reliability) helps
Start reading

It's twenty to ten on New Year's Eve in Gurugram. Aditi is at a friend's flat in Sector 56, the party is bigger than anyone planned, and the food has run out. She opens Zomato, scrolls past a banner warning that demand is high tonight, and finds a biryani place about six kilometres away, which we'll call Noor's Biryani. She adds four chicken dum biryanis, pays by UPI, India's instant bank-to-bank payment system, and the app says they'll arrive in 38 minutes. Across town, a tablet on Noor's counter starts ringing. A quarter of an hour later, Sunil, a delivery partner on an electric scooter, gets a ping on his phone. At 10:17 Aditi's phone buzzes: Sunil is downstairs.

It looks like Uber with food in the back, and part of it is. Zomato still has to find a nearby rider, and our Uber case study (chapter 50) already covered how to do that among millions of moving phones. But a ride has two parties, and an order has three. Aditi, Noor's kitchen and Sunil each hold one piece of the order on their own device, on their own network, and any of them can tap a button twice, lose signal, or not answer. There's a step in the middle that no server controls: someone in a hot kitchen has to cook four biryanis, and nobody knows exactly how long that will take. Send Sunil too early and he stands at the counter for fifteen minutes, earning nothing. Send him too late and the biryani goes cold under the heat lamp. And Aditi isn't alone: on New Year's Eve 2023, Zomato's own engineers wrote that orders peaked at about 8,400 a minute.

In this case study we'll design the system behind Aditi's order the way an engineer would: start with the obvious design, find exactly where it breaks, and fix it. The question we'll keep coming back to is this: when Aditi taps Place Order, how does Zomato make a promise about when her food will arrive, and then keep it, when the kitchen and the rider aren't machines it controls and the whole country is ordering at once? We'll follow her order from the restaurant list to the doorstep: the state the order moves through, the list of restaurants that can deliver to her, the clock that predicts the kitchen, the moment Sunil is sent, the dot that moves on her map, the refund if it all goes wrong, and the night that tests all of it.

01What we're building, and how big

1.1What it has to do

Getting Aditi her biryani comes down to a short list:

  1. Show the restaurants that can deliver to her right now, with their menus and what's available, and a delivery time for each.
  2. Take the order and the payment, exactly once, however many times she taps.
  3. Get the restaurant to accept it and cook it.
  4. Predict when the food will be ready and when it will arrive, and show that as a promise.
  5. Send a delivery partner to the restaurant at the right moment, sometimes with a second order on the same trip.
  6. Track the order live, including the rider's position on a map, until it's handed over.
  7. Handle cancellations and refunds when something goes wrong, and decide who pays for the mistake.

And the qualities it needs while doing that:

  • Never lose or double an order. A paid order that the restaurant never hears about is the worst failure in the business: Aditi has paid, nobody is cooking.
  • Honest times. A delivery time on the restaurant list is a promise, and missing it costs trust, refunds and coupons.
  • Work on bad networks. The restaurant's tablet sits on shop Wi-Fi, Sunil's phone moves between towers on a scooter, and Aditi might be on a congested network at a party.
  • Survive the peaks. Demand isn't smooth: it jumps at dinner, on rainy evenings, during a cricket final, and most of all on New Year's Eve.

This list differs from Uber's in one important way. Uber's hard question was about space: which of the moving cars is closest, by road. Zomato inherits that question, but its hardest problems are about time and agreement: when will the food be ready, when should the rider leave, and how do three people's phones agree on what state the order is in.

1.2How big is it?

Zomato is the food-delivery business of a listed company that used to be called Zomato Ltd. In February 2025 its board approved renaming the company Eternal Ltd, effective 20 March 2025, keeping "Zomato" as the name of the app; Eternal also owns Blinkit (quick commerce), District (going out) and Hyperpure (supplies for restaurants). Eternal's annual report for the year to March 2026 gives the scale of the food business:

Metric (food delivery)FY24FY25FY26
Orders delivered753 million853 million979 million
Average monthly transacting customers18.4 million20.6 million24.3 million
Average order value (net)₹365₹385₹388

The financial year runs April to March, so FY26 is April 2025 to March 2026. In that year the average month had about 330,000 restaurants and 552,000 delivery partners active, across more than 900 cities. By the quarter ending June 2026 the delivery partners had grown to an average of 638,000 a month. Spread evenly, 979 million orders a year is roughly 2.7 million a day, but nothing about food is spread evenly.

New Year's Eve is the night that shapes the design. Zomato's engineering team described 31 December 2023 in a post in February 2024: orders peaked at about 8,400 a minute, the day passed 3 million orders, and at the same peak the platform served about 81 million requests a minute. On New Year's Eve 2025, Zomato's founder, Deepinder Goyal, posted that Zomato and Blinkit together delivered more than 7.5 million orders that day, through more than 450,000 delivery partners; he didn't give a per-minute peak or a split between the two apps.

Your turn: design it before reading on

At the 2023 peak, about 8,400 orders a minute arrived and the platform served about 81 million requests a minute. How many requests is that per order placed? Where could they all be coming from?

02Version 1: one table and a nearest rider

2.1The obvious design

Start with the most obvious design: a single order service and a database. Restaurants and their menus live in tables. When Aditi opens the app, the server lists every restaurant within, say, seven kilometres. When she taps Place Order, it inserts a row into an orders table with status PLACED. Noor's tablet asks the server every few seconds whether it has new orders, and when the cook taps Accept, the row becomes ACCEPTED. At that moment, the server finds the nearest free rider, the way chapter 50 did, and assigns them. The rider's app updates the row as they pick up and deliver, and Aditi's app asks every few seconds for the status.

Version 1: three apps, one order service, one table
place, pollpoll, acceptpicked, deliveredAditi's appNoor's tabletpolls every 5 sSunil's appOrder serviceorders tableid, status, rider
Step 1. Aditi taps Place Order. The service inserts a row with status PLACED.
1 / 3

On a quiet Tuesday afternoon in one city, this works. On New Year's Eve it breaks in four places, and each of them is a section of this chapter:

  • The list lies. It shows Noor's because Noor's is within seven kilometres. But there may be no free riders anywhere near Noor's tonight, or Noor's kitchen may already have forty orders queued. Aditi orders, and then waits an hour and a half.
  • The rider waits. Sunil is assigned the moment Noor's accepts. He's six minutes away and the biryani takes eighteen to pack, so he stands at the counter for twelve minutes, on the busiest night of the year, when he could have delivered another order. Multiply that by every order in the country.
  • The order isn't safe. Aditi taps Place Order, the network stalls, she taps again: two orders, two payments. Noor's tablet taps Accept, the reply is lost, the tablet retries. Sunil's phone has no signal in the basement car park when he taps Picked Up.
  • The night overwhelms it. One service and one table, polled by every app every few seconds, can't take 81 million requests a minute.

We'll fix them in an order that follows the order's own life, but we'll start with the third, because everything else is built on top of the order being right.

03The order, a state machine three people share

3.1Who may move the order, and where to

Look at the life of Aditi's order and you'll see it moves through a fixed set of stages, and only some moves between them make sense. It can go from placed to accepted, but not from placed to delivered. It can be cancelled before the restaurant starts cooking, but cancelling after Sunil has handed it over makes no sense. A thing that is always in exactly one of a fixed set of states, with a fixed list of allowed moves between them, is called a state machine. Each allowed move is a transition.

What makes a food order's state machine unusual is that the transitions belong to different people. Aditi can place and, within limits, cancel. Only Noor's kitchen can accept, or mark the food ready. Only Sunil can say he's picked it up and delivered it. And some moves belong to the system itself: assigning a rider, or cancelling an order the restaurant hasn't answered in time. Zomato hasn't published its internal states, so the ones below are our design, but every food platform needs something of this shape.

Aditi's order and who moves it
place, cancelaccept, readypicked up, deliveredassignAditicustomer appNoor's kitchenrestaurant tabletSunilpartner appOrder servicechecks each transitionOrder storestate + versionOrder eventsKafkaDispatch
Step 1. Aditi places the order. It's PLACED once her payment is confirmed. Only her app can make this move.
1 / 5

That diagram also shows where the order lives once version 1 is taken apart. One service owns the order's state and checks every transition. Everything else that cares about the order, dispatch, notifications, tracking, the restaurant's analytics, learns about changes by reading events, small messages saying "order 7781 is now ACCEPTED", from a log like Kafka (chapter 23). Zomato's February 2024 post says its Kafka clusters carried more than 450 million messages a minute at the New Year's Eve 2023 peak.

?Why not let each app write the status directly?

Because the apps disagree, and late. Suppose Aditi's screen is forty seconds old and still says PLACED, so it shows a Cancel button with no fee. In those forty seconds, Noor's cook accepted and started frying onions. If Aditi's app could write CANCELLED straight into the row, the kitchen would be cooking an order nobody will collect, and nobody would pay for the wasted food. A single owner that checks "is this move allowed from the state the order is in now?" turns that race into a clear answer: the cancel is refused or turned into a different cancel, one with a charge, as section 8 will show.

3.2Taps that arrive twice, late or out of order

Now the network. Every one of the three apps sometimes sends the same request twice, and sometimes sends requests in the wrong order. Let's walk through the cases one at a time, on Aditi's order:

  • Aditi double-taps Place Order, or taps it, sees a spinner, and taps again. Two requests arrive. If each makes an order, she pays twice and Noor's cooks eight biryanis.
  • Noor's tablet taps Accept and never hears back. The server got it and accepted the order, but the reply was lost on the shop's Wi-Fi. The tablet, reasonably, retries.
  • Aditi cancels from a stale screen, the case from 3.1: her request is about a state that no longer exists.
  • Sunil's phone has no signal in a basement car park. He taps Picked Up, then rides out and taps Delivered at Aditi's door. When the phone reconnects, it sends both, and the network delivers Delivered first.

Each case has a standard fix. For repeats, every request carries an idempotency key, a unique ID the app creates once per action and resends unchanged with every retry. On its side, the server remembers which keys it has already applied, and a repeat gets the original answer instead of being applied again. This is the same idea the Stripe case study (chapter 52) uses for payments. For stale requests, every order carries a version number that goes up by one on every transition. A request says which version it was looking at, and if the order has moved on since, the server refuses it and sends back the current state, so the app can show Aditi what happened. That's called optimistic concurrency: nobody locks the order, but a write based on an out-of-date view fails. For out-of-order events, the transition table itself does the work: Delivered from FOOD_READY isn't an allowed move, so it's refused, and the partner app keeps it in its outbox and retries after Picked Up has gone through.

This short program applies all three rules to a sequence of events like the ones above. ALLOWED is the transition table, seen is the set of idempotency keys already applied, and each event carries an ID, the state it wants, and the version its sender last saw:

An order state machine with idempotency keys and versions
python
Python
# Which state may follow which. Anything else is refused.
ALLOWED = {
    "PLACED":           {"ACCEPTED", "CANCELLED"},
    "ACCEPTED":         {"FOOD_READY", "CANCELLED"},
    "FOOD_READY":       {"PICKED_UP"},
    "PICKED_UP":        {"DELIVERED"},
}
 
order = {"id": "ORD-7781", "state": "PLACED", "version": 1}
seen = set()          # event ids already applied
 
def apply(event_id, new_state, expected_version):
    if event_id in seen:
        return "duplicate, ignored"
    if expected_version != order["version"]:
        return f"stale (sender saw v{expected_version}, order is v{order['version']})"
    if new_state not in ALLOWED.get(order["state"], set()):
        return f"refused: {order['state']} -> {new_state}"
    order["state"] = new_state
    order["version"] += 1
    seen.add(event_id)
    return f"ok, now {new_state} v{order['version']}"
 
events = [
    ("r-1", "ACCEPTED", 1),    # restaurant taps Accept
    ("r-1", "ACCEPTED", 1),    # its tablet retries: the reply was lost
    ("c-9", "CANCELLED", 1),   # customer cancels from a screen that was 40 s old
    ("r-2", "FOOD_READY", 2),
    ("p-4", "DELIVERED", 3),   # partner app, offline, replays out of order
    ("p-3", "PICKED_UP", 3),
    ("p-4", "DELIVERED", 4),   # the replay, retried in the right order
]
for e in events:
    print(f"{e[0]:4} {e[1]:10} ->", apply(*e))
output
C++
r-1  ACCEPTED   -> ok, now ACCEPTED v2
r-1  ACCEPTED   -> duplicate, ignored
c-9  CANCELLED  -> stale (sender saw v1, order is v2)
r-2  FOOD_READY -> ok, now FOOD_READY v3
p-4  DELIVERED  -> refused: FOOD_READY -> DELIVERED
p-3  PICKED_UP  -> ok, now PICKED_UP v4
p-4  DELIVERED  -> ok, now DELIVERED v5

The tablet's retry of r-1 is recognised and ignored, so the order is accepted once. Aditi's cancel, sent while she was looking at version 1, is refused as stale, and her app would now show "the restaurant has started preparing your order" with the terms for cancelling at that stage. Sunil's Delivered, arriving before his Picked Up, is refused without harm, because the table doesn't allow it yet; once Picked Up is applied, the same event goes through. Notice the order: the duplicate check comes first, so a retry of an event that already succeeded gets a friendly answer instead of a confusing "stale". A real service would also return the stored result for a duplicate.

Sunil's phone needs one more rule, because it spends a lot of its life offline. It records each tap locally, with its own ID and the time it happened, and sends them in order when it can. On the server, the time Sunil tapped is kept for the record (so the delivery time is right), while the version check and transition table keep the state correct whatever order the network delivered things in.

3.3Writing the state and telling everyone

There's one more way to lose an order, and it hides between two writes. When Noor's accepts, the order service has to do two things: save ACCEPTED in the order store, and publish an event so that Dispatch starts planning a rider. If it saves and then crashes before publishing, the order is accepted and no rider ever comes. If it publishes first and then fails to save, Dispatch sends a rider for an order the store says was never accepted.

A standard fix is called the transactional outbox. In the same database transaction that changes the order's state, the service also inserts the event into an outbox table. Either both are saved or neither is. A separate process reads the outbox and publishes each event to Kafka, marking it sent, and retries until it succeeds. That process might publish an event twice if it crashes between publishing and marking, so every consumer has to cope with duplicates, which they can, by the same idempotency trick: each event carries an ID, and Dispatch ignores one it has already handled.

Decision

Who drives an order from one state to the next?

Choreography
Each service reacts to events and emits its own: payment confirmed → restaurant notified → dispatch plans.
  • No central component
  • Services stay loosely coupled
  • Nobody holds the whole flow in one place
  • Timeouts and stuck orders are hard to see
chosen
Orchestration
One workflow per order owns the state, calls each step, and sets timers for what should happen next.
  • The whole life of an order in one place
  • Timers built in: 'restaurant hasn't accepted in 3 minutes'
  • The orchestrator must be durable and highly available
  • More coupling to one component

For an order, orchestration wins, mostly because of timers. Half the interesting logic is about things that didn't happen: the restaurant didn't accept, the rider didn't arrive, the food wasn't marked ready. With choreography, every one of those is a separate timer in a separate service. With one durable workflow per order, the order owns its own deadlines. Zomato hasn't published how its order flow is built. DoorDash has: in February 2021 it described rebuilding checkout so that each order is saved to Cassandra, announced on Kafka, and then driven by a workflow in Cadence, an open-source engine for durable workflows. The steps (fraud check, payment, creating the delivery, sending the order to the merchant) were split into retryable, idempotent steps tracked by an order state machine, and when a step fails for good, the workflow undoes the earlier steps to leave the order in a "clean failed state". DoorDash said the rebuild saved about 1% of orders a day that the old flow would have left stuck.

Those timers are where the system acts on its own. If Noor's doesn't accept within a few minutes, the tablet rings louder, an operations agent may call the restaurant, and eventually the order is cancelled and Aditi refunded, without her having to ask. Zomato's exact timings and steps aren't public, but every platform needs a ladder like it, because an unanswered order with a paid customer is exactly the failure from section 1.1.

Aditi's order is now safe. But she only placed it because the list told her Noor's could deliver to her, in 38 minutes. That's the first promise the system makes, before any order exists, and it's where version 1's list lied.

04Which restaurants can deliver to Aditi, right now

4.1Near isn't enough

Version 1 listed every restaurant within seven kilometres. That answers "is it close?", but Aditi's question is "if I order from here, will it reach me hot, in a reasonable time?" Three things decide that, and only the first is about distance:

  • Distance by road. Noor's is six kilometres away in a straight line, but if the only way across is a congested flyover, the ride may take forty minutes. Platforms usually draw each restaurant's delivery area as a shape on the map, not a circle.
  • Riders nearby. If every rider in Noor's part of Gurugram is already carrying an order, Aditi's biryani will sit on the counter until one frees up.
  • The kitchen's load. If Noor's already has forty orders in the queue, the forty-first won't be cooked in eighteen minutes, whatever the rider does.

Whether a restaurant can deliver to an address right now, in reasonable time, is called serviceability. Version 1 checked only the first condition, and checked it as a circle. The other two change minute by minute, and on New Year's Eve they change faster than anything else in the system.

Swiggy's engineers described their version of this in 2018: a restaurant is serviceable if a delivery from it is possible within some stipulated time, they show roughly 300 to 500 restaurants to each customer while considering thousands more, and the list has to appear within half a second, or the customer leaves the app. Zomato hasn't published how its serviceability works, so what follows is a reasonable design built from those requirements.

4.2Compute it per area, ahead of time

A naive way to check serviceability is to do it when Aditi opens the app: for each of a few thousand restaurants near her, compute the road distance, count free riders near it, and look up its queue. That's thousands of road-distance calculations per app open, and on New Year's Eve, tens of thousands of people open the app every minute. Swiggy wrote in 2021 that at an assumed 100,000 app visits a minute during peak demand, its serviceability checks came to nearly 200 million evaluations a minute, with a target of under 200 milliseconds for the slowest 1% of requests.

But most of the answer doesn't depend on Aditi. Rider supply around Noor's and the length of Noor's queue are the same for everyone looking at Noor's this minute. Only the last leg, from Noor's to her particular building, is hers. So we split the work:

  1. Per area, every minute or so: cut the city into cells, as chapter 50 did with H3 hexagons. A background job reads the rider positions and order events and computes, for each cell, how many riders are free or about to be free, and how many orders are waiting for one. Another computes each restaurant's kitchen stress: orders accepted but not yet ready, compared with what that kitchen normally handles at this hour. Swiggy's 2023 post on predicting delivery time lists this signal among its model's inputs: "stress at the restaurant", derived from orders placed and orders prepared in the recent past.
  2. Per restaurant, from those numbers: decide its delivery radius for the next minute. With plenty of riders and a calm kitchen, Noor's might deliver up to seven kilometres. With riders scarce, the radius shrinks to four; with the kitchen overwhelmed, it might be switched off for new orders altogether.
  3. Per request, cheaply: when Aditi opens the app, look up her cell, fetch the restaurants whose current delivery area covers it, and estimate the last leg from a precomputed table of cell-to-cell travel times. Swiggy's serviceability platform caches distances keyed by map cells in this way, because computing a fresh route (from Google's directions service or a search over an OpenStreetMap road graph) is too slow and too costly to do for every pair.
Building Aditi's restaurant list
who covers my cell?ranktimesAditi's appSector 56Listing serviceServiceabilityper cell, per minuteSearch indexSolr, sharded by areaTime servicedelivery estimateArea jobsriders, kitchen stressOrder + location events
Step 1. In the background, jobs read rider positions and order events and recompute, every minute or so, how many riders each cell has and how stressed each kitchen is.
1 / 5

Zomato has published a lot about the ranking step. Its 2023 series on search, which it said handles 100 million queries a day, describes an index built on Solr, the open-source search engine, with popularity stored per delivery area. Storing a separate popularity field for every area made the index's field list explode and run the servers out of memory, so the team moved to child documents per area instead. And because nearly every search is about one small area, they split the index into shards, separate pieces on separate machines, along geographic lines, so that about 95% of queries were answered by a single shard. That's the same insight as chapter 50's location index: data for a hyperlocal business should be split by place, because almost every question is about one place.

Decision

When should a restaurant's serviceability be decided?

On every request
Compute road distance, rider supply and kitchen load for each restaurant when the customer opens the app.
  • Always the freshest possible answer
  • Thousands of calculations per app open
  • Collapses exactly when traffic peaks
chosen
Per area, every minute
Background jobs compute supply, stress and each restaurant's delivery area; requests only look them up.
  • A request is a cheap lookup
  • The same answer for everyone in the cell, so it caches well
  • Can be degraded gracefully under load
  • Up to a minute stale
  • A sudden rush can overshoot before the next update

A minute of staleness costs little here, because rider supply and kitchen queues don't change much in a minute, while the per-request version would fall over on precisely the nights that matter. Overshooting is a real risk, though: if a thousand people order from one popular kitchen in the same minute, the next update arrives too late. That's why kitchen stress also has to be enforced at order time, by the order service itself, with a hard limit on open orders per kitchen; section 9 comes back to it.

4.3Menus that are mostly the same, and the part that isn't

Aditi taps Noor's and the menu opens: a hundred-odd items, photos, prices, options like "extra raita". Menus are the easiest thing in the system to cache, because they barely change. One menu is read by everyone in Gurugram who opens Noor's tonight, and changed perhaps a few times a week. So the menu is stored as one document per restaurant with a version number, and cached close to the app: in memory caches, and for the photos on a CDN, a network of caching servers spread close to users. Zomato's New Year's Eve 2023 post gives a sense of how much of the platform is cache reads: its in-memory caches held about 35 TB and served about a billion reads a minute at the peak.

Availability is the part that isn't stable. At 9:30 Noor's runs out of mutton, and the cook taps "out of stock" on the tablet. If that flag lives inside the cached menu, either the whole menu has to be thrown out of every cache, or customers keep ordering mutton biryani for the next half hour and Noor's keeps rejecting orders. So we split the two: the menu document is cached for a long time, and a small availability overlay, a list of item IDs that are off right now, is stored separately, kept for seconds, and applied on top when the menu is shown. The order service checks availability once more when Aditi taps Place Order, because the overlay she saw may be a few seconds old.

Aditi's menu: a long-lived document and a seconds-fresh overlay
Aditi's appCDNdish photosMenu servicemenu + overlayOrder servicerechecks on orderMenu cacheversioned, hoursStock overlayoff now, secondsMenu storedoc per restaurantNoor's tabletmarks out of stock
Step 1. Aditi opens Noor's. The photos come from the CDN, and the menu document comes from memory, cached for hours under its version number. Only a miss reaches the menu store.
1 / 4

Every restaurant on Aditi's list showed a delivery time, and Noor's said 38 minutes. That number decides whether she orders, and once she has, it's a promise. Where does it come from?

05Predicting the kitchen, and the promise

5.1What 38 minutes is made of

An ETA, an estimated time of arrival, is a sum of pieces, and for food it has more pieces than for a ride. Zomato's 2020 post on prep time described its delivery ETA as three parts: food preparation time, the rider's pickup time, and the rider's drop time. Swiggy's 2018 post wrote down how they combine:

Delivery Time = Max (Assignment Delay + First Mile Time, Prep Time) + Last Mile Time

Unpacked: the first mile is the rider's trip to the restaurant and the last mile is the trip from the restaurant to the customer. Before the first mile can start, the system has to find and assign a rider, which takes time of its own. The food leaves the restaurant when both the food and the rider are there, so the pickup moment is the later of "rider arrives" and "food is ready". Then the last mile, plus a few minutes to park, find the building and hand over. For Aditi's order it might look like this:

PieceMinutesWhat decides it
Assign a rider and ride to Noor's3 + 6Riders near Noor's, traffic
Cook and pack four biryanis18The kitchen, its queue, the dish
Pickup momentmax(9, 18) = 18Whichever is later
Ride to Sector 5615Distance, traffic, a two-wheeler's route
Park, find the flat, hand over3The building, the gate, the stairs
Total36, shown as 38Plus a buffer for uncertainty

Look at which number dominates. The rider's legs are a routing problem, the kind chapter 50 solved for Uber with a road graph and a learned correction. Zomato's 2022 posts on delivery ETA describe moving that part from a map-graph-based model to a tree-based machine-learning model (LightGBM), trained per city, served in about 2 to 3 milliseconds, at thousands of requests a second. Swiggy's 2018 post adds a detail that matters in India: mapping services didn't predict travel times well for two-wheelers, which slip through traffic that holds cars up. But the biggest piece, and the least certain, is the kitchen.

5.2Predicting a kitchen you can't see

The food preparation time, or prep time, is the time from the restaurant accepting the order to the food being packed and ready. It's hard to predict for reasons that have nothing to do with computers. Four biryanis from a big pot already on the stove take minutes; four from a kitchen with forty orders ahead of them take much longer. A tandoor holds a fixed number of naan at once. One restaurant can be quick at 4 pm and slow at 9 pm on a Saturday.

Zomato has written about this twice. In its February 2020 post, the model represented each dish with a learned list of numbers called an embedding (there were about 3.5 million distinct dishes on the platform), each restaurant with another, and fed in the restaurant's current running orders (up to five) and last five completed orders through a neural network layer designed for sequences, an LSTM. That cut the average error from 4.64 to 4.13 minutes.

That post also describes a problem that's more interesting than the model: the training data lied. Its only signal of "food ready" was the rider picking it up, so if the rider arrived late, the prep time looked long; if the rider arrived early and waited, it looked about right. The data mixed the kitchen's behaviour with the rider's. Zomato's fix was a product change, a Food Order Ready button on the restaurant's app, so the kitchen itself says when the food is ready, and it reported a 9% improvement in predictions landing within five minutes of the truth. DoorDash ran into the same trap from the other side in 2020: when its rider arrived after the food was ready, the true prep time was hidden (statisticians call such data censored), and it adjusted the training data to account for that before fitting an ordinary model.

Where the prep time comes from
trainNoor's tabletestimate, ready tapOrder eventsOrder historyready taps = labelsKitchen stateopen + last 5Prep modelembeddings, LSTMTravel modelLightGBM, per cityTime serviceto list + dispatch
Step 1. Noor's accepts with its own estimate and, when the food is packed, taps Food Order Ready. Both become order events.
1 / 4

?Why does the kitchen accept with its own estimate?

Because the kitchen knows things no model does: that the mutton has just gone in, or that the cook is alone tonight. That's why our state machine in section 3 had Noor's accept "with an estimate". The platform's model and the kitchen's estimate can be combined, and over time the model learns how far to trust each restaurant's own number.

5.3Late is worse than early

Noor's 38 minutes is a promise in two directions at once, and the two directions don't cost the same.

If the promise is too long, Aditi sees 55 minutes and orders from somewhere else, or not at all. If it's too short, she gets her biryani late, and a late order on New Year's Eve generates complaints, refunds and coupons. Swiggy's 2023 post states the trade-off plainly: delays make customers unhappy, but a conservative promise reduces the chance they order. DoorDash's 2021 post on ETAs says the same and then does something about it: it trained its model with a loss function that explicitly treats a late delivery as worse than an early one by a chosen factor.

A loss function is the score a model is trained to make small. An ordinary one punishes being two minutes early and two minutes late equally. An asymmetric loss punishes the late side more, so the trained model leans towards slightly longer estimates, by exactly the amount the business chooses. Zomato's 2022 prep-time post describes the same choice: a modified quantile loss that penalises underestimates. The ETA shown to Aditi then adds a buffer that grows with uncertainty: Zomato's 2022 ETA post lists real-time buffers for demand, traffic, weather and rider supply. On New Year's Eve all four are high at once.

Zomato's 2022 ETA post also mentions something it doesn't do: it doesn't show its ETAs to riders or give them a countdown. The promise is to the customer; pressuring the rider with a clock on a crowded road is a different matter. Zomato said the same in 2022 about its ten-minute delivery pilot, where riders weren't told the promised time.

And the prediction has a second use, which turns out to be just as important. Once we know Noor's biryani will be ready in about 18 minutes, we know when Sunil needs to be there.

06Sending the rider

6.1Timing the rider to the food

Version 1 assigned Sunil the moment Noor's accepted. Sunil was six minutes away and the biryani took eighteen, so he waited twelve minutes at the counter. Every food platform has written about this tension in nearly the same words. DoorDash's 2021 post on its dispatch system puts it this way: dispatch a rider too early and they wait for the order; dispatch too late and the food sits and gets cold. What we want is a rider who walks in just as the bag is sealed.

So we change the rule: dispatch just in time. Instead of sending a rider when the order is accepted, send one so that they arrive when the food is predicted to be ready: send time = predicted ready time − travel time to the restaurant. How well that works depends entirely on the prediction from section 5. This program compares the two rules over 10,000 simulated orders whose real prep time averages 18 minutes but varies a lot from order to order, for a rider who is always 6 minutes away. noise is how wrong our prediction is (as a standard deviation, so "±3 min" means about two-thirds of guesses are within three minutes), and early lets the rider leave a few minutes before the plan says:

When to send the rider: at the order, or just in time
python
Python
import random
random.seed(7)
 
# 10,000 biryani orders. Real prep time averages 18 minutes but varies a lot.
# The rider is 6 minutes' ride from the restaurant when assigned.
TRAVEL = 6
orders = [max(5, random.gauss(18, 5)) for _ in range(10_000)]
 
def simulate(policy, noise=0.0, early=0.0):
    rider_wait = food_wait = 0.0
    for prep in orders:
        if policy == "at order":
            send_at = 0                                  # assign the moment the order is placed
        else:
            guess = prep + random.gauss(0, noise)        # our prediction of prep time
            send_at = max(0, guess - TRAVEL - early)     # leave so we arrive as food is ready
        arrive = send_at + TRAVEL
        rider_wait += max(0, prep - arrive)              # rider standing at the counter
        food_wait += max(0, arrive - prep)               # food sitting on the counter
    n = len(orders)
    return rider_wait / n, food_wait / n
 
print(f"{'policy':34} rider waits  food waits")
for label, policy, noise, early in [
        ("assign at order", "at order", 0, 0),
        ("just in time, perfect guess", "jit", 0, 0),
        ("just in time, guess +-3 min", "jit", 3, 0),
        ("just in time, guess +-6 min", "jit", 6, 0),
        ("guess +-6 min, leave 3 min early", "jit", 6, 3)]:
    r, f = simulate(policy, noise, early)
    print(f"{label:34} {r:5.1f} min   {f:5.1f} min")
output
C++
policy                             rider waits  food waits
assign at order                     12.0 min     0.0 min
just in time, perfect guess          0.0 min     0.0 min
just in time, guess +-3 min          1.2 min     1.2 min
just in time, guess +-6 min          2.3 min     2.3 min
guess +-6 min, leave 3 min early     3.7 min     1.2 min

The first line is version 1: twelve minutes of waiting per order, every order. On a night with hundreds of thousands of orders, that's probably tens of thousands of rider-hours spent standing at counters, which means tens of thousands of hours of deliveries not made. Just-in-time dispatch with a perfect prediction removes it entirely. With a realistic prediction, the waits come back, split between the rider (when the food is late) and the food (when the rider is late), and they grow in step with how wrong the prediction is. In the last line, a few minutes' head start is the knob: leaving a few minutes early trades rider time for hotter food. Where to set it is a business decision, and it's the same asymmetry as section 5.3, seen from the rider's side.

Predict before you read on

Just-in-time dispatch decides at 9:52 that Sunil should leave at 9:58. What should the system do with Sunil in those six minutes?

6.2Who to send: the matching problem again

Once we know when a rider should leave, choosing which rider is the problem chapter 50 solved for Uber: build a table of how long each candidate rider takes to reach each restaurant, and pick the pairing that's best overall instead of greedily giving each order its nearest rider. That's the assignment problem, and chapter 50 shows with a small example why solving it for a batch beats greedy matching.

Food adds two twists. First, the riders worth considering include ones who aren't free yet: a rider who will drop off an order around the corner from Noor's in four minutes may be the best choice for an order whose food is ready in eighteen. Swiggy's 2019 post calls this next order assignment. Second, the cost of a pairing isn't just travel time; it includes the rider's waiting, the food's waiting and the chance that the rider declines the offer, because riders are independent workers and can say no. DoorDash's 2021 post describes models that predict exactly these quantities for each possible offer (when the order will be ready, travel times including parking, and how likely the rider is to accept), feeding an optimiser that makes the final decision.

6.3One rider, two orders

On a busy night there's a better trick than sending the right rider: sending one rider for two orders. Suppose Noor's has Aditi's order and, two minutes later, another for a flat in Sector 57, a kilometre from Aditi. Two riders could each make a trip from Noor's, or one rider could take both bags and drop them one after the other. Carrying more than one order on a single trip is called batching. (This is a different "batch" from chapter 50's, which meant collecting requests for a few seconds before matching them.)

Batching saves riders, which on New Year's Eve are the scarcest thing in the system. It costs time: someone's food now takes a detour. Swiggy's 2019 post put the prize in a sentence, that batching all orders would double the capacity of the fleet, and the catch in another: finding the best batches is NP-hard, meaning no known algorithm solves it exactly in reasonable time as the number of orders grows, while each order has to be assigned within minutes. So platforms set a limit on how much extra time batching may add to any customer's delivery, and batch only when the pair fits under it.

This program shows the trade-off. It scatters 40 orders over a busy 4 km × 4 km area in ten minutes, each with its own restaurant and customer, and pairs orders greedily, cheapest first, as long as neither customer's delivery gets more than max_extra minutes longer than a solo trip would be. The times are counted from when the rider sets off with the first order, at 3 minutes per kilometre plus 2 minutes per handover, and for simplicity both orders in a pair are ready at once:

Batching: riders saved against minutes added
python
Python
import math, random
from itertools import combinations
random.seed(3)
 
MIN_PER_KM = 3          # a scooter in evening traffic: about 20 km/h
HANDOVER = 2            # minutes to park, climb stairs, hand over
 
# 40 orders in one busy 4 km x 4 km area over ten minutes: (restaurant, customer)
pt = lambda: (random.uniform(0, 4), random.uniform(0, 4))
orders = [(pt(), pt()) for _ in range(40)]
km = lambda a, b: math.dist(a, b)
 
def solo(o):
    r, c = o
    return km(r, c) * MIN_PER_KM + HANDOVER
 
def pair(o1, o2):
    """One rider: pick up both, drop the nearer customer first.
    Returns minutes until each customer has their food."""
    (r1, c1), (r2, c2) = o1, o2
    best = None
    for first, second in [(c1, c2), (c2, c1)]:
        t_pick = km(r1, r2) * MIN_PER_KM                      # ride between restaurants
        t1 = t_pick + km(r2, first) * MIN_PER_KM + HANDOVER   # first drop
        t2 = t1 + km(first, second) * MIN_PER_KM + HANDOVER   # second drop
        times = (t1, t2) if first is c1 else (t2, t1)
        if best is None or sum(times) < sum(best):
            best = times
    return best
 
def plan(max_extra):
    """Greedily pair orders when neither customer waits more than
    max_extra minutes longer than a solo delivery would take."""
    left, trips, times = set(range(len(orders))), 0, []
    cands = []
    for i, j in combinations(range(len(orders)), 2):
        a, b = pair(orders[i], orders[j])
        extra = max(a - solo(orders[i]), b - solo(orders[j]))
        if extra <= max_extra:
            cands.append((extra, i, j, a, b))
    for extra, i, j, a, b in sorted(cands):
        if i in left and j in left:
            left -= {i, j}; trips += 1; times += [a, b]
    for i in left:
        trips += 1; times.append(solo(orders[i]))
    return trips, sum(times) / len(times), max(times)
 
print("extra wait allowed   riders needed   avg delivery   worst")
for limit in [0, 3, 6, 10, 20]:
    trips, avg, worst = plan(limit)
    print(f"{limit:>8} min        {trips:>6}           {avg:5.1f} min    {worst:5.1f} min")
output
C++
extra wait allowed   riders needed   avg delivery   worst
       0 min            40             8.0 min     15.0 min
       3 min            35             8.4 min     15.0 min
       6 min            25             9.9 min     17.4 min
      10 min            22            10.8 min     17.8 min
      20 min            20            11.9 min     24.0 min

With no extra wait allowed, nothing is batched: forty orders, forty riders. Allowing six extra minutes pairs up most of the orders that sit near each other, cutting the riders needed from 40 to 25 while the average delivery grows by roughly two minutes. After that the returns shrink: going from 10 to 20 minutes saves only two more riders, and the unluckiest customer's wait jumps from about 18 minutes to 24. Look at the shape more than the exact numbers: the first few minutes of allowed detour buy a lot of riders; after that, each rider saved costs more and more customer time, concentrated on a few unlucky customers.

A street grid with a central depot marked D and three coloured routes, each leaving the depot and visiting three or four stops before returning
The textbook version of the problem: a vehicle routing problem, where vehicles leave a depot and each visits several stops. Food batching is a harder cousin: the 'depots' are many restaurants, each stop has a deadline, and new orders keep arriving while the routes are being driven.Image: Zootos, CC BY-SA 4.0, via Wikimedia Commons

Academic work on meal delivery agrees with the small program. A 2018 study of meal delivery by Reyes and colleagues at Georgia Tech, built on real order data from Grubhub, found that allowing bundles of orders and committing riders sensibly were the key to completing as many deliveries as possible; with bundling switched off, the share of orders left undelivered almost tripled, from 0.25% to 0.69%. It also suggested that what drives how much batching happens isn't how many orders there are but the balance between orders and scheduled rider-hours: the algorithm used bundles to make up for a shortage of riders, and bundled less when riders were plentiful. That's probably one reason batching pays most at dinner peaks.

Decision

How much should batching be allowed to slow a delivery down?

Never batch
One rider, one order.
  • Fastest delivery for every customer
  • Simplest to promise
  • Needs the most riders exactly when riders are scarce
  • Lower earnings per hour for riders
chosen
Batch under a delay limit that moves with supply
Batch when no customer's delivery grows by more than a few minutes; allow more when riders are short.
  • Saves many riders for a small delay
  • Responds to the night's conditions
  • Some customers wait a few minutes longer
  • The promise must account for possible batching
Batch aggressively
Pack as many orders per rider as routes allow.
  • Fewest riders
  • Long, uneven delays
  • Late, cold food for the last stop

DoorDash's 2021 post describes its optimiser accepting a slightly later delivery to batch orders "without violating our on-time promise to our customers", and one of its authors, answering a reader, put the rule plainly: a bonus for the efficiency a batch adds, a penalty for delaying delivery, and no batch when the penalty wins. The scoring also adds a penalty for uncertainty, because a batched route stacks two kitchens' and two buildings' worth of unpredictability. Swiggy's 2019 post adds a subtle point about the promise: don't add a fixed batching buffer to every order's ETA, only to the orders that are likely to be batched. Zomato hasn't published how it batches.

6.4Putting dispatch together

Combining the three ideas gives a dispatcher that runs in a loop. Every few seconds, for each area of the city, it takes the orders that will need a rider soon, the riders who are free or about to be, and the predictions of ready times and travel times. It proposes possible offers, including pairs of orders, scores them, solves for the best set, and sends out only the offers whose moment has come; the rest are planned and recomputed next round.

Dispatch: when, who, and with which other order
offeracceptOrder eventsaccepted, readyRider locationslatest per riderPredictionsprep, travel, acceptOffer generationcandidates + pairsOptimiserper area, every few sSunil's appofferOrder service
Step 1. Noor's accepts Aditi's order with an 18-minute estimate. A minute later it accepts the Sector 57 order. Both arrive as events.
1 / 5

DoorDash has described its dispatcher in more detail than anyone else. Its 2020 post says its old dispatcher assigned one delivery at a time, as a bipartite matching solved by the Hungarian algorithm (the algorithm chapter 50 mentions), and that this was too slow for large areas and couldn't express routes with two or more deliveries. It reformulated dispatch as a mixed-integer program, an optimisation where some variables must be whole numbers, here yes-or-no "this rider gets this order" choices, and solved it with Gurobi, a commercial solver it found up to ten times faster than the Hungarian algorithm on the matching problem. The optimiser runs on several instances, split by region, multiple times a minute.

Sunil accepts. He's on his way to Noor's, and Aditi's app now shows a scooter on a map.

07The dot on Aditi's map

7.1From polling to a held-open connection

Aditi now watches a small scooter icon crawl towards Noor's, wait, and then head for Sector 56. That icon is Sunil's phone reporting its position, the server passing it on, and Aditi's app drawing it. Version 1 did the passing-on by polling: Aditi's app asked every few seconds "where is my rider?", and most of the time the answer was the same as before.

Zomato described its own move away from this in a June 2026 post. It began with periodic HTTP polling, and found two problems. Every update is a tiny payload, but each HTTP request carries its own connection setup and headers, which often outweighed the payload. And tracking felt jerky, with rider markers that jumped instead of moving. The fix was MQTT, a publish/subscribe protocol designed for devices on unreliable networks. A device keeps one connection open to a server called a broker. It can publish small messages to a named topic, and subscribe to topics to receive whatever others publish there. MQTT's message headers can be as small as two bytes, and the connection is kept alive by a cheap ping at an agreed interval, so a dead link is noticed without the app constantly asking.

Sunil's position, from his phone to Aditi's map
positionpublishpushSunil's phoneGPS, adaptive rateLocation ingestvalidate, snap to roadLatest positionsone per rider, in memoryMQTT brokertopic per orderDispatch + ETAAditi's appsubscribed
Step 1. Sunil's phone sends its position over its own held-open connection, more often while he's moving, less while he's standing at Noor's counter.
1 / 4

A tweet goes to millions of followers (chapter 65) and a cricket ball to millions of viewers (chapter 69), but Sunil's position matters to one customer, perhaps two when he's carrying a batch, and occasionally a support agent. So a topic per order, with one or two subscribers, is a natural fit. What's hard is keeping a few hundred thousand of these thin connections alive at once on phones that keep losing signal.

7.2Phones, batteries and patchy networks

Sunil's phone is the most demanding device in the system. Zomato's post notes that delivery partner apps often run for several hours at a stretch, while networks drop or switch between Wi-Fi and mobile data, phone makers' battery optimisations get in the way, and Android limits what apps may do in the background. Its library, Pulse, which it has open-sourced, sits on the Eclipse Paho MQTT client and adds what those conditions demand: it queues connect, subscribe and publish commands so they run in order after a reconnect, reconnects automatically and restores subscriptions, retries failed operations with exponential or sequential retry policies, and checks the connection's health with Android's scheduled-work APIs (WorkManager or AlarmManager). Zomato says it runs in its delivery partner, customer and Hyperpure apps.

Inside Sunil's app: staying connected on a bad network
ON THE PHONERider appGPS, offers, tapsHealth checkWorkManagerPulsequeue, retriesOffline bufferfixes while offlinePaho clientMQTTMQTT brokerIngestsnap to road
Step 1. The app hands every connect, subscribe and publish to Pulse, which queues them so they run in order, on top of the Paho MQTT client.
1 / 4

How often should Sunil's phone report? Zomato hasn't published its rate, so here is how we'd reason about it. Uber's drivers report roughly every four seconds (chapter 50). A scooter at 20 km/h covers roughly 22 metres in four seconds, which is fine for a smooth map. But GPS is the most power-hungry thing a phone does besides the screen, and Sunil's shift is long. So the rate should depend on what Sunil is doing: every few seconds while riding with an order, much less often while he stands at Noor's counter or waits for the next offer, and a burst of fixes as he approaches Aditi's gate, where the "arriving" notification depends on it. When the network drops, the phone keeps recording fixes and sends a compressed batch when it reconnects; the customer's map can skip the gap, but the record of the route, used for pay and disputes, needs every point.

Then there's the question of how wrong a position can be. In a dense city, GPS can be off by tens of metres between tall buildings, and a raw trace jumps across rows of houses. The ingest service snaps each point to the road network and discards points that would need an impossible speed.

Decision

How should the customer's app get the rider's position?

Polling over HTTP
The app asks every few seconds.
  • Simplest to build and cache
  • Works through any proxy
  • Connection overhead per tiny update
  • Jerky markers; mostly 'nothing new'
WebSocket
A held-open, two-way connection; the server pushes.
  • Instant updates
  • Works from browsers
  • You build your own topics, acknowledgements and offline delivery
chosen
MQTT
Publish/subscribe over a held-open connection, with a broker.
  • Tiny headers, built-in keep-alive
  • Delivery guarantees and topics built in
  • Made for flaky networks
  • A broker fleet to run at very high connection counts
  • Needs a robust client on every phone

Zomato chose MQTT for its Android apps and described it in June 2026. MQTT's delivery guarantees come in three levels: at most once (fire and forget), at least once (resend until acknowledged), and exactly once (a four-step handshake), so each kind of message can pick its cost. Zomato's post doesn't say which level it uses for what, but the choice follows from the message. A position update is fine at most once, because a newer one is a few seconds away; an order offer to a rider wants at least once, and the order service's idempotency from section 3 handles the duplicates. Uber made the same polling-to-push move for the same reasons, as chapter 50 described.

Sunil reaches Sector 56 at 10:17 and hands over the biryanis. Most orders end like this. When an order doesn't end like this, the money gets complicated.

08Payment, cancellations and who pays for mistakes

8.1Pay first, then place

Aditi paid by UPI before the order existed for the kitchen, and that order of events is deliberate. If Noor's started cooking first and the payment then failed, someone would be left with four biryanis and no money. So the order starts life in a state before PLACED, say PAYMENT_PENDING, invisible to the restaurant, and moves to PLACED only when the payment is confirmed.

Our UPI case study (chapter 70) explains why "confirmed" isn't always a quick yes or no: a UPI payment can come back as pending when a bank doesn't answer in time, and its real outcome may be known only after status checks, minutes later. Nobody wants to hold a hungry customer for minutes, so there are two reasonable designs: wait a short time for the status to resolve and, if it doesn't, cancel the order and let UPI's own reversal return the money if it was taken; or accept a pending payment for trusted customers and take the risk. Zomato hasn't published which it does, and it probably does both for different customers. Either way, the payment is a separate step with its own idempotency key, so that Aditi's double tap from section 3.2 can't charge her twice.

8.2When an order is cancelled

Orders get cancelled for many reasons, and the right response depends on who caused it and how far the order had got. Zomato's terms of service, last updated in February 2026, set out the principle. If Aditi cancels after placing the order, she's liable for a charge equal to the parts of the order value that are attributable to the restaurant and the delivery partner, because the kitchen may have cooked and the rider may have ridden. If the cancellation is caused by Zomato, the restaurant or the delivery partner, she isn't charged. Zomato doesn't publish a table of who gets what at each stage, so the one below is our design, applying that principle to the order's state at the moment of cancelling:

Cancelled whenCaused byAditiNoor'sSunil
Before the restaurant acceptsAditiFull refundNothing cooked, nothing owedNot assigned
After accepting, before pickupAditiCharged for food preparedPaid for the foodPaid for the trip if he was on his way
Restaurant rejects or runs outNoor'sFull refundNot paid; may be penalisedPaid for the trip if he was on his way
No rider found in timeZomatoFull refundPaid if the food was madeNot assigned
Rider has an accident with the foodDelivery sideFull refund or a fresh orderPaidNot penalised for safety

Each row is a small saga, the pattern from chapter 31: a sequence of steps already committed in different places (payment taken, kitchen notified, rider dispatched), and on failure, a matching sequence of compensating steps that undo or pay for each one. The refund goes back through the payment provider, keyed by the order's ID so that a retried refund can't pay Aditi twice; the restaurant's and rider's payouts are adjusted in their ledgers, which are settled in batches later. Driven by the order's workflow from section 3.3, a crash halfway through resumes where it stopped.

A cancellation as a saga: Noor's runs out of mutton
Noor's tabletrejects the orderOrder serviceCANCELLEDSagaa step per partyPaymentsrefund by order IDNoor's ledgerno payoutSunil's ledgertrip payAditifull refundPayoutssettled later
Step 1. Noor's rejects the order. The order service moves it to CANCELLED and records who caused it, because the cause decides who pays.
1 / 4

There's a neat addition to the cancelled-in-transit row. In November 2024 Zomato wrote that roughly 4 lakh (400,000) perfectly good orders already on their way to customers were being cancelled on the platform each month, and launched Food Rescue: when an order is cancelled while a rider is carrying it, it pops up, at a lower price, for customers within 3 km of that rider. In state-machine terms, a cancelled order in a rider's bag gets one more allowed transition: back onto the market, with a new customer and a new destination, while the original customer's cancellation and refund rules still apply.

So far every piece works on a busy night. Section 9 is about the one night a year that pushes all of them at once.

09New Year's Eve

9.1Getting ready for a spike you can see coming

New Year's Eve is unusual among peaks: everyone knows when it's coming. Zomato's engineering team wrote up its preparation for 31 December 2023 in February 2024, and it reads like a checklist for any known peak:

  • Find the limits before the night does. The team ran about 14,000 minutes of load tests across more than 100 services, reaching a peak test throughput of 50 million requests a minute across a platform of more than 300 microservices.
  • Break things on purpose. It injected very slow responses from upstream services through its internal service mesh, built with Kuma, flipped kill switches to check the degraded modes worked, and injected cache misses, to see what happens when a dependency gets slow or a cache goes cold.
  • Build switches to turn things off. Kill switches and degraded modes let operators disable non-essential features under pressure, so the order path keeps its capacity.
  • Scale up ahead of the peak. The team scaled up before the night, not during it: about 11,000 EC2 virtual machines running about 40,000 containers (ECS tasks) on AWS, over a mix of DynamoDB, Aurora and MongoDB databases, with Kafka carrying over 450 million messages a minute.

That last item is about timing. Autoscaling that reacts to rising load starts adding machines after the load has arrived, and new machines take minutes to become useful. For a peak at a known hour, it's better to scale on the calendar. DoorDash's 2023 post on failure mitigation, which introduced its use of Aperture, an open-source load-management tool, recommends exactly this: predictive autoscaling on a schedule (it names KEDA's cron scaler) instead of reactive scaling. Zomato's post adds that 60% of that day's orders were delivered within 30 minutes.

9.2When demand still outruns supply

More servers fix the software's limits. They don't fix the real constraint, which is the number of riders and the number of kitchens' hands. At 9:40 on New Year's Eve, demand for riders can outrun supply however much compute is running, and pretending otherwise produces 90-minute deliveries, cold food and cancellations. So the system needs levers on the physical side too. Here they are, gentlest first:

  1. Batch more. Section 6.3's delay limit can loosen when riders are short, which is what the Reyes study's finding about the order-to-rider balance suggests.
  2. Shrink delivery areas. Section 4.2's per-minute serviceability shrinks each restaurant's radius as rider supply drops, so fewer long trips are promised. Ten minutes later, Aditi probably wouldn't have seen Noor's at all.
  3. Throttle kitchens. A restaurant whose open orders pass a limit stops taking new ones for a few minutes, so that the orders it has accepted come out on time. That limit is enforced in the order service at checkout, not just in the list, so a rush can't overshoot it.
  4. Raise the price of delivery. Zomato's terms of service allow a delivery surge charge based on distance, time, demand, traffic, weather and seasonal peaks, and a surge fee on Zomato Gold orders during high demand or heavy rain, which the terms say is passed to delivery partners. A higher fee turns some marginal orders away and, if it reaches riders, draws more of them onto the road, the same logic as Uber's surge in chapter 50.
  5. Show the truth. The "high demand" banner Aditi scrolled past, and a longer ETA, are part of the same machinery: honest promises keep the cancellations down.

Swiggy wrote in 2021 that unless riders are moved between zones ahead of demand, typically 15 minutes ahead, its stress system ends up gradually taking fewer new orders to prevent delays and cancellations. That's the last resort of every food platform: when supply runs out, taking fewer orders beats taking orders it can't deliver.

9.3Protecting the order path

Software has one more lever, for the moment when a service is overloaded anyway. Load shedding means a service refuses some requests on purpose, quickly, when it has more in flight than it can finish, instead of accepting everything and getting slower for everyone. DoorDash's Aperture post describes shedding at a service's entrance once concurrent requests exceed a limit that adapts to the latency the service sees, to keep the useful throughput up, and names one of the failures overload causes, a death spiral: some nodes fail, traffic piles onto the healthy ones, and they fail too.

What to shed is the design decision, and Zomato's three buckets from section 1.2 answer it. Browsing is the bulk of the traffic and the least valuable per request, so recommendations, rich images and personalised rows go first; the list falls back to simpler, cached versions. Tracking can drop to slower updates. The order path, placing, accepting, assigning, delivering and refunding, is the last thing to be shed, because an order in progress is money already taken and food already cooking.

Shedding from the bottom up
SHED FIRSTApps81M a minuteAPI gatewaytags each requestBrowsingrecs, rich rowsTrackingrider positionsOrder pathplace to refundCached listsimple fallback
Step 1. Every request is tagged at the edge as browsing, tracking or the order path, so overloaded services know what to refuse first.
1 / 4

10The whole system

10.1Every box, and why it's there

Zomato's order flow, end to end
MUST BE EXACTLY RIGHTAditi's appNoor's tabletSunil's appAPI gatewaypriority, sheddingListing + searchserviceability, menusOrder servicestate machine, workflowMQTT brokerstracking, offersPaymentsUPI, refundsOrder eventsKafkaDispatchtiming, batchingTime serviceprep + travel ETA
Step 1. Aditi opens the app. The listing service reads per-area serviceability, ranks restaurants from a search index sharded by area, and asks the time service for each one's delivery estimate.
1 / 5
ComponentWhat it doesAdded because
Order service and storeOwns each order's state; checks every transition; idempotency keys and versionsDouble taps, lost replies and stale screens corrupted orders (§3)
Outbox and event logPublishes every state change exactly as it was savedSaving and announcing separately could lose an order (§3.3)
Order workflow with timersDrives each order; acts when a step never comesHalf the logic is about things that didn't happen (§3.3)
Serviceability, per area per minuteDelivery areas from rider supply and kitchen stressA list by distance promised what couldn't be delivered (§4)
Menu cache plus availability overlayLong-lived menus, seconds-fresh stock flagsOne expiry can't fit both (§4.3)
Time servicePrep time and travel predictions, asymmetric loss, buffersThe list's time is a promise; late is worse than early (§5)
Dispatch optimiserWhen to send a rider, which one, with which other orderRiders waited at counters; riders are scarce at peaks (§6)
MQTT brokers and location ingestRider positions to one customer; offers to ridersPolling was heavy and jerky on mobile networks (§7)
Payment and refund sagasPay before placing; compensate each party on cancellationCancellations need defined, retryable money flows (§8)
Priority and shedding at the edgeSheds browsing before ordersPeaks overwhelm whatever isn't protected (§9)

10.2From top to bottom

zoomZomatoOrder serviceOrder state machinetransition table + version + idempotency keys
LevelThe choiceData structure or algorithm
SystemSort work into browsing, tracking and ordersCaches for the first, push streams for the second, transactions for the third
OrderOne owner checks every moveState machine as a transition table; version number for optimistic concurrency; set of applied idempotency keys
Order eventsSave and announce togetherOutbox table in the same transaction; idempotent consumers
ServiceabilityPer cell, per minuteMap from cell to covering restaurants; cell-to-cell travel-time table
MenusSplit by rate of changeVersioned document per restaurant; small overlay set of unavailable item IDs
Prep timeLearn the kitchenDish and restaurant embeddings; sequence model over running and recent orders; asymmetric loss
Dispatch timingArrive as food is readysend time = predicted ready − travel; recomputed every few seconds
Matching and batchingOptimise per area, oftenAssignment problem (chapter 50) generalised to a mixed-integer program with batch variables
TrackingPush to one subscriberMQTT topic per order; latest-position map per rider

11What goes wrong, and what it costs

11.1Failures this design has to survive

What happensWhat the user seesWhat the design does
Aditi taps Place Order twiceOne order, one chargeThe idempotency key makes the second tap return the first result
Noor's tablet loses its reply and retries AcceptNothingDuplicate key, ignored
The restaurant never acceptsA refund within minutes, without askingThe order workflow's timer escalates, then cancels and refunds
The kitchen is slower than predictedSunil waits; the ETA growsDispatch timing absorbs some of it; the ETA is re-predicted and the app updated
Sunil's phone loses signalThe scooter freezes on the mapThe app records taps and fixes offline and replays them; the transition table keeps the order consistent
A sudden rush on one kitchen"Currently not accepting orders"Kitchen throttling at checkout
New Year's EveA banner, a surge fee, longer ETAs, fewer frillsPre-scaling, batching, shrinking areas, shedding browsing before orders

11.2The tradeoffs, in one table

DecisionChosenGiven upWhy it was worth it
Who changes an orderOne owner checks every transitionApps writing their own statusThree parties and stale screens disagree
How an order movesA durable workflow per orderFully decoupled servicesTimers for steps that never come
ServiceabilityPer area, every minutePer-request freshnessCheap lookups that survive peaks
The promiseLean late-averse, with buffersThe shortest possible number on the listA late order costs more than a lost one
When to send a riderJust in timeRiders standing by earlyRider time is the scarcest thing at peaks
BatchingUnder a delay limit that moves with supplyA few minutes for some customersMany riders saved when they're short
TrackingPush over MQTTThe simplicity of pollingSmooth maps, less overhead, flaky networks handled
OverloadShed browsing firstRich pages on the busiest nightOrders in progress are protected

12Summary

  1. A food order has three parties, each on their own device and network, so its state must have one owner that checks every move.
  2. The order is a state machine, with transitions owned by the customer, the restaurant, the rider or the system, and refused if they don't fit the current state.
  3. Idempotency keys, versions and the transition table turn double taps, lost replies, stale screens and offline replays into harmless refusals.
  4. Save the state and the event together with an outbox, and drive each order with a durable workflow whose timers handle the steps that never come.
  5. Serviceability is about riders and kitchens as well as distance, and is computed per area every minute so a request is a lookup.
  6. Cache menus for hours and availability for seconds, as separate pieces.
  7. An ETA is a sum dominated by the kitchen, predicted from dishes, restaurants and their running orders, and trained to treat late as worse than early.
  8. Send the rider so they arrive as the food is ready, not when the order is placed; the better the prep-time prediction, the less anyone waits.
  9. Batching saves riders at the cost of minutes; the first few minutes of allowed detour buy the most, and the limit should move with rider supply.
  10. Live tracking is a push to one subscriber, over a protocol built for patchy mobile networks, with the phone adapting its reporting rate to save battery.
  11. On the busiest night, protect the order path: scale ahead of time, batch more, shrink delivery areas, throttle kitchens, and shed browsing before orders.

13Build this

A one-evening food delivery simulator.

  • Generate a city of 50 restaurants and 2,000 customers on a 10 km × 10 km grid. Give each restaurant a prep time that grows with its number of open orders, and generate orders at a rate that rises to a peak at 9 pm and falls after.
  • Implement the order state machine from section 3 with idempotency keys, and drive it from simulated restaurant, rider and customer events, including random duplicates and out-of-order delivery. Check that no order ever ends in an impossible state.
  • Add 150 riders. Compare dispatching at acceptance with dispatching just in time, using a prep-time prediction with adjustable noise. Measure rider idle time, food wait time and delivery time.
  • Add batching with a delay limit, and plot riders needed against average and worst delivery time as the limit grows. Then cut the riders to 100 and see where the best limit moves.

14Interview questions

beginnerWhy can't a food delivery app just assign the nearest rider as soon as the order is placed?›

Because the food isn't ready yet. If the kitchen needs 18 minutes and the nearest rider is 6 minutes away, the rider waits 12 minutes at the counter, doing nothing on the busiest nights. Platforms predict when the food will be ready and send a rider so they arrive at that moment. Waiting to decide also leaves room to pick a better rider who becomes free in the meantime, or to pair the order with another one.

beginnerWhat goes into the delivery time shown on a restaurant listing?›

The time to find a rider and get them to the restaurant, the kitchen's preparation time, whichever of those two is later, then the ride to the customer and the handover. Preparation time is usually the largest and least certain piece. The total is trained to lean towards being slightly long, because a late order costs more than a customer who chose another restaurant, and a buffer is added for demand, traffic, weather and rider supply.

intermediateA customer double-taps Place Order on a slow network. How do you make sure only one order and one payment happen?›

The app generates an idempotency key once for the checkout and sends it with every retry. The order service records keys it has applied and returns the original result for a repeat instead of creating a second order. Each payment call carries its own key derived from the order, so a retried charge or refund can't happen twice. That mechanism also protects against restaurant tablets and rider apps retrying after a lost reply.

intermediateHow would you decide which restaurants to show a customer on a busy night?›

Distance alone overpromises. A restaurant is serviceable if a delivery from it can arrive in reasonable time, which depends on road travel time, free riders near the restaurant and how loaded its kitchen is. Compute rider supply and kitchen stress per area every minute in the background, and from them each restaurant's current delivery area; at request time, look up which restaurants cover the customer's cell. Enforce kitchen limits again at checkout, because a sudden rush can overshoot a minute-old list.

deepExplain the trade-off in batching orders, and how you'd set the limit.›

One rider carrying two orders saves a rider at the cost of a detour for at least one customer. Allowing a few minutes of extra delivery time pairs up the orders that are naturally close, saving many riders; beyond that, each rider saved costs more customer time, concentrated on the last stop. Set a cap on extra delay per customer, score batches with an efficiency bonus against a delay and uncertainty penalty, and loosen the cap when riders are short relative to orders, which research on meal delivery found to be the real driver of when batching pays.

deepThe order service saves the new state and then publishes an event. What can go wrong, and how do you fix it?›

If it crashes between the two, the state and the rest of the system disagree: an order accepted with no rider ever planned, or a rider planned for an order that was never accepted. Write the event into an outbox table in the same database transaction as the state change, and have a separate relay publish outbox rows to the event log, retrying until it succeeds. The relay may publish twice, so every consumer must be idempotent, using the event ID to ignore repeats.

15Go deeper

check yourself
An order is in FOOD_READY. The rider's app sends DELIVERED before PICKED_UP. What should the order service do?›

Refuse DELIVERED, because it isn't an allowed move from FOOD_READY, without changing anything. The app keeps the event and retries after PICKED_UP has been applied, and the recorded tap times keep the delivery time correct.

Prep time is predicted at 20 minutes and the rider is 8 minutes away. When should the rider be sent, and what happens if the prediction is 5 minutes short?›

At minute 12, to arrive at minute 20. If the food is ready at 25, the rider waits 5 minutes at the counter. Leaving earlier would make that worse; leaving later risks the food waiting instead.

Why is a rider's position delivered at most once, but an order offer at least once?›

A newer position replaces a lost one within seconds, so resending is wasted effort. A lost offer means an order with no rider, so it's resent until acknowledged, and the order service ignores duplicates by key.

'A Tale of Scale: Behind the Scenes at Zomato Tech for NYE 2023' (Zomato blog, Feb 2024)

Peak orders and requests per minute, load tests, chaos testing through the service mesh, kill switches, and the size of the fleet, caches, databases and Kafka on the night.

Zomato on prep time (2020, 2022) and delivery ETA (2022)

Dish and restaurant embeddings, sequence models over running orders, the Food Order Ready button, quantile loss against underestimates, and the move to tree-based ETA models served in milliseconds.

'Pulse: Building resilient MQTT infrastructure for Android at Zomato' (Eternal blog, Jun 2026)

Why Zomato moved tracking from polling to MQTT, and what a client needs to survive flaky networks and battery savers. The library is open source.

'Using ML and Optimization to Solve DoorDash's Dispatch Problem' (DoorDash, 2021)

DeepRed: offer generation, ML predictions of ready time, travel and acceptance, and a mixed-integer optimiser that batches and delays dispatch. The 2020 post on moving from the Hungarian algorithm to Gurobi pairs with it.

'The Swiggy Delivery Challenge', parts one and two (Swiggy Bytes, 2018–2019)

Serviceability, the delivery-time equation, just-in-time and next-order assignment, and why batching is both valuable and NP-hard.

Reyes et al., 'The Meal Delivery Routing Problem' (2018)

A rolling-horizon dispatch algorithm on Grubhub data, and what bundling, commitment and the order-to-rider balance do to delivery times.

Eternal annual report 2025–26 and shareholder letters

Orders, customers, restaurants and delivery partners by year and quarter.

Designing Uber

Geo-indexing with H3, finding nearby drivers, and the assignment problem that food dispatch builds on. Chapter 50.

Designing UPI

How Aditi's payment moves between banks, and what "pending" means. Chapter 70.

Designing Stripe

Idempotency keys and ledgers for payments and refunds. Chapter 52.

Distributed Transactions

Sagas and compensating steps, the shape of every cancellation. Chapter 31.

Reliability

Timeouts, retries, load shedding and degraded modes. Chapter 40.

Kafka & Logs

The event log under the order flow. Chapter 23.