KnowSys

Designing JioHotstar: Live Cricket at Scale

It's the IPL final, a wicket falls, Dhoni walks out to bat, and millions of people open the app in the same few minutes. We'll design the system that gets every one of them the same ball, a few seconds behind the stadium: live ingest, segment length and latency, request collapsing at the CDN, multi-CDN steering across India's mobile networks, scaling ahead of the crowd, panic mode, server-side ads and the scoreboard.

⏱ 50 min read◆ IntermediateAssumes: the YouTube case study (chapter 51), chapter 25 (caching), chapter 42 (queueing and capacity) helps
Start reading

It's the IPL final. Ritu is at home in Kanpur, half watching on her phone over 4G while she cooks. Chennai are chasing, the required rate is climbing, and in the sixteenth over a wicket falls. Up on the screen, the camera finds the dressing room stairs, and Dhoni walks out to bat. Her phone buzzes with a notification from the app, and so do a few million other phones across India. Within a couple of minutes, millions of people who weren't watching are.

A full cricket stadium at dusk with the floodlights on, seen from the stands
The Narendra Modi Stadium in Ahmedabad, full for the 2023 World Cup final. It holds more than a hundred thousand people. That night 59 million more watched on one streaming app, and they all wanted the same ball at the same moment.Photo: Ishi6181, CC0, via Wikimedia Commons

From Ritu's side it's one tap: the app opens, the stream starts, an ad plays at the end of the over, the score in the corner ticks over. Behind that tap, a lot is less simple than it looks. Ritu wants to see a ball that was bowled a few seconds ago and didn't exist before that, so no cache anywhere can have it ready in advance. Everyone else wants exactly the same few seconds of video at exactly the same moment. An audience like this doesn't grow smoothly; it jumps by millions when something happens on the field. And the neighbour's TV, a wall away, cheers a boundary before Ritu's picture shows the ball leaving the bat.

Our YouTube case study (chapter 51) already built the machinery for video on demand: renditions, segments, manifests, CDN caches, the player's bitrate choice. This chapter starts from that design and asks what breaks when the video is live and the audience is a whole country. So the question we'll keep coming back to is this: when Dhoni walks out and millions more people tap Play in the same few minutes, how does every one of them get the same ball, a few seconds behind the stadium, without anything falling over? We'll follow the video from the stadium to Ritu's phone, then the crowd from the notification to the play button, and finish with the ads, the score and the checks that decide who may watch.

01What we're building, and how big

1.1What it has to do

The live-cricket part of JioHotstar comes down to a short list:

  1. Start playing quickly when someone taps the match, wherever they are and whatever network they're on.
  2. Stay close to live, so the picture isn't far behind the stadium (and the neighbour's TV).
  3. Keep playing through the whole match, adapting to the viewer's connection.
  4. Show ads in the breaks between overs, and count that they were shown.
  5. Show the score and the match's key moments alongside the video.
  6. Let in only the people allowed to watch: the right subscription, the right country, a licensed device.
  7. Carry many versions of the match: commentary in several languages, extra camera angles, and quality levels from a phone on a weak signal up to a 4K TV.

And the qualities it needs while doing that:

  • The video must play. Everything else on the screen can fail before the video does. Hotstar's engineering posts come back to this rule again and again: nobody misses the extras if the match itself won't play.
  • Surge-proof: the audience can grow by millions in a couple of minutes, at moments nobody can schedule.
  • Live: a delay of tens of seconds is noticeable, and a delay of a minute means the score notification arrives before the ball.

1.2How big is it?

India's cricket streams have broken the world record for simultaneous viewers several times, and each record was set by a single service in a single country. What matters is peak concurrency: the most people watching at the same instant.

DateMatchServicePeak concurrent viewers
Jun 2017Champions Trophy finalHotstar4.7 million
May 2019IPL finalHotstar18.6 million
10 Jul 2019World Cup semi-final, India v New ZealandHotstar25.3 million
29 May 2023IPL final, Chennai v GujaratJioCinema32.1 million
19 Nov 2023World Cup final, India v AustraliaDisney+ Hotstar59 million
9 Mar 2025Champions Trophy final, India v New ZealandJioHotstar61.2 million
5 Mar 2026T20 World Cup semi-final, India v EnglandJioHotstar65.2 million
8 Mar 2026T20 World Cup final, India v New ZealandJioHotstar72.5 million

You've probably noticed that two of those names have since merged. Reliance's Viacom18, which ran JioCinema, and Disney's Star India combined into a joint venture called JioStar in November 2024, and on 14 February 2025 JioCinema and Disney+ Hotstar became one app, JioHotstar. Most of the engineering history in this chapter comes from Hotstar's team, whose blog has kept running under the new name.

Two more facts shape the design. In March 2026 JioHotstar's engineers wrote that a live event needs upwards of 60–80 terabits per second of bandwidth for streaming. And the viewers are on phones: JioStar's CTO described a region "where 85% of the viewership is on mobile" in an April 2026 talk, and at the end of 2025 India's telecom regulator counted about 983 million wireless internet subscribers against 45 million wired ones.

Your turn: design it before reading on

Suppose 50 million people are watching, on phones, at an average of about 1.5 megabits per second each. How much bandwidth is that? And if each player fetches one 4-second segment and re-reads the playlist every 4 seconds, how many requests a second hit the delivery network?

02Version 1: the YouTube design, pointed at a live match

2.1Reuse what works for video on demand

The obvious first design is the one from chapter 51 with a camera at the front. An encoder takes the match feed and produces a ladder of renditions, the same picture at several resolutions and bitrates (chapter 51, §4.2). A packager cuts each rendition into short segments of a few seconds and lists them in a playlist (HLS's name for the manifest), and both go to origin storage. A CDN, a network of cache servers close to viewers, sits in front of the origin. Ritu's player reads the playlist, fetches segments, and picks a rendition for each one with its adaptive-bitrate algorithm (chapter 51, §8).

Version 1: the on-demand design with a live feed at the front
Match feedfrom the stadiumEncoderladder of renditionsPackagersegments + playlistOriginnewest segmentsCDN edgecachesRitu's phone4G, Kanpur
Step 1. The production truck at the stadium sends out the match. The encoder compresses it into several renditions, from a small phone picture to a sharp TV one.
1 / 4

HLS has a way to say "this playlist is still growing". Here is a live playlist from the HLS specification, RFC 8216:

C++
#EXTM3U
#EXT-X-VERSION:3
#EXT-X-TARGETDURATION:8
#EXT-X-MEDIA-SEQUENCE:2680
 
#EXTINF:7.975,
https://priv.example.com/fileSequence2680.ts
#EXTINF:7.941,
https://priv.example.com/fileSequence2681.ts
#EXTINF:7.975,
https://priv.example.com/fileSequence2682.ts

Compare it with the on-demand playlist in chapter 51 and one tag is missing: EXT-X-ENDLIST. Without it, the player knows more segments are coming and must reload the playlist to find them. EXT-X-MEDIA-SEQUENCE:2680 says the first segment listed is number 2,680; the packager keeps only the last few segments in the list and drops older ones off the top, so the playlist is a window that slides along with the match.

2.2Where it breaks

This design works for a small live stream. For the IPL final, it breaks in four places, and each one is something Ritu would notice:

  • The ingest is a single thread. One feed, one encoder, one packager. If any of them hiccups, every viewer's picture freezes at the same instant, and there's no older copy to fall back on.
  • It's too far behind. With 8-second segments, Ritu's picture trails the stadium by half a minute or more. Her neighbour cheers, her phone buzzes with "SIX!", and then she sees the ball.
  • The CDN gets hammered in a way it wasn't built for. Millions of players ask for segment 2,683 within the same second or two of it appearing, and the caches don't have it yet.
  • The crowd arrives faster than servers can be added. When Dhoni walks out, the video bytes come from CDNs that were sized in advance, but the app's own servers, which log people in, build the home page and hand out the stream, see millions of new arrivals in a couple of minutes.

We'll take these in order, starting where the video starts: in the stadium.

03From the stadium to the cloud

3.1Cameras, a truck and a contribution feed

The pictures start with the broadcast cameras around the ground. Their signals run by cable to an outside broadcast van (OB van), a production truck parked at the stadium, where a director picks which camera is live, adds replays and graphics, and mixes in the commentators' audio. What comes out is one finished programme, the match as you'd see it on TV.

A broadcast television camera with a long zoom lens at a football stadium
Where the video begins: a broadcast camera with a long lens, here at a football match at Wembley. Each camera feeds the production truck with an uncompressed or nearly uncompressed picture; everything after this point is about squeezing that picture and getting it to people.Photo: Richard Yeowat, CC BY-SA 3.0, via Wikimedia Commons
The inside of a broadcast production truck, with rows of monitors above a vision mixing desk
The production gallery inside an outside broadcast truck at Wimbledon. The wall of monitors shows every camera; the director cuts between them, and the output is the single programme that becomes the contribution feed.Photo: Richard Yeowart, CC BY-SA 3.0, via Wikimedia Commons

That programme has to leave the stadium for wherever the encoders are. What carries it out of the stadium is called the contribution feed: a high-quality, lightly compressed signal sent over dedicated fibre or a satellite uplink to a broadcast centre. It's called contribution to tell it apart from distribution, the heavily compressed copies sent to viewers. A contribution feed carries far more bits than any viewer will receive, because every later step loses some quality and the encoder needs a clean picture to start from.

Two satellite uplink trucks with dish antennas parked at a sports event
Satellite uplink trucks at the 2005 World Athletics Championships in Helsinki. A dish like these is one way a contribution feed leaves a venue; fibre is the other, and an event can use both so that losing one path doesn't take the match off air.Photo: Timo Newton-Syms, CC BY-SA 2.0, via Wikimedia Commons

A cricket stream is many programmes at once. In April 2026, JioHotstar's engineers described a single live event carried "in 12 languages, 5 camera angles", in variants for TVs and phones and for formats such as Dolby Vision and 4K, and said they "spin up 100s of encoders for every match". At peak they have run close to 55 live streams at the same time. Each language is a different audio track (and often a different commentary team), so a Tamil viewer and a Hindi viewer get the same pictures with different sound.

3.2Encoding and packaging, a few seconds at a time

From here the pipeline is chapter 51's, run continuously. Each programme gets compressed into a ladder of renditions, and the packager cuts them into segments, every rendition at the same instants so the player can switch between them. What's different is the clock. For a film, the transcoding pipeline can take an hour. For a match, the encoder has to keep up with real time forever, and the packager publishes a new segment for every rendition, language and angle every few seconds, then rewrites each playlist to announce it.

Because the playlist is rewritten so often, it's the most-requested file in the system. Every player re-reads it every few seconds, and RFC 8216 sets the rhythm: after a reload that found something new, the player waits at least one target duration before reloading; if nothing changed, it tries again after half a target duration.

3.3Two of everything

One encoder failing for thirty seconds during a final is the kind of failure every viewer notices at once. So the usual design runs two complete pipelines in parallel: two contribution paths out of the stadium, two sets of encoders, two packagers, each producing the same segments with the same numbers and timestamps.

The match from the stadium to the origin, twice
AT THE STADIUMfibresatelliteCamerasaround the groundOB vanprogramme + commentaryEncoders Aregion 1×100Encoders Bregion 2×100Packager APackager BOriginpicks a good copy
Step 1. The director in the OB van cuts between cameras and mixes in commentary. Each language is its own audio track.
1 / 4

JioHotstar hasn't published how its pipelines are duplicated. Netflix has, for its own live events: its December 2025 post on its Live Origin describes two independent pipelines in separate cloud regions, with separate contribution feeds, encoders and packagers, and an origin that, when asked for a segment, checks the candidates from both pipelines in a fixed order and serves the first valid one. It works because both encoders cut segments at identical moments, so segment 4,817 from either pipeline covers exactly the same two seconds of the match.

What JioHotstar has described is how it checks the result. Its 2026 "watchdog" post lists the failures its automated validators catch before viewers do: "manifest 404s" where a regional feed hadn't reached the CDN edge, ad markers that would have produced a black screen, and "media sequence drift", where the segment numbers of different bitrates fall out of step and a player switching between them would buffer forever.

Decision

How many copies of the live pipeline should run?

One pipeline, fast restart
A single chain from stadium to origin; restart a failed piece quickly.
  • Half the cost
  • Nothing to keep in step
  • Any failure freezes every viewer at once
  • A restart takes longer than a viewer's buffer
chosen
Two pipelines, origin picks
Two independent chains in different regions, clock-locked, with the origin choosing a good segment per request.
  • One pipeline can fail without viewers noticing
  • Can be tested by switching one off
  • Double the encoding cost
  • Both must cut identical segments

For an event watched by tens of millions, the cost of a duplicate encoding chain is small next to the cost of the whole country seeing a frozen frame. Keeping the two in lockstep is the harder engineering, and drift between streams is exactly what JioHotstar's watchdog checks for.

So the pipeline produces segments reliably. Next: how long they take to reach Ritu, and why her neighbour hears the six first.

04How far behind the stadium

4.1Where the seconds go

The delay between something happening in front of a camera and it appearing on a viewer's screen is called glass-to-glass latency: from the camera's lens to the screen's glass. Let's add it up for Ritu, from the moment the ball leaves the bat.

  • Production and contribution. The OB van adds a little, and the trip out of the stadium adds more. A satellite in geostationary orbit is about 36,000 km up, so a hop up and down covers over 70,000 km and takes about a quarter of a second at the speed of light.
  • Encoding. Encoders look ahead a little to compress well, so this costs around a second or two.
  • Waiting for the segment to finish. The packager can't publish segment 4,817 until the last frame of it has been encoded. With 6-second segments, the first frame of the ball waits up to 6 seconds.
  • The CDN and the network, usually well under a second.
  • The player's safety margin. This is the big one. RFC 8216 says that for a live playlist the client "SHOULD NOT choose a segment that starts less than three target durations from the end of the Playlist file. Doing so can trigger playback stalls." So a player using 6-second segments starts about 18 seconds behind the newest segment.

Add those up with 6-second segments and Ritu is roughly half a minute behind the stadium. Hotstar wrote in 2018 that it had cut its live latency over the previous year "from being roughly 55s behind broadcast, to approximately 15–20s behind broadcast", by measuring how long each step of its encoding workflow took and tuning the encoder settings. And broadcast TV is itself several seconds behind the stadium, which is why the neighbour still hears it first.

?Why does the player stay three segments behind on purpose?

Because in live video the buffer can't grow. For a film, Ritu's player could download a minute ahead and coast through a dead spot in the network. For a match, the video a minute ahead hasn't happened yet. At best the player can be at the newest segment, so the distance it chooses to stay behind the live edge is the only buffer it has. Starting three segments back means a network stall of up to that long empties the buffer without freezing the picture. Every second of safety is a second of delay, and every second of delay removed is a second less of protection.

4.2Smaller pieces: chunks and partial segments

An obvious fix is shorter segments. Halve the segment length and you halve both the wait for the segment to finish and the three-segment safety margin. But every segment has to start with a keyframe, a picture that can be decoded without any earlier one, and keyframes are expensive, so very short segments cost compression (chapter 51, §7.4). They also multiply requests: 1-second segments mean four times the playlist reloads and segment fetches of 4-second ones.

A way out is to keep segments a reasonable length for compression and caching, but let the player fetch each one in pieces while it's still being made. CMAF (the Common Media Application Format), the fragmented-MP4 packaging that both DASH and HLS players can read, allows a segment to be written as a series of small chunks, each a few hundred milliseconds of video that can be sent as soon as it's encoded. Low-Latency HLS exposes these to the player as partial segments, listed in the playlist with an EXT-X-PART tag. This is the end of a low-latency playlist from the second edition of the HLS specification (May 2026 draft):

C++
#EXTM3U
#EXT-X-TARGETDURATION:4
...
#EXTINF:4.00008,
fileSequence270.mp4
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="filePart271.0.mp4"
#EXT-X-PART:DURATION=2.00004,URI="filePart271.1.mp4"
#EXTINF:4.00008,
fileSequence271.mp4
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="filePart272.0.mp4"
#EXT-X-PART:DURATION=0.50001,URI="filePart272.1.mp4"
#EXTINF:2.50005,
fileSequence272.mp4
#EXT-X-DISCONTINUITY
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="midRoll273.0.mp4"
#EXT-X-PART:DURATION=2.00004,URI="midRoll273.1.mp4"
#EXTINF:4.00008,
midRoll273.mp4
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="midRoll274.0.mp4"
#EXT-X-PRELOAD-HINT:TYPE=PART,URI="midRoll274.1.mp4"

Read it from the bottom. Segment 274 is still being encoded: its first part, midRoll274.0.mp4, exists, and the EXT-X-PRELOAD-HINT tells the player the name of the next part before it exists, so the player can ask for it now and the server answers the moment it's ready. INDEPENDENT=YES marks the parts that start with a keyframe, where a player can begin. (The midRoll names and the EXT-X-DISCONTINUITY tag are an ad break; section 10 comes back to them.)

Two more mechanisms stop the playlist itself from adding delay. With blocking playlist reload, a player asks for "the playlist once part 274.1 exists", and the server holds the request open until it does, instead of the player polling and guessing. And the safety margin shrinks with the parts: the specification's PART-HOLD-BACK, the distance a low-latency player keeps from the live edge, "MUST be at least twice the Part Target Duration" and "SHOULD be at least three times" it.

That specification is blunt about the price: "A shorter Target Duration reduces latency but also reduces available buffer, handicaps adaption and increases delivery overhead, increasing the likelihood of playback stall." So how much stall risk are we buying with each second of latency? Let's measure it.

4.3Trading latency for stalls

Predict before you read on

Ritu's player keeps three chunks of the match in hand. Going from 4-second segments to 2-second segments cuts her delay behind the stadium from about 20 seconds to about 12. What happens to the number of viewers who see at least one freeze during a three-hour match?

This program models a viewer on a mobile network for a three-hour match. Every ten minutes or so the phone hits a dead spot that lasts about two seconds on average, and now and then a longer one of several seconds, such as a hand-off between cell towers. Each player starts the match holding as much video as its safety margin and can never hold more, because the rest of the match hasn't happened. During a dead spot the buffer drains one second per second; if it runs out, the picture freezes. When the network comes back, the player refills at one and a half times real time. A fixed 4 seconds of encoding, packaging and delivery, the same for every row, stands in for the rest of the pipeline.

Segment length against latency and stalls, for 2,000 simulated viewers
python
Python
import random
random.seed(7)
 
FIXED = 4.0       # seconds: capture, encode, package, CDN (same for all)
MATCH = 3 * 3600  # a three-hour match
VIEWERS = 2000
 
def one_viewer(hold):
    # The buffer starts at the hold-back and can't grow past it,
    # because the next piece of the match hasn't been played yet.
    buf, t, stalls, frozen = hold, 0.0, 0, 0.0
    while t < MATCH:
        if random.random() < 1 / 600:           # a dead spot every ~10 min
            gap = random.expovariate(1 / 2)      # ~2 s on average
            if random.random() < 0.05:
                gap += random.uniform(3, 10)     # now and then, a tower hand-off
            if gap > buf:
                stalls += 1
                frozen += gap - buf
            buf = max(0.0, buf - gap)
            t += gap
        buf = min(hold, buf + 0.5)               # refill at 1.5x real time
        t += 1
    return stalls, frozen
 
print("chunk       hold-back  behind stadium  stalls/match  frozen s  viewers who stall")
for name, chunk, hold in [("6 s seg", 6, 18), ("4 s seg", 4, 12),
                          ("2 s seg", 2, 6), ("1 s part", 1, 3),
                          ("0.5 s part", 0.5, 1.5)]:
    runs = [one_viewer(hold) for _ in range(VIEWERS)]
    stalls = sum(s for s, _ in runs) / VIEWERS
    frozen = sum(f for _, f in runs) / VIEWERS
    share = sum(s > 0 for s, _ in runs) / VIEWERS
    print(f"{name:11} {hold:5.1f} s   {FIXED + chunk + hold:8.1f} s   "
          f"{stalls:8.2f}    {frozen:6.1f}     {share:5.0%}")
output
C++
chunk       hold-back  behind stadium  stalls/match  frozen s  viewers who stall
6 s seg      18.0 s       28.0 s       0.01       0.0        1%
4 s seg      12.0 s       20.0 s       0.13       0.3       12%
2 s seg       6.0 s       12.0 s       1.58       4.3       80%
1 s part      3.0 s        8.0 s       4.75      12.8       99%
0.5 s part    1.5 s        6.0 s       8.96      22.4      100%

Read the rows from the top. Six-second segments keep Ritu 28 seconds behind and almost nobody freezes. Four-second segments buy 8 seconds of latency for a freeze that one viewer in eight sees once. Two-second segments buy another 8 seconds, and now most viewers freeze at least once, because a 6-second buffer is shorter than the occasional tower hand-off. Below that, every viewer freezes several times a match. So the curve is lopsided: the first seconds of latency are cheap to remove and the last ones are very expensive.

This model is deliberately crude. Real dead spots are probably rarely total, so a real player would drop to a lower rendition and keep downloading slowly instead of stopping, and a player can also widen its hold-back after a stall. But the shape holds, and it's the shape the HLS specification warns about. It's why "lowest possible latency" is the wrong goal for a mobile-first audience, and why services aim for a delay that's a little behind broadcast TV and stable.

Decision

How far behind live should the player sit?

Long segments (6 s+)
Classic HLS: three long segments of hold-back.
  • Almost no stalls on a bad network
  • Fewest requests per viewer
  • Half a minute behind the stadium
  • Score alerts arrive before the ball
chosen
Short segments (2–4 s)
Shorter segments, three of them held back.
  • 10–20 s behind, close to broadcast TV
  • Nothing special needed in CDNs or players
  • More stalls on mobile networks
  • More requests, slightly worse compression
Low-latency parts
CMAF chunks as LL-HLS partial segments, blocking reload, preload hints.
  • A few seconds behind the stadium
  • Tiny buffer: many stalls on a weak signal
  • Every hop must pass chunks on as they arrive

JioHotstar's segment length isn't published. Large mobile audiences probably land on the marked option: Netflix's 2025 description of its live platform uses 2-second segments, and Hotstar's 2018 target was 15–20 seconds behind broadcast. Low-latency modes make the most sense where the network is good, such as a TV on home broadband, and a service can offer different hold-backs to different devices from the same segments.

So now Ritu's player asks for a small segment every couple of seconds. So does everyone else's, and they all ask for the same one.

05Everyone wants the newest segment

5.1Why live breaks the CDN's assumptions

A CDN for on-demand video works because popular files stay popular. In Kanpur, the first viewer of a film causes a miss, the edge server fetches it and keeps it, and the next ten thousand viewers over the following week get hits. Chapter 51 showed how well that works when popularity is skewed.

Live video breaks this in three ways at once.

  • The file is brand new. Segment 4,817 didn't exist two seconds ago, so every cache in the country starts empty.
  • Everyone wants it in the same second or two. Ten minutes from now, almost nobody will ask for it again. Its life as a popular object lasts roughly as long as it takes to play it.
  • The playlist changes every few seconds. It can only be cached for a moment, or players would never learn about the new segment.

Put the first two together and you get a burst: in the instant segment 4,817 is announced, millions of requests land on edge servers that don't have it. Each edge server that sends every one of those misses to the origin turns one popular file into a flood. This is the same thundering herd, or cache stampede, that chapter 25 met when one hot key expired (§8), except that in live video it happens on purpose, every two to six seconds, for three hours.

Rows of server racks in a data centre
An origin is a few racks like these. It can comfortably serve each new segment a few hundred times; it can't serve it a few hundred thousand times in one second, which is what an unprotected CDN would ask of it.Photo: BalticServers.com, CC BY-SA 3.0, via Wikimedia Commons

Playlists have their own version of the problem, and Hotstar's 2019 post warned about it under the heading "TTL Cloudbursts": with several layers of caches, if you lose track of how long each one keeps each object, "large number of objects expiring around the same time" can send "a request cloudburst at your origin". A playlist cached for 2 seconds at every edge expires at every edge every 2 seconds.

5.2Request collapsing

Fixing it starts at each edge server. When the first request for segment 4,817 misses, the edge starts one fetch from the tier above. Every request for the same segment that arrives while that fetch is in flight doesn't start another; it waits in line, and when the one fetch returns, the edge answers all of them from it. This is called request collapsing (or request coalescing). Fastly's documentation describes it in those terms: on a miss, concurrent requests share one origin fetch, and the requests that arrive meanwhile join a "waiting list".

Collapsing at the edge turns "every miss goes upstream" into "one request per edge server goes upstream". That's a big improvement, but a large CDN has thousands of edge servers, and each one still fetches its own copy. So the second half of the fix is a layer between the edges and the origin: a small number of big caches, often one per region, that every edge fetches from. AWS calls its version Origin Shield, and its documentation describes exactly the live case: requests for content not yet in the shield's cache "are consolidated with other requests for the same object, resulting in as few as one request going to your origin", and it lists "origins that provide just-in-time packaging for live streaming" and "workloads that use multiple content delivery networks" as the uses it's for.

Segment 4,817 is announced: collapsing at the edge and at the shield
Ritu + 40,000 otherssame edgeViewers elsewhereother edgesEdge, Kanpurwaiting listOther edges×50Shieldregional cacheOriginsegment 4,817
Step 1. The playlist announces segment 4,817. Within a second or two, Ritu and tens of thousands of others on the same edge server ask for it. The first request misses.
1 / 5

How much does each layer save? This program sends two million viewers at one region's 50 edge servers, all asking for segment 4,817 at random moments in a two-second window. Fetching a copy from the tier above takes 200 milliseconds. It counts how many requests reach the origin under four policies.

A thundering herd for one live segment, with and without collapsing
python
Python
import random
random.seed(42)
 
VIEWERS = 2_000_000   # watching in one region
EDGES = 50            # edge servers they're spread across
WINDOW = 2.0          # seconds over which players ask for the new segment
FETCH = 0.200         # seconds to fetch one copy from the tier above
 
# every viewer asks one edge for segment 4817 at some moment in the window
arrivals = [[] for _ in range(EDGES)]
for _ in range(VIEWERS):
    arrivals[random.randrange(EDGES)].append(random.uniform(0, WINDOW))
for a in arrivals:
    a.sort()
 
def misses(times):
    """Requests that find the cache empty: they arrive before the
    first fetch (started by the first request) has come back."""
    ready = times[0] + FETCH
    return [t for t in times if t < ready]
 
# 1. edges forward every miss to the origin
no_collapse = sum(len(misses(a)) for a in arrivals)
# 2. each edge sends one fetch; later requests wait for it
edge_collapse = sum(1 for a in arrivals if a)
# 3. those edge fetches go to a shield, which forwards each of its misses
first_fetches = sorted(a[0] for a in arrivals if a)
shield_no_collapse = len(misses(first_fetches))
# 4. the shield collapses them too: the first one fetches, the rest wait
shield_collapse = 1
 
print(f"{VIEWERS:,} viewers ask for segment 4817 within {WINDOW:.0f} s")
print(f"edges forward every miss          : {no_collapse:>7,} origin requests")
print(f"edges collapse                    : {edge_collapse:>7,} origin requests")
print(f"edges collapse, shield forwards   : {shield_no_collapse:>7,} origin requests")
print(f"edges collapse, shield collapses  : {shield_collapse:>7,} origin request")
print(f"requests that waited on a fetch   : {no_collapse:,} "
      f"({no_collapse / VIEWERS:.0%} of all), none longer than {FETCH*1000:.0f} ms")
output
C++
2,000,000 viewers ask for segment 4817 within 2 s
edges forward every miss          : 200,364 origin requests
edges collapse                    :      50 origin requests
edges collapse, shield forwards   :      50 origin requests
edges collapse, shield collapses  :       1 origin request
requests that waited on a fetch   : 200,364 (10% of all), none longer than 200 ms

Without collapsing, 200,364 requests reach the origin for one segment, a tenth of all viewers, because a tenth of the requests arrive during the 200 ms that each edge's first fetch is in flight. That happens again for the next segment two seconds later, and for every rendition and language. Collapsing at the edges cuts it to 50, one per edge. Line three is the trap: a shield that doesn't collapse is just another cache that misses, and all 50 edge fetches arrive inside its own fetch window, so it forwards all 50. Only when the shield collapses too does the origin see one request per segment per rendition. You can see the cost on the last line: a tenth of viewers waited up to 200 ms for their copy, a bit of delay that no viewer would notice next to a multi-second buffer.

5.3The details that bite

Collapsing has sharp edges, and live streaming finds all of them.

  • An uncacheable reply breaks the line. If the origin's response can't be cached (an error, or a header that forbids it), every request on the waiting list was waiting for nothing. Fastly's documentation says the queued requests are then sent to the origin one after another, and warns that "in some cases this can create extreme response times of several minutes", and suggests failing such requests quickly instead.
  • Asking too early. A player that requests segment 4,818 a moment before it exists gets a 404. If the CDN caches that 404 for even a second, everyone behind it gets a 404 too, for a file that now exists. Low-Latency HLS's blocking reload and preload hints exist partly so that a request for something about to exist is held open instead of failed.
  • Tiny lifetimes. HTTP's cache lifetimes are counted in whole seconds, which is coarse for a playlist that changes every two seconds. Netflix's Live Origin post describes adding millisecond-grain caching to its nginx servers for this reason, and, during a surge, telling Open Connect to cache identical requests for five seconds and returning HTTP 503 to low-priority traffic such as rewinding.
Decision

How do you stop the newest segment from flattening the origin?

A bigger origin
Serve every edge miss directly; add origin servers until it copes.
  • Nothing clever in the CDN
  • Needs capacity for hundreds of thousands of identical requests per segment
  • Fails the moment the audience grows past the plan
Collapse at the edge
One upstream fetch per edge per object; the rest wait.
  • Cuts origin load to one request per edge
  • Thousands of edges still means thousands of fetches
chosen
Collapse at edge and shield
A regional tier between edges and origin that collapses again.
  • About one origin request per segment per rendition
  • Lets several CDNs share one shield
  • An extra hop of latency on every miss
  • The shield is now critical and must be redundant

With collapsing at both tiers, the origin's load depends on the number of renditions and languages, not on the number of viewers. So the audience can grow from 25 to 70 million while the origin's load stays flat. Hotstar made the same point in 2019 when it said its scaling wasn't "the CDN (alone)": a CDN is a "shock absorber" only if you "architect your CDN to sit in front of your origin" and know its limits.

So the origin is safe. But one region's edges are no longer enough when the audience needs 60–80 terabits a second, so where does all that bandwidth come from?

06Getting the bytes to 50 million phones

6.1More than one CDN, and caches inside the ISPs

No single CDN has spare capacity on the scale of an Indian cricket final. Hotstar's 2024 infrastructure post says it outright: when the team modelled bandwidth for 50 million concurrent streams, "the demand was far beyond what our CDN providers could handle." So the stream is spread across several CDNs at once, a design called multi-CDN. JioHotstar's April 2026 watchdog post names three whose edges it checks against each other: "Akamai, Cloud front, JIO".

That third name matters. Cheapest and fastest of all are the bytes that never leave the viewer's own internet provider. Chapter 51 described Google Global Cache, servers Google places inside ISPs' networks; Netflix does the same with its Open Connect appliances. For a viewer on Jio's mobile network, a cache inside Jio's network is a few hops away and never touches a link between providers. An independent analysis by the magazine Voice&Data during IPL 2024 found that Jio's users were served almost entirely from JioCinema's own CDN servers inside the Jio network, while some users of other mobile operators were served from outside India. How JioHotstar's caches are placed today isn't published.

Network switches and fibre cabling in a cage at an internet exchange point
Where networks meet: the core of the AMS-IX internet exchange in Amsterdam. Bytes from a CDN that isn't inside the viewer's ISP have to cross a link like these, at an exchange or a private interconnect, and those links can fill up during a final.Photo: Fabienne Serriere, CC BY-SA 3.0, via Wikimedia Commons

6.2The last mile is a cell tower

Most of those bytes end on a mobile network, and a mobile network shares its capacity per cell: everyone connected to the same tower splits that tower's radio capacity. A busy market or a train during a final can probably put hundreds of phones on one cell, all streaming the same match. Each player's bitrate algorithm handles this viewer by viewer, dropping to a smaller rendition when downloads slow down (chapter 51 covers how).

A mobile phone tower standing in farm fields in Punjab
The last hop for most viewers: a mobile tower in Punjab. Every phone near it shares its radio capacity, so a crowd watching the same match on one cell competes for the same bandwidth whatever the CDNs have in reserve.Photo: Sahil Dhiman, CC BY-SA 4.0, via Wikimedia Commons

But the service can also act for everyone at once. Hotstar's 2017 post on surviving traffic spikes described progressive degradation: the key levers, it said, are "the bit-rates at which we offer the content, which must degrade as concurrency grows", because the bandwidth in the infrastructure is a hard limit. Removing the top rung of the ladder for phones when the audience passes some size costs each viewer a little sharpness and frees a large slice of total bandwidth. A 2026 JioHotstar post puts it with an image of a bridge: everyone wants to drive a wide luxury bus, but the lanes are finite, and during a peak surge there isn't room for everyone to drive one at once.

6.3Which CDN should Ritu use?

With several CDNs, something must decide which one each viewer fetches from, and the decision changes during the match. A CDN can be healthy across India and struggling for one mobile operator in one city. JioHotstar described its answer in March 2026, a service it calls the QoS Routing Manager.

zoomJioHotstarDeliveryCDN steeringCohort score

It works in four steps.

  1. Group viewers into cohorts. A cohort is a combination of network and place: "ASN-Country-State-City-UserType". (An ASN, autonomous system number, identifies a network on the internet, such as one mobile operator.) Ritu is in the cohort "Jio, India, Uttar Pradesh, Kanpur, mobile".
  2. Score each CDN for each cohort from the players' own reports, which arrive as heartbeats. Three measurements are each mapped to a score: playback failure rate (how often a stream fails to start), rebuffering, and round-trip time. A cumulative score weighs them in that order, failures most heavily, "to ensure that 'reachability' is the absolute priority".
  3. Smooth the scores with an exponentially weighted moving average, so one bad ten-second window doesn't send a whole city to another CDN. Each new score counts for a fraction α of the average and the old average for the rest, with α about 0.6, which keeps roughly the last five scores significant.
  4. Shift traffic weights, with brakes. Before a CDN reaches its capacity, its share is throttled even if it scores well. During a surge the logic switches to making all CDNs run out of capacity at the same time. A CDN's share can only change by a capped amount each minute, and no working CDN ever drops below a floor, "typically 5%", so its caches stay warm in case it's suddenly needed.

Here's step 3 on a small example. One CDN's score for Ritu's cohort is steady around 90, then a ten-second window comes in at 30 because of a tower hand-off in one neighbourhood:

WindowRaw scoreSmoothed (α = 0.6)
19090.0
28888.8
33053.5
48974.8
59083.9

The dip shows, but it's about 40% shallower than the raw one, and two windows later the score is most of the way back. Combined with the cap on how far a share can move per minute, a blip moves a little traffic for a minute instead of the whole cohort. JioHotstar tested the manager against the old routing during a T20 match, with cohorts split only by ASN and state, and in that A/B test the treatment group's playback failure rate improved by 11%, and rebuffering and round-trip time by 2% each.

Steering Ritu's cohort between CDNs
heartbeatsRitu's phoneJio · KanpurPlayback APIreturns CDN URLsQoS Routing Managerweights per cohortPlayer telemetryheartbeats, OLAPCDN ACache in Jio's networkCDN C
Step 1. Ritu taps the match. The playback API looks up her cohort, Jio in Kanpur on mobile, and the current CDN weights for it.
1 / 5
Decision

How should a live match get to tens of millions of viewers?

Own CDN, filled in advance
Netflix's Open Connect for on-demand: appliances inside ISPs, filled overnight with what members will watch.
  • Almost no traffic crosses provider links at peak
  • Fill happens when networks are quiet
  • Needs to know the content in advance: impossible for a match
  • One network: no fallback if it fills up
Own CDN, pulled live
Netflix's approach for live events: the same appliances fetch new segments as they're made.
  • One system to run
  • Reuses a huge existing footprint
  • All the live risk sits on one network
  • Built and tuned for a different traffic shape
chosen
Several CDNs plus ISP caches, steered
JioHotstar: multiple commercial CDNs and in-network caches, weighted per cohort from player telemetry.
  • More total capacity than any one CDN
  • Routes around a CDN struggling in one city
  • A steering system to build and trust
  • Every CDN must be configured and tested identically

Netflix's Open Connect overview says the program was designed as "a proactive, directed caching solution", and that Netflix can deploy "the majority of our content and software updates proactively during off-peak fill windows" because "we can predict with high accuracy what our members will watch". None of that applies to a ball bowled two seconds ago. When Netflix streamed the Paul–Tyson fight in November 2024, it reported 65 million concurrent streams at the peak, and many viewers reported buffering and dropped streams. For JioHotstar, where a single match can need more bandwidth than any one provider has in India, spreading across many CDNs is a necessity, and steering is what makes many CDNs behave like one.

So the video reaches the phones. But the video was never where the crowd hurt most. When Dhoni walks out, what breaks first is the app's own servers.

07The surge: toss, first ball, wickets

7.1The shape of a match

Watch the number of viewers across one match and it rises fairly smoothly from the toss to the death overs. Watch the number of people arriving each minute and you see something else. JioHotstar's engineers plotted new sessions per minute and called it the onboarding rate: how fast people are joining, as opposed to concurrency, how many are watching. Their December 2025 post lists what the curve shows in every match: "the mini-spike at toss", "the explosion at first ball", "the middle overs lull", a second-innings spike, "the pre-death overs ramp", and "the sudden drop at match end".

Those numbers are steep. At the first ball, they wrote, "6–8 million users joining in the first 2 minutes", the home page's API traffic jumping 300–400% and the watch page's (the page with the player) 400–500%. "Then, strangely, 10 minutes later the same services hum along happily as concurrency climbs toward 50 or 60 million."

A crowd gathered at night around a small television showing a cricket match, one man cheering with his arms raised
A street crowd in India gathers round a TV when the match gets tense. On a streaming app the same moment shows up as a rush of people opening the app at once. Concurrency counts the people watching; the onboarding rate counts the people arriving, and arriving is the expensive part.Photo: Arvind Jain, CC BY-SA 2.0, via Wikimedia Commons

Unplanned moments are worse than the first ball, because nobody knows when they'll come. A wicket brings a new batter, and a famous one brings a crowd. Hotstar wrote about the first time this hurt, during the India–Pakistan opener of the 2017 Champions Trophy: "Virat Kohli walks in to bat — a notification is sent to a large segment and boom! Our login API that was cruising at 10K TPS, shot up to 125 K TPS in a matter of seconds." (TPS is transactions per second.) That's a twelvefold jump. Hotstar had scaled its servers for 7 million viewers, but the database behind the login service hadn't been scaled for the extra traffic all those servers brought on, and it "started to melt down".

Ritu's notification about Dhoni is the same event. Here the app's own push notification creates the surge it then has to absorb.

7.2What the surge hits

Why do the app's servers suffer more than the CDNs? Because a viewer who's already watching costs the backend almost nothing: their player talks to the CDN, plus an occasional heartbeat. A viewer who's arriving hits every service on the way in, as the StepSequence below shows for Ritu.

Ritu taps the notification: one arrival, many backend calls
Ritu's appAPI gatewayBackend servicesCDNauthprofile, homewatch pageplaybackplaylist, segments
Step 1. Is this session still valid? The login service checks her token.
1 / 5

So the services on the arrival path, login, profile, home page, watch page and playback, scale with the onboarding rate. Services used while watching, heartbeats, the scoreboard and ads, scale with concurrency. The December 2025 post maps each service to the signal it scales on: authentication, profile selection and the watch page on onboarding rate plus throughput, ads on concurrency plus throughput, and the home page on a mix of both, for a reason section 8 explains.

A surge like this lasts a few minutes. Can the system add servers fast enough to catch it?

08Scaling before the crowd arrives

8.1Why autoscaling is too slow

The usual answer to rising load is an autoscaler: watch a metric such as CPU use, and when it crosses a threshold, start more servers. It works well when load rises over tens of minutes. In Hotstar's 2019 re:Invent talk, as summarised by attendees, the arrival rate reached more than a million new users per minute, while the autoscaler took about 90 seconds to react and the application about a minute to boot. That talk also listed the cloud's own limits: errors when the provider had no more instances of a type to give, one instance type per scaling group, steps too coarse.

Hotstar had reached the same conclusion the year before. Its 2018 post put it in one rule, "No Auto-Scaling — Give yourself headroom", and an analogy: "Do you end up ordering food as guests start arriving for your party. I hope you don't." The strategy was "to estimate the peak concurrency, then scale up ahead of the event".

This program puts numbers on it. Ten million people are watching when Dhoni walks out. Three million join in each of the next two minutes, a million a minute for three more, then a trickle. Each arrival makes six backend calls as the app opens and the stream starts (an illustration); each viewer already watching makes about one a minute. One API server handles 1,000 requests a second. The autoscaler keeps servers 70% busy and orders more when they get busier, which arrive 150 seconds later. Our pre-scaled fleet sits on a fixed rung, sized beforehand for a forecast of 25 million viewers.

Dhoni walks out: an autoscaled fleet against a pre-scaled one
python
Python
# Dhoni walks out at minute 0 and a push notification goes out.
# How many viewers join in each minute after that:
joining = [3_000_000, 3_000_000, 1_000_000, 1_000_000, 1_000_000] + [200_000] * 5
 
CALLS_PER_JOIN = 6   # API calls as the app opens and the stream starts
PER_SERVER = 1_000   # requests per second one API server handles
DELAY = 150          # seconds: autoscaler notices (~90) + server boots (~60)
LADDER = 1_250       # servers on the rung for a 25M-viewer forecast
 
def demand(viewers, join_per_min):
    # joiners hit the start-up path; everyone watching makes a call a minute
    return join_per_min / 60 * CALLS_PER_JOIN + viewers / 60
 
viewers = 10_000_000
auto = demand(viewers, 0) / PER_SERVER / 0.7   # 70% busy before the spike
orders = []                                    # (ready at second, servers)
lost = {"auto": 0, "ladder": 0}
used = {"auto": 0, "ladder": 0}
 
print("min  joining/min  demand req/s  autoscaled  served  pre-scaled  served")
for minute, j in enumerate(joining):
    for half in range(2):                       # a decision every 30 s
        t = minute * 60 + half * 30
        viewers += j / 2
        need = demand(viewers, j)
        auto += sum(n for at, n in orders if at == t)
        coming = sum(n for at, n in orders if at > t)
        target = need / PER_SERVER / 0.7
        if target > auto + coming:              # order the shortfall
            orders.append((t + DELAY, target - auto - coming))
        for name, servers in (("auto", auto), ("ladder", LADDER)):
            lost[name] += max(0, need - servers * PER_SERVER) * 30
            used[name] += servers / 2
        if half == 0:
            print(f"{minute:3}  {j:11,}  {need:12,.0f}  {auto:8,.0f}   "
                  f"{min(1, auto * PER_SERVER / need):5.0%}  {LADDER:8,}   "
                  f"{min(1, LADDER * PER_SERVER / need):5.0%}")
 
for name in ("auto", "ladder"):
    print(f"{name:6}: {lost[name] / 1e6:5.1f} M requests turned away, "
          f"{used[name]:,.0f} server-minutes")
output
C++
min  joining/min  demand req/s  autoscaled  served  pre-scaled  served
  0    3,000,000       491,667       238     48%     1,250    100%
  1    3,000,000       541,667       238     44%     1,250    100%
  2    1,000,000       375,000       238     63%     1,250    100%
  3    1,000,000       391,667       738    100%     1,250    100%
  4    1,000,000       408,333       810    100%     1,250    100%
  5      200,000       338,333       810    100%     1,250    100%
  6      200,000       341,667       810    100%     1,250    100%
  7      200,000       345,000       810    100%     1,250    100%
  8      200,000       348,333       810    100%     1,250    100%
  9      200,000       351,667       810    100%     1,250    100%
auto  :  39.0 M requests turned away, 6,560 server-minutes
ladder:   0.0 M requests turned away, 12,500 server-minutes

Look at the first three rows of the autoscaled fleet. Demand roughly triples in the first minute, and for two and a half minutes the fleet stays the size it was before Dhoni walked out, serving between 44% and 63% of what's asked. By the time new servers arrive, in minute 3, the worst of the rush is over. Over those minutes it turns away 39 million requests, and each one is a person staring at a spinner, who will probably tap again and make it worse. Meanwhile the pre-scaled fleet serves everything, at the cost of nearly twice the server-minutes. It was sized for a forecast of 25 million and the audience stopped at about 20 million, so once the rush is over, well over half of it sits idle. That's the trade, paid in money to buy minutes.

8.2Ladders

Scaling ahead needs a plan for how big to be at each audience size. Hotstar called it a ladder: a table, worked out from load tests, that says for each level of concurrency how many servers each service needs. In 2018 the team "built pessimistic traffic models" for each of its three pillars, subscriptions, metadata and streaming, "basis which we came up with ladders that controlled server farms depending on the estimated concurrency". During a match, engineers stepped the fleet up a rung whenever the headroom above current concurrency fell below a threshold.

zoomJioHotstarBackendScalingLadder rung

Hotstar's ladders evolved in three steps, each fixing the cost of the last.

  • 2019: automation, in shadow first. Hand-stepped ladders worked, but "required human oversight and was costly". After moving to Kubernetes, the team built its own autoscaling engine "that took into account multiple variables that mattered to our system", and "ran auto-scaling in shadow" (deciding, but not acting) until it trusted it. That year it supported "2x more concurrency in 2019 with 10x less compute overall".
  • 2023: the platform's own limits. Preparing for 50 million viewers at the 2023 World Cup, the team found that the Kubernetes control plane started returning errors when asked to add over 400 nodes at once, so pre-scaling now proceeds in automated steps of 100–300 nodes. A subnet ran out of IP addresses at about 350 nodes because each node reserved 35 addresses in advance. And services with more than 1,000 pods hit a Kubernetes limit: the older Endpoints API lists at most 1,000 addresses per service, and the team's API gateway relied on it, so those services were given bigger pods instead of more of them.
  • 2025: scale on arrivals, not just viewers. As scaling became more automatic, incidents at lower concurrencies showed that concurrency alone missed the first-ball rush, which is why the onboarding rate from section 7 became a scaling signal. An old split between a business-as-usual mode and a live mode gave way to one model: high onboarding rate scales the arrival path, high concurrency scales playback, both scale everything, and both low leaves "an efficient, cost-effective base ladder" that autoscales on live traffic.

One surprise came at the end of a match. Concurrency drops and nobody new arrives, yet "homepage traffic spikes 200–300% within seconds", because everyone who was watching goes back to the home page at once, like a cinema at the interval. They named it the off-boarding rate and scale the home page on onboarding rate plus concurrency times a post-match coefficient. Hotstar has filed a patent on it.

Decision

How should the backend grow for a match?

Reactive autoscaling
Add servers when a metric such as CPU crosses a threshold.
  • Pay only for what's used
  • No forecast needed
  • Minutes too slow for a first-ball rush
  • Depends on the cloud having capacity at that moment
Ladders on forecast concurrency
Scale to a rung for the expected peak before the match; step up on headroom.
  • Capacity is there before the crowd
  • Cloud capacity reserved in advance
  • Pays for idle servers
  • Misses spikes driven by arrivals, not totals
chosen
Ladders on concurrency and onboarding rate
Each service scales on the signal that stresses it, with a base ladder and live metrics on top.
  • Arrival paths ready for the rush, playback paths for the peak
  • One model for quiet days and finals
  • Needs forecasts per service and per match
  • More signals to get right

Notice the progression: Hotstar started by rejecting autoscaling, then rebuilt it on signals that matched its traffic. Even so, the cloud's own limits never went away, so the big steps still happen before the match: Hotstar's 2017 checklist included having the cloud load balancers specially "warmed up" and making sure the right instance types would be available.

Even a good ladder can be wrong. If the audience overshoots the forecast, what gives way?

09Panic mode: protect the video

9.1Jettison what you can

In 2019, ahead of a World Cup game where "breaching the 30M barrier was not un-thinkable", Hotstar asked its team what it would take to serve 50 million. They went through every service: its rated limits, and "what would be the tsunami, that took the service down". Then they agreed an order in which to turn services off: "a sequencing of which service to jettison (turn off), as the concurrency marched upwards, so that customers could keep coming in and the video would keep playing." Hotstar compared it to Apollo 13, where ground control kept only what the crew needed to get home.

A system with these switches built in has a panic mode: a set of pre-planned levers that shed non-essential work under overload. Hotstar's 2019 talk named the first things to go, recommendations and personalisation, to protect video streaming and payments. This Scene plays it out for Ritu's screen as the audience climbs past the forecast.

Panic levers as the audience overshoots
Running normallyDegradedcached, slower, simplerSwitched offvideoadsscorecardfreshhome pagepersonalisedrecommendationssocial feed
Step 1. Before Dhoni walks out, everything on Ritu's screen is live: the video, ads, the scorecard, personalised rows, the social feed under the player.
1 / 5

That order is illustrative; JioHotstar's current sequence isn't published. What Hotstar has published are the individual levers: recommendations and personalisation (2019), refresh rates for "scorecards, Watch-more, live feeds, and key moments" (2024), and bitrate (2017). Hotstar's 2017 post adds a distinction that's easy to miss. For Tier 1 actions, login, subscriptions and watching, the lever is a positive panic: the action appears to succeed and the backend work is deferred, so a viewer whose subscription check can't complete right now is let in and checked later. For Tier 2 actions, such as account settings and logout, it's a negative panic: show a polite failure message and move on.

9.2Clients that back off

Panic mode's other half lives in the app. When a backend struggles, the worst thing tens of millions of clients can do is retry immediately, all at once; every failure then doubles the load that caused it. Hotstar's 2018 post: "Retries will make the problems worse. Clients must be smart about inferring when things don't look right, and add 'jitter' to the requests they make." Its apps understand panic switches from the server, which tell them "to ease off momentarily, either exponential back-off or sometimes a custom back-off depending on the interaction". Exponential back-off means waiting twice as long after each failure; jitter means adding a random amount to each wait, so that a million clients that failed at the same instant don't all retry at the same instant too.

Netflix found the same thing with live events: in its 2025 account, viewers restarting streams after interruptions as short as 30 seconds caused a "10x increase in traffic load", and the fixes included server-guided backoff for devices and prioritised shedding of traffic at its cloud gateway.

So the video is protected, and the ads are on the protected list because they pay for it. How do they get into the stream?

10Ads in the break between overs

10.1Every viewer reaches the break at once

At the end of the over, the broadcast cuts to an ad break. Ritu should see ads chosen for her, or at least for people like her, stitched in so the picture never stutters, and then be back in time for the first ball of the next over.

There are two ways to put an ad into a stream. With client-side ad insertion (CSAI), the stream carries a marker, the player sees it, pauses the match, asks an ad server for ads, plays them with a second player, then switches back. With server-side ad insertion (SSAI), the ads are stitched into the stream itself: the playlist Ritu's player reads lists ad segments where the match segments would be, and the player just keeps playing.

Both methods start from a marker in the broadcast. When the director goes to a break, the production system inserts an SCTE-35 cue, a standard message meaning "an ad break starts here, and lasts this long". In HLS it can appear as a date-range tag; this one is from the HLS specification:

C++
#EXT-X-DATERANGE:ID="splice-6FFFFFF0",\
   START-DATE="2014-03-05T11:15:00Z",PLANNED-DURATION=59.993,\
   SCTE35-OUT=0xFC002F000000000000FF000014056FFFFFF000E081622DCAFF0\
   00052636200000000000A0008029896F50000008700000000

SCTE35-OUT marks the start of a 60-second break and a matching SCTE35-IN its end. In the low-latency playlist from section 4, the ad segments (midRoll273.mp4) follow an EXT-X-DISCONTINUITY tag, which the specification requires whenever the encoding or timestamps change, as they do when one piece of video ends and an unrelated one begins.

Now picture the cue arriving for 50 million players. With CSAI, every one of them calls the ad server in the same second, a thundering herd from section 5 aimed at a service that can't be cached, because every answer is personalised. With SSAI, the decision moves to the server and can be made once per group of similar viewers. JioHotstar's 2026 post on its ad system says SSAI at its scale "often happens at a Cohort level where we are grouping similar users", and the stitched ad segments are ordinary video files that the CDN caches and collapses like any other.

An ad break for Ritu's cohort, stitched on the server
PackagerAd stitcherAd serverCDNRitu's playerSCTE-35 outads for cohort?pod: 3 adscohort playlistnext segmentsSCTE-35 in
Step 1. The over ends; the cue says an ad break starts at segment 4,903 and how long it lasts.
1 / 6

10.2Choosing the ads in under 100 milliseconds

Choosing the ads is a busy, spiky job of its own. JioHotstar's July 2026 post on it shows its traffic peaking at every ad break, and dropping "more predictable and lower during the innings break" (20:40 to 21:10 in its graph). For each break it picks "2–3 ads out of thousands for a 30-second ad 'pod'", all in under 100 ms "even during massive concurrency spikes like the final over of an IPL match". It filters ads by targeting, checks frequency caps so nobody sees the same ad too often, uses pacing algorithms (it names PID and SHALE) so a campaign's impressions are spread over its whole run, and fills the break through priority tiers, with leftover time going to in-house promos. Ads sold through open programmatic auctions are the risky part: their fill rate "can plummet to 1–2%", against 80–100% for pre-negotiated deals.

That post also describes a third method, server-guided ad insertion (SGAI): the server tells the player what to play and when, and the player does the final insertion, which allows one-to-one personalisation "at a live event scale where we do not have the cue markers in advance".

Decision

How should ads get into a live stream?

Client-side (CSAI)
The player sees the cue, calls the ad server and plays the ads itself.
  • Fully personal ads
  • Simple servers
  • Every player calls the ad server in the same second
  • Switching players can stutter; easy to block
chosen
Server-side (SSAI)
Ads stitched into the playlist per cohort.
  • One decision per cohort, not per viewer
  • Ad segments cached on the CDN
  • Seamless on every device
  • Personal only to the cohort
  • Measuring views needs extra reporting
Server-guided (SGAI)
The server schedules ads per viewer; the player inserts them.
  • One-to-one personalisation
  • Needs capable players
  • Per-viewer decisions at the break

For live sport, SSAI fits the traffic: the break arrives for everyone at the same moment, and turning tens of millions of personalised requests into a few thousand cohort decisions is the same move as request collapsing in section 5. JioHotstar uses SSAI at cohort level and has described SGAI for more personal targeting; its watchdog validates SSAI markers on every live stream because a broken cue means a black screen in the break.

11The scoreboard, a system of its own

11.1Cacheable, and separate from the video

The score in the corner of Ritu's screen doesn't come from the video. It's data: runs, wickets, overs, the batters' scores, and a list of key moments (wickets, boundaries, milestones) that she can tap to jump back to. Where JioHotstar's ball-by-ball data comes from isn't published; on any cricket broadcast it starts with scorers at the ground entering each ball as it happens.

The large hand-operated scoreboard at Eden Gardens, showing batting and bowling figures for India and England
The manual scoreboard at Eden Gardens, Kolkata. The app's scoreboard carries the same information, a few small records per ball. It's tiny next to the video, but every viewer asks for it, all the time.Photo: JokerDurden, CC BY-SA 4.0, via Wikimedia Commons

That data is small, but everybody wants it, constantly. An obvious way to serve it is a score API the app polls every few seconds. At 50 million viewers polling every 5 seconds, that's 10 million requests a second, nearly half as many as the video itself, for a few hundred bytes each. Hotstar's 2024 post describes the answer: the scorecard, the concurrency count and key moments are "highly cacheable", identical for everyone, so they were moved to a separate CDN domain with "leaner but optimal security checks and rate controls". Serving them from CDN caches with a short lifetime turns 10 million requests a second into a trickle at the score service. Keeping them on their own domain isolates them, so a misconfigured rule for the scoreboard can't break video, and the refresh rate becomes one of the panic levers from section 9.

Not everything around the match is cacheable. Hotstar's social feed under the player, launched for IPL 2019, carried content tied to each moment of play, where, the team wrote, "caching at any level is also not possible". For that they built a push service: phones keep a persistent connection open (they chose the MQTT protocol and EMQX broker after benchmarks), designed to "broadcast content to 50 Million simultaneously connected clients". That same system carried live emoji reactions, about 5 billion of them during the 2019 World Cup, aggregated on the server so phones receive the popular ones instead of every tap.

Decision

Should the scoreboard be pushed or polled?

Push over open connections
Phones hold a connection; the server sends each ball.
  • Updates arrive the moment they're published
  • No wasted requests
  • Tens of millions of open connections to run
  • A broker fan-out on every ball
chosen
Poll a cached endpoint
Phones fetch the scorecard every few seconds from the CDN, with a short cache lifetime.
  • The CDN absorbs nearly all the load
  • Refresh rate is a dial for panic mode
  • Up to one polling interval stale
  • Requests even when nothing changed

Since the scorecard is the same for everyone, the CDN can answer it, and staleness of a few seconds is harmless because the video is behind anyway, as the next part shows. Push is worth its cost for per-moment social features that can't be cached; Hotstar built both and used each where it fitted.

11.2Don't spoil the ball

Section 4 left the scoreboard one more problem. Ritu's video is about 15 seconds behind the stadium. But the score service isn't: the scorer enters a six as it lands.

Predict before you read on

Ritu's video is 15 seconds behind the stadium. Her app polls a fresh scorecard every 5 seconds. A batter hits a six. What does she see?

One fix is to give every score event the stadium time at which it happened, and let the player decide when to show it. And the player knows the stadium time of the frame on screen: HLS's EXT-X-PROGRAM-DATE-TIME tag ties the first frame of a segment to an absolute date and time. So the app can hold each score update until the video reaches it, and show the six on the scoreboard as the ball crosses the rope on Ritu's screen. Whether JioHotstar synchronises its scorecard this way isn't published, but any design has to solve this, because a scoreboard that's more up to date than the picture spoils the match.

12Who may watch: entitlement and DRM

12.1Checked once, enforced everywhere

Before Ritu's player gets a playlist URL, the playback service decides whether she may watch: is her subscription tier one that includes this match, is she in a country where JioHotstar holds the rights, is she under her plan's device limit? That's an entitlement check, and it sits on the arrival path from section 7, so it surges with the onboarding rate.

Then there's protecting the video itself. Every segment is encrypted, and a phone can only decrypt them with keys from a licence server, through the phone's built-in DRM (digital rights management) system: Widevine on Android and in Chrome, FairPlay on Apple devices, PlayReady on Windows and many TVs. Each new viewer's player asks the licence server for a licence as it starts (and again only if the keys are rotated during the stream), and then decrypts every segment locally.

For live, the design rule is to do the expensive checks once per session and nothing per segment. The CDN can't call a database for each of 25 million requests a second, so the playback service hands the player URLs carrying a signed token that the edge can verify with a key, without asking anyone. So the licence server and entitlement service see each viewer once, at the start, which makes them onboarding-rate services, scaled and pre-warmed like the rest of the arrival path.

And when they're overloaded, the 2017 rule applies: login, subscription and watch are Tier 1 actions, and a positive panic lets the viewer in and settles the check afterwards. Letting a few people watch an over they weren't entitled to is far cheaper than turning millions of paying viewers away from a final. JioHotstar hasn't published the details of these services, though its watchdog lists "a DRM mismatch" among the failures that can leave a TV spinning while the manifest looks perfect.

13The whole system

13.1Every box, and why it's there

JioHotstar's live path, from the stadium to Ritu's phone
VIDEO PATHscorecardStadium + OB vancontribution ×2Encoders + packagerstwo pipelinesOrigin + ad stitcherpicks a copy, per cohortShieldcollapsesCDNs + ISP cachesedges collapseRitu's phoneplayer + appAPI gatewaypanic switchesArrival servicesauth, home, playback, DRMScore servicecacheableCDN steeringcohort scores
Step 1. Dhoni walks out. The OB van's programme goes out over two contribution paths to two encoding pipelines in different regions, which produce identical segments.
1 / 6
ComponentWhat it doesAdded because
Duplicate pipelinesTwo clock-locked chains from stadium to originOne failure would freeze every viewer (§3)
Short segments, hold-back2–4 s pieces, a few held in reserveLong segments put viewers half a minute behind (§4)
Collapsing at edge and shieldOne upstream fetch per new object per tierMillions ask for each new segment at once (§5)
Multi-CDN + ISP cachesCapacity from several providers, close to viewersNo single CDN has 60–80 Tbps in India (§6)
CDN steeringPer-cohort scores from player telemetryCDNs struggle city by city, network by network (§6)
Ladders on concurrency and onboarding rateScale before the crowd arrivesAutoscaling is minutes too slow (§7, §8)
Panic switches, client back-offShed non-video work; spread out retriesForecasts can be wrong (§9)
SSAI per cohortAds stitched into the playlistEvery player reaches the break at once (§10)
Score service on its own CDN domainCacheable scorecard and key momentsPolling from every viewer, isolated from video (§11)
Entitlement + DRM once per sessionSigned URLs, one licence per startThe CDN can't check a database per segment (§12)

13.2From top to bottom

LevelThe choiceData structure or algorithm
SystemProtect the video above everythingAn ordered list of panic levers, each with a switch
IngestTwo pipelines, origin picksSegments numbered and timestamped identically; first-valid selection
PackagingSliding live playlistMedia sequence numbers, a window of recent segments; LL-HLS parts and preload hints
PlayerHold back from the live edgeA buffer capped by the live edge; hold-back of about three target durations
CDNCollapse misses at every tierA per-object waiting list behind one in-flight fetch
DeliverySteer each cohort to a CDNCohort key (ASN, state, city, user type); weighted score; EWMA with α ≈ 0.6; capped weight changes and a 5% floor
BackendLaddersA table from concurrency and onboarding rate to servers per service
AdsStitch per cohortSCTE-35 cues; discontinuity tags; pacing (PID, SHALE) and tiered fill
ScoreCache and synchroniseShort-lived cached JSON; events stamped with stadium time, shown at the video's program date-time

14What goes wrong, and what it costs

14.1Failures this design has to survive

What happensWhat Ritu seesWhat the design does
One encoding pipeline failsNothingThe origin serves segments from the other pipeline
A regional feed never reaches the CDNNothing, if caught earlyValidators flag manifest 404s before the match is published
Millions ask for a new segment at onceUp to a fraction of a second's waitCollapsing at edge and shield; one origin fetch per segment
One CDN degrades in one cityA brief stall, then normalThe cohort's weight shifts to better CDNs, a capped amount per minute
Her cell tower is crowdedSofter pictureHer player drops a rendition; the service may cap the top rung for everyone
A famous batter walks outNormal startArrival services were scaled up their ladders before the match
The audience overshoots the forecastPlain home page, slower scorecardPanic levers shed features in a planned order; the video keeps playing
A backend error during the surgeA short wait, then successThe app backs off with jitter instead of retrying at once
The ad stitcher gets a bad cueWithout validation, a black breakSSAI markers are checked continuously on every stream

14.2The tradeoffs, in one table

DecisionChosenGiven upWhy it was worth it
IngestTwo pipelinesHalf the encoding costOne failure mustn't freeze the country
Latency2–4 s segments, a few held backThe last few seconds of delayA tiny buffer stalls on mobile networks
Origin protectionCollapsing at edge and shieldA little wait on a missOrigin load stops growing with the audience
DeliveryMany CDNs, steeredOne simple providerNo single provider has the bandwidth
ScalingLadders ahead of the matchMoney for idle serversAutoscaling can't catch a first-ball rush
OverloadPanic levers in a set orderFeatures during peaksThe video must play
AdsServer-side, per cohortOne-to-one targetingNo ad-server herd at every break
ScoreboardCached pollingA few seconds of freshnessThe CDN carries it; the video is behind anyway

15Summary

  1. Live breaks the on-demand design in four places: the pipeline is a single point of failure, the stream lags, the CDN gets a herd for every new segment, and the backend gets a crowd faster than it can grow.
  2. The pipeline runs twice, from two contribution paths to two clock-locked encoders, so the origin can pick a good copy of every segment.
  3. Latency is mostly the player's own safety margin: about three target durations behind the live edge, the only buffer a live player can have.
  4. Shorter segments and LL-HLS parts trade stalls for seconds, and the trade gets steep: below a few seconds of buffer, a mobile viewer freezes several times a match.
  5. Every new segment is a thundering herd, and collapsing at the edge and at a shield brings the origin's load down to about one request per segment per rendition.
  6. No single CDN carries an Indian final, so the stream is spread over several CDNs and caches inside ISPs, steered per cohort by smoothed scores from the players' own telemetry.
  7. The surge is about arrivals, not viewers: the onboarding rate drives login, home page and playback load, and a famous batter's walk to the crease can multiply it in seconds.
  8. Autoscaling is minutes too slow, so the backend climbs ladders sized from load tests before the crowd arrives, on concurrency and onboarding rate.
  9. Panic mode protects the video by switching off features in a planned order and telling clients to back off with jitter.
  10. Ads are stitched on the server per cohort, so the break doesn't send every player to the ad server in the same second.
  11. The scoreboard is cacheable data on its own CDN domain, and has to be held back to match the video, or it spoils the ball.

16Build this

A live stream and its herd, on one laptop.

  • Use ffmpeg to turn a looping video file into a live HLS stream: -f hls -hls_time 2 -hls_list_size 5 -hls_flags delete_segments. Serve the directory with any static web server and play it in a browser with hls.js. Watch the playlist slide.
  • Put an nginx cache in front of the server, with proxy_cache_lock on (nginx's request collapsing). Write a small load generator that, every time the playlist gains a segment, requests it from 2,000 concurrent connections. Count origin requests with the lock off and on.
  • Set the playlist's cache lifetime to 1, 2 and 4 seconds, and measure how far behind the source each player falls.
  • Change -hls_time to 6, 4, 2 and 1. Measure the delay with a clock burned into the video (ffmpeg's drawtext with the time) against the real clock, and use your system's network link conditioner to add two-second dropouts. Compare stall counts with the TryIt in section 4.
  • Add a /score endpoint that returns an event with a timestamp every few seconds. Make the page show each event only when the video's EXT-X-PROGRAM-DATE-TIME reaches it.

17Interview questions

beginnerWhy is a live stream behind the stadium, and why not make the segments tiny?›

The delay adds up from encoding, waiting for each segment to finish, delivery, and above all the player's hold-back: HLS players start about three target durations behind the newest segment, because in live video the distance from the live edge is the only buffer the player can have. Tiny segments shrink that buffer, so a phone hitting a few seconds of dead signal freezes. They also cost compression, because every segment starts with a keyframe, and they multiply requests. Most services settle on a few seconds of segment and a delay close to broadcast TV.

beginnerWhat is request collapsing, and why does live video need it?›

When many requests for the same uncached object arrive together, the cache sends one request upstream and makes the others wait for its answer. Live video needs it because every new segment is requested by millions of players within a second or two of appearing, when no cache has it yet. Without collapsing, each edge server would forward thousands of misses to the origin for every segment; with collapsing at the edges and at a shield tier, the origin sees about one request per segment per rendition.

intermediateWhy doesn't autoscaling handle a cricket match, and what do you do instead?›

Arrivals come in minutes: millions in the first two minutes after the first ball or a famous batter's walk-out, while an autoscaler takes a minute or more to notice and new servers another minute to boot, if the cloud has them. So you scale ahead: ladders, worked out from load tests, that set each service's size for the expected concurrency and arrival rate, applied before the match and stepped up as headroom shrinks. Hotstar also scales arrival-path services on the onboarding rate, and plans an order of features to switch off if the audience overshoots.

intermediateWhy server-side ad insertion for live sport?›

Every viewer reaches the ad break at the same moment. With client-side insertion, every player calls the ad server at once with a personalised request that can't be cached, a herd aimed at a service that must answer in milliseconds. Server-side insertion makes one decision per cohort of similar viewers and stitches ordinary video segments into that cohort's playlist, so the CDN caches the ads like any other segment and every device plays them seamlessly. In exchange, targeting is per cohort, not per person, which server-guided insertion tries to recover.

deepDesign how a multi-CDN live service decides which CDN each viewer uses.›

Group viewers into cohorts by network and place (JioHotstar uses ASN, country, state, city and user type), because a CDN can be fine nationally and failing for one operator in one city. Score each CDN per cohort from the players' own telemetry, weighting playback failures above rebuffering above round-trip time, and smooth the scores (an EWMA with α ≈ 0.6) so a ten-second blip doesn't move a city. Then turn scores into traffic weights with brakes: throttle a CDN before its capacity limit, during surges aim for all CDNs to fill together, cap the change per minute, and keep every working CDN above a floor so its caches stay warm for failover. Hard constraints come first: never send a cohort to a CDN with no presence on its network.

deepThe score on screen keeps spoiling the video. Why, and how do you fix it?›

The video is 10–20 seconds behind the stadium because of encoding, segments and the player's hold-back, but the scorecard is published as each ball happens. Stamp every score event with the stadium time it happened, and have the app show it only when the video on screen reaches that time, which the player knows from the program date-time of each segment. Notifications have the same problem, and need either a delay or a choice to accept the spoiler.

18Go deeper

check yourself
Why can't a live player keep a minute of video in its buffer, as an on-demand player can?›

Because the next minute of the match hasn't happened yet. A live player can only be as far ahead as the newest segment, so the hold-back it chooses is the whole of its buffer, and every second of buffer is a second of delay.

Edges collapse their misses, and a shield sits in front of the origin but doesn't collapse. How many origin requests per segment, with 50 edges?›

About 50. All the edges' first fetches reach the shield within its own fetch time, so each one misses and is forwarded. So the shield has to collapse too to get the origin down to one request per segment.

Concurrency is falling and no one is joining, yet the home page service is overloaded. What happened?›

A match, or an innings, just ended. Everyone who was watching returns to the home page at once: Hotstar calls it the off-boarding rate and scales the home page on concurrency times a post-match coefficient.

Hotstar: 'Scaling Is Not An Accident' (2018) and 'Scaling the Hotstar Platform for 50M' (2019)

Ladders and headroom instead of autoscaling, clients that back off with jitter, cutting latency from 55 s to 15–20 s behind broadcast, and the jettison order for 50 million viewers.

Hotstar: 'T For Tsunami: Dealing with traffic spikes' (2017)

The Kohli walk-in that took the login API from 10K to 125K TPS, progressive degradation of bitrates, and positive and negative panics.

JioHotstar: 'Scaling Tales: Discovering Onboarding Rate' (2025)

The first-ball rush, ladders on arrivals as well as viewers, and the post-match storm on the home page.

JioHotstar: 'Orchestrating JioHotstar Traffic' (2026)

The QoS Routing Manager: cohorts, scores from player telemetry, EWMA smoothing and capacity steering across CDNs.

HTTP Live Streaming, RFC 8216 and the 2nd edition draft

The live playlist rules, the three-target-duration hold-back, partial segments, blocking reload, SCTE-35 date ranges and content steering.

Netflix: 'Behind the Streams: Live at Netflix' (2025) and the Live Origin post (2025)

Live from a VOD company: 2-second segments, two pipelines with an origin that picks, millisecond caching and retry storms. Read with the Open Connect overview for the on-demand contrast.

Designing YouTube

Renditions, segments, manifests, CDN caches and the bitrate algorithm that this chapter builds on. Chapter 51.

Caching

Stampedes and request coalescing inside one service, the small version of section 5. Chapter 25.

Queueing and Capacity

Why headroom matters and how utilisation turns into waiting. Chapter 42.

Designing Zoom

The other end of the latency scale: real-time video over UDP for a meeting, not HTTP segments for a country. Chapter 53.

Reliability

Load shedding, back-off and graceful degradation in general. Chapter 40.