It's the IPL final. Ritu is at home in Kanpur, half watching on her phone over 4G while she cooks. Chennai are chasing, the required rate is climbing, and in the sixteenth over a wicket falls. Up on the screen, the camera finds the dressing room stairs, and Dhoni walks out to bat. Her phone buzzes with a notification from the app, and so do a few million other phones across India. Within a couple of minutes, millions of people who weren't watching are.

From Ritu's side it's one tap: the app opens, the stream starts, an ad plays at the end of the over, the score in the corner ticks over. Behind that tap, a lot is less simple than it looks. Ritu wants to see a ball that was bowled a few seconds ago and didn't exist before that, so no cache anywhere can have it ready in advance. Everyone else wants exactly the same few seconds of video at exactly the same moment. An audience like this doesn't grow smoothly; it jumps by millions when something happens on the field. And the neighbour's TV, a wall away, cheers a boundary before Ritu's picture shows the ball leaving the bat.
Our YouTube case study (chapter 51) already built the machinery for video on demand: renditions, segments, manifests, CDN caches, the player's bitrate choice. This chapter starts from that design and asks what breaks when the video is live and the audience is a whole country. So the question we'll keep coming back to is this: when Dhoni walks out and millions more people tap Play in the same few minutes, how does every one of them get the same ball, a few seconds behind the stadium, without anything falling over? We'll follow the video from the stadium to Ritu's phone, then the crowd from the notification to the play button, and finish with the ads, the score and the checks that decide who may watch.
01What we're building, and how big
1.1What it has to do
The live-cricket part of JioHotstar comes down to a short list:
- Start playing quickly when someone taps the match, wherever they are and whatever network they're on.
- Stay close to live, so the picture isn't far behind the stadium (and the neighbour's TV).
- Keep playing through the whole match, adapting to the viewer's connection.
- Show ads in the breaks between overs, and count that they were shown.
- Show the score and the match's key moments alongside the video.
- Let in only the people allowed to watch: the right subscription, the right country, a licensed device.
- Carry many versions of the match: commentary in several languages, extra camera angles, and quality levels from a phone on a weak signal up to a 4K TV.
And the qualities it needs while doing that:
- The video must play. Everything else on the screen can fail before the video does. Hotstar's engineering posts come back to this rule again and again: nobody misses the extras if the match itself won't play.
- Surge-proof: the audience can grow by millions in a couple of minutes, at moments nobody can schedule.
- Live: a delay of tens of seconds is noticeable, and a delay of a minute means the score notification arrives before the ball.
1.2How big is it?
India's cricket streams have broken the world record for simultaneous viewers several times, and each record was set by a single service in a single country. What matters is peak concurrency: the most people watching at the same instant.
| Date | Match | Service | Peak concurrent viewers |
|---|---|---|---|
| Jun 2017 | Champions Trophy final | Hotstar | 4.7 million |
| May 2019 | IPL final | Hotstar | 18.6 million |
| 10 Jul 2019 | World Cup semi-final, India v New Zealand | Hotstar | 25.3 million |
| 29 May 2023 | IPL final, Chennai v Gujarat | JioCinema | 32.1 million |
| 19 Nov 2023 | World Cup final, India v Australia | Disney+ Hotstar | 59 million |
| 9 Mar 2025 | Champions Trophy final, India v New Zealand | JioHotstar | 61.2 million |
| 5 Mar 2026 | T20 World Cup semi-final, India v England | JioHotstar | 65.2 million |
| 8 Mar 2026 | T20 World Cup final, India v New Zealand | JioHotstar | 72.5 million |
You've probably noticed that two of those names have since merged. Reliance's Viacom18, which ran JioCinema, and Disney's Star India combined into a joint venture called JioStar in November 2024, and on 14 February 2025 JioCinema and Disney+ Hotstar became one app, JioHotstar. Most of the engineering history in this chapter comes from Hotstar's team, whose blog has kept running under the new name.
Two more facts shape the design. In March 2026 JioHotstar's engineers wrote that a live event needs upwards of 60–80 terabits per second of bandwidth for streaming. And the viewers are on phones: JioStar's CTO described a region "where 85% of the viewership is on mobile" in an April 2026 talk, and at the end of 2025 India's telecom regulator counted about 983 million wireless internet subscribers against 45 million wired ones.
Suppose 50 million people are watching, on phones, at an average of about 1.5 megabits per second each. How much bandwidth is that? And if each player fetches one 4-second segment and re-reads the playlist every 4 seconds, how many requests a second hit the delivery network?
02Version 1: the YouTube design, pointed at a live match
2.1Reuse what works for video on demand
The obvious first design is the one from chapter 51 with a camera at the front. An encoder takes the match feed and produces a ladder of renditions, the same picture at several resolutions and bitrates (chapter 51, §4.2). A packager cuts each rendition into short segments of a few seconds and lists them in a playlist (HLS's name for the manifest), and both go to origin storage. A CDN, a network of cache servers close to viewers, sits in front of the origin. Ritu's player reads the playlist, fetches segments, and picks a rendition for each one with its adaptive-bitrate algorithm (chapter 51, §8).
HLS has a way to say "this playlist is still growing". Here is a live playlist from the HLS specification, RFC 8216:
#EXTM3U
#EXT-X-VERSION:3
#EXT-X-TARGETDURATION:8
#EXT-X-MEDIA-SEQUENCE:2680
#EXTINF:7.975,
https://priv.example.com/fileSequence2680.ts
#EXTINF:7.941,
https://priv.example.com/fileSequence2681.ts
#EXTINF:7.975,
https://priv.example.com/fileSequence2682.tsCompare it with the on-demand playlist in chapter 51 and one tag is missing: EXT-X-ENDLIST. Without it, the player knows more segments are coming and must reload the playlist to find them. EXT-X-MEDIA-SEQUENCE:2680 says the first segment listed is number 2,680; the packager keeps only the last few segments in the list and drops older ones off the top, so the playlist is a window that slides along with the match.
2.2Where it breaks
This design works for a small live stream. For the IPL final, it breaks in four places, and each one is something Ritu would notice:
- The ingest is a single thread. One feed, one encoder, one packager. If any of them hiccups, every viewer's picture freezes at the same instant, and there's no older copy to fall back on.
- It's too far behind. With 8-second segments, Ritu's picture trails the stadium by half a minute or more. Her neighbour cheers, her phone buzzes with "SIX!", and then she sees the ball.
- The CDN gets hammered in a way it wasn't built for. Millions of players ask for segment 2,683 within the same second or two of it appearing, and the caches don't have it yet.
- The crowd arrives faster than servers can be added. When Dhoni walks out, the video bytes come from CDNs that were sized in advance, but the app's own servers, which log people in, build the home page and hand out the stream, see millions of new arrivals in a couple of minutes.
We'll take these in order, starting where the video starts: in the stadium.
03From the stadium to the cloud
3.1Cameras, a truck and a contribution feed
The pictures start with the broadcast cameras around the ground. Their signals run by cable to an outside broadcast van (OB van), a production truck parked at the stadium, where a director picks which camera is live, adds replays and graphics, and mixes in the commentators' audio. What comes out is one finished programme, the match as you'd see it on TV.


That programme has to leave the stadium for wherever the encoders are. What carries it out of the stadium is called the contribution feed: a high-quality, lightly compressed signal sent over dedicated fibre or a satellite uplink to a broadcast centre. It's called contribution to tell it apart from distribution, the heavily compressed copies sent to viewers. A contribution feed carries far more bits than any viewer will receive, because every later step loses some quality and the encoder needs a clean picture to start from.

A cricket stream is many programmes at once. In April 2026, JioHotstar's engineers described a single live event carried "in 12 languages, 5 camera angles", in variants for TVs and phones and for formats such as Dolby Vision and 4K, and said they "spin up 100s of encoders for every match". At peak they have run close to 55 live streams at the same time. Each language is a different audio track (and often a different commentary team), so a Tamil viewer and a Hindi viewer get the same pictures with different sound.
3.2Encoding and packaging, a few seconds at a time
From here the pipeline is chapter 51's, run continuously. Each programme gets compressed into a ladder of renditions, and the packager cuts them into segments, every rendition at the same instants so the player can switch between them. What's different is the clock. For a film, the transcoding pipeline can take an hour. For a match, the encoder has to keep up with real time forever, and the packager publishes a new segment for every rendition, language and angle every few seconds, then rewrites each playlist to announce it.
Because the playlist is rewritten so often, it's the most-requested file in the system. Every player re-reads it every few seconds, and RFC 8216 sets the rhythm: after a reload that found something new, the player waits at least one target duration before reloading; if nothing changed, it tries again after half a target duration.
3.3Two of everything
One encoder failing for thirty seconds during a final is the kind of failure every viewer notices at once. So the usual design runs two complete pipelines in parallel: two contribution paths out of the stadium, two sets of encoders, two packagers, each producing the same segments with the same numbers and timestamps.
JioHotstar hasn't published how its pipelines are duplicated. Netflix has, for its own live events: its December 2025 post on its Live Origin describes two independent pipelines in separate cloud regions, with separate contribution feeds, encoders and packagers, and an origin that, when asked for a segment, checks the candidates from both pipelines in a fixed order and serves the first valid one. It works because both encoders cut segments at identical moments, so segment 4,817 from either pipeline covers exactly the same two seconds of the match.
What JioHotstar has described is how it checks the result. Its 2026 "watchdog" post lists the failures its automated validators catch before viewers do: "manifest 404s" where a regional feed hadn't reached the CDN edge, ad markers that would have produced a black screen, and "media sequence drift", where the segment numbers of different bitrates fall out of step and a player switching between them would buffer forever.
How many copies of the live pipeline should run?
- Half the cost
- Nothing to keep in step
- Any failure freezes every viewer at once
- A restart takes longer than a viewer's buffer
- One pipeline can fail without viewers noticing
- Can be tested by switching one off
- Double the encoding cost
- Both must cut identical segments
For an event watched by tens of millions, the cost of a duplicate encoding chain is small next to the cost of the whole country seeing a frozen frame. Keeping the two in lockstep is the harder engineering, and drift between streams is exactly what JioHotstar's watchdog checks for.
So the pipeline produces segments reliably. Next: how long they take to reach Ritu, and why her neighbour hears the six first.
04How far behind the stadium
4.1Where the seconds go
The delay between something happening in front of a camera and it appearing on a viewer's screen is called glass-to-glass latency: from the camera's lens to the screen's glass. Let's add it up for Ritu, from the moment the ball leaves the bat.
- Production and contribution. The OB van adds a little, and the trip out of the stadium adds more. A satellite in geostationary orbit is about 36,000 km up, so a hop up and down covers over 70,000 km and takes about a quarter of a second at the speed of light.
- Encoding. Encoders look ahead a little to compress well, so this costs around a second or two.
- Waiting for the segment to finish. The packager can't publish segment 4,817 until the last frame of it has been encoded. With 6-second segments, the first frame of the ball waits up to 6 seconds.
- The CDN and the network, usually well under a second.
- The player's safety margin. This is the big one. RFC 8216 says that for a live playlist the client "SHOULD NOT choose a segment that starts less than three target durations from the end of the Playlist file. Doing so can trigger playback stalls." So a player using 6-second segments starts about 18 seconds behind the newest segment.
Add those up with 6-second segments and Ritu is roughly half a minute behind the stadium. Hotstar wrote in 2018 that it had cut its live latency over the previous year "from being roughly 55s behind broadcast, to approximately 15–20s behind broadcast", by measuring how long each step of its encoding workflow took and tuning the encoder settings. And broadcast TV is itself several seconds behind the stadium, which is why the neighbour still hears it first.
?Why does the player stay three segments behind on purpose?
Because in live video the buffer can't grow. For a film, Ritu's player could download a minute ahead and coast through a dead spot in the network. For a match, the video a minute ahead hasn't happened yet. At best the player can be at the newest segment, so the distance it chooses to stay behind the live edge is the only buffer it has. Starting three segments back means a network stall of up to that long empties the buffer without freezing the picture. Every second of safety is a second of delay, and every second of delay removed is a second less of protection.
4.2Smaller pieces: chunks and partial segments
An obvious fix is shorter segments. Halve the segment length and you halve both the wait for the segment to finish and the three-segment safety margin. But every segment has to start with a keyframe, a picture that can be decoded without any earlier one, and keyframes are expensive, so very short segments cost compression (chapter 51, §7.4). They also multiply requests: 1-second segments mean four times the playlist reloads and segment fetches of 4-second ones.
A way out is to keep segments a reasonable length for compression and caching, but let the player fetch each one in pieces while it's still being made. CMAF (the Common Media Application Format), the fragmented-MP4 packaging that both DASH and HLS players can read, allows a segment to be written as a series of small chunks, each a few hundred milliseconds of video that can be sent as soon as it's encoded. Low-Latency HLS exposes these to the player as partial segments, listed in the playlist with an EXT-X-PART tag. This is the end of a low-latency playlist from the second edition of the HLS specification (May 2026 draft):
#EXTM3U
#EXT-X-TARGETDURATION:4
...
#EXTINF:4.00008,
fileSequence270.mp4
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="filePart271.0.mp4"
#EXT-X-PART:DURATION=2.00004,URI="filePart271.1.mp4"
#EXTINF:4.00008,
fileSequence271.mp4
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="filePart272.0.mp4"
#EXT-X-PART:DURATION=0.50001,URI="filePart272.1.mp4"
#EXTINF:2.50005,
fileSequence272.mp4
#EXT-X-DISCONTINUITY
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="midRoll273.0.mp4"
#EXT-X-PART:DURATION=2.00004,URI="midRoll273.1.mp4"
#EXTINF:4.00008,
midRoll273.mp4
#EXT-X-PART:DURATION=2.00004,INDEPENDENT=YES,URI="midRoll274.0.mp4"
#EXT-X-PRELOAD-HINT:TYPE=PART,URI="midRoll274.1.mp4"Read it from the bottom. Segment 274 is still being encoded: its first part, midRoll274.0.mp4, exists, and the EXT-X-PRELOAD-HINT tells the player the name of the next part before it exists, so the player can ask for it now and the server answers the moment it's ready. INDEPENDENT=YES marks the parts that start with a keyframe, where a player can begin. (The midRoll names and the EXT-X-DISCONTINUITY tag are an ad break; section 10 comes back to them.)
Two more mechanisms stop the playlist itself from adding delay. With blocking playlist reload, a player asks for "the playlist once part 274.1 exists", and the server holds the request open until it does, instead of the player polling and guessing. And the safety margin shrinks with the parts: the specification's PART-HOLD-BACK, the distance a low-latency player keeps from the live edge, "MUST be at least twice the Part Target Duration" and "SHOULD be at least three times" it.
That specification is blunt about the price: "A shorter Target Duration reduces latency but also reduces available buffer, handicaps adaption and increases delivery overhead, increasing the likelihood of playback stall." So how much stall risk are we buying with each second of latency? Let's measure it.
4.3Trading latency for stalls
Ritu's player keeps three chunks of the match in hand. Going from 4-second segments to 2-second segments cuts her delay behind the stadium from about 20 seconds to about 12. What happens to the number of viewers who see at least one freeze during a three-hour match?
This program models a viewer on a mobile network for a three-hour match. Every ten minutes or so the phone hits a dead spot that lasts about two seconds on average, and now and then a longer one of several seconds, such as a hand-off between cell towers. Each player starts the match holding as much video as its safety margin and can never hold more, because the rest of the match hasn't happened. During a dead spot the buffer drains one second per second; if it runs out, the picture freezes. When the network comes back, the player refills at one and a half times real time. A fixed 4 seconds of encoding, packaging and delivery, the same for every row, stands in for the rest of the pipeline.
import random
random.seed(7)
FIXED = 4.0 # seconds: capture, encode, package, CDN (same for all)
MATCH = 3 * 3600 # a three-hour match
VIEWERS = 2000
def one_viewer(hold):
# The buffer starts at the hold-back and can't grow past it,
# because the next piece of the match hasn't been played yet.
buf, t, stalls, frozen = hold, 0.0, 0, 0.0
while t < MATCH:
if random.random() < 1 / 600: # a dead spot every ~10 min
gap = random.expovariate(1 / 2) # ~2 s on average
if random.random() < 0.05:
gap += random.uniform(3, 10) # now and then, a tower hand-off
if gap > buf:
stalls += 1
frozen += gap - buf
buf = max(0.0, buf - gap)
t += gap
buf = min(hold, buf + 0.5) # refill at 1.5x real time
t += 1
return stalls, frozen
print("chunk hold-back behind stadium stalls/match frozen s viewers who stall")
for name, chunk, hold in [("6 s seg", 6, 18), ("4 s seg", 4, 12),
("2 s seg", 2, 6), ("1 s part", 1, 3),
("0.5 s part", 0.5, 1.5)]:
runs = [one_viewer(hold) for _ in range(VIEWERS)]
stalls = sum(s for s, _ in runs) / VIEWERS
frozen = sum(f for _, f in runs) / VIEWERS
share = sum(s > 0 for s, _ in runs) / VIEWERS
print(f"{name:11} {hold:5.1f} s {FIXED + chunk + hold:8.1f} s "
f"{stalls:8.2f} {frozen:6.1f} {share:5.0%}")chunk hold-back behind stadium stalls/match frozen s viewers who stall
6 s seg 18.0 s 28.0 s 0.01 0.0 1%
4 s seg 12.0 s 20.0 s 0.13 0.3 12%
2 s seg 6.0 s 12.0 s 1.58 4.3 80%
1 s part 3.0 s 8.0 s 4.75 12.8 99%
0.5 s part 1.5 s 6.0 s 8.96 22.4 100%Read the rows from the top. Six-second segments keep Ritu 28 seconds behind and almost nobody freezes. Four-second segments buy 8 seconds of latency for a freeze that one viewer in eight sees once. Two-second segments buy another 8 seconds, and now most viewers freeze at least once, because a 6-second buffer is shorter than the occasional tower hand-off. Below that, every viewer freezes several times a match. So the curve is lopsided: the first seconds of latency are cheap to remove and the last ones are very expensive.
This model is deliberately crude. Real dead spots are probably rarely total, so a real player would drop to a lower rendition and keep downloading slowly instead of stopping, and a player can also widen its hold-back after a stall. But the shape holds, and it's the shape the HLS specification warns about. It's why "lowest possible latency" is the wrong goal for a mobile-first audience, and why services aim for a delay that's a little behind broadcast TV and stable.
How far behind live should the player sit?
- Almost no stalls on a bad network
- Fewest requests per viewer
- Half a minute behind the stadium
- Score alerts arrive before the ball
- 10–20 s behind, close to broadcast TV
- Nothing special needed in CDNs or players
- More stalls on mobile networks
- More requests, slightly worse compression
- A few seconds behind the stadium
- Tiny buffer: many stalls on a weak signal
- Every hop must pass chunks on as they arrive
JioHotstar's segment length isn't published. Large mobile audiences probably land on the marked option: Netflix's 2025 description of its live platform uses 2-second segments, and Hotstar's 2018 target was 15–20 seconds behind broadcast. Low-latency modes make the most sense where the network is good, such as a TV on home broadband, and a service can offer different hold-backs to different devices from the same segments.
So now Ritu's player asks for a small segment every couple of seconds. So does everyone else's, and they all ask for the same one.
05Everyone wants the newest segment
5.1Why live breaks the CDN's assumptions
A CDN for on-demand video works because popular files stay popular. In Kanpur, the first viewer of a film causes a miss, the edge server fetches it and keeps it, and the next ten thousand viewers over the following week get hits. Chapter 51 showed how well that works when popularity is skewed.
Live video breaks this in three ways at once.
- The file is brand new. Segment 4,817 didn't exist two seconds ago, so every cache in the country starts empty.
- Everyone wants it in the same second or two. Ten minutes from now, almost nobody will ask for it again. Its life as a popular object lasts roughly as long as it takes to play it.
- The playlist changes every few seconds. It can only be cached for a moment, or players would never learn about the new segment.
Put the first two together and you get a burst: in the instant segment 4,817 is announced, millions of requests land on edge servers that don't have it. Each edge server that sends every one of those misses to the origin turns one popular file into a flood. This is the same thundering herd, or cache stampede, that chapter 25 met when one hot key expired (§8), except that in live video it happens on purpose, every two to six seconds, for three hours.

Playlists have their own version of the problem, and Hotstar's 2019 post warned about it under the heading "TTL Cloudbursts": with several layers of caches, if you lose track of how long each one keeps each object, "large number of objects expiring around the same time" can send "a request cloudburst at your origin". A playlist cached for 2 seconds at every edge expires at every edge every 2 seconds.
5.2Request collapsing
Fixing it starts at each edge server. When the first request for segment 4,817 misses, the edge starts one fetch from the tier above. Every request for the same segment that arrives while that fetch is in flight doesn't start another; it waits in line, and when the one fetch returns, the edge answers all of them from it. This is called request collapsing (or request coalescing). Fastly's documentation describes it in those terms: on a miss, concurrent requests share one origin fetch, and the requests that arrive meanwhile join a "waiting list".
Collapsing at the edge turns "every miss goes upstream" into "one request per edge server goes upstream". That's a big improvement, but a large CDN has thousands of edge servers, and each one still fetches its own copy. So the second half of the fix is a layer between the edges and the origin: a small number of big caches, often one per region, that every edge fetches from. AWS calls its version Origin Shield, and its documentation describes exactly the live case: requests for content not yet in the shield's cache "are consolidated with other requests for the same object, resulting in as few as one request going to your origin", and it lists "origins that provide just-in-time packaging for live streaming" and "workloads that use multiple content delivery networks" as the uses it's for.
How much does each layer save? This program sends two million viewers at one region's 50 edge servers, all asking for segment 4,817 at random moments in a two-second window. Fetching a copy from the tier above takes 200 milliseconds. It counts how many requests reach the origin under four policies.
import random
random.seed(42)
VIEWERS = 2_000_000 # watching in one region
EDGES = 50 # edge servers they're spread across
WINDOW = 2.0 # seconds over which players ask for the new segment
FETCH = 0.200 # seconds to fetch one copy from the tier above
# every viewer asks one edge for segment 4817 at some moment in the window
arrivals = [[] for _ in range(EDGES)]
for _ in range(VIEWERS):
arrivals[random.randrange(EDGES)].append(random.uniform(0, WINDOW))
for a in arrivals:
a.sort()
def misses(times):
"""Requests that find the cache empty: they arrive before the
first fetch (started by the first request) has come back."""
ready = times[0] + FETCH
return [t for t in times if t < ready]
# 1. edges forward every miss to the origin
no_collapse = sum(len(misses(a)) for a in arrivals)
# 2. each edge sends one fetch; later requests wait for it
edge_collapse = sum(1 for a in arrivals if a)
# 3. those edge fetches go to a shield, which forwards each of its misses
first_fetches = sorted(a[0] for a in arrivals if a)
shield_no_collapse = len(misses(first_fetches))
# 4. the shield collapses them too: the first one fetches, the rest wait
shield_collapse = 1
print(f"{VIEWERS:,} viewers ask for segment 4817 within {WINDOW:.0f} s")
print(f"edges forward every miss : {no_collapse:>7,} origin requests")
print(f"edges collapse : {edge_collapse:>7,} origin requests")
print(f"edges collapse, shield forwards : {shield_no_collapse:>7,} origin requests")
print(f"edges collapse, shield collapses : {shield_collapse:>7,} origin request")
print(f"requests that waited on a fetch : {no_collapse:,} "
f"({no_collapse / VIEWERS:.0%} of all), none longer than {FETCH*1000:.0f} ms")2,000,000 viewers ask for segment 4817 within 2 s
edges forward every miss : 200,364 origin requests
edges collapse : 50 origin requests
edges collapse, shield forwards : 50 origin requests
edges collapse, shield collapses : 1 origin request
requests that waited on a fetch : 200,364 (10% of all), none longer than 200 msWithout collapsing, 200,364 requests reach the origin for one segment, a tenth of all viewers, because a tenth of the requests arrive during the 200 ms that each edge's first fetch is in flight. That happens again for the next segment two seconds later, and for every rendition and language. Collapsing at the edges cuts it to 50, one per edge. Line three is the trap: a shield that doesn't collapse is just another cache that misses, and all 50 edge fetches arrive inside its own fetch window, so it forwards all 50. Only when the shield collapses too does the origin see one request per segment per rendition. You can see the cost on the last line: a tenth of viewers waited up to 200 ms for their copy, a bit of delay that no viewer would notice next to a multi-second buffer.
5.3The details that bite
Collapsing has sharp edges, and live streaming finds all of them.
- An uncacheable reply breaks the line. If the origin's response can't be cached (an error, or a header that forbids it), every request on the waiting list was waiting for nothing. Fastly's documentation says the queued requests are then sent to the origin one after another, and warns that "in some cases this can create extreme response times of several minutes", and suggests failing such requests quickly instead.
- Asking too early. A player that requests segment 4,818 a moment before it exists gets a 404. If the CDN caches that 404 for even a second, everyone behind it gets a 404 too, for a file that now exists. Low-Latency HLS's blocking reload and preload hints exist partly so that a request for something about to exist is held open instead of failed.
- Tiny lifetimes. HTTP's cache lifetimes are counted in whole seconds, which is coarse for a playlist that changes every two seconds. Netflix's Live Origin post describes adding millisecond-grain caching to its nginx servers for this reason, and, during a surge, telling Open Connect to cache identical requests for five seconds and returning HTTP 503 to low-priority traffic such as rewinding.
How do you stop the newest segment from flattening the origin?
- Nothing clever in the CDN
- Needs capacity for hundreds of thousands of identical requests per segment
- Fails the moment the audience grows past the plan
- Cuts origin load to one request per edge
- Thousands of edges still means thousands of fetches
- About one origin request per segment per rendition
- Lets several CDNs share one shield
- An extra hop of latency on every miss
- The shield is now critical and must be redundant
With collapsing at both tiers, the origin's load depends on the number of renditions and languages, not on the number of viewers. So the audience can grow from 25 to 70 million while the origin's load stays flat. Hotstar made the same point in 2019 when it said its scaling wasn't "the CDN (alone)": a CDN is a "shock absorber" only if you "architect your CDN to sit in front of your origin" and know its limits.
So the origin is safe. But one region's edges are no longer enough when the audience needs 60–80 terabits a second, so where does all that bandwidth come from?
06Getting the bytes to 50 million phones
6.1More than one CDN, and caches inside the ISPs
No single CDN has spare capacity on the scale of an Indian cricket final. Hotstar's 2024 infrastructure post says it outright: when the team modelled bandwidth for 50 million concurrent streams, "the demand was far beyond what our CDN providers could handle." So the stream is spread across several CDNs at once, a design called multi-CDN. JioHotstar's April 2026 watchdog post names three whose edges it checks against each other: "Akamai, Cloud front, JIO".
That third name matters. Cheapest and fastest of all are the bytes that never leave the viewer's own internet provider. Chapter 51 described Google Global Cache, servers Google places inside ISPs' networks; Netflix does the same with its Open Connect appliances. For a viewer on Jio's mobile network, a cache inside Jio's network is a few hops away and never touches a link between providers. An independent analysis by the magazine Voice&Data during IPL 2024 found that Jio's users were served almost entirely from JioCinema's own CDN servers inside the Jio network, while some users of other mobile operators were served from outside India. How JioHotstar's caches are placed today isn't published.

6.2The last mile is a cell tower
Most of those bytes end on a mobile network, and a mobile network shares its capacity per cell: everyone connected to the same tower splits that tower's radio capacity. A busy market or a train during a final can probably put hundreds of phones on one cell, all streaming the same match. Each player's bitrate algorithm handles this viewer by viewer, dropping to a smaller rendition when downloads slow down (chapter 51 covers how).

But the service can also act for everyone at once. Hotstar's 2017 post on surviving traffic spikes described progressive degradation: the key levers, it said, are "the bit-rates at which we offer the content, which must degrade as concurrency grows", because the bandwidth in the infrastructure is a hard limit. Removing the top rung of the ladder for phones when the audience passes some size costs each viewer a little sharpness and frees a large slice of total bandwidth. A 2026 JioHotstar post puts it with an image of a bridge: everyone wants to drive a wide luxury bus, but the lanes are finite, and during a peak surge there isn't room for everyone to drive one at once.
6.3Which CDN should Ritu use?
With several CDNs, something must decide which one each viewer fetches from, and the decision changes during the match. A CDN can be healthy across India and struggling for one mobile operator in one city. JioHotstar described its answer in March 2026, a service it calls the QoS Routing Manager.
It works in four steps.
- Group viewers into cohorts. A cohort is a combination of network and place: "ASN-Country-State-City-UserType". (An ASN, autonomous system number, identifies a network on the internet, such as one mobile operator.) Ritu is in the cohort "Jio, India, Uttar Pradesh, Kanpur, mobile".
- Score each CDN for each cohort from the players' own reports, which arrive as heartbeats. Three measurements are each mapped to a score: playback failure rate (how often a stream fails to start), rebuffering, and round-trip time. A cumulative score weighs them in that order, failures most heavily, "to ensure that 'reachability' is the absolute priority".
- Smooth the scores with an exponentially weighted moving average, so one bad ten-second window doesn't send a whole city to another CDN. Each new score counts for a fraction α of the average and the old average for the rest, with α about 0.6, which keeps roughly the last five scores significant.
- Shift traffic weights, with brakes. Before a CDN reaches its capacity, its share is throttled even if it scores well. During a surge the logic switches to making all CDNs run out of capacity at the same time. A CDN's share can only change by a capped amount each minute, and no working CDN ever drops below a floor, "typically 5%", so its caches stay warm in case it's suddenly needed.
Here's step 3 on a small example. One CDN's score for Ritu's cohort is steady around 90, then a ten-second window comes in at 30 because of a tower hand-off in one neighbourhood:
| Window | Raw score | Smoothed (α = 0.6) |
|---|---|---|
| 1 | 90 | 90.0 |
| 2 | 88 | 88.8 |
| 3 | 30 | 53.5 |
| 4 | 89 | 74.8 |
| 5 | 90 | 83.9 |
The dip shows, but it's about 40% shallower than the raw one, and two windows later the score is most of the way back. Combined with the cap on how far a share can move per minute, a blip moves a little traffic for a minute instead of the whole cohort. JioHotstar tested the manager against the old routing during a T20 match, with cohorts split only by ASN and state, and in that A/B test the treatment group's playback failure rate improved by 11%, and rebuffering and round-trip time by 2% each.
How should a live match get to tens of millions of viewers?
- Almost no traffic crosses provider links at peak
- Fill happens when networks are quiet
- Needs to know the content in advance: impossible for a match
- One network: no fallback if it fills up
- One system to run
- Reuses a huge existing footprint
- All the live risk sits on one network
- Built and tuned for a different traffic shape
- More total capacity than any one CDN
- Routes around a CDN struggling in one city
- A steering system to build and trust
- Every CDN must be configured and tested identically
Netflix's Open Connect overview says the program was designed as "a proactive, directed caching solution", and that Netflix can deploy "the majority of our content and software updates proactively during off-peak fill windows" because "we can predict with high accuracy what our members will watch". None of that applies to a ball bowled two seconds ago. When Netflix streamed the Paul–Tyson fight in November 2024, it reported 65 million concurrent streams at the peak, and many viewers reported buffering and dropped streams. For JioHotstar, where a single match can need more bandwidth than any one provider has in India, spreading across many CDNs is a necessity, and steering is what makes many CDNs behave like one.
So the video reaches the phones. But the video was never where the crowd hurt most. When Dhoni walks out, what breaks first is the app's own servers.
07The surge: toss, first ball, wickets
7.1The shape of a match
Watch the number of viewers across one match and it rises fairly smoothly from the toss to the death overs. Watch the number of people arriving each minute and you see something else. JioHotstar's engineers plotted new sessions per minute and called it the onboarding rate: how fast people are joining, as opposed to concurrency, how many are watching. Their December 2025 post lists what the curve shows in every match: "the mini-spike at toss", "the explosion at first ball", "the middle overs lull", a second-innings spike, "the pre-death overs ramp", and "the sudden drop at match end".
Those numbers are steep. At the first ball, they wrote, "6–8 million users joining in the first 2 minutes", the home page's API traffic jumping 300–400% and the watch page's (the page with the player) 400–500%. "Then, strangely, 10 minutes later the same services hum along happily as concurrency climbs toward 50 or 60 million."

Unplanned moments are worse than the first ball, because nobody knows when they'll come. A wicket brings a new batter, and a famous one brings a crowd. Hotstar wrote about the first time this hurt, during the India–Pakistan opener of the 2017 Champions Trophy: "Virat Kohli walks in to bat — a notification is sent to a large segment and boom! Our login API that was cruising at 10K TPS, shot up to 125 K TPS in a matter of seconds." (TPS is transactions per second.) That's a twelvefold jump. Hotstar had scaled its servers for 7 million viewers, but the database behind the login service hadn't been scaled for the extra traffic all those servers brought on, and it "started to melt down".
Ritu's notification about Dhoni is the same event. Here the app's own push notification creates the surge it then has to absorb.
7.2What the surge hits
Why do the app's servers suffer more than the CDNs? Because a viewer who's already watching costs the backend almost nothing: their player talks to the CDN, plus an occasional heartbeat. A viewer who's arriving hits every service on the way in, as the StepSequence below shows for Ritu.
So the services on the arrival path, login, profile, home page, watch page and playback, scale with the onboarding rate. Services used while watching, heartbeats, the scoreboard and ads, scale with concurrency. The December 2025 post maps each service to the signal it scales on: authentication, profile selection and the watch page on onboarding rate plus throughput, ads on concurrency plus throughput, and the home page on a mix of both, for a reason section 8 explains.
A surge like this lasts a few minutes. Can the system add servers fast enough to catch it?
08Scaling before the crowd arrives
8.1Why autoscaling is too slow
The usual answer to rising load is an autoscaler: watch a metric such as CPU use, and when it crosses a threshold, start more servers. It works well when load rises over tens of minutes. In Hotstar's 2019 re:Invent talk, as summarised by attendees, the arrival rate reached more than a million new users per minute, while the autoscaler took about 90 seconds to react and the application about a minute to boot. That talk also listed the cloud's own limits: errors when the provider had no more instances of a type to give, one instance type per scaling group, steps too coarse.
Hotstar had reached the same conclusion the year before. Its 2018 post put it in one rule, "No Auto-Scaling — Give yourself headroom", and an analogy: "Do you end up ordering food as guests start arriving for your party. I hope you don't." The strategy was "to estimate the peak concurrency, then scale up ahead of the event".
This program puts numbers on it. Ten million people are watching when Dhoni walks out. Three million join in each of the next two minutes, a million a minute for three more, then a trickle. Each arrival makes six backend calls as the app opens and the stream starts (an illustration); each viewer already watching makes about one a minute. One API server handles 1,000 requests a second. The autoscaler keeps servers 70% busy and orders more when they get busier, which arrive 150 seconds later. Our pre-scaled fleet sits on a fixed rung, sized beforehand for a forecast of 25 million viewers.
# Dhoni walks out at minute 0 and a push notification goes out.
# How many viewers join in each minute after that:
joining = [3_000_000, 3_000_000, 1_000_000, 1_000_000, 1_000_000] + [200_000] * 5
CALLS_PER_JOIN = 6 # API calls as the app opens and the stream starts
PER_SERVER = 1_000 # requests per second one API server handles
DELAY = 150 # seconds: autoscaler notices (~90) + server boots (~60)
LADDER = 1_250 # servers on the rung for a 25M-viewer forecast
def demand(viewers, join_per_min):
# joiners hit the start-up path; everyone watching makes a call a minute
return join_per_min / 60 * CALLS_PER_JOIN + viewers / 60
viewers = 10_000_000
auto = demand(viewers, 0) / PER_SERVER / 0.7 # 70% busy before the spike
orders = [] # (ready at second, servers)
lost = {"auto": 0, "ladder": 0}
used = {"auto": 0, "ladder": 0}
print("min joining/min demand req/s autoscaled served pre-scaled served")
for minute, j in enumerate(joining):
for half in range(2): # a decision every 30 s
t = minute * 60 + half * 30
viewers += j / 2
need = demand(viewers, j)
auto += sum(n for at, n in orders if at == t)
coming = sum(n for at, n in orders if at > t)
target = need / PER_SERVER / 0.7
if target > auto + coming: # order the shortfall
orders.append((t + DELAY, target - auto - coming))
for name, servers in (("auto", auto), ("ladder", LADDER)):
lost[name] += max(0, need - servers * PER_SERVER) * 30
used[name] += servers / 2
if half == 0:
print(f"{minute:3} {j:11,} {need:12,.0f} {auto:8,.0f} "
f"{min(1, auto * PER_SERVER / need):5.0%} {LADDER:8,} "
f"{min(1, LADDER * PER_SERVER / need):5.0%}")
for name in ("auto", "ladder"):
print(f"{name:6}: {lost[name] / 1e6:5.1f} M requests turned away, "
f"{used[name]:,.0f} server-minutes")min joining/min demand req/s autoscaled served pre-scaled served
0 3,000,000 491,667 238 48% 1,250 100%
1 3,000,000 541,667 238 44% 1,250 100%
2 1,000,000 375,000 238 63% 1,250 100%
3 1,000,000 391,667 738 100% 1,250 100%
4 1,000,000 408,333 810 100% 1,250 100%
5 200,000 338,333 810 100% 1,250 100%
6 200,000 341,667 810 100% 1,250 100%
7 200,000 345,000 810 100% 1,250 100%
8 200,000 348,333 810 100% 1,250 100%
9 200,000 351,667 810 100% 1,250 100%
auto : 39.0 M requests turned away, 6,560 server-minutes
ladder: 0.0 M requests turned away, 12,500 server-minutesLook at the first three rows of the autoscaled fleet. Demand roughly triples in the first minute, and for two and a half minutes the fleet stays the size it was before Dhoni walked out, serving between 44% and 63% of what's asked. By the time new servers arrive, in minute 3, the worst of the rush is over. Over those minutes it turns away 39 million requests, and each one is a person staring at a spinner, who will probably tap again and make it worse. Meanwhile the pre-scaled fleet serves everything, at the cost of nearly twice the server-minutes. It was sized for a forecast of 25 million and the audience stopped at about 20 million, so once the rush is over, well over half of it sits idle. That's the trade, paid in money to buy minutes.
8.2Ladders
Scaling ahead needs a plan for how big to be at each audience size. Hotstar called it a ladder: a table, worked out from load tests, that says for each level of concurrency how many servers each service needs. In 2018 the team "built pessimistic traffic models" for each of its three pillars, subscriptions, metadata and streaming, "basis which we came up with ladders that controlled server farms depending on the estimated concurrency". During a match, engineers stepped the fleet up a rung whenever the headroom above current concurrency fell below a threshold.
Hotstar's ladders evolved in three steps, each fixing the cost of the last.
- 2019: automation, in shadow first. Hand-stepped ladders worked, but "required human oversight and was costly". After moving to Kubernetes, the team built its own autoscaling engine "that took into account multiple variables that mattered to our system", and "ran auto-scaling in shadow" (deciding, but not acting) until it trusted it. That year it supported "2x more concurrency in 2019 with 10x less compute overall".
- 2023: the platform's own limits. Preparing for 50 million viewers at the 2023 World Cup, the team found that the Kubernetes control plane started returning errors when asked to add over 400 nodes at once, so pre-scaling now proceeds in automated steps of 100–300 nodes. A subnet ran out of IP addresses at about 350 nodes because each node reserved 35 addresses in advance. And services with more than 1,000 pods hit a Kubernetes limit: the older Endpoints API lists at most 1,000 addresses per service, and the team's API gateway relied on it, so those services were given bigger pods instead of more of them.
- 2025: scale on arrivals, not just viewers. As scaling became more automatic, incidents at lower concurrencies showed that concurrency alone missed the first-ball rush, which is why the onboarding rate from section 7 became a scaling signal. An old split between a business-as-usual mode and a live mode gave way to one model: high onboarding rate scales the arrival path, high concurrency scales playback, both scale everything, and both low leaves "an efficient, cost-effective base ladder" that autoscales on live traffic.
One surprise came at the end of a match. Concurrency drops and nobody new arrives, yet "homepage traffic spikes 200–300% within seconds", because everyone who was watching goes back to the home page at once, like a cinema at the interval. They named it the off-boarding rate and scale the home page on onboarding rate plus concurrency times a post-match coefficient. Hotstar has filed a patent on it.
How should the backend grow for a match?
- Pay only for what's used
- No forecast needed
- Minutes too slow for a first-ball rush
- Depends on the cloud having capacity at that moment
- Capacity is there before the crowd
- Cloud capacity reserved in advance
- Pays for idle servers
- Misses spikes driven by arrivals, not totals
- Arrival paths ready for the rush, playback paths for the peak
- One model for quiet days and finals
- Needs forecasts per service and per match
- More signals to get right
Notice the progression: Hotstar started by rejecting autoscaling, then rebuilt it on signals that matched its traffic. Even so, the cloud's own limits never went away, so the big steps still happen before the match: Hotstar's 2017 checklist included having the cloud load balancers specially "warmed up" and making sure the right instance types would be available.
Even a good ladder can be wrong. If the audience overshoots the forecast, what gives way?
09Panic mode: protect the video
9.1Jettison what you can
In 2019, ahead of a World Cup game where "breaching the 30M barrier was not un-thinkable", Hotstar asked its team what it would take to serve 50 million. They went through every service: its rated limits, and "what would be the tsunami, that took the service down". Then they agreed an order in which to turn services off: "a sequencing of which service to jettison (turn off), as the concurrency marched upwards, so that customers could keep coming in and the video would keep playing." Hotstar compared it to Apollo 13, where ground control kept only what the crew needed to get home.
A system with these switches built in has a panic mode: a set of pre-planned levers that shed non-essential work under overload. Hotstar's 2019 talk named the first things to go, recommendations and personalisation, to protect video streaming and payments. This Scene plays it out for Ritu's screen as the audience climbs past the forecast.
That order is illustrative; JioHotstar's current sequence isn't published. What Hotstar has published are the individual levers: recommendations and personalisation (2019), refresh rates for "scorecards, Watch-more, live feeds, and key moments" (2024), and bitrate (2017). Hotstar's 2017 post adds a distinction that's easy to miss. For Tier 1 actions, login, subscriptions and watching, the lever is a positive panic: the action appears to succeed and the backend work is deferred, so a viewer whose subscription check can't complete right now is let in and checked later. For Tier 2 actions, such as account settings and logout, it's a negative panic: show a polite failure message and move on.
9.2Clients that back off
Panic mode's other half lives in the app. When a backend struggles, the worst thing tens of millions of clients can do is retry immediately, all at once; every failure then doubles the load that caused it. Hotstar's 2018 post: "Retries will make the problems worse. Clients must be smart about inferring when things don't look right, and add 'jitter' to the requests they make." Its apps understand panic switches from the server, which tell them "to ease off momentarily, either exponential back-off or sometimes a custom back-off depending on the interaction". Exponential back-off means waiting twice as long after each failure; jitter means adding a random amount to each wait, so that a million clients that failed at the same instant don't all retry at the same instant too.
Netflix found the same thing with live events: in its 2025 account, viewers restarting streams after interruptions as short as 30 seconds caused a "10x increase in traffic load", and the fixes included server-guided backoff for devices and prioritised shedding of traffic at its cloud gateway.
So the video is protected, and the ads are on the protected list because they pay for it. How do they get into the stream?
10Ads in the break between overs
10.1Every viewer reaches the break at once
At the end of the over, the broadcast cuts to an ad break. Ritu should see ads chosen for her, or at least for people like her, stitched in so the picture never stutters, and then be back in time for the first ball of the next over.
There are two ways to put an ad into a stream. With client-side ad insertion (CSAI), the stream carries a marker, the player sees it, pauses the match, asks an ad server for ads, plays them with a second player, then switches back. With server-side ad insertion (SSAI), the ads are stitched into the stream itself: the playlist Ritu's player reads lists ad segments where the match segments would be, and the player just keeps playing.
Both methods start from a marker in the broadcast. When the director goes to a break, the production system inserts an SCTE-35 cue, a standard message meaning "an ad break starts here, and lasts this long". In HLS it can appear as a date-range tag; this one is from the HLS specification:
#EXT-X-DATERANGE:ID="splice-6FFFFFF0",\
START-DATE="2014-03-05T11:15:00Z",PLANNED-DURATION=59.993,\
SCTE35-OUT=0xFC002F000000000000FF000014056FFFFFF000E081622DCAFF0\
00052636200000000000A0008029896F50000008700000000SCTE35-OUT marks the start of a 60-second break and a matching SCTE35-IN its end. In the low-latency playlist from section 4, the ad segments (midRoll273.mp4) follow an EXT-X-DISCONTINUITY tag, which the specification requires whenever the encoding or timestamps change, as they do when one piece of video ends and an unrelated one begins.
Now picture the cue arriving for 50 million players. With CSAI, every one of them calls the ad server in the same second, a thundering herd from section 5 aimed at a service that can't be cached, because every answer is personalised. With SSAI, the decision moves to the server and can be made once per group of similar viewers. JioHotstar's 2026 post on its ad system says SSAI at its scale "often happens at a Cohort level where we are grouping similar users", and the stitched ad segments are ordinary video files that the CDN caches and collapses like any other.
10.2Choosing the ads in under 100 milliseconds
Choosing the ads is a busy, spiky job of its own. JioHotstar's July 2026 post on it shows its traffic peaking at every ad break, and dropping "more predictable and lower during the innings break" (20:40 to 21:10 in its graph). For each break it picks "2–3 ads out of thousands for a 30-second ad 'pod'", all in under 100 ms "even during massive concurrency spikes like the final over of an IPL match". It filters ads by targeting, checks frequency caps so nobody sees the same ad too often, uses pacing algorithms (it names PID and SHALE) so a campaign's impressions are spread over its whole run, and fills the break through priority tiers, with leftover time going to in-house promos. Ads sold through open programmatic auctions are the risky part: their fill rate "can plummet to 1–2%", against 80–100% for pre-negotiated deals.
That post also describes a third method, server-guided ad insertion (SGAI): the server tells the player what to play and when, and the player does the final insertion, which allows one-to-one personalisation "at a live event scale where we do not have the cue markers in advance".
How should ads get into a live stream?
- Fully personal ads
- Simple servers
- Every player calls the ad server in the same second
- Switching players can stutter; easy to block
- One decision per cohort, not per viewer
- Ad segments cached on the CDN
- Seamless on every device
- Personal only to the cohort
- Measuring views needs extra reporting
- One-to-one personalisation
- Needs capable players
- Per-viewer decisions at the break
For live sport, SSAI fits the traffic: the break arrives for everyone at the same moment, and turning tens of millions of personalised requests into a few thousand cohort decisions is the same move as request collapsing in section 5. JioHotstar uses SSAI at cohort level and has described SGAI for more personal targeting; its watchdog validates SSAI markers on every live stream because a broken cue means a black screen in the break.
11The scoreboard, a system of its own
11.1Cacheable, and separate from the video
The score in the corner of Ritu's screen doesn't come from the video. It's data: runs, wickets, overs, the batters' scores, and a list of key moments (wickets, boundaries, milestones) that she can tap to jump back to. Where JioHotstar's ball-by-ball data comes from isn't published; on any cricket broadcast it starts with scorers at the ground entering each ball as it happens.

That data is small, but everybody wants it, constantly. An obvious way to serve it is a score API the app polls every few seconds. At 50 million viewers polling every 5 seconds, that's 10 million requests a second, nearly half as many as the video itself, for a few hundred bytes each. Hotstar's 2024 post describes the answer: the scorecard, the concurrency count and key moments are "highly cacheable", identical for everyone, so they were moved to a separate CDN domain with "leaner but optimal security checks and rate controls". Serving them from CDN caches with a short lifetime turns 10 million requests a second into a trickle at the score service. Keeping them on their own domain isolates them, so a misconfigured rule for the scoreboard can't break video, and the refresh rate becomes one of the panic levers from section 9.
Not everything around the match is cacheable. Hotstar's social feed under the player, launched for IPL 2019, carried content tied to each moment of play, where, the team wrote, "caching at any level is also not possible". For that they built a push service: phones keep a persistent connection open (they chose the MQTT protocol and EMQX broker after benchmarks), designed to "broadcast content to 50 Million simultaneously connected clients". That same system carried live emoji reactions, about 5 billion of them during the 2019 World Cup, aggregated on the server so phones receive the popular ones instead of every tap.
Should the scoreboard be pushed or polled?
- Updates arrive the moment they're published
- No wasted requests
- Tens of millions of open connections to run
- A broker fan-out on every ball
- The CDN absorbs nearly all the load
- Refresh rate is a dial for panic mode
- Up to one polling interval stale
- Requests even when nothing changed
Since the scorecard is the same for everyone, the CDN can answer it, and staleness of a few seconds is harmless because the video is behind anyway, as the next part shows. Push is worth its cost for per-moment social features that can't be cached; Hotstar built both and used each where it fitted.
11.2Don't spoil the ball
Section 4 left the scoreboard one more problem. Ritu's video is about 15 seconds behind the stadium. But the score service isn't: the scorer enters a six as it lands.
Ritu's video is 15 seconds behind the stadium. Her app polls a fresh scorecard every 5 seconds. A batter hits a six. What does she see?
One fix is to give every score event the stadium time at which it happened, and let the player decide when to show it. And the player knows the stadium time of the frame on screen: HLS's EXT-X-PROGRAM-DATE-TIME tag ties the first frame of a segment to an absolute date and time. So the app can hold each score update until the video reaches it, and show the six on the scoreboard as the ball crosses the rope on Ritu's screen. Whether JioHotstar synchronises its scorecard this way isn't published, but any design has to solve this, because a scoreboard that's more up to date than the picture spoils the match.
12Who may watch: entitlement and DRM
12.1Checked once, enforced everywhere
Before Ritu's player gets a playlist URL, the playback service decides whether she may watch: is her subscription tier one that includes this match, is she in a country where JioHotstar holds the rights, is she under her plan's device limit? That's an entitlement check, and it sits on the arrival path from section 7, so it surges with the onboarding rate.
Then there's protecting the video itself. Every segment is encrypted, and a phone can only decrypt them with keys from a licence server, through the phone's built-in DRM (digital rights management) system: Widevine on Android and in Chrome, FairPlay on Apple devices, PlayReady on Windows and many TVs. Each new viewer's player asks the licence server for a licence as it starts (and again only if the keys are rotated during the stream), and then decrypts every segment locally.
For live, the design rule is to do the expensive checks once per session and nothing per segment. The CDN can't call a database for each of 25 million requests a second, so the playback service hands the player URLs carrying a signed token that the edge can verify with a key, without asking anyone. So the licence server and entitlement service see each viewer once, at the start, which makes them onboarding-rate services, scaled and pre-warmed like the rest of the arrival path.
And when they're overloaded, the 2017 rule applies: login, subscription and watch are Tier 1 actions, and a positive panic lets the viewer in and settles the check afterwards. Letting a few people watch an over they weren't entitled to is far cheaper than turning millions of paying viewers away from a final. JioHotstar hasn't published the details of these services, though its watchdog lists "a DRM mismatch" among the failures that can leave a TV spinning while the manifest looks perfect.
13The whole system
13.1Every box, and why it's there
| Component | What it does | Added because |
|---|---|---|
| Duplicate pipelines | Two clock-locked chains from stadium to origin | One failure would freeze every viewer (§3) |
| Short segments, hold-back | 2–4 s pieces, a few held in reserve | Long segments put viewers half a minute behind (§4) |
| Collapsing at edge and shield | One upstream fetch per new object per tier | Millions ask for each new segment at once (§5) |
| Multi-CDN + ISP caches | Capacity from several providers, close to viewers | No single CDN has 60–80 Tbps in India (§6) |
| CDN steering | Per-cohort scores from player telemetry | CDNs struggle city by city, network by network (§6) |
| Ladders on concurrency and onboarding rate | Scale before the crowd arrives | Autoscaling is minutes too slow (§7, §8) |
| Panic switches, client back-off | Shed non-video work; spread out retries | Forecasts can be wrong (§9) |
| SSAI per cohort | Ads stitched into the playlist | Every player reaches the break at once (§10) |
| Score service on its own CDN domain | Cacheable scorecard and key moments | Polling from every viewer, isolated from video (§11) |
| Entitlement + DRM once per session | Signed URLs, one licence per start | The CDN can't check a database per segment (§12) |
13.2From top to bottom
| Level | The choice | Data structure or algorithm |
|---|---|---|
| System | Protect the video above everything | An ordered list of panic levers, each with a switch |
| Ingest | Two pipelines, origin picks | Segments numbered and timestamped identically; first-valid selection |
| Packaging | Sliding live playlist | Media sequence numbers, a window of recent segments; LL-HLS parts and preload hints |
| Player | Hold back from the live edge | A buffer capped by the live edge; hold-back of about three target durations |
| CDN | Collapse misses at every tier | A per-object waiting list behind one in-flight fetch |
| Delivery | Steer each cohort to a CDN | Cohort key (ASN, state, city, user type); weighted score; EWMA with α ≈ 0.6; capped weight changes and a 5% floor |
| Backend | Ladders | A table from concurrency and onboarding rate to servers per service |
| Ads | Stitch per cohort | SCTE-35 cues; discontinuity tags; pacing (PID, SHALE) and tiered fill |
| Score | Cache and synchronise | Short-lived cached JSON; events stamped with stadium time, shown at the video's program date-time |
14What goes wrong, and what it costs
14.1Failures this design has to survive
| What happens | What Ritu sees | What the design does |
|---|---|---|
| One encoding pipeline fails | Nothing | The origin serves segments from the other pipeline |
| A regional feed never reaches the CDN | Nothing, if caught early | Validators flag manifest 404s before the match is published |
| Millions ask for a new segment at once | Up to a fraction of a second's wait | Collapsing at edge and shield; one origin fetch per segment |
| One CDN degrades in one city | A brief stall, then normal | The cohort's weight shifts to better CDNs, a capped amount per minute |
| Her cell tower is crowded | Softer picture | Her player drops a rendition; the service may cap the top rung for everyone |
| A famous batter walks out | Normal start | Arrival services were scaled up their ladders before the match |
| The audience overshoots the forecast | Plain home page, slower scorecard | Panic levers shed features in a planned order; the video keeps playing |
| A backend error during the surge | A short wait, then success | The app backs off with jitter instead of retrying at once |
| The ad stitcher gets a bad cue | Without validation, a black break | SSAI markers are checked continuously on every stream |
14.2The tradeoffs, in one table
| Decision | Chosen | Given up | Why it was worth it |
|---|---|---|---|
| Ingest | Two pipelines | Half the encoding cost | One failure mustn't freeze the country |
| Latency | 2–4 s segments, a few held back | The last few seconds of delay | A tiny buffer stalls on mobile networks |
| Origin protection | Collapsing at edge and shield | A little wait on a miss | Origin load stops growing with the audience |
| Delivery | Many CDNs, steered | One simple provider | No single provider has the bandwidth |
| Scaling | Ladders ahead of the match | Money for idle servers | Autoscaling can't catch a first-ball rush |
| Overload | Panic levers in a set order | Features during peaks | The video must play |
| Ads | Server-side, per cohort | One-to-one targeting | No ad-server herd at every break |
| Scoreboard | Cached polling | A few seconds of freshness | The CDN carries it; the video is behind anyway |
15Summary
- Live breaks the on-demand design in four places: the pipeline is a single point of failure, the stream lags, the CDN gets a herd for every new segment, and the backend gets a crowd faster than it can grow.
- The pipeline runs twice, from two contribution paths to two clock-locked encoders, so the origin can pick a good copy of every segment.
- Latency is mostly the player's own safety margin: about three target durations behind the live edge, the only buffer a live player can have.
- Shorter segments and LL-HLS parts trade stalls for seconds, and the trade gets steep: below a few seconds of buffer, a mobile viewer freezes several times a match.
- Every new segment is a thundering herd, and collapsing at the edge and at a shield brings the origin's load down to about one request per segment per rendition.
- No single CDN carries an Indian final, so the stream is spread over several CDNs and caches inside ISPs, steered per cohort by smoothed scores from the players' own telemetry.
- The surge is about arrivals, not viewers: the onboarding rate drives login, home page and playback load, and a famous batter's walk to the crease can multiply it in seconds.
- Autoscaling is minutes too slow, so the backend climbs ladders sized from load tests before the crowd arrives, on concurrency and onboarding rate.
- Panic mode protects the video by switching off features in a planned order and telling clients to back off with jitter.
- Ads are stitched on the server per cohort, so the break doesn't send every player to the ad server in the same second.
- The scoreboard is cacheable data on its own CDN domain, and has to be held back to match the video, or it spoils the ball.
16Build this
A live stream and its herd, on one laptop.
- Use
ffmpegto turn a looping video file into a live HLS stream:-f hls -hls_time 2 -hls_list_size 5 -hls_flags delete_segments. Serve the directory with any static web server and play it in a browser with hls.js. Watch the playlist slide. - Put an nginx cache in front of the server, with
proxy_cache_lock on(nginx's request collapsing). Write a small load generator that, every time the playlist gains a segment, requests it from 2,000 concurrent connections. Count origin requests with the lock off and on. - Set the playlist's cache lifetime to 1, 2 and 4 seconds, and measure how far behind the source each player falls.
- Change
-hls_timeto 6, 4, 2 and 1. Measure the delay with a clock burned into the video (ffmpeg'sdrawtextwith the time) against the real clock, and use your system's network link conditioner to add two-second dropouts. Compare stall counts with the TryIt in section 4. - Add a
/scoreendpoint that returns an event with a timestamp every few seconds. Make the page show each event only when the video'sEXT-X-PROGRAM-DATE-TIMEreaches it.
17Interview questions
beginnerWhy is a live stream behind the stadium, and why not make the segments tiny?›
The delay adds up from encoding, waiting for each segment to finish, delivery, and above all the player's hold-back: HLS players start about three target durations behind the newest segment, because in live video the distance from the live edge is the only buffer the player can have. Tiny segments shrink that buffer, so a phone hitting a few seconds of dead signal freezes. They also cost compression, because every segment starts with a keyframe, and they multiply requests. Most services settle on a few seconds of segment and a delay close to broadcast TV.
beginnerWhat is request collapsing, and why does live video need it?›
When many requests for the same uncached object arrive together, the cache sends one request upstream and makes the others wait for its answer. Live video needs it because every new segment is requested by millions of players within a second or two of appearing, when no cache has it yet. Without collapsing, each edge server would forward thousands of misses to the origin for every segment; with collapsing at the edges and at a shield tier, the origin sees about one request per segment per rendition.
intermediateWhy doesn't autoscaling handle a cricket match, and what do you do instead?›
Arrivals come in minutes: millions in the first two minutes after the first ball or a famous batter's walk-out, while an autoscaler takes a minute or more to notice and new servers another minute to boot, if the cloud has them. So you scale ahead: ladders, worked out from load tests, that set each service's size for the expected concurrency and arrival rate, applied before the match and stepped up as headroom shrinks. Hotstar also scales arrival-path services on the onboarding rate, and plans an order of features to switch off if the audience overshoots.
intermediateWhy server-side ad insertion for live sport?›
Every viewer reaches the ad break at the same moment. With client-side insertion, every player calls the ad server at once with a personalised request that can't be cached, a herd aimed at a service that must answer in milliseconds. Server-side insertion makes one decision per cohort of similar viewers and stitches ordinary video segments into that cohort's playlist, so the CDN caches the ads like any other segment and every device plays them seamlessly. In exchange, targeting is per cohort, not per person, which server-guided insertion tries to recover.
deepDesign how a multi-CDN live service decides which CDN each viewer uses.›
Group viewers into cohorts by network and place (JioHotstar uses ASN, country, state, city and user type), because a CDN can be fine nationally and failing for one operator in one city. Score each CDN per cohort from the players' own telemetry, weighting playback failures above rebuffering above round-trip time, and smooth the scores (an EWMA with α ≈ 0.6) so a ten-second blip doesn't move a city. Then turn scores into traffic weights with brakes: throttle a CDN before its capacity limit, during surges aim for all CDNs to fill together, cap the change per minute, and keep every working CDN above a floor so its caches stay warm for failover. Hard constraints come first: never send a cohort to a CDN with no presence on its network.
deepThe score on screen keeps spoiling the video. Why, and how do you fix it?›
The video is 10–20 seconds behind the stadium because of encoding, segments and the player's hold-back, but the scorecard is published as each ball happens. Stamp every score event with the stadium time it happened, and have the app show it only when the video on screen reaches that time, which the player knows from the program date-time of each segment. Notifications have the same problem, and need either a delay or a choice to accept the spoiler.
18Go deeper
Why can't a live player keep a minute of video in its buffer, as an on-demand player can?›
Because the next minute of the match hasn't happened yet. A live player can only be as far ahead as the newest segment, so the hold-back it chooses is the whole of its buffer, and every second of buffer is a second of delay.
Edges collapse their misses, and a shield sits in front of the origin but doesn't collapse. How many origin requests per segment, with 50 edges?›
About 50. All the edges' first fetches reach the shield within its own fetch time, so each one misses and is forwarded. So the shield has to collapse too to get the origin down to one request per segment.
Concurrency is falling and no one is joining, yet the home page service is overloaded. What happened?›
A match, or an innings, just ended. Everyone who was watching returns to the home page at once: Hotstar calls it the off-boarding rate and scales the home page on concurrency times a post-match coefficient.
Ladders and headroom instead of autoscaling, clients that back off with jitter, cutting latency from 55 s to 15–20 s behind broadcast, and the jettison order for 50 million viewers.
The Kohli walk-in that took the login API from 10K to 125K TPS, progressive degradation of bitrates, and positive and negative panics.
The first-ball rush, ladders on arrivals as well as viewers, and the post-match storm on the home page.
The QoS Routing Manager: cohorts, scores from player telemetry, EWMA smoothing and capacity steering across CDNs.
The live playlist rules, the three-target-duration hold-back, partial segments, blocking reload, SCTE-35 date ranges and content steering.
Live from a VOD company: 2-second segments, two pipelines with an origin that picks, millisecond caching and retry storms. Read with the Open Connect overview for the on-demand contrast.
19Related chapters
Renditions, segments, manifests, CDN caches and the bitrate algorithm that this chapter builds on. Chapter 51.
Stampedes and request coalescing inside one service, the small version of section 5. Chapter 25.
Why headroom matters and how utilisation turns into waiting. Chapter 42.
The other end of the latency scale: real-time video over UDP for a meeting, not HTTP segments for a country. Chapter 53.
Load shedding, back-off and graceful degradation in general. Chapter 40.
