How would I design Netflix in 45 Minutes: A System Design Walkthrough
How do you stream 100M+ hours of video a day to 260M+ subscribers worldwide without buffering? The answer isn't one magic server. It's a layered system of micro-frontends, a global control plane, a massive content-delivery network pushed to the edge of the internet, and some seriously smart adaptive streaming. Let's design it from scratch.
Series: System Design Walkthroughs (1 article)
- 1How would I design Netflix in 45 Minutes: A System Design Walkthrough This article
Why Netflix is a hard question
When the interviewer said "design Netflix," my first instinct was to start drawing databases. That's the trap. Netflix is really two systems stuck together, and if you treat it as one you'll get lost.
There's the part that handles browsing, search, login, recommendations, and figuring out whether you're allowed to watch something. That's mostly normal request-response traffic. Then there's the part that actually moves the video, which is a completely different beast measured in terabits per second.
The thing that made the interview go well was separating those two early and being clear about which one I was talking about at any given moment. I'll call them the control plane and the data plane.
Figuring out what we're building
I spent the first few minutes pinning down scope, mostly by asking the interviewer and getting nods. We landed on:
- People can browse and search a catalog of movies and shows
- People can stream video on demand on phones, TVs, and the web
- The system recommends things based on what you've watched
- It remembers where you left off so you can resume
- Video quality adapts to your network
I also said what I was leaving out, because that matters as much as what you include. No billing, no DRM internals, no content production pipeline, no live streaming. Just video on demand. The interviewer seemed happy I drew that line myself instead of trying to boil the ocean.
The numbers that actually shaped the design
Before touching architecture I did some rough math, and this is where the design really got decided.
- Storage first. Say there are around ten thousand titles. Each one gets encoded into something like a thousand different files once you account for every codec, resolution, and audio track. A single movie in all its forms can run into hundreds of gigabytes. Add it all up and you're in petabyte territory.
- Then bandwidth, which is the scary one. If roughly a hundred million streams are going at peak and each averages about 5 megabits per second, you're looking at something on the order of 500 terabits per second leaving the system. I said that number out loud and then said the obvious conclusion: there is no single building on earth that serves that. The video has to live close to people, not in one place.
Those two numbers, petabytes of storage and 500 terabits of egress, decided almost everything that followed.
The shape of the system
Here's roughly what I drew on the whiteboard.
- The client, whatever it is, talks to the control plane first. That's where it figures out what to show you and, when you hit play, where to actually get the video from. The control plane is a set of services behind a gateway: login, browse and search, recommendations, and a playback authorization service that acts as the gatekeeper for streaming.
- Those services sit on top of caches and databases. The important move is that when you press play, the control plane doesn't hand you video. It hands you a signed URL pointing at a nearby edge server, and your player goes and talks to that server directly.
So the whole thing in one line: the client asks the control plane what to play and where to get it, then pulls the actual bytes straight from the closest edge box. The control plane never touches the video itself. Once I said that, the interviewer visibly relaxed, and I knew I was on the right track.
The control plane
I kept this part deliberately boring, because it mostly is. It's a collection of stateless services behind an API gateway. The gateway is the front door. It routes requests, checks auth, and is also the first place you can push back when traffic gets ugly. Behind it:
- A login service that validates tokens and known devices
- A browse and search service that serves the home page rows and search results, cached heavily because the same rows get requested constantly
- A recommendation service that does the heavy machine learning work offline, overnight, and just serves precomputed results fast at request time
- A playback authorization service, which is the one that matters for streaming. It checks whether you're entitled to watch, whether you've hit your plan's concurrent stream limit, and then picks the best edge server for you and signs a URL
For storage I didn't try to use one database for everything. Viewing history and playback progress are write heavy and need to survive a region going down, so I'd put those in Cassandra and accept eventual consistency. Hot metadata like title details goes in a memory cache in front of everything so reads come back in under a millisecond. Search gets its own search index. Recommendations are precomputed and read out of a key-value store.
The theme here is cache a lot, store the durable stuff carefully, and lean toward staying available rather than being perfectly consistent. If your recommendations are a few minutes stale nobody notices. If playback breaks, people cancel.
How the video actually moves
This is the part I almost got wrong, and it's the part the interviewer cared about most.
- Netflix does not stream video out of AWS. AWS runs the brains. The video comes off Netflix's own network of servers called Open Connect (I read about it recently a paper, pretty impressive if you ask me). These are physical boxes they build and then place inside internet providers' networks and at the big internet exchange points. So when you watch something, the bytes often never even leave your own ISP. The box is sitting right there in their data center.
- A few reasons this is clever. The path to you is as short as it physically gets, so it's fast and reliable. The ISPs are happy to rack these boxes for free because it saves them a fortune on the traffic they'd otherwise have to carry. And because Netflix builds the hardware, they can tune the whole thing for one job: reading video off disk in order and shoving it down the wire.
- The other trick I brought up is that Netflix doesn't wait for you to ask. It predicts what's going to be popular tomorrow and copies those files onto the edge boxes overnight when the network is quiet. So by the time the new season drops and everyone hits play at once, the file is already sitting a few feet from them. That's a very different model from a normal CDN that only caches something after the first person requests it.
- Then there's the thing that makes streaming feel smooth, which is adaptive bitrate. The video isn't one file. During encoding it gets chopped into small segments, a few seconds each, and every segment is encoded at a bunch of different qualities from very low up to 4K. A small manifest file lists all the options. Your player grabs the manifest and then makes its own decisions, segment by segment. It starts with a low quality one so the video begins almost instantly, then as it measures your actual bandwidth and how full its buffer is, it climbs up to higher quality. If your wifi stumbles, it quietly steps back down instead of freezing. That's why Netflix starts fast and sags gracefully instead of spinning.
- I also mentioned the encoding pipeline since it's where that pile of files comes from. The source master gets validated, split into chunks, and encoded in parallel into all those bitrates and codecs on AWS, then stored and pushed out to the edge boxes. Netflix goes further and analyzes each title, even each shot, to decide how many bits it needs. A dark, still scene needs far fewer bits than an explosion, so spending them unevenly saves a huge amount of bandwidth for the same picture quality.
Scaling it
- For the control plane, scaling is the usual story. The services are stateless, so you add more copies behind the gateway and autoscale on load. Run it active-active across multiple AWS regions so that if one region has a bad day you can shift traffic to another one quickly. Netflix actually practices this by evacuating regions on purpose. The caches are tiered within and across regions so a miss is rare and failover is quick, and Cassandra spreads writes across nodes and regions.
- For the data plane, you scale by adding more edge boxes, not by making the origin bigger. Popular content gets copied onto more boxes closer to more people, while the long tail of rarely watched stuff lives on fewer, bigger boxes. Requests go to the nearest box, fall back to one at an exchange point, and only in rare cases reach all the way back to origin.
- The last point I made on scaling is that resilience is really what lets you scale. Netflix is famous for deliberately killing servers in production with tools like Chaos Monkey. It sounds reckless, but it forces every service to be built to survive failure, and that's the reason the whole thing holds together at a hundred million streams.
Keeping it from falling over
The interviewer pushed on how you protect a system this size, so I laid out a few layers.
- The gateway does rate limiting per user and per device, which stops a buggy client that's hammering retries from taking a service down with it.
- The playback authorization service enforces the concurrent stream limit, which is both a product rule and a quiet way to cap load.
- When a service starts getting slow, adaptive load shedding kicks in and starts dropping the least important traffic first so the critical playback path keeps working.
- Circuit breakers wrap calls to things like recommendations, so if that service is struggling the gateway just serves a generic cached row instead of failing the whole page.
- Finally, there's backpressure between services so a slow component downstream doesn't pile up and topple everything upstream.
The rule I kept coming back to is that a slightly worse experience always beats an outage. You drop the fancy personalized row long before you ever drop someone's video.
Walking through a single play
To tie it together I traced what happens when you press play.
- You click a title.
- The client goes through the gateway to the playback authorization service, which checks you're allowed, checks you're under your stream limit, and figures out which edge boxes are closest to you.
- It sends back a manifest with signed URLs for the best couple of boxes for your connection.
- Your player reads the manifest, grabs a low quality first segment so playback starts right away, then streams the rest straight from the local edge box, moving the quality up and down as your network changes.
- Every so often it reports your progress back to the control plane so that resume works no matter what device you pick up next.
The thing worth noticing is that after you get that manifest, AWS is out of the picture. All the heavy lifting is the edge box talking to your player. That split is the single idea the whole design hangs on.
What I'd tell someone walking into this interview
If I could hand my past self a note, it would say: split the control plane from the data plane in the first two minutes, do the bandwidth math early because it makes every later decision for you, remember that the video comes off a custom edge network and not AWS, lean toward staying available over being consistent on the playback path, and have a real answer ready for how you throttle and shed load when things go wrong. (AND STILL AWS HAS NETFLIX AS ONE OF THEIR BIGGEST CUSTOMERS!)
That's day one - next I'll pull apart another one. If you'd have designed any of this differently, especially the edge caching or the throttling, I'd genuinely like to hear it, because the fun of these questions is that there's never just one answer.
A quick note on reading further: Netflix's own tech blog has good deep dives on Open Connect and their encoding work. 10/10 recommend this read!
Series: System Design Walkthroughs (1 article)
- 1How would I design Netflix in 45 Minutes: A System Design Walkthrough This article
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article