SRE capacity planning is the work of checking how much demand a service can handle while still meeting its response-time and correctness goals, and making sure enough capacity is in place before that demand arrives. Overload happens when more work arrives than the service can complete. Good planning covers expected busy periods and the failures that are likely during them, such as losing a computer in the middle of a sale.
TicketDesk, our fictional booking app, needs enough capacity for people to buy tickets when demand rises. It also needs a way to limit incoming work when extra capacity cannot arrive in time, because at some point every service hits a limit.
The arithmetic in this lesson is deliberately simple. If you can divide and round up, you can follow all of it.

- 01Estimate busy demand
- 02Plan for a lost copy
- 03Limit waiting and retries
What demand should you plan for?
You should plan for the busiest realistic mix of operations the service will face, not the average. A quiet afternoon of browsing does not resemble the first minute of a popular ticket sale, and saving a reservation requires more work than displaying a page that was already prepared.
Test a realistic mix of operations, including the busy bursts. Check where demand concentrates. One event, or one part of the storage system, can be overloaded while other parts have room to spare, and an overall average will hide that.
Adding capacity takes time. Autoscaling means automatically increasing or decreasing the amount of running software in response to demand. If it takes ten minutes to add usable capacity, autoscaling will miss a sale whose peak lasts five minutes. Forecast far enough ahead to allow preparation, then check whether actual demand matched the assumptions so the next forecast is better.
How many running copies does TicketDesk need?
TicketDesk needs six instances in this example: five to carry the forecast peak and one spare to cover losing a copy. An instance is one running copy of the software handling requests. A request is an attempt asking it to do something, such as book a ticket. A representative test shows one instance can handle 200 requests per second within the chosen response-time and correctness limits.
Peak demand is forecast at 900 requests per second. The team wants to handle that demand even after losing one instance:
Required surviving instances = ceiling(900 / 200) = 5
Initial fleet = 5 + 1 = 6 instances
Capacity after one loss = 5 × 200 = 1,000 requests/second
Dividing 900 by 200 gives 4.5. You cannot run half an instance, so round up to five surviving copies. Then add the one copy that may fail. A fleet is the group of instances running the service.
Six is the minimum under this simplified model, before any extra margin. It assumes demand is spread evenly and that each surviving instance keeps its tested capacity. The services TicketDesk relies on, called dependencies, must support the total work too. If confirmations are limited to 700 per second, adding more booking instances does not remove that limit. Losing several copies together needs a different estimate. The testing lesson uses this six-instance model in a controlled failure exercise.
Can a queue absorb all the extra work?
A queue can absorb a short burst, but it cannot absorb a sustained overload, because if work keeps arriving faster than it finishes the waiting list grows without limit. A queue is a waiting list of tasks. It buys time while the service catches up; it does not add capacity.
Suppose confirmation work arrives at 1,000 items per second and finishes at 800:
Backlog growth = 1,000 - 800 = 200 items/second
Growth in one minute = 200 × 60 = 12,000 items
With 12,000 items ahead of it and 800 completing each second, a new item faces roughly 15 seconds of waiting before its own processing even starts. Actual scheduling and work size can change that estimate, but it already warns you that customers will notice the delay.
Set limits on waiting time and queue size. Decide in advance which work can be declined or deferred, and make the result visible to the customer. Once a booking is accepted, quietly dropping it breaks the promise you made.
How do repeat attempts make overload worse?
Repeat attempts make overload worse because each retry adds demand at exactly the moment the service is least able to handle it. A customer waits too long and tries again. Software can also retry automatically. Either way, the struggling service now has more to do.
A dependency that slows down keeps more work unfinished. More callers retry, pressure increases, and the failure spreads to parts of the system that were healthy a minute ago. This is one way a cascading failure develops: trouble in one part causes further trouble elsewhere.
The defences are simple to state. Limit the number of retries and the total waiting time. Backoff means waiting longer between repeat attempts. Jitter adds random variation to those waits so that many callers do not all return at the same instant. A retry budget caps the additional work that retries are allowed to create.
Idempotency means repeating the same operation does not create an extra side effect, such as a second booking. It keeps retries correct, which matters a great deal, but repeated attempts still cost processing time. Correct and cheap are different properties.
What can the service temporarily give up?
A service can temporarily give up optional features to protect its essential ones. TicketDesk could turn off suggested events during a sale to preserve resources for bookings. This is a degraded mode: the service offers less functionality while protecting the task customers came for.
Agree on those choices before overload happens, not during it. Record what customers will receive and how the team will measure it. Reporting a successful reservation that was never saved is not a reduced service; it is a hidden failure.
Also check what happens after work is declined. If every caller immediately retries, the rejected work comes straight back and recreates the pressure. The service's limits and the callers' behaviour need to fit together.
Why can more workers or redirected traffic fail to help?
More workers or redirected traffic fail to help when the real constraint is somewhere else, such as a shared dependency that is already saturated. Workers are pieces of software that perform tasks. Measurements from the golden signals help you locate the limit before you change capacity. Adding workers helps when their processing capacity is the constraint. If a shared dependency is already overloaded, more workers simply send it even more work.
Likewise, moving requests to another region helps only if that destination has spare usable capacity. A bottleneck is the part that limits the overall rate. Identify it before assuming that every capacity increase will help; otherwise you may spend money making the problem worse.
Exercise: “Can't we just autoscale?”
Use the example with 1,000 arrivals and 800 completions per second. Explain when extra capacity would arrive and whether it would relieve the bottleneck.
While you investigate, reduce optional work, limit repeat attempts, and avoid admitting work that cannot meet its deadline. Accepted bookings still need a safe plan for completion.
Your reasoning should distinguish immediate harm reduction from the lasting capacity fix. In an interview, state your workload assumptions and explain what test would show that the service can handle launch demand after the expected failure.
Quick answers
How do you calculate how many servers you need?
Divide the forecast peak demand by the tested capacity of one instance and round up, then add the number of instances you expect to lose. For 900 requests per second at 200 per instance, that is five surviving instances plus one spare, so six in total.
What is a cascading failure?
A cascading failure is one where trouble in one part of a system causes further trouble elsewhere. A slow dependency causes callers to retry, the retries add load, and the overload spreads to components that were healthy.
Why does autoscaling not always prevent overload?
Autoscaling takes time to add usable capacity, so a burst that peaks faster than that delay will overwhelm the service before new instances are ready. It also cannot help when the bottleneck is a shared dependency that the new instances would only load further.