This SRE interview preparation capstone brings together everything in the track: reliability goals, capacity, safe changes, incident response, and data recovery, in one guided design exercise. You will plan a busy launch for TicketDesk, our fictional booking service, and then explain how you would respond when the plan starts to fail. It mirrors the shape of a real reliability design interview, where you are given a scenario and asked to reason through it out loud.

A capstone is a final exercise that combines earlier lessons. If a term feels unfamiliar, go back to its lesson using the track navigation; that is what it is there for. No previous workplace experience is required. Use the given assumptions to work through the decisions.

If this feels like a lot, that is normal. Nobody is expected to hold all of it in their head at once, which is why we work through it one section at a time.

Tux comic illustrating reliability interview capstone.
  1. 01Define the customer promise
  2. 02Plan for failure and recovery
  3. 03Explain the evidence
Connect the promised booking experience to practical decisions, calculations, and checks for recovery.

What is the launch scenario?

The launch scenario is a busy ticket sale on TicketDesk, with fixed assumptions about demand, capacity, the reliability target, and a risky late release. Customers browse an event, reserve a ticket, and receive confirmation. Some try again if a result is delayed. Suggested events are optional. Correct reservations and confirmations are essential.

Checkout receives booking requests. A request is an attempt asking the software to do something. The later confirmation work is performed by worker software and uses shared stored records.

Use these invented practice assumptions:

  • Peak checkout demand is forecast at 900 valid requests per second.
  • One serving instance, a running copy of the software, handles 200 requests per second within the tested response-time and correctness limits.
  • The service must handle losing one serving instance.
  • The proposed success SLO, or service level objective, is 99.9% over a rolling 30-day window.
  • The measured window contains 30 million valid requests and 18,000 failures.
  • Confirmed bookings must not be silently lost or duplicated.
  • A new release changes confirmation behaviour shortly before launch.

Allow about 30 to 45 minutes, or work through one section at a time. Download the blank worksheet and attempt the questions before reading the worked discussion. This is original practice material, not a question reported from any employer.

What are you promising the customer?

You are promising the customer a defined success rate for valid booking attempts and a confirmation that arrives within a stated time. Start by defining a valid booking attempt and a successful result. How should sold-out seats and repeated attempts be counted? How soon must confirmation arrive? What should customers see while they wait?

An SLI, or service level indicator, measures one aspect of service performance. Propose a success measurement and a timely-completion objective. State where each is collected and how accepted work is followed through to its final outcome.

Measuring only the first checkout response can miss confirmations that never arrive. Optional suggestions may be disabled under pressure; reservations must still produce an honest result. Explain how your thresholds fit the customer's needs. The 99.9% target is an assumption for this exercise, not a universal target to quote in every interview.

How much failure budget and capacity remain?

In this scenario 60% of the failure budget is already used, and the service needs six instances to carry the peak after losing one. An error budget is the failure allowance within the reliability target. Calculate the allowed failures, the fraction used, and how many running copies are needed after one fails:

Allowed bad requests = 30,000,000 × 0.001 = 30,000
Consumed fraction = 18,000 / 30,000 = 60%
Remaining allowance = 12,000 requests, or 40%
Required surviving instances = ceiling(900 / 200) = 5
Initial fleet = 5 + 1 = 6 instances before extra margin

The permitted failure fraction is 0.1%, written as 0.001. Divide 900 by 200 and round up to get five surviving copies, then add the one copy that may be lost.

State what the estimate assumes: evenly distributed work, unchanged per-copy performance, and enough capacity in shared dependencies. Then ask about bursts, how long it takes to add capacity, and failures that could remove several copies at once. Interviewers like hearing the assumptions named before the number.

This budget counts requests. The 43.2-minute allowance for a simple 30-day time-based 99.9% target does not describe the remaining 12,000 requests here; the two are different ways of measuring the same objective, and you should not mix them.

Should you expand the new version?

No, you should pause expansion, because the canary's customer results are ten times worse than the control's. A canary is the small group receiving the new version first. The control uses the unchanged version at the same time. Here the canary has 200 failures out of 20,000 valid requests, or 1%. The control has 180 out of 180,000, or 0.1%. Confirmations are waiting longer. The installation tool reports success.

Explain why these customer results justify pausing under the agreed release checks. A successful installation does not establish successful bookings; it only shows the software started.

Check whether the canary exercised the risky booking types and stayed within absolute limits. Describe recovery. Rollback means restoring the earlier software, but it cannot automatically undo payments or new record formats. If the old software cannot read the changed records, rollback may itself be unsafe.

Name who makes the decision and what evidence would allow expansion to resume.

What do you do about the growing queue?

For the growing queue, you limit incoming work, reduce optional work, bound retries, and find the bottleneck, while making sure accepted bookings still complete. A queue is a waiting list of work. Confirmation work arrives at 1,000 items per second and finishes at 800. The backlog grows by 200 per second, or 12,000 in a minute. Customers retry, adding demand on top.

Limit incoming work and waiting time, reduce optional work, and bound repeat attempts while you investigate the bottleneck, the part that limits completion. Extra workers help only if that part can use them. Accepted bookings still need a safe completion plan.

Explain when you would declare an incident, meaning a problem requiring a coordinated response. Assign technical and communication responsibilities. Record changes and coordinate actions that interact. Give an update describing confirmed impact, current action, uncertainties, and the next update time.

Verify recovery through the affected bookings, response time, waiting confirmations, and correct records under normal demand. Falling errors alone may just mean fewer customers are reaching the app.

Which assumptions could break your plan?

The assumptions most likely to break the plan are independent dependencies, even per-copy performance, and a recovery that stops at copying files. A dependency is another service TicketDesk needs. Suppose two independent dependencies must both work in sequence. Each is 99.9% available:

0.999 × 0.999 = 0.998001 = 99.8001%

Multiplication gives the joint availability in that simplified model. Shared failures, alternative paths, or different measurement periods can change the result. Explain those limits before applying the estimate to a real service.

For data recovery, choose a recovery point objective (RPO), the maximum acceptable lost-data interval, and a recovery time objective (RTO), the maximum acceptable recovery duration. Include checking the restored bookings, reconciling external actions such as payments, and reopening safely. A reachable app with duplicate reservations has not recovered correctly.

What should improve after the incident?

After the incident, propose two actions with owners and acceptance tests, meaning checks that demonstrate the intended improvement. One should strengthen a protection that failed; the other should reduce repeated manual work.

A release check for confirmation delay can be tested by replaying the failing workload. A recovery process can be judged by fewer manual repairs, correct bookings, and the effort needed to maintain it. Explain which action comes first, using the risk you observed rather than personal preference.

How should you review your reasoning?

Review your reasoning by marking each area as missing, mentioned, or explained with evidence: the customer promise, the calculations, the failure assumptions, the measurements, the response responsibilities, the recovery, and the follow-up.

This is a self-review exercise, not a hiring score. Revisit any area where you named a technique but could not explain its purpose; that gap is exactly what an interviewer will probe. Practise a two-minute answer to “What would you improve first?” The listener should be able to follow your choice from the risk, to the change, to the evidence you would check.

Quick answers

How do you prepare for an SRE system design interview?

Practise reasoning out loud through a scenario: state the customer promise, work the capacity and error-budget sums with named assumptions, decide how to release safely, describe how you would respond to failure, and explain how you would verify recovery. Interviewers reward a connected chain of reasoning over memorised tools.

How do you calculate an error budget from an SLO?

Multiply the number of valid requests in the window by the permitted failure fraction. For 30 million requests and a 99.9% objective, the allowance is 30,000,000 × 0.001 = 30,000 failures; 18,000 observed failures means 60% of the budget is used.

What questions should you ask in a reliability design interview?

Ask about peak demand and its shape, how much work one instance can handle, which failures the service must survive, what the reliability target is and how it is measured, and what must never be lost or duplicated. Those answers determine every later decision.