An SRE incident response interview usually gives you a failure in pieces and watches how you decide with incomplete information. The interviewer is not checking whether you guess the cause. They are checking whether you establish impact first, state what each piece of evidence rules out, choose actions whose effects you can predict, and verify the result. This capstone runs that format from start to finish, with a stop point at each stage so you can decide before reading on.

The scenario is new, but every skill comes from the earlier lessons in this series, starting with what to do when an alert fires. TicketDesk is a fictional service; the times, numbers, and people are invented for practice. Download the capstone worksheet if you want to write your answers before reading each discussion. Treat each “Stop and decide” as if the interviewer had just paused and asked, “What do you do?”

Tux prepares tickets under a lamp beside a sand timer as a coordinator holds a stadium gate for a waiting crowd.
Prepare for the surge, keep the critical path safe, and agree who can delay the on-sale.

The setting

It is a Saturday. At 10:00, TicketDesk opens sales for a stadium tour that is expected to bring five times the usual traffic for the first fifteen minutes. The on-sale has a coordinator, Mei, who can delay it if engineering asks. You are on call. Search, the part of the app customers use to find events, normally handles 1,000 requests a second and answers 92% of them from a cache, a fast store of recent results. The remaining 8% go to the main database, which also serves checkout. The database has been tested to about 1,200 read queries a second before response times climb.

Stage one: an alert at 09:41

The pager says: search p99 latency above 2 seconds for five minutes, currently 3.8 seconds. The p99 is the value that 99% of requests finish within, so one request in a hundred is taking almost four seconds or more. The alert covers all regions.

Stop and decide. What do you check first, and what do you tell Mei?

Establish impact before cause. Checkout success is 99.9%, its normal value, so customers can still buy tickets. Search is returning results, with errors at 0.2%, but slowly; the customer experience is a sluggish event page, not a failed booking. The impact is real but bounded, and it has a deadline: at 10:00 search will receive five times this load.

Acknowledge the page, open the search runbook, and start the incident note with the detection time. Then tell Mei in one message: “Search is slow since about 09:36. Bookings are working. We are investigating and will give you a go or hold recommendation for the on-sale by 09:52.” You have not promised a cause or a fix; you have promised a decision at a time that leaves her room to act.

Writing the decision time down commits you to it. A coordinator with no information at 09:58 will make the on-sale decision without you.

Stage two: evidence at 09:47

Three charts and one log come back. The cache hit rate fell from 92% to 40% at 09:30 and is now at 55%. Database read queries rose from about 80 a second to 600 a second over the same period, and database CPU is at 85%. The cache dashboard shows two of its four nodes marked unhealthy. The deploy log shows search service version 3.2 rolled out to all hosts between 09:28 and 09:31.

Stop and decide. What are your hypotheses, and which single check separates them?

Write the candidates with their predictions. H1: version 3.2 changed how cache keys are built, so every existing entry misses and the cache is warming from empty. If true, the hit rate fell at the deploy and is climbing steadily as new entries fill in. H2: two cache nodes failed, taking half the cached entries with them. If true, the nodes became unhealthy around 09:30 and the hit rate dropped to roughly half, which 40% is close to. H3: a wave of unusual searches ahead of the on-sale is missing the cache. If true, request volume or query variety rose.

The separating check is the cache node history. Noor pulls it: both nodes have been marked unhealthy since Thursday, carry no traffic, and the other two have held the full working set since then. That rules out H2 as a cause of today's change, while leaving a stale alert nobody cleaned up as a finding for later. Request volume is a flat 1,000 a second, which rules out H3. The 3.2 release notes say “normalised search keys: lowercase and trimmed,” which means every key written before 09:30 is now unreachable under its new spelling. H1 fits all of it.

The arithmetic confirms the mechanism. At a 92% hit rate, 1,000 searches a second send 80 queries to the database; at 40%, they send 600. The database went from idle to 85% CPU because of a key format, and checkout, which shares that database, is one bad decision away from being dragged in.

Stage three: the decision at 09:52

The hit rate is 62% and climbing about one point a minute as the cache refills. The on-sale is eight minutes away. Five thousand searches a second against a database that tolerates 1,200 queries a second needs a hit rate of at least 76%:

Required hit rate = 1 - 1,200 / 5,000 = 0.76 = 76%
Hit rate at 09:52 = 62%, rising about 1 point per minute
Expected to reach 76% at about 10:06, after the on-sale starts

Stop and decide. What do you recommend to Mei, what action do you take in the meantime, and what is your stop condition?

Lay out the options with what each one does to the number that matters. Rolling back to 3.1 restores the old key format, but entries written in that format expire after 30 minutes, so by 09:55 almost all of them are gone; rollback produces a second cold cache right at 10:00. That is the worst available action, and it is the one an interviewer will often offer you. Adding database read replicas helps, but provisioning takes about 20 minutes. Delaying the on-sale is a business decision that only Mei can make, and she needs a time and a reason.

Two actions do help before 10:00. Pre-warming the cache, by running the two thousand most common searches through version 3.2 so their new keys exist, takes about five minutes, adds nothing dangerous, and raises the hit rate for exactly the queries the on-sale will produce. And preparing a protective switch: the search service has a cached-only mode that answers from the cache and fails uncached queries quickly instead of querying the database. It degrades search but shields checkout.

The recommendation to Mei at 09:52: “We recommend holding the on-sale until search's cache hit rate is above 76%; we are pre-warming now and expect that by 09:58. If we are not there by 10:00, we ask for a ten-minute hold.” The stop condition, written in the note: if database CPU exceeds 90% at any point, switch search to cached-only mode to protect checkout, and tell Mei immediately.

Notice the shape of that answer. It names the metric that decides, the action that moves it, the time at which you will know, and the protective step if the prediction fails. That shape is what the interviewer is listening for; the specific numbers are secondary.

Stage four: the on-sale at 10:00

The pre-warm finishes at 09:57 and the hit rate on live traffic reads 81%. Mei goes ahead at 10:00. Between 10:00 and 10:05, search receives 4,800 requests a second with an 88% hit rate, because on-sale searches are highly repetitive; the database sees about 576 queries a second and its CPU settles at 70%. Search p99 latency is 900 milliseconds. Checkout success holds at 99.8%, within its objective.

Stop and decide. Has the service recovered, and what is still open?

Recovered means the customer experience is back within its objectives under the demand that matters, and it is: search is fast, checkout is healthy, and the on-sale is proceeding. Say so in the note and to Mei, with the numbers.

Then list what recovery has not settled. Version 3.2 is still running with a cache that would be cold again after any rollback, so a pre-warm step must be part of any future rollback of this service. The two stale unhealthy cache nodes need cleaning up or replacing, and the alert that has been quietly red since Thursday needs a reason to be trusted again. Search and checkout share a database, so a search problem can become a booking problem; that coupling is a design finding, not a Saturday fix. And the on-sale rehearsal never included a cache warm-up, which is why a routine deploy on a Saturday morning became an incident.

Those go into the postmortem with owners, as the previous lesson describes. The incident closes when the on-sale traffic subsides and the objectives have held for the agreed period, not at 10:05 because everyone is relieved.

How do you narrate an incident like this in an interview?

Tell it in the order you decided things, with the evidence attached to each decision, and finish with what you changed afterwards. A clear structure for a five-minute answer:

Impact first: what customers experienced, measured, and what they did not experience. Evidence next: the three or four observations that mattered, and what each one ruled out. Then the decision, with the calculation that drove it and the alternative you rejected and why. Then verification: the numbers that showed recovery, under the load that mattered. Finally, the follow-up: the finding behind each action item.

The common mistakes are the mirror image. Leading with the cause and skipping impact. Listing every dashboard you would open without saying what each would tell you. Choosing rollback because “it is usually safe.” Declaring victory when one chart improves. And ending the story at recovery, as if the incident had nothing to teach.

If your experience comes from practice scenarios like this one rather than a production incident, say so, as the interview capstone in the beginner guide advises. An honest account of a well-reasoned exercise is a better answer than a vague account of a real incident you did not understand.

Exercise: run the scenario with one fact changed

Suppose the pre-warm had finished at 09:57 but the live hit rate had read 70%, not 81%. Decide what you recommend to Mei at 09:58, what you would say about the cached-only switch, and how you would explain the trade-off between a delayed on-sale and degraded search. Write it out, then compare with this version:

At 70% the database would receive about 1,500 queries a second at on-sale load, above its tested limit, so I would ask Mei for a ten-minute hold and tell her why in one sentence: going at 10:00 risks checkout, not only search. During the hold I would keep the pre-warm running on the next most common queries and watch the hit rate climb. If she cannot hold, I would go ahead with cached-only mode enabled from the start, which keeps bookings safe and makes some searches fail fast; I would tell support what customers will see. Either way the stop condition stays: database CPU above 90% switches search to cached-only.

Explain your reasoning for which option you present as the recommendation and which as the fallback. The strongest answers keep checkout safe in every branch and make the business trade-off explicit for the person who owns it, rather than quietly deciding it for them.

Quick answers

What do interviewers look for in an SRE incident response question?

Impact assessment before cause, hypotheses with evidence that rules them in or out, actions chosen for predictable effect and reversibility, explicit verification of recovery, and follow-up that addresses why the failure was possible.

Is rolling back always the safest first action in an incident?

No. Rollback is unsafe when the previous version cannot read data the new version wrote, when it would repeat a cold start or migration, or when it does not address the actual cause. State the compatibility check before recommending it.

How should you answer if you have never handled a real production incident?

Use a practice scenario or a real project failure and say clearly that it was practice. Walk through impact, evidence, decision, and verification in order; the reasoning is what is being assessed, not the size of the outage.