# Incident response capstone: a failure decided in stages

An original TicketDesk practice scenario. Use it alongside the lesson:
https://everythingsre.com/incident-response/incident-response-capstone

Write your answer at each stop before reading the next stage in the lesson. Timing yourself to three minutes per stop makes it feel closer to an interview.

## The setting

It is Saturday. At 10:00 TicketDesk opens an on-sale expected to bring five times the usual traffic for fifteen minutes. The coordinator, Mei, can delay it if engineering asks. Search normally handles 1,000 requests a second with a 92% cache hit rate; misses go to the main database, which also serves checkout and has been tested to about 1,200 read queries a second.

## Stage one: 09:41

Alert: search p99 latency above 2 seconds for five minutes, currently 3.8 seconds, all regions. Checkout success is 99.9%. Search errors are 0.2%.

1. What is the customer impact right now, and what is not affected?
2. Write the first message to Mei. Include a time by which you will give a recommendation.

## Stage two: 09:47

Cache hit rate fell from 92% to 40% at 09:30 and is now 55%. Database reads rose from 80 to 600 queries a second; database CPU is 85%. Two of four cache nodes show as unhealthy. Search version 3.2 deployed to all hosts between 09:28 and 09:31.

3. Write three hypotheses, each with the evidence that would confirm or rule it out.
4. Name the single check that separates them, and say what result you expect.
5. Show the arithmetic that connects the hit rate to the database load.

## Stage three: 09:52

Hit rate is 62%, rising about one point a minute. The on-sale brings 5,000 searches a second. Cache entries expire after 30 minutes. A pre-warm of the two thousand most common searches takes about five minutes. Read replicas take about 20 minutes. Search has a cached-only mode that protects the database but fails uncached queries fast.

6. Calculate the minimum hit rate that keeps the database under 1,200 queries a second.
7. For rollback, pre-warm, replicas, delay, and cached-only mode, state what each does to that number and by when.
8. Write your recommendation to Mei, the action you take now, and your stop condition.

## Stage four: 10:00

Pre-warm finished at 09:57; live hit rate 81%. From 10:00 to 10:05: 4,800 searches a second, 88% hit rate, database at 70% CPU, search p99 900 milliseconds, checkout success 99.8%.

9. Has the service recovered? Give the numbers that justify your answer.
10. List what is still open for the postmortem, with the finding behind each item.

## Variation

Repeat stage three assuming the pre-warm finished with a live hit rate of 70%. What changes in your recommendation, and what stays the same?

## Self-review

Mark each area as missing, mentioned, or explained with evidence: impact before cause, hypotheses with separating checks, a calculation that drove the decision, a rejected alternative with the reason, verification under realistic load, and follow-up items with findings.

This is a learning checklist, not an employer score. If you discuss this scenario in an interview, present it as practice material you worked through, not an incident you handled.
