# Incident response practice: the first alert

An original TicketDesk practice scenario. Use it alongside the lesson:
https://everythingsre.com/incident-response/first-alert

Try every question on paper before you read the worked answer in the lesson. Your first attempt will have gaps; that is what the self-review at the end is for.

## Evidence supplied

At 14:20, the West region has 1,200 failures in 20,000 valid checkout requests over five minutes. East has 80 in 80,000. Both regions previously ran at a 0.1% failure rate. A new confirmation behaviour was enabled in West at 14:18. Arrival counts are similar to the preceding sample. Some customers report paying without receiving a confirmation. The oldest waiting confirmation is ten minutes old; the normal expectation in this exercise is two minutes.

The practice policy declares an incident when a region exceeds 5% booking failures for five minutes, or when accepted paid bookings miss their confirmation deadline. The tested disabling setting affects only future work, preserves accepted records, can be reversed, and should reach the affected copies within one minute.

## Attempt before reading the worked answer

1. Describe the known impact. Which of the numbers count requests rather than customers? Which accepted bookings need more investigation?
2. Write one observation and two possible explanations for it. What single check would tell the two explanations apart?
3. Assign the response lead, technical investigation, confirmation recovery, and communication responsibilities. Combine roles only if the workload allows it.
4. Propose a mitigation. State its owner, the safety assumptions behind it, the expected result, the stop condition, and the fallback if it does not work.
5. Write a support update covering known impact, current action, what is uncertain, and the next update time. Do not state a cause or a repair estimate you cannot support.
6. At 14:26, West's new failure rate is back to 0.1%, but the old confirmations are still waiting. Explain what has improved and what is still unfinished.
7. List the customer recovery checks you would run, and the evidence you would preserve for the follow-up review.

## Self-review

Mark each area as missing, mentioned, or explained with evidence: customer impact, uncertainty, coordination, mitigation safety, communication, and recovery verification.

This is a learning checklist, not an employer score. If you discuss this scenario in an interview, present it as what it is: practice material you worked through, not an incident you handled.
