When an urgent SRE alert fires, the first steps are to acknowledge it, establish the customer impact, and get the right people working on reducing that harm. You then use the evidence in front of you to choose a safe action and check whether the affected customers actually recover. You can, and usually should, start responding before you know the full cause.

This lesson puts you in charge of the first response to a TicketDesk booking failure. TicketDesk is a fictional service; the times, numbers, people, and operating rules below are invented for practice. You will decide what to check and what to tell others while some important facts are still missing, which is exactly how real incidents feel.

If the vocabulary is new, keep the SRE glossary open in another tab. The beginner lessons on monitoring and on-call troubleshooting cover the foundations, but you can follow this lesson without them.

Tux examines a jammed ticket kiosk with a magnifying glass while a customer waits. The restart lever is untouched.
The alarm draws your attention. Find out what happened to the waiting customer before choosing a repair.

What does the alert actually tell you?

An alert tells you that a measured condition crossed a threshold for a period of time. It does not tell you why. At 14:20, you receive an urgent notification: valid checkout requests in TicketDesk's West region have exceeded a 5% failure rate for five minutes. A region is a group of servers and supporting infrastructure that serves part of the customer demand. A request is a message asking the software to do something, such as booking a seat.

So the alert identifies a symptom and a time window. It does not yet establish why bookings failed, whether payments completed, or whether East-region customers are quietly affected by a related confirmation problem.

Acknowledge the notification first, so the alerting system and your teammates know someone is on it. Then open the runbook, the written recovery guidance for this alert, and check that you have access to the measurements and controls it describes. If you do not have access, use the agreed escalation route to get someone who does. Waiting silently never repairs a service.

Start a short incident note. Write down the detection time, your name as the initial responder, the known symptom, and the next thing you plan to check. You will add detail as you learn more.

Which customers are affected?

Before touching anything, find out who is hurting and how badly. The dashboard, a page of live measurements, shows this five-minute sample. A failure here means a valid checkout request ended with an internal error or a timeout. A timeout means the result did not arrive within the allowed waiting time.

MeasurementWestEast
Valid checkout requests20,00080,000
Failed requests1,20080
Failure rate6%0.1%
Failure rate before this window0.1%0.1%

The West calculation is 1,200 divided by 20,000, which is 6%. Across both regions together, 1,280 of 100,000 requests failed, which is 1.28%. Notice how the combined number makes West's situation look far less severe than it is. Always keep the affected group's own result visible.

A request count is not a count of distinct customers. One frustrated person may have tried five times. Report failed attempts, and only estimate customers once you have evidence to support it.

Customer support has also received reports of people who paid but are still waiting for confirmations. The oldest waiting confirmation is ten minutes old; in this exercise, the usual expectation is two minutes. Some of that confirmation work runs on software shared between regions. East's healthy checkout percentage therefore does not prove that every East customer has received a ticket.

Check when each symptom began and whether it is spreading. Give priority to the payment and reservation records that could be wrong or unfinished, alongside the new failed attempts.

Do you need to declare an incident?

Yes, in this scenario. For the exercise, TicketDesk's operating policy calls for an incident response when valid booking failures exceed 5% for five minutes in a region, or when accepted paid bookings miss their confirmation deadline. Both conditions are now true, so this needs a coordinated investigation.

These are practice rules for TicketDesk. Severity names and thresholds vary between teams, so do not present this threshold as a universal SRE standard in an interview.

Declare the incident and name a response lead. Ask a checkout owner to investigate the new failures and a confirmation owner to check the accepted work that is still waiting. Someone also needs to keep the timeline and send updates. In a small response, one person can hold more than one role if the workload allows it.

The incident-management lesson explains these responsibilities in detail. The practical question is simple: does every urgent task have an owner? How many job titles appear in the chat matters much less.

What is a fact, and what is still a hypothesis?

A fact is something you observed; a hypothesis is an explanation you still need to test. Keeping the two separate is one of the most valuable habits in incident response.

You find that a new confirmation behaviour was enabled in West at 14:18. The number of arriving checkout requests is similar to the previous five-minute sample.

Your observations are that West's failure rate rose, accepted confirmations are waiting too long, and a software change happened shortly before the alert. A hypothesis is a proposed explanation to test: perhaps the new behaviour creates more confirmation work per booking and overwhelms the shared processing.

Test the prediction. Compare the work generated by changed and unchanged bookings. Check how many tasks arrive and how many finish, and whether delays began after the setting changed. A queue is a waiting list of work; it grows whenever arrivals outpace completions.

You may also find that a dependency, another service TicketDesk relies on, slowed down at the same time. Similar request counts do not rule that out. Keep each explanation tied to evidence that could tell it apart from the others.

In your notes, write “new behaviour may be contributing” while that is all the evidence supports. A recent change is always worth investigating, but timing alone is not proof of cause.

Which action would reduce harm safely?

The safest action is one whose benefit, risks, and verification you understand, and that can be reversed if it makes things worse. The release owner proposes stopping expansion of the new behaviour. The runbook also describes a setting that disables it for future work. For this exercise, that setting has been tested, leaves accepted reservations intact, and can be reversed. Its effect should reach all affected software copies within one minute.

Agree on the action and its owner before applying it. Record the time and the expected result: fewer new failures and a lower rate of new confirmation work. The confirmation owner checks whether waiting bookings can finish safely. Define a stop condition too: if duplicate reservations appear or another region gets worse, reassess immediately.

This is a mitigation, an action intended to reduce the harm happening right now. It is reasonable to apply while the cause is still under investigation, precisely because its benefit, risks, and verification are understood.

Compare that with restarting all of the confirmation software. You do not yet know which paid bookings are halfway through processing or whether a restart would repeat their actions. “It worked last time” does not answer those questions. A restart could be appropriate with supporting evidence and a safe recovery procedure; it should never be an unexplained default.

Rolling back to an earlier version also needs a compatibility check. As the release lesson explains, old software may not understand records already written by the new version. And changing the software does not reverse a payment that already happened.

What should you tell customer support?

Give support enough confirmed information to help customers, and be clear about what remains uncertain. After stopping expansion, an update could read:

Booking attempts in West are failing more often than usual. Some customers are also waiting longer for confirmations. We have paused expansion of a recent change and are checking affected bookings and payments. We have not yet confirmed the status of every delayed reservation. Our next update is at 14:35.

This update does not promise a repair time and does not advise customers to pay again. Neither would be supported by the current evidence, and the second could cause duplicate charges.

Keep the technical investigation moving while an assigned person handles updates. At 14:35, send another update even if the investigation is still ongoing. Say what changed, what is still affected, and when people will hear more. Silence makes everyone assume the worst.

When can you say the service has recovered?

A service has recovered when the affected customers can complete what they came to do, not when a chart looks better. At 14:26, West's recent checkout failure rate falls back to 0.1% after the setting is disabled. The oldest waiting confirmation is still ten minutes old.

New bookings are improving, but the earlier affected work is still unfinished. Check whether accepted reservations produce the correct tickets, whether payments match those reservations, and whether delayed work finishes without creating duplicates. A shrinking queue is only good news if tasks are completing rather than being discarded.

Watch the affected regions under the demand the service must handle. Use the agreed observation period and confirm that failures and delays do not creep back. Give every temporary mitigation an owner so it is reviewed after the immediate response ends.

Record detection, mitigation, and customer recovery as separate milestones. The later postmortem will use that timeline to investigate the cause and propose lasting improvements.

Exercise: give your first response aloud

Spend two minutes answering, out loud: “You receive the 14:20 alert. What do you do next?” Use the evidence supplied above, and leave uncertain details uncertain. Download the response worksheet if you want to make notes before reading the discussion.

A defensible answer would acknowledge the alert, check customer impact in both checkout and confirmations, and start a coordinated response under the stated policy. It would pause expansion, compare the new behaviour with other explanations, and use the tested disabling setting if its safety assumptions still hold. It would also name who checks the accepted bookings and who sends the next update.

Explain your reasoning for the recovery check. Fewer new errors do not complete the bookings that were already waiting. In an interview, a command or technique is only as good as your explanation of what you expect it to change, what could go wrong, and which result would make you keep or abandon it.

Check your understanding

Try these four questions before opening the answers. They use the scenario's stated assumptions, not a secret employer scoring rubric.

1. What should you report from the West measurement?

  • A. Exactly 1,200 customers have lost tickets.
  • B. 1,200 of 20,000 valid checkout requests failed; distinct customer impact still needs investigation.
  • C. Exactly 6% of all TicketDesk customers were charged twice.
  • D. East's checkout result proves all its confirmations are healthy.
Check answer 1

B. The sample counts requests, and one customer can make several attempts. It also says nothing about lost payments or duplicate charges. Report the measured failure count while you investigate accepted bookings, customer reports, and confirmation completion.

2. What does the change at 14:18 establish?

  • A. It proves the new behaviour is the only cause.
  • B. It means investigation should wait until the incident ends.
  • C. It is a plausible contributor worth testing against other explanations.
  • D. It proves a customer surge caused the failure.
Check answer 2

C. Timing makes the change worth examining first, but timing is not proof. Compare work per booking, arrivals, completions, and dependency behaviour. Evidence that distinguishes those explanations is stronger than assuming the latest change must explain every symptom.

3. Which immediate action has the strongest support here?

  • A. Coordinate stopping expansion and use the tested disabling setting, with an owner and checks for new failures and accepted work.
  • B. Restart all confirmation software without checking unfinished payments.
  • C. Return to old software without checking whether it can read the new records.
  • D. Expand the new behaviour into East to gather more data.
Check answer 3

A. The scenario gives you evidence about the setting's reversibility, its effect, and its protection of accepted records. Confirm those assumptions still apply, coordinate the change, and observe the result. Each of the other actions introduces unresolved risks or exposes more customers to harm.

4. Errors fall at 14:26, but old confirmations still wait. What should you say?

  • A. Every affected customer has recovered.
  • B. The incident is over because the home page opens.
  • C. The waiting work can be discarded to make the queue smaller.
  • D. New checkout results improved; delayed bookings still need recovery and correctness checks.
Check answer 4

D. The lower error rate only describes recent requests. Earlier accepted bookings remain part of the incident until their outcome is understood and handled. Verify that confirmations complete and that reservations and payments are correct, then watch for sustained recovery under realistic demand.