SRE on-call troubleshooting is the work of responding when a service needs urgent attention, investigating what went wrong, testing possible explanations, and choosing actions that help the service recover safely. Being on call means you are the person responsible for that response during your shift. Troubleshooting is the method you use once the notification arrives.

At TicketDesk, our fictional booking app, an urgent notification might arrive because customers cannot receive their tickets. The responder needs to understand what those customers are experiencing before deciding what to change. That order matters: understand first, then act.

If the idea of being woken up to fix something you did not build sounds intimidating, that is normal. This lesson walks through the sequence one step at a time, and none of it requires memorising commands.

Tux comic illustrating on-call and troubleshooting.
  1. 01Who is affected?
  2. 02Test an explanation
  3. 03Reduce harm and check results
Start with the customer’s problem. Use evidence to choose a safe action and verify the booking experience afterward.

What happens when an urgent alert arrives?

When an urgent alert arrives, the first step is to acknowledge it, establish who is leading the response, and work out what customers are trying to do and which part of that is failing. A page is an urgent notification asking someone to respond now. Acknowledging it tells everyone else that someone is looking, so nobody duplicates effort or assumes the problem is being ignored.

Next, find the start time, the affected attempts, and any recent changes. A request is an attempt asking the software to perform an operation, such as booking a seat. Are all booking requests failing, or only a particular group? These few checks help you judge urgency without opening every chart you have.

If the problem is spreading, or several teams need to be involved, ask for a coordinated incident response. One person cannot investigate, repair software, keep notes, and inform customers all at once, and nobody expects you to. The incident-management lesson explains how those responsibilities can be shared.

What should be ready before someone goes on call?

Before someone goes on call, they need working access to the systems, knowledge of the service, a current runbook, and a clear escalation route. A runbook is a written guide to known problems and their recovery actions. It should describe the software as it is today, not as it was a year ago, and its important steps should have been practised at least once.

An escalation route tells you who to ask for more help, and how to reach them. The schedule itself also needs to be sustainable. Frequent overnight interruptions leave people unable to work effectively the next day, and a tired responder makes worse decisions.

Discovering during a failure that your account cannot disable a risky change creates avoidable delay. Check your access and the recovery procedures before your responsibility starts, not after the first page.

What makes a useful explanation to test?

A useful explanation is a hypothesis: a proposed cause that evidence can support or contradict. “The app is slow” describes a symptom and cannot be tested. “The new version creates extra confirmation work for every booking, filling the shared waiting list” can be tested, because it predicts something you can go and measure.

A queue is a list of waiting work. If the new version really creates extra work, changed bookings should add more queue items than unchanged bookings. If a surge in customers is the cause instead, the number of booking attempts should rise. Look for evidence that separates those two explanations rather than evidence that fits both.

Keep observations separate from guesses in your notes. Record what changed and when, along with the result of each check. Where practical, change one meaningful thing at a time. If you make several repairs together and the service recovers, you will not know which one helped, and you will not know which to keep.

What does the evidence suggest at TicketDesk?

In this invented example, the evidence points to a new software version that adds confirmation work. TicketDesk introduces the version to one group at 10:00. Errors rise in that group at 10:04, while unchanged bookings still work. By 10:06, confirmations have been waiting longer in both groups.

Why both groups? The confirmation queue is shared. Extra work from the changed group could slow confirmations for everyone, although a customer surge remains another possibility. To separate the two, compare incoming attempts with the amount of confirmation work generated per booking.

Suppose switching off the new behaviour is a tested action that can be reversed. Taking it may reduce harm before the full investigation finishes. Before you do, state the expected result: less new work, fewer errors, and a waiting list that starts shrinking. Writing the prediction down is what lets you check whether it came true.

If errors fall but confirmation age keeps rising, the action helped one symptom and not the other. Some customers still lack tickets. Give that remaining work attention before anyone declares recovery.

How do you choose a safe immediate action?

A safe immediate action, called a mitigation, is one that reduces the harm customers are experiencing now, can be undone if it does not help, and whose effect you can observe quickly. Examples include stopping the expansion of a new version, disabling optional recommendations, or sending work to computers with spare capacity.

Compare the likely benefit with the action's risks. Which customers or records could it affect? Can it be undone? How quickly will the results show whether it helped? You will not always have perfect answers, but asking the questions stops you from reaching for the first idea that comes to mind.

Some actions look helpful and are not. Adding more software workers, the pieces of software that perform waiting tasks, helps only if the limited part of the service can use them. More workers can put extra pressure on a dependency, meaning another service TicketDesk relies on. Restarting software can interrupt unfinished work or repeat an action such as a payment.

Sometimes the safest action is to stop something. If new updates are corrupting booking records, temporarily stopping those updates may prevent more damage. Keeping the booking page reachable would offer little benefit if it kept creating incorrect reservations.

When has the service recovered?

The service has recovered when the affected customer journey works under normal demand, and you have watched it long enough to be confident the problem is not returning. Look at failed bookings, response time, pending confirmations, and record correctness in the groups that had trouble. Observe long enough to notice delayed effects or work building up again.

Be suspicious of numbers that improve for the wrong reason. A lower error count could mean fewer customers are reaching checkout at all. A shorter queue could mean tasks were discarded rather than completed. Account for accepted bookings and their final outcome, not just the shape of a chart.

Be precise when you report time to recover. Teams use MTTR, mean time to recovery, for different intervals, including time to mitigate, time to restore, and time to fully resolve. State what the clock measures rather than relying on the acronym. A timeline with clear milestones is much easier for other people to interpret.

Exercise: should you restart immediately?

A colleague says, “A restart fixed it last time.” Explain what you would check in the first five minutes and whether you would restart now.

A previous repair is useful history, but today's cause may differ. You might still choose the restart if its effects are understood and waiting would cause more harm. Either way, explain the result you expect and how you would verify it.

Your reasoning should connect the customer's problem to the evidence, the immediate action, and the recovery check. Include the point at which you would ask for help. In an interview, that connected chain is what the interviewer is listening for; the specific commands matter far less.

Quick answers

What does an SRE do when on call?

An on-call SRE acknowledges urgent alerts, works out what customers are experiencing, forms a testable explanation, takes a safe action to reduce harm, and verifies that the customer journey has recovered. They also know when to escalate to a coordinated incident response.

What is the difference between mitigation and root cause?

Mitigation reduces the immediate harm, for example by switching off a new behaviour, even before you know exactly why it failed. Finding the root cause is the deeper investigation into why the failure happened, which usually continues after customers are no longer affected.

How do you know an incident is resolved?

Check the affected customer journey under normal demand, including accepted bookings and their final outcome, and watch long enough to catch delayed effects. A falling error count alone is not proof, because it can also mean fewer customers are reaching the service.