Incident investigation in SRE means turning what you can see into explanations you can test, then choosing the check that rules the most explanations out. You are not trying to be right on the first guess. You are trying to shrink the list of possible causes quickly, while the mitigation you already applied keeps customers safe. Good investigators write down what each piece of evidence proves, what it rules out, and what it leaves open.

This lesson continues the TicketDesk incident from what to do when an alert fires. TicketDesk is a fictional service; every time, number, name, and rule below is invented for practice. If you have not read the first lesson, here is where things stand: it is 14:26, you disabled a new confirmation behaviour that was enabled in the West region at 14:18, the West checkout failure rate has dropped from 6% back to 0.1%, and paid customers are still waiting too long for their confirmations.

If any term is unfamiliar, the SRE glossary has short definitions, and the beginner lesson on on-call troubleshooting covers the basic investigation loop this lesson builds on.

Tux studies a notebook beside a cracked pipe and a pipe crowded with tickets feeding a sorting bin.
Two symptoms can have different causes. Choose the next check by the evidence it can give you.

Why is the incident not over when the error rate drops?

Because a falling error rate only tells you that new requests are succeeding; it says nothing about the customers already harmed or about causes you have not found yet. At 14:26, new West checkouts work again. But the confirmation owner reports that the oldest unfinished confirmation is still about ten minutes old, and support keeps receiving messages from people who paid and have no ticket.

That combination should make you suspicious. You disabled a change and one symptom vanished within a minute, while the other symptom did not move. If the change were the whole story, both symptoms would improve together. Something else is going on, and your job now is to find it without guessing.

Write a one-line status in the incident note: “14:27. Checkout failures mitigated by disabling the West confirmation setting. Confirmation delay unchanged. Cause of delay unknown. Investigating.” The last two words matter: they tell everyone that the incident is still open.

What do you actually know at this point?

You know four things, and it helps to list them as plain observations before you reason about them.

First, West checkout failures rose from 0.1% to 6% between 14:18 and 14:26, and fell back to 0.1% within a minute of disabling the new setting. Second, the new setting was enabled in West at 14:18. Third, confirmations in both regions are completing late, with the oldest unfinished one roughly ten minutes old. Fourth, checkout request volume has been steady; there is no traffic surge in either region.

Notice what is not on that list. You do not know why checkout waited on confirmation work at all, you do not know when the confirmation delay started, and you do not know whether anything else changed around the same time. Those gaps are where the hypotheses come from.

A hypothesis is a proposed explanation that makes a prediction you can check. “The new setting caused everything” is a hypothesis. So is “something unrelated to the setting slowed confirmations before 14:18.” Your evidence so far fits both, which is exactly why you cannot stop here.

How do you turn symptoms into hypotheses you can test?

Write each candidate explanation as a sentence that predicts something specific, then ask what evidence would make it wrong. An explanation that no evidence could contradict is not useful during an incident.

For TicketDesk, three candidates fit the facts so far:

H1: the new setting doubled confirmation work and overloaded a shared confirmation system. If true, the confirmation queue should have started growing at 14:18, not before, and the arrival rate of confirmation tasks should have jumped at the same moment.

H2: an unrelated loss of confirmation capacity began before 14:18, and the setting made it worse. If true, the queue should already have been growing before 14:18, and something in the confirmation system itself, such as the number of working hosts, should have changed earlier.

H3: a downstream dependency, such as the email or SMS provider, slowed down. If true, individual confirmation tasks should be taking longer to finish, and the provider's own status or response times should show a change.

Each hypothesis points at a different chart. That is the sign you have written them well. If two hypotheses predicted the same evidence, you would need a sharper question to separate them.

Which check rules out the most explanations?

Start with the queue history, because one chart can distinguish H1 from H2 in a single look. A queue is a waiting list of work; it grows whenever tasks arrive faster than they complete. Ask the confirmation owner for arrivals and completions per minute since 14:00.

PeriodTasks arriving per minuteTasks completing per minuteQueue change per minute
Before 14:082702700
14:08 to 14:18270200+70
14:18 to 14:26324200+124
After 14:26270200+70

Read the table slowly, because it answers two questions at once. Completions dropped from 270 to 200 per minute at 14:08, ten minutes before anyone touched the setting. That is strong evidence for H2 and against “the setting caused everything.” Arrivals then jumped from 270 to 324 at 14:18 and fell back at 14:26, which is exactly what H1 predicted for the setting's contribution. Both hypotheses are partly right: the capacity loss came first, and the setting made a bad situation worse.

The arithmetic also tells you how big the problem is. Between 14:08 and 14:18 the queue grew by 70 tasks a minute, so about 700 tasks were waiting when the setting was enabled. Between 14:18 and 14:26 it grew by 124 a minute for eight minutes, another 992 tasks, for a total near 1,692. At 70 a minute since then, it reaches roughly 1,972 by 14:30. At the current completion rate of 200 a minute, that is about ten minutes of waiting for the newest task, which matches what support is seeing.

West's share of bookings explains the 14:18 jump. West handles about 20% of checkouts, so 20% of 270 is 54 tasks a minute. Doubling West's confirmation work takes that to 108, and 270 plus 54 more is 324. The numbers fit the mechanism, which is more convincing than the timing alone.

What changed at 14:08?

Something reduced completions by about a third, so look at the confirmation workers themselves. A worker is a program that takes a task from the queue, performs it, and records the result. TicketDesk runs three confirmation hosts, each completing about 100 tasks a minute.

The platform engineer checks the host list and finds that one of the three stopped reporting at 14:08. Nobody was alerted, because the only alert on that system fires when the queue exceeds a fixed size, and the queue grew slowly enough to stay under that limit until the setting pushed it over. Two hosts remain, which explains the 200 completions a minute almost exactly.

This is worth pausing on. The lost host would have been found eventually, but it was found now because you refused to accept “the change did it” as a complete answer. In the postmortem lesson you will turn this gap into an action item. For now, record it as a confirmed contributing cause.

Why did a confirmation setting break checkout?

This is the remaining puzzle, and it matters because the answer decides whether re-enabling the setting later is safe. Confirmations happen after checkout, so a slow confirmation queue should delay tickets, not fail bookings.

The checkout owner reads the code path and finds the mechanism. When the new behaviour is on, checkout does not simply hand the confirmation task to the queue and return. It waits for the queue service to acknowledge the richer confirmation request, with a three-second timeout, before reporting success. Under the overloaded queue, some acknowledgements took longer than three seconds, so checkout returned an error even though the reservation and the payment had already been saved.

That explains the worst customer reports. Those customers were charged, their seats were held, and the app showed them a failure. Many of them tried again. Whether those retries created duplicate reservations depends on whether checkout treats a repeated attempt for the same basket as the same booking. That question goes straight to the top of the list for the mitigation lesson.

Notice that H3, the slow provider, has quietly dropped out. Task durations have been steady and the provider shows no change. You do not need to prove H3 false in a courtroom sense; you need evidence that makes it unlikely enough to stop spending time on it, and the stable task durations do that.

How do you keep investigating without stalling the response?

Separate the questions by who needs the answer and when. Some questions decide the next action in the next five minutes: how many bookings are paid but unconfirmed, and can confirmation sends be safely repeated? Others improve the eventual fix but can wait an hour: why was there no alert on worker count?

Give each open question an owner and write it in the incident note with the time it was raised. A question with no owner is a question nobody is answering. The lead, which is you right now, should not be the person running every check; your job is to keep the list honest and make sure results come back to the note.

Record what each result rules out, not only what it suggests. “Provider latency flat since 13:00, rules out H3 as a cause of the delay” is a line the next responder can trust. “Provider looks fine” is a line they will have to re-check.

Exercise: write the investigation log

Using only the evidence in this lesson, write the 14:35 investigation summary for the incident note. Include the confirmed causes, what each piece of evidence ruled out, the open questions, and who owns each one. Then compare with this version:

14:35. Two contributing causes confirmed. (1) Confirmation host 3 stopped at 14:08; completions fell from 270 to 200 per minute; no alert fired. (2) The West setting enabled at 14:18 doubled West confirmation tasks and made checkout wait up to three seconds for queue acknowledgement, causing the 6% failure rate. Disabling the setting at 14:26 fixed checkout; it did not restore capacity. Provider latency has been flat since 13:00, which rules out a downstream slowdown. Queue is near 2,000 tasks and still growing by about 70 a minute. Open: count of paid-but-unconfirmed bookings (Priya, checkout); whether repeated confirmation sends are safe (Dev, confirmations); whether failed checkouts created duplicate reservations (Priya).

Explain your reasoning for the order of the open questions. The safety of repeated sends decides whether any recovery can start, so it belongs first. In an interview, saying which evidence ruled an explanation out is worth more than listing every dashboard you would open.

Quick answers

What is the difference between a symptom and a cause in an incident?

A symptom is what you can observe, such as a rising error rate or a growing queue. A cause is the mechanism that produced it. Incidents often have more than one contributing cause, and a mitigation can remove a symptom while a cause remains.

Should you assume the most recent change caused the incident?

Treat the most recent change as the first hypothesis to test, not as the answer. Check whether the symptom started before the change; if it did, something else is involved, even if the change made it worse.

How do you know when to stop investigating during an incident?

Stop when the remaining uncertainty no longer changes your next action. Once you know enough to choose a safe mitigation and to verify recovery, deeper questions about why the failure was possible belong in the postmortem.