Incident coordination means making sure every urgent task has exactly one owner, everyone works from the same picture of what is happening, and every decision is written down with its expected result. It is a job in its own right. When the person coordinating also tries to run the queries and apply the fixes, both jobs suffer, and the response starts to depend on whoever is loudest in the channel.

This lesson continues the TicketDesk incident from mitigating safely. TicketDesk is a fictional service; the people, times, and policies below are invented for practice. It is 14:55. Three extra confirmation hosts are coming online, the orphaned confirmations are being re-enqueued in batches, and your on-call shift ends at 15:00. Five people are doing technical work, forty people are watching, and someone senior has just asked whether this is “a Sev1.”

The beginner lesson on incident management introduces the roles used here. This lesson is about using them when the response has grown beyond what one person can hold in their head.

Tux passes a handover clipboard to a penguin in a scarf while responders work at separate stations and observers wait behind a rope.
Give each task an owner, keep one shared record, and make the handover explicit.

Who is doing what right now?

Start by writing the roster, because a response you cannot describe is a response you cannot coordinate. At 14:55 the TicketDesk roster looks like this.

You are the incident lead, sometimes called the incident commander. Your job is to hold the overall picture, make or delegate decisions, and keep the incident note current. Priya owns checkout, including the orphaned-booking enqueue. Dev owns confirmations and the delivery provider. Noor owns the platform and the new hosts. Sam owns communications: customer updates, support briefings, and the stakeholder channel. Ravi, who enabled the setting at 14:18, is in the channel and available to answer questions about it.

Everyone else is an observer. Observers are welcome, but they do not have tasks, and the response does not wait for them.

In a small incident one person can hold two of these roles. The test is whether each urgent task still has an owner who is not also doing something else urgent. When Priya was running the duplicate check and the orphan query at the same time, the orphan query waited. Splitting the work is not bureaucracy; it is throughput.

What does the incident lead actually do?

The lead keeps the picture whole and the queue of decisions moving; the lead does not run the investigation personally. In practice that means four repeating actions.

First, maintain the incident note. It holds the current status, customer impact, confirmed causes, active mitigations with their stop conditions, open questions with owners, and the time of the next update. A reader who joins at 15:10 should be able to read the note and understand the incident without scrolling the channel.

Second, make the timing decisions. Should the orphan enqueue wait for the new hosts? Is the observation period long enough? These are judgment calls that technical owners should not have to make alone while they are also watching a dashboard.

Third, protect the responders from interruptions. When a product manager asks Dev in the channel how many customers are affected, the lead answers, or redirects the question to the note. Dev keeps working.

Fourth, decide when the incident changes state: when it escalates, when mitigation is complete, when it closes. Those moments need one person to say them out loud.

If you find yourself typing a database query, stop and ask who should be doing it instead. The exception is a tiny response where you are the only engineer, and even then, narrate what you are doing in the note so a second person can pick it up.

How do you keep a crowded channel useful?

Give observers a place to get answers that is not the channel where the work happens. Forty people watching an incident produce questions, suggestions, and well-meaning “have you tried” messages. Each one costs a responder a context switch.

Sam opens a stakeholder channel and posts a short summary there every fifteen minutes, with a link to the incident note. The response channel gets a pinned message: “If you do not have an assigned task, please read the incident note rather than asking here. Questions go to the stakeholder channel.” It feels blunt. It is also the single most effective thing a lead can do for the people actually fixing the service.

Suggestions still arrive. Treat them as hypotheses: write them in the note under open questions if they are plausible, assign an owner if they are worth checking, and otherwise thank the person and move on. You do not owe the channel an argument.

How do you record decisions so they can be checked later?

Write every decision with the time, the owner, the expected result, and the condition that would make you stop. You have been doing this since the first lesson; now the habit pays off, because the note already holds the mitigation decisions in a form the next person can act on.

By 15:00 the decision log reads, in part:

14:26 Disabled West confirmation setting (Ravi). Expected: checkout failures stop. Result: West failures 0.1% by 14:27. Confirmed.

14:43 Add three confirmation hosts, six total (Noor). Expected: queue falls about 330 a minute from roughly 14:51. Stop if provider errors exceed 1% or duplicate-ticket reports appear.

14:52 Enqueue 1,140 orphaned confirmations in batches of 200 (Priya). Expected: delivery receipts for each batch within two minutes. Stop if any batch shows duplicate sends.

Notice that each entry predicts something measurable. That is what makes the log useful during the incident and not only afterwards. If the queue does not fall from 14:51, you know immediately that something about the host plan is wrong, instead of discovering it at 15:30.

When should you escalate, and to whom?

Escalate when the response needs a person, a permission, or a decision that nobody currently involved can provide. Escalation is not an admission of failure; it is routing.

Three escalations come up at TicketDesk. The provider's delivery receipts are arriving slowly for one batch, and Dev wants the provider's support line opened; that is an external escalation, and Sam owns it because it is communication. Noor needs approval to run six hosts beyond the usual budget for the rest of the day; that goes to whoever owns the platform budget, and the lead asks for it with the drain calculation attached. And the senior observer's question about severity needs an answer.

For this exercise, TicketDesk's practice policy defines severity by customer impact: paid bookings without confirmation for more than thirty minutes is severity 2, and severity 1 is reserved for checkout being unavailable or records being lost. Checkout is working and no records are lost, so this is a severity 2 incident, and the lead says so in both channels with the reason attached. Severity labels differ between companies; what matters in an interview is that you tie the label to the stated policy and the measured impact, rather than to how stressful the afternoon feels.

How do you hand over an incident at a shift change?

Hand over by giving the next lead the same picture you have, in writing, with a short overlap, and then stepping back completely. Your shift ends at 15:00 and Lena is taking over. A handover done badly loses more time than any single technical mistake in this incident.

Do not hand over in the middle of an action. The orphan enqueue is running in batches; finish the current batch, or make Priya the clear owner of its completion, before you leave. Then write the handover from the incident note. A good one fits on one screen:

Handover 15:00, from you to Lena. Status: mitigations applied, recovery in progress, not yet verified. Customer impact: confirmations delayed by up to 17 minutes for bookings since 14:08; 1,140 West bookings charged without confirmation between 14:18 and 14:26, enqueue in progress; 23 customers with second bookings referred to support. Causes: confirmation host 3 failed 14:08; West setting at 14:18 doubled tasks and made checkout block on the queue. Active: six hosts (Noor), stop if provider errors above 1%; orphan enqueue batch 4 of 6 (Priya), stop on duplicate sends; setting stays disabled. Open: three undelivered confirmations from batch 2 (Dev); budget approval for extra hosts (lead). Next customer update 15:15 (Sam). Done means: queue under one minute old for 30 minutes, paid-but-unconfirmed query near zero with remaining rows explained, duplicates zero.

Read it aloud with Lena for five minutes, answer questions, then post “15:05. Lena is incident lead” in both channels and update the note. After that, you are an observer. Answering questions from the sidelines for the next hour is tempting; it also means two people think they are leading.

What does progress look like by 15:20?

Progress is each predicted result arriving on time, and Lena checks them one by one against the decision log. The queue began falling at 14:51 and reached zero at 15:03; the oldest task has been under a minute old since then. All six orphan batches have delivery receipts: 1,137 delivered, 3 rejected by the provider because the customer's email address is invalid, and those three go to support with the booking details. The same-basket duplicate check still returns zero. West checkout failures have been 0.1% for nearly an hour.

That is every prediction confirmed, which means the incident can move toward closing. It does not mean it is closed. The agreed observation period is thirty minutes of healthy results, so the earliest the incident ends is 15:33, and the final customer update waits for that. Writing those updates, and turning what you have learned into changes, is the subject of the next lesson.

Exercise: write the handover for a different moment

Suppose your shift had ended at 14:40 instead, before the host decision was made and while the duplicate check was still running. Write the handover note Lena would need at that moment. Then compare it with the 15:00 version above and explain your reasoning about what changed: which items move from “active” to “open,” and which decisions you would make before leaving rather than leaving for Lena.

The strongest answers hand over decisions that are ready to be made, rather than deferring them, and are explicit about which checks are incomplete. In an interview, describing a handover you ran well is a strong story, because it shows you understand that an incident is a shared piece of work rather than a personal performance.

Quick answers

What does an incident commander do?

The incident commander, or incident lead, keeps the overall picture, assigns owners, makes timing and escalation decisions, and maintains the shared incident record. They coordinate the technical work rather than doing it themselves.

How many people should be in an incident channel?

As many as have tasks, plus a scribe and a communications owner. Everyone else should read a shared incident note and ask questions in a separate stakeholder channel so responders are not interrupted.

How do you hand over an incident to the next on-call engineer?

Finish or clearly assign any in-progress action, write a one-screen summary covering status, impact, causes, active mitigations with stop conditions, open questions with owners, and the definition of done, walk through it together, then announce the new lead and step back.