A good incident status update says what customers are experiencing, what you have done about it, what you have not yet confirmed, and when the next update will arrive. It does not guess at causes in public, promise a repair time the evidence does not support, or shrink the impact with words like “a small number.” After the incident, the same discipline turns into a postmortem whose action items someone can finish and someone else can verify.

This lesson finishes the TicketDesk incident that began in what to do when an alert fires. TicketDesk is a fictional service; the numbers, names, and policies are invented for practice. It is 15:15. Lena is the incident lead, the confirmation queue has been healthy since 15:03, and the orphaned confirmations have been delivered. Sam is about to send the next customer update.

The beginner lesson on blameless postmortems explains why the review avoids personal blame. This lesson shows what the review looks like for an incident you have followed from the first alert.

Tux posts a notice for waiting customers while another penguin records follow-up actions inside the ticket office.
Tell customers what the evidence supports, then turn the findings into actions someone can verify.

What should the 15:15 update say?

It should describe the current customer experience, the actions taken, the checks still running, and the next update time, in that order. Here is the version Sam sends to customer support and posts on the status page:

15:15. Bookings are confirming normally again. Customers who booked between 14:08 and 15:03 may have received their confirmation later than usual, in some cases by up to 17 minutes. Customers who saw an error in the West region between 14:18 and 14:26 after paying have now been sent their confirmations; nobody was charged twice for the same booking. We are watching the service for the next 30 minutes before we consider this resolved. Next update by 15:45.

Compare it with the 14:35 update from the first lesson. That one said “we have not yet confirmed the status of every delayed reservation.” This one can say the reservations are confirmed because the delivery receipts and the duplicate check exist. The confidence in the wording tracks the evidence, which is the whole skill.

Three things are deliberately absent. There is no explanation of the host failure or the setting, because customers need to know what happened to their booking, not to your infrastructure, and because the postmortem has not been written yet. There is no promise that the incident is over. And “some cases by up to 17 minutes” replaces any vaguer phrase, because support will be asked “how late?” and deserves a number.

How do internal updates differ from customer updates?

Internal updates carry the causes, the decisions, and the open questions; customer updates carry the experience and the next step. Both come from the same incident note, which is why keeping it current all afternoon was worth the effort.

The 15:15 stakeholder update is Lena's, not Sam's, and it reads differently:

15:15 internal. Severity 2, recovery in progress. Queue healthy since 15:03; six hosts running, three temporary, budget approved until 18:00. Orphaned confirmations: 1,137 delivered, 3 rejected for invalid email, with support. Duplicates: none. Causes: host 3 disk failure at 14:08 with no alert; West setting at 14:18 doubled confirmation tasks and blocked checkout on queue acknowledgement. Setting remains disabled pending a checkout change. Observation period ends 15:33. Postmortem owner: you.

Note that the postmortem has an owner before the incident is closed. If that assignment waits until tomorrow, the people with the clearest memory will have moved on.

When is the incident actually over?

The incident is over when every closing condition you wrote down earlier has held for the agreed observation period, not when the dashboards look calm. TicketDesk's conditions, written at 14:55, were: the oldest queued task under one minute old for thirty minutes, the paid-but-unconfirmed query near zero with any remaining rows explained, and the same-basket duplicate check at zero.

At 15:33 Lena checks each one. The oldest task has been under a minute since 15:03. The paid-but-unconfirmed query returns three rows, which are the three invalid-email bookings now with support. Duplicates are zero. West checkout failures have been at 0.1% since 14:27. Every condition holds, so Lena declares the incident resolved at 15:35, posts it in both channels, and Sam sends the closing update:

15:40. Resolved. Bookings and confirmations have been working normally for more than 30 minutes. If you booked this afternoon and have not received a confirmation, or if you believe you were charged for a booking you did not make, please contact support with your booking reference.

Resolved is not the same as finished. The temporary hosts still need removing at 18:00, the setting is still disabled, and a checkout change is still required before it can return. Each of those gets an owner and a date in the note before anyone goes home.

How do you state customer impact honestly?

State impact in terms customers would recognise, give counts with their definitions, and separate what you measured from what you estimated. TicketDesk's impact section is written at 15:50 while the numbers are fresh.

Measured: 1,200 West checkout attempts returned an error between 14:18 and 14:26. Of those, 1,140 had captured payment and held a reservation, and received their confirmation between 14:55 and 15:05, roughly 30 to 45 minutes late. 60 attempts failed before payment and were not charged. 23 customers completed a second, separate booking for the same event after seeing the error; support is contacting them about refunds under the usual policy. 3 confirmations could not be delivered because of invalid email addresses.

Estimated: about 13,000 confirmations were delivered later than the two-minute expectation between roughly 14:14, when the queue delay first passed two minutes, and 15:03. The estimate multiplies the arrival rate of 270 tasks a minute by the affected period; the exact count would need a query against delivery timestamps.

Not affected: East checkout, reservation and payment records, and ticket validity. Saying what was not affected is as useful to a reader as saying what was.

Resist rounding 1,140 down to “around a thousand” or 13,000 up to “tens of thousands.” Both are easier to type than to defend.

What does the postmortem need to explain?

A postmortem needs to explain why each contributing cause was possible, not only what happened, because the action items come from the “why it was possible” layer. TicketDesk's timeline is already in the incident note; the analysis adds five findings.

The host failure at 14:08 went unnoticed for more than twenty minutes because the only confirmation alert watched queue size against a fixed limit, and the queue grew too slowly to reach it. Nothing watched the number of healthy hosts or compared completions with arrivals.

Three hosts at 90% utilisation left no room to lose one. The capacity lesson calls this the missing spare: the service was sized for the normal day, not for the normal day minus a host.

Checkout blocked on a queue acknowledgement with a three-second timeout when the new setting was on. A slow confirmation path should delay tickets; it should never fail a booking whose payment has already been taken.

The setting was enabled for all of West at 14:18 without a limited rollout and without checking confirmation capacity first. The release lesson would have started with a small group and a comparison.

The first signal that confirmations were late came from customer support, not from monitoring. That ordering is itself a finding.

None of these findings names a person as the cause, and none needs to. Ravi enabling the setting was the trigger, and the review treats it as the moment a weak system was tested, not as a mistake to punish. If the analysis had stopped at “Ravi should have been more careful,” every other finding would have stayed hidden.

What makes an action item worth writing?

An action item is worth writing when a named owner can finish it by a date and someone else can check that it had the intended effect. “Add more monitoring” fails that test. “Alert when the oldest queued confirmation exceeds two minutes for five minutes, owner Dev, by 24 October, verified by a drill that stops one host” passes.

TicketDesk's list, with the finding each one addresses:

Alert on confirmation health: page when the oldest queued task exceeds two minutes for five minutes, or when fewer than three hosts report healthy. Owner Dev. Verified by stopping one host in a drill and confirming the page arrives.

Capacity with a spare: run four confirmation hosts as the normal configuration and document the N+1 rule in the runbook. Owner Noor. Verified by the same drill showing no queue growth with one host stopped.

Remove the blocking call: checkout hands the confirmation task to the queue without waiting for the richer acknowledgement, and a reconciliation job every five minutes enqueues any paid booking still missing a confirmation. Owner Priya. Verified by a test that delays the queue and confirms checkout still succeeds and the booking still confirms.

Rollout rule for critical-path settings: any setting that changes checkout or confirmation behaviour is enabled for a small group first with a stated capacity check. Owner Ravi. Verified by review of the next two such changes.

Re-enable the setting only after the checkout change ships and the limited rollout rule is used. Owner Ravi, dependent on Priya's item.

Five items, each with an owner, a verification, and a finding behind it. A postmortem with twenty items and no owners produces nothing; this one can be checked in a month, and the reliability testing lesson describes the drill that does the checking.

Exercise: write the closing update and one action item

Suppose the duplicate check at 15:33 had found four customers with two confirmed reservations for the same basket. Rewrite the 15:40 customer update, and write one action item that addresses the finding. Then compare with this version:

15:40. Resolved. Bookings and confirmations have been working normally for more than 30 minutes. Four customers received two reservations for one booking; we are contacting each of them directly and will refund the duplicate. If you booked this afternoon and have not received a confirmation, please contact support with your booking reference.

Action item: identify why the basket idempotency key did not prevent four duplicate reservations, fix the gap, and add a daily duplicate-reservation report. Owner Priya, by 24 October, verified by replaying the four affected baskets in a test environment and confirming one reservation each.

Explain your reasoning for naming the duplicates in the public update rather than handling them quietly. Customers compare notes, and a refund that arrives before the complaint is cheaper than one that arrives after. In an interview, being asked to write an update is common; the answer they want states the impact plainly and never promises what the evidence cannot support.

Quick answers

What should an incident status update include?

Current customer impact, actions taken so far, what is still being checked, and the time of the next update. Leave out speculation about causes and any repair time you cannot support with evidence.

When should you declare an incident resolved?

When the closing conditions you wrote down during the response have all held for an agreed observation period. A quiet dashboard for a few minutes is not a closing condition.

What makes a good postmortem action item?

A specific change with a named owner, a due date, the finding it addresses, and a way for someone else to verify it worked. If it cannot be verified, it is a wish rather than an action.