Incident mitigation means choosing the action that reduces customer harm fastest, with the smallest chance of causing new harm, and then checking that it worked. Rolling back, restarting, adding capacity, and repairing records are all mitigations, and none of them is automatically the safe choice. The safe choice is the one whose effect you can predict, whose risks you have checked, and which you can reverse or stop if the prediction is wrong.

This lesson continues the TicketDesk incident from investigating without guessing. TicketDesk is a fictional service; the numbers, names, and rules below are invented for practice. It is 14:38. You know that one of three confirmation hosts died at 14:08, that a West setting doubled confirmation work and caused checkout failures until you disabled it at 14:26, and that the confirmation queue is still growing by about 70 tasks a minute.

The beginner lessons on safe releases and data integrity explain rollback compatibility and correct records in a calmer setting. Here you will apply both under time pressure.

Tux directs helpers bringing an extra ticket sorter beside a roped-off reverse lever and a protected stack of envelopes.
Add capacity where it helps, protect accepted work, and check the risks before reaching for rollback.

What harm is happening right now, and to whom?

Before comparing actions, write down each group of affected customers and what they need, because different actions help different groups. At 14:38, TicketDesk has three groups.

The largest group is everyone whose booking succeeded but whose confirmation is sitting in the queue. The queue is near 2,500 tasks and growing. These customers are waiting, not harmed permanently, but every minute of waiting produces more support contacts and more retries.

The second group is the customers behind the 1,200 failed West checkouts between 14:18 and 14:26. Priya, the checkout owner, has run a query: 1,140 of those attempts saved a reservation and captured payment before checkout timed out, and 60 failed before payment was taken. The 1,140 have been charged and shown an error. Nothing in the system is going to confirm them on its own, because checkout never recorded that it handed the task to the queue.

The third group is anyone who retried after seeing that error. Checkout uses the customer's basket identifier as an idempotency key, which means a repeated attempt for the same basket attaches to the same reservation instead of creating a second one. Priya checks: zero duplicate reservations for the same basket. She also finds 23 accounts that completed a second, separate booking for the same event within the window. Those are not duplicates in the system's sense, but a customer may not see it that way, so the list goes to support for follow-up rather than to engineering for a repair.

That last distinction is worth remembering for interviews. “Are there duplicates?” has a technical answer and a customer answer, and you need both.

Which recovery options are on the table?

List every option you can think of, including the bad ones, because explaining why an option is unsafe is how you show your reasoning. For TicketDesk, five options come up in the incident channel.

Roll back the confirmation service to last week's version. The richer confirmation behaviour shipped in version 2.14 at 13:50, behind the setting you disabled. Rolling back to 2.13 would remove the code entirely.

Restart all confirmation workers. A familiar reflex when something is slow.

Add confirmation hosts. Replace the host that died and possibly add more.

Discard the queue and rebuild it from the database. Start fresh by querying every booking without a confirmation and enqueueing it again.

Re-enqueue the 1,140 orphaned confirmations. Create confirmation tasks for the paid bookings that checkout never handed to the queue.

Several of these can be combined. The question is which ones reduce harm, which ones merely feel decisive, and which ones could make a customer's day worse.

Why is rollback unsafe here, even though a change caused the problem?

Rollback is unsafe when the new version has already written records the old version cannot read, and that is exactly the situation at TicketDesk. Between 14:18 and 14:26, West bookings made under the new setting stored their confirmation requests in the richer 2.14 format. Version 2.13 does not understand that format. If you roll back, every one of those tasks fails when a worker picks it up, and the customers who have already waited longest wait longer still.

Rollback also does nothing about capacity. The queue is growing because two hosts cannot keep up with normal demand, and the version of the software running on those two hosts does not change that arithmetic. A rollback here would be effort spent on the cause that is already mitigated, with a real chance of new harm.

Ask the compatibility question every time: can the old version read what the new version wrote? When the answer is no, or unknown, rollback is not a quick fix. It is a migration performed in a hurry.

Why do restarts and queue rebuilds feel safe but are not?

Restarting all workers does not add capacity. The same two hosts come back and process at the same 200 tasks a minute. Worse, each host holds a lease on the tasks it is working on, and a restart abandons those leases mid-task. Dev, the confirmation owner, confirms that confirmation sends are keyed by booking identifier and that the delivery provider ignores a repeated send for the same key within 24 hours, so a restart probably would not produce duplicate tickets. “Probably would not” is a poor reason to take an action with no benefit.

Discarding the queue and rebuilding it from the database sounds thorough. It has two problems. The rebuild query would take several minutes during which nothing is processed at all, and the queue currently holds tasks for bookings whose records are consistent only because the task exists. If the rebuild query misses a case, those customers lose their place entirely. Rebuilding is a recovery tool for a corrupted queue, not a cure for a slow one.

How do you calculate whether adding capacity is enough?

Divide the backlog by the difference between completion rate and arrival rate; that gives the time to drain. Noor, the platform engineer, reports that host 3 failed a disk, that a replacement takes about eight minutes through the standard provisioning path, and that the same path can add more hosts in parallel. Each host completes about 100 tasks a minute, and arrivals are steady at 270.

By 14:42, the queue will hold about 2,800 tasks. Compare two choices:

Four hosts:  capacity 400, arrivals 270, net drain 130 per minute
             2,800 / 130 ≈ 22 minutes to clear the backlog

Six hosts:   capacity 600, arrivals 270, net drain 330 per minute
             2,800 / 330 ≈ 8.5 minutes to clear the backlog

Twenty-two minutes is acceptable; eight and a half is clearly better for the customers already waiting. Before choosing six, check the next limit downstream. Dev confirms that the delivery provider accepts up to 800 sends a minute on TicketDesk's plan, so 600 completions a minute will not simply move the bottleneck. If the provider limit had been 400, six hosts would have been wasted money and a new source of errors.

Write the decision with its expected result and its stop condition: “14:43. Adding three confirmation hosts, six total, owner Noor. Expect queue to fall by roughly 330 a minute from about 14:51. Stop and reassess if provider error rate rises above 1% or if duplicate-ticket reports appear.” The three extra hosts are temporary; give them an owner and a removal date before the incident closes, or they will still be running next month.

How do you protect the work that was already accepted?

Accepted work needs its own recovery plan, because capacity alone will not touch the 1,140 orphaned bookings. Checkout never enqueued their confirmations, so no number of hosts will ever send them.

The safe repair is to enqueue those confirmations now, and the reason it is safe is the idempotency key Dev checked earlier. Each send is keyed by booking identifier. If a handful of those bookings turn out to have a task in the queue after all, the provider will ignore the second send. Repeating a safe operation is cheap; missing an unsafe one is expensive.

Sequence it deliberately. Enqueue the 1,140 after the new hosts are processing, not before, so they do not sit behind the existing backlog for another ten minutes. Priya runs the enqueue in batches of 200 and checks the first batch delivers before continuing. And leave the setting disabled: until checkout stops waiting on the queue acknowledgement, re-enabling it would recreate the 14:18 failure under load.

How do you know the mitigation worked?

A mitigation worked when the customer-facing result improves in the way you predicted, for every affected group, and stays improved. At 14:55, check each group separately.

For the waiting customers: the queue should be falling by about 330 a minute and the age of the oldest task should be dropping. If the queue is flat, the new hosts are not taking work. If completions rose but the oldest task is not getting younger, a stuck task is holding the front of the queue.

For the 1,140 orphaned bookings: confirmations delivered, measured by the provider's delivery receipts rather than by “enqueued.” Re-run the paid-but-unconfirmed query after the batches finish; it should return a number close to zero, with any remaining rows explained.

For duplicates: re-run the same-basket duplicate check and watch support for duplicate-ticket reports. A prediction you did not test is a hope.

Only when all three hold, and hold for an agreed observation period, does recovery belong in the next customer update. The coordination lesson picks up at 14:55 with the response now spread across several people and a shift change approaching.

Exercise: defend your mitigation choice

Suppose the provider limit had been 400 sends a minute instead of 800. Decide how many hosts to add, calculate the time to drain the backlog, and say what you would tell the response lead about the trade-off. Then compare with this reasoning:

With a provider limit of 400 a minute, capacity above four hosts is wasted and may cause provider errors. Four hosts give 400 completions against 270 arrivals, a net 130 a minute, so about 22 minutes to clear 2,800 tasks. I would add one host, not three, and tell the lead that the provider limit, not our hosts, sets the recovery time. If 22 minutes is too long, the next lever is reducing arrivals, for example by pausing non-urgent confirmation types, and that needs a product decision.

Explain your reasoning for excluding rollback in one sentence; an interviewer will usually ask. The strongest version names the compatibility problem and the fact that rollback does not add capacity, rather than saying rollback is “risky” in general.

Quick answers

What is the difference between mitigation and a fix?

A mitigation reduces customer harm now, often without removing the cause; a fix removes the cause so the failure cannot recur. During an incident you mitigate first and fix later, unless the fix is the fastest safe mitigation available.

When is rolling back a deployment the wrong choice?

Rolling back is wrong when the old version cannot read data the new version has already written, when the rollback does not address the actual cause, or when you cannot predict its effect. Check compatibility before every rollback, not after.

How do you estimate how long a queue backlog will take to clear?

Divide the backlog size by the net drain rate, which is completion rate minus arrival rate. If the net rate is zero or negative, the backlog will never clear without more capacity or fewer arrivals.