SRE release engineering is the work of delivering new software versions in a repeatable, controlled way. A safe release limits how many customers a change can affect, checks their results against a group that did not get the change, and includes a tested way to recover if something goes wrong.

Imagine TicketDesk, our fictional booking app, changing how it sends booking confirmations. Trying that version with a small group first lets the team learn about failures before every customer receives it. This lesson walks through that process from the first build to the moment you decide whether to stop.

Tux comic illustrating safe changes and releases.
  1. 01Try with a small group
  2. 02Compare booking results
  3. 03Continue or stop
A small first group limits exposure. Compare its bookings before giving the new version to more customers.

What changes when an app is released?

When an app is released, a new version of its instructions starts handling real customers, which is why the team must know exactly which version did what. Software is the set of instructions computers follow. Developers change those instructions to fix problems or add features. A build packages a particular version for use; deployment puts it into the environment where it will run. Production means the environment serving real users.

The team needs to know which version and settings handled a booking. That makes a later failure much easier to connect with a change. Repeatable build and deployment steps reduce accidental differences between versions.

Review effort should follow the consequences. Changing a page label is different from changing stored reservations. Both need an owner and observation, while a reservation change needs stronger checks for correctness and recovery.

What is a canary release?

A canary release gives a limited group of requests or users the new version before more people receive it, so problems show up in a small group first. Compare their outcomes with an unchanged group running at the same time, called the control.

Also check absolute limits. Two groups failing equally badly would not be a reason to continue. A comparison shows whether the change made results worse, while an absolute check asks whether those results are acceptable at all.

Choose the group carefully. A small amount of browsing traffic might never exercise the booking feature being changed. TicketDesk needs evidence from creating reservations, sending confirmations, correctly declining unavailable seats, and handling repeat attempts.

Before starting, agree on the observation period and the stop conditions. The person responsible can then stop expansion when a check fails, rather than negotiating a rule while customers are affected.

Would you expand this new version?

No: in this example the new version fails ten times as often as the unchanged one, and the agreed rule says to stop. For this fictional exercise, the new confirmation version handles 20,000 valid requests and records 200 internal failures. The unchanged version handles 180,000 requests with 180 failures in the same period.

Canary error rate = 200 / 20,000 = 1%
Control error rate = 180 / 180,000 = 0.1%
Observed canary error rate is ten times the control rate

A regression means performance became worse. Assume the agreed rule stops expansion for a well-sampled regression of this size or a failure of the absolute limit. The release owner should pause and investigate.

These results show worse outcomes without proving the cause. Preserve the comparison data and version details. Use the tested recovery plan when appropriate, then check affected bookings and any confirmations still waiting. There is no need to wait for the full 30-day reliability objective to be missed before acting on a visible release problem.

Does rollback undo everything?

No. Rollback returns to an earlier software version, but it does not automatically undo changes to stored information or actions already performed elsewhere. It can help when the new version caused trouble, and that is the limit of what it does.

Suppose today's software writes reservations in a format yesterday's software cannot read. Restoring yesterday's version leaves those new records in place. A payment already sent to an external service also remains a payment.

Plan compatibility between versions. One approach introduces a format both versions can read, changes usage gradually, and removes the old format only after it is no longer needed for recovery.

A feature flag is a setting that switches a behaviour on or off. It can limit further harm, but only if disabling it actually stops the risky work. Test how the setting reaches all running copies and who can change it. Sometimes repairing the affected records or releasing a fix is safer than restoring the old version. The data-recovery lesson explains how to check those records before reopening.

What makes a launch ready?

A launch is ready when the team has evidence that the service can handle expected demand, its dependencies and monitoring are in place, and a responder can recover from the likely failures. A launch is a planned introduction of a service or feature. Check whether tested capacity can handle the expected demand, including the relevant failure case. Check the services it depends on, the monitoring, and the responder's ability to recover.

A production readiness review brings that evidence together. Important gaps need owners and decisions. Some gaps may be acceptable for a limited launch; an untested change to critical booking records may block it.

Additional safety mechanisms also need maintenance. Be able to explain which failure each one handles and how a responder will know it is working. A complicated plan that nobody can operate safely is a weak recovery plan.

Exercise: installation succeeds, confirmations do not

The deployment tool reports success, and the software is running. Confirmation completion falls by 8%. What should happen next?

Explain your reasoning using both observations. The tool reports that installation steps finished; the booking measurement shows missing customer outcomes. Stop expansion, identify affected bookings, and use the planned recovery action if it is safe for the records already written.

Verify that delayed confirmations finish without duplicates. In an interview, walk through the sequence from identifying the change to limiting exposure, comparing results, stopping, and recovering. Include data compatibility so the listener can see how recovery would work after the change has already begun.

Quick answers

What is the difference between a canary and a blue-green deployment?

A canary sends a small share of real traffic to the new version and compares it with the old one before expanding. A blue-green deployment runs two complete environments and switches all traffic from one to the other at once, which gives a fast switch back but no gradual comparison.

When should you roll back a release?

Roll back when the new version is causing harm, the earlier version is known to be safe, and restoring it will not strand records or actions the new version has already created. If stored data has changed, repairing forward may be safer.

What is a production readiness review?

It is a structured check before launch that gathers evidence on capacity, dependencies, monitoring, and recovery, assigns owners to important gaps, and decides whether the launch can proceed as planned or in a limited form.