An SRE error budget is the amount of failure a service can have while still meeting its reliability target, called a service level objective (SLO). If the target is 99.9% successful booking attempts, the remaining 0.1% is the failure allowance for the agreed period. The budget turns an abstract target into a quantity you can spend, track, and run out of.

Teams compare actual failures with that allowance when deciding whether to introduce more changes or spend more time fixing reliability problems. The budget helps guide a decision; it does not make every kind of customer harm acceptable. We will come back to that limit at the end, because it is the part people most often get wrong in interviews.

This lesson builds on the service-level lesson. If you are comfortable with what an SLO is, you have everything you need.

Tux comic illustrating error budgets and risk.
  1. 0130,000 allowed failures
  2. 0218,000 used
  3. 0312,000 remaining
A 99.9% success target allows 30,000 failures in 30 million attempts. Using 18,000 consumes 60% of that allowance.

Why allow any failure?

Allowing some failure is what makes it possible to keep improving the service without freezing it. A booking app needs updates. People fix defects, improve features, and change how the software works. A release is a new version made available to users. Even a carefully tested release can have a problem that did not appear during testing.

Trying to avoid every possible failure would take ever-increasing effort and could delay useful improvements. An SLO makes the reliability goal explicit. The error budget tells the team how much failure fits within that goal.

Unused budget does not mean the team should deliberately break the service. It can leave room for a limited, carefully observed change. If repeated failures use most of the allowance, that is a signal to give reliability repairs more attention.

How do you calculate a request-based budget?

To calculate a request-based error budget, multiply the total valid requests by the permitted failure fraction, then compare actual failures with that allowance. A request is an attempt asking software to do something, such as book a ticket. For this fictional TicketDesk exercise, count valid checkout requests over 30 days. The target is 99.9% success. There are 30,000,000 valid requests and 18,000 failures.

Allowed bad fraction = 1 - 0.999 = 0.001
Allowed bad requests = 30,000,000 × 0.001 = 30,000
Budget consumed = 18,000 / 30,000 = 60%
Budget remaining = 12,000 requests, or 40%

To get the allowance, multiply the total attempts by the permitted failure fraction. To find how much was used, divide actual failures by that allowance.

The observed success rate is 99.94%. Only 0.06% of requests failed, but those failures used 60% of the budget. If those two percentages feel like they disagree, look at what each one divides by: all attempts in the first case, allowed failures in the second. Different denominators, different meanings.

With a rolling window, the calculation always looks back over the latest 30 days. Older attempts and failures fall out as time moves forward. Future traffic can change too, so the remaining 12,000 requests are a result for this measured window, not a guaranteed allowance for tomorrow.

Is the budget a number of minutes?

Sometimes, but only when the SLO is measured in time rather than requests. A time-based target counts how long a service is working. In a simple model where it is either fully working or fully unavailable, a 99.9% target over 30 days permits 43.2 minutes of bad time:

30 × 24 × 60 × 0.001 = 43.2 minutes

The first example counts requests, so it cannot automatically be converted into those minutes. Ten minutes of failure during a busy ticket sale may affect far more attempts than ten minutes overnight. Some bookings could also fail while others succeed, which a simple up-or-down model cannot describe.

State what the SLO measures before doing the arithmetic. This avoids a common interview mistake: treating every 99.9% objective as the same downtime allowance.

What should the team do when the budget runs low?

When the budget runs low, the team should follow a policy it agreed in advance, rather than negotiating under pressure. Agree on that policy before a launch deadline makes decisions harder. A policy describes actions, who decides, and what evidence allows ordinary changes to resume.

TicketDesk might continue normal releases while its objective is met and recent failures are under control. As failures increase, it could expose fewer customers to each new version. If the objective is missed, it could pause optional risky changes while repairing the service.

Urgent security fixes and changes that restore reliability need an agreed exception process. A policy that says “stop everything” without explaining recovery criteria can block the very work needed to improve the situation.

Look at what used the budget. If a service TicketDesk depends on repeatedly fails, postponing unrelated features each month will not repair that weakness. The policy should create time for a lasting improvement, such as the engineering work described in toil and automation.

What does an error budget leave out?

An error budget leaves out any harm that a percentage cannot make acceptable, such as duplicate charges, lost reservations, or leaked private information. A booking success target does not settle those questions. They need protections suited to their consequences. An acceptable failure percentage cannot justify charging the same person twice.

A total percentage can also hide a smaller group. Suppose a small event fails every time while a much larger event works. TicketDesk might still meet its overall objective, but the small event's customers cannot buy tickets at all. Investigate their experience rather than dismissing it because budget remains.

Exercise: would you approve both changes?

TicketDesk has 40% of its budget left. One proposed change updates suggested events and can be switched off quickly. Another changes how confirmed reservations are stored.

Explain your reasoning before deciding. For the suggestions, trying the change with a small group may give useful evidence with limited impact. For reservation records, ask how the team will check existing bookings and recover if some records change before the process stops.

The release lesson explains a canary, a rollout to a small group before more customers receive the change. For now, compare the affected customers, current failures, ability to detect trouble, and recovery options. State who makes the decision and what result would make them stop. Remaining budget is useful context alongside those checks, not a substitute for them.

Quick answers

What is an error budget in simple terms?

It is the amount of failure your reliability target still allows. If the target is a certain success percentage, everything short of that is the budget, and you track how much of it recent failures have used.

What happens when the error budget is exhausted?

The team follows its agreed error budget policy, which usually means pausing optional risky changes and spending the time on reliability repairs, while keeping an exception process for urgent security and recovery fixes.

Can you convert an error budget into downtime minutes?

Only if the SLO is time-based. A request-based budget counts failed attempts, and the same number of minutes can contain very different numbers of attempts, so the two cannot be swapped without stating the assumption.