SLI vs SLO vs SLA comes down to three layers: a service level indicator (SLI) measures how well a service is working, a service level objective (SLO) is the target you want that measurement to meet, and a service level agreement (SLA) is an agreement with a customer about service levels and what happens if the commitment is missed. The SLI is the number, the SLO is the goal for that number, and the SLA is the promise built on top of it.
For TicketDesk, our fictional booking app, these could mean measuring successful bookings, aiming for a chosen success rate, and agreeing on a remedy if customers receive less than the promised service. Start with the measurement. The other two depend on knowing exactly what you are counting, so we spend most of this lesson there.
If you have not read the introduction to SRE yet, that is a good place to start. You do not need any of the arithmetic in this lesson memorised; what matters is being able to explain each step.

- 01Measure bookings
- 02Set the target
- 03Agree the commitment
What should an SLI measure?
A good service level indicator measures what the customer actually experiences, not just whether a computer is switched on. Imagine trying to buy a ticket. The page opens, you select a seat, and then payment hangs. A computer may still be running, but your booking has failed. A useful SLI describes that experience.
A request is a message asking the software to do something. Pressing the booking button sends a request. TicketDesk could measure the percentage of valid booking requests that finish correctly, without a software failure or a timeout. A timeout means the result did not arrive within the allowed waiting time.
“Valid” needs a definition. Asking for an available seat with acceptable payment details is different from asking for a seat that has already sold out. A correctly explained sold-out result need not count as a software failure. The team must agree on these rules before anyone uses the numbers, otherwise two people can look at the same day and report different results.
You also need to decide where to collect the measurement, what happens when a measurement is missing, and how repeat attempts are counted. Those choices change the answer, so write them down alongside the indicator itself.
Is an attempt the same as a completed booking?
No, and keeping the two separate is one of the most useful habits you can build. You press Book, wait, try again twice, and finally receive a ticket. Counting messages might give two failed requests and one successful request. Counting your whole journey might give one completed booking, with a long delay.
Both views describe something useful. Keep their names and rules separate. Otherwise the team could make a bad result appear better simply by switching what it counts, and nobody would notice until a customer complained.
A successful checkout response can also arrive before the final confirmation. If confirmation matters to customers, measure that step too. Checking only the first response could miss people who paid but never received their tickets, which is exactly the failure that upsets customers most.
How do you write an SLO?
A service level objective states which attempts count, what a good result looks like, how many must be good, and over what period. For this exercise, assume TicketDesk chooses:
At least 99.9% of valid checkout requests finish correctly without an internal failure over a rolling 30-day window, measured where checkout receives the request.
“Rolling” means looking back 30 days whenever you calculate the result. Tomorrow's window moves forward by a day, so the oldest day drops out and a new one comes in.
Latency is the time taken to get a response. TicketDesk could add a separate objective for it:
At least 99% of valid checkout requests finish correctly within 800 milliseconds over the same window.
A millisecond is one thousandth of a second, so 800 milliseconds is 0.8 seconds. These targets are invented practice assumptions. Real targets should reflect what customers need and what the service can support. Notice that the latency objective still says “finish correctly”: a quick error does not meet an objective that requires a correct result.
Does TicketDesk meet its targets?
TicketDesk meets its success objective in this example but misses its latency objective, and working through both shows why the two cannot be treated as one. In 30 days, TicketDesk receives 2,000,000 valid requests and 1,600 fail. Subtract the failures, divide by all valid attempts, and express the result as a percentage:
Good requests = 2,000,000 - 1,600 = 1,998,400
Success SLI = 1,998,400 / 2,000,000 = 99.92%
That meets the 99.9% success objective. It does not yet tell us whether people waited too long.
Suppose 25,000 requests took longer than 800 milliseconds. Of those slow requests, 1,000 also failed. When counting everything that was either slow or failed, subtract that overlap so the same requests are not counted twice:
Bad requests = 25,000 + 1,600 - 1,000 = 25,600
Good requests = 2,000,000 - 25,600 = 1,974,400
Correct-and-fast SLI = 1,974,400 / 2,000,000 = 98.72%
TicketDesk misses the 99% correct-and-fast objective. The two objectives overlap, so their failure allowances cannot simply be added together. If that overlap step felt fiddly, you are not alone; it is the part interviewers most often probe. The error-budget lesson explains how to turn one objective into a failure allowance.
Why can an average hide slow customers?
An average hides slow customers because a few very slow bookings barely move the mean when most bookings are quick. If most bookings finish quickly but a few take several seconds, the average may still look good. A percentile describes how much of the measured group falls below a particular value. For example, p99 is a boundary below which roughly 99% of the measured response times fall.
Percentiles are good for investigating delays. For an SLO, counting requests that finish within an agreed limit gives a simple pass-or-fail measurement that anyone can recalculate.
Be careful when combining results. Averaging p99 values from separate computers does not give the p99 for all requests together. An overall percentage can also hide customers in one area who are having a much worse experience. The monitoring lesson works through a case where one failing region is hidden by the overall result.
Where does an SLA fit?
A service level agreement records an agreed service commitment and the consequences of missing it, such as a remedy specified by the agreement. Its rules, time period, and exclusions may differ from the team's internal SLO, so never assume they match.
The internal objective can be stricter, leaving room to react before the customer commitment is missed. That is a choice the team makes, not a rule. You cannot infer the exact customer promise from an internal chart; read the agreement.
Exercise: make “fast and available” measurable
Write an objective for booking a ticket. Say which requests count, what success means, how long a customer may wait, the percentage required, the period, and where you would measure it.
Check your reasoning against a sold-out seat, a payment timeout, and a checkout response followed by a missing confirmation. Could another person calculate your result without asking what you meant? If the answer is no, tighten the definition until it is. Explaining those decisions is useful interview practice, even before you know which monitoring tool would collect the numbers.
Quick answers
What is the difference between an SLI and an SLO?
An SLI is the measurement, such as the percentage of valid booking requests that succeed. An SLO is the target for that measurement over a stated period, such as a required success percentage across a rolling window.
Is an SLA the same as an SLO?
No. An SLO is an internal target the team sets for itself. An SLA is an external agreement with a customer that includes what happens if the commitment is missed, and its rules and exclusions can differ from the internal objective.
Why use percentiles instead of averages for latency?
An average can look healthy while a small group of customers waits many seconds. A percentile such as p99 shows the boundary that the slowest customers fall beyond, which an average hides.