Keep this nearby while you learn

SRE terms, in plain language.

Use these definitions when a word interrupts your reading. Each links to a lesson that explains the idea through TicketDesk, our fictional booking app. Start with software and service if technology is new to you.

32 terms · Updated October 10, 2026 · No login needed

What is Software?

Instructions that tell a computer what to do. TicketDesk's software checks seats, saves reservations, and sends confirmations.

Worked example: The SRE role

What is Service?

Something people use over the internet, such as a ticket-booking app. Reliability concerns whether they can finish the task they came to do.

Worked example: The SRE role

What is Production?

The environment serving real customers. A change that works in a separate test environment can still fail under production demand or conditions.

Worked example: Safe changes and releases

What is Request?

A message asking software to perform an operation. One person may send several booking requests by trying again, so a request count is not necessarily a customer count.

Worked example: SLIs, SLOs, and SLAs

What is Availability?

Whether a service can be used successfully under the measurement's agreed rules. A reachable home page does not establish that customers can book tickets.

Worked example: SLIs, SLOs, and SLAs

What is Latency?

The time an operation takes. A booking response taking 800 milliseconds has 0.8 seconds of latency. A fast error is still a failed booking.

Worked example: Monitoring and observability

What is Percentile?

A boundary describing part of a measured group. A p99 response time is a value at or below which roughly 99% of measured response times fall; it is not a maximum.

Worked example: SLIs, SLOs, and SLAs

What is SLI?

Service level indicator: a measurement of service performance. TicketDesk might count the fraction of valid booking attempts that finish correctly.

Worked example: SLIs, SLOs, and SLAs

What is SLO?

Service level objective: the target for an SLI over a stated period. It needs clear counting rules, a definition of a good result, and a measurement window.

Worked example: SLIs, SLOs, and SLAs

What is SLA?

Service level agreement: a customer agreement describing service commitments and the consequences of missing them. Its rules can differ from an internal SLO.

Worked example: SLIs, SLOs, and SLAs

What is Error budget?

The failure allowance within an SLO. For 99.9% success across 30 million valid requests, 0.1%, or 30,000 requests, may fail while still meeting that target.

Worked example: Error budgets and risk

What is Burn rate?

The observed failure rate divided by the SLO's allowed failure rate. A 1% failure rate against a 0.1% allowance is a burn rate of 10, not 10% of the budget consumed.

Worked example: Actionable alerting

What is Toil?

Recurring operational work that could be automated and leaves no lasting improvement. Repeating the same manual confirmation repair is an example; investigating a new failure can produce lasting understanding.

Worked example: Toil and automation

What is Monitoring?

Collecting information about a service's behaviour to notice and investigate problems. Measure customer results as well as the computers performing the work.

Worked example: Monitoring and observability

What is Metric?

A numerical measurement recorded over time, such as failed booking attempts per minute. Show the number of attempts behind a percentage when interpreting it.

Worked example: Monitoring and observability

What is Log?

A record of an event in the software. Logs can help explain an individual failure, but should avoid unnecessary private customer information.

Worked example: Monitoring and observability

What is Trace?

A record following one request through different software steps. It helps show where time was spent, although the location of a delay does not by itself prove its cause.

Worked example: Monitoring and observability

What is Saturation?

How close a resource is to its useful limit. A growing waiting list can reveal a limit even when the computers' processor use looks ordinary.

Worked example: Monitoring and observability

What is Dependency?

Another service or component the app relies on. If TicketDesk needs a payment service, that service's failures can affect whether bookings finish.

Worked example: Capacity and overload

What is Queue?

A waiting list of work. It can absorb a brief burst, but keeps growing when work arrives faster than it finishes. Accepted bookings still need a safe completion plan.

Worked example: Capacity and overload

What is Instance?

One running copy of the software. Several instances can share demand, provided the workload can be distributed and shared dependencies can keep up.

Worked example: Capacity and overload

What is Canary?

A limited group receiving a new version before wider rollout. Compare representative customer outcomes with an unchanged group and check absolute safety limits.

Worked example: Safe changes and releases

What is Rollback?

Returning to an earlier software version. It does not automatically undo stored-record changes or external actions, such as payments that already happened.

Worked example: Safe changes and releases

What is Runbook?

Written guidance for investigating known problems and recovering safely. It should explain current context, stop conditions, escalation, and how to check the result.

Worked example: On-call and troubleshooting

What is Mitigation?

An action that reduces current customer harm while investigation continues. Pausing a faulty release may help before the team knows the full cause.

Worked example: On-call and troubleshooting

What is Incident?

A service problem requiring coordinated response. Wrong booking records can qualify even when the website remains reachable.

Worked example: Incident response

What is Postmortem?

A review of an incident's impact, timeline, decisions, and contributing conditions. Useful follow-up actions have owners and checks that show the intended protection works.

Worked example: Postmortems and learning

What is Idempotency?

Repeating the same operation does not create an extra side effect. It can prevent a retry from creating a duplicate booking, but the repeated attempt still consumes capacity.

Worked example: Capacity and overload

What is Invariant?

A rule that must remain true, including during failure and recovery. For TicketDesk, one confirmed reservation should refer to one valid booking.

Worked example: Data integrity and recovery

What is RPO?

Recovery point objective: the maximum acceptable lost-data interval. It concerns how much recent work may be missing after recovery, not how long copying files takes.

Worked example: Data integrity and recovery

What is RTO?

Recovery time objective: the maximum acceptable time to restore usable service, using an agreed start and finish. Include checks and safe reopening when they are part of that definition.

Worked example: Data integrity and recovery

What is Acceptance test?

A check showing that a proposed improvement does what was intended. Replaying a failure and observing the new stop check is stronger evidence than merely creating a chart.

Worked example: Postmortems and learning
Explore the beginner's guide