SRE reliability testing checks how a service behaves when it is busy, when something fails, and while it recovers. A useful reliability test states a specific claim and then measures whether customers still get the intended result under those conditions. It answers a different question from ordinary feature testing, which asks only whether the code works when nothing goes wrong.
For TicketDesk, our fictional booking app, completing one ordinary purchase in a test is a start. The team also needs evidence that bookings stay correct when customers retry, or when part of the software stops in the middle of a sale.
You do not need any testing tools to follow this lesson. The skill we are practising is deciding what to test and what a result does and does not prove.

- 01Choose a claim to test
- 02Limit the experiment
- 03Check bookings and recovery
What can an ordinary test leave unanswered?
An ordinary functional test leaves unanswered what happens when a response is lost, a dependency slows down, or a booking is interrupted part-way through. A functional test checks whether a feature produces the expected result, such as confirming one reservation. A dependency is another service the app relies on.
Those failure situations change what work is left unfinished, and unfinished work is where lost and duplicate reservations come from. TicketDesk needs to avoid both while still meeting its response-time and success goals.
Choose reliability tests from those requirements. If a repeated attempt could create another ticket, test that path. If accepted confirmation work can stop before it finishes, test whether the stop is noticed and the work is recovered.
Code coverage measures which parts of the software's instructions ran during testing. A recovery path can be fully covered and still never have been exercised after the particular interruption that matters. Coverage alone cannot establish that the important failure cases work.
How do you make a reliability claim testable?
You make a reliability claim testable by stating the exact load, the exact failure, and the exact outcome you expect, so that the result can only be pass or fail. A hypothesis is a proposed claim to test. Building on the capacity lesson, TicketDesk could propose:
At 900 representative checkout requests per second, losing one serving instance will keep success and response time within the agreed limits. Confirmed bookings will not be lost or duplicated.
An instance is one running copy of the serving software. Representative requests resemble the work real customers perform. A test made mostly of simple page views cannot show whether the more expensive booking steps survive that same rate, so the mix matters as much as the number.
Specify the observation period, the mix of operations, where measurements are taken, and the correctness checks. Record the differences between the test setup and the service real users use, called production. That record tells the team where the result applies and where it does not.
How would you test the loss of one instance?
You would test the loss of one instance by establishing a baseline under load in an isolated environment, removing one copy on purpose, watching customer outcomes, and then restoring and checking that delayed work completes correctly. Start somewhere the experiment cannot unexpectedly affect customers. A plan for this fictional exercise is:
- Run six serving instances with the chosen workload and record normal results as a baseline.
- Check that measurements, alerts, and recovery instructions work.
- Remove one instance using the environment's test mechanism.
- Watch completed requests, response times, failures, pending confirmations, and booking correctness.
- Stop if the agreed impact limits are crossed and use the recovery plan.
- Restore the baseline and check that delayed work finishes without duplicates.
Decide who watches and who can stop the experiment before you start. Testing against real customer traffic requires the responsible team's authorisation and operating process; never improvise that.
If checkout stays within target but confirmations wait too long, part of the claim has failed. The customer's journey includes receiving the ticket, so a checkout-only chart would miss that weakness entirely.
What should you do when the test fails?
When a reliability test fails, use the evidence to locate the mechanism, fix that mechanism, and repeat the same test under comparable conditions. The surviving instances may lack capacity. Demand may be spread unevenly. A recovery action may be overloading a shared confirmation service.
Make an improvement tied to the mechanism you found, then re-run the relevant test. Be honest with yourself here: lowering the demand can produce a pass without solving the original weakness, and that pass is worth nothing.
Check for new costs too. A change that helps recovery might slow normal operation or create more repeated work. Record a specific result, such as confirmation delay crossing its objective after one instance disappeared, rather than only noting that a failure drill failed. Specific results are what the next engineer can act on.
Does a passing test mean the service is ready?
A passing test supports the claim under the conditions tested, and nothing more; readiness is a broader judgement that includes people, process, and recovery. Real customers may bring different demand, data, settings, or combinations of failures. Explain those limits when deciding how much confidence to place in the result.
A production readiness review checks whether the service is ready to be operated for real users. It considers reliability goals, responsibilities, measurements, alerts, capacity, safe releases, incident response, and data recovery.
Important gaps need owners and evidence that they are closed, or an explicit decision to accept the risk. The people responding to urgent problems also need working access, practice, and a route to ask for help. A recovery procedure can be technically possible and still too slow if nobody on duty knows how to perform it.
Exercise: challenge “testing passed, so the sale is safe”
A teammate says the test environment passed, so the launch can proceed. What would you ask?
Find out which operations, how much demand, and which failure were tested. Ask how customer outcomes were measured, how the environment differs from production, and whether alert delivery and recovery were actually exercised.
Prioritise the remaining gaps by expected customer harm. A missing test for losing confirmed bookings may matter more than a minor environment difference. Explain your reasoning so an interviewer can see how you use evidence without asking a test to prove more than it demonstrated.
Quick answers
What is chaos engineering in SRE?
Chaos engineering is the practice of deliberately introducing failures, such as removing an instance or slowing a dependency, under controlled conditions to check that the service behaves as claimed. The key words are controlled and claimed: you state what should happen, limit the impact, and verify the result.
What is a production readiness review?
A production readiness review is a structured check, before a service serves real users, that its reliability goals, ownership, monitoring, alerts, capacity, release process, incident response, and data recovery are in place. Gaps get owners or an explicit decision to accept the risk.
Why is code coverage not enough for reliability?
Code coverage shows which instructions ran during tests, not whether the service behaves correctly after a lost response, a slow dependency, or an interrupted booking. A recovery path can be fully covered and still fail in the one situation that matters.