The library
SRE guides
Browse guides to production debugging, reliability fundamentals, and SRE interview preparation.
Start here · A Beginner's Guide to SRE14 ordered lessons, worked scenarios, and 56 practice questions →- Incident Response and Management10 min read
SRE incident response: what to do when an alert fires
Learn the first steps of SRE incident response with a worked booking outage: assess customer impact, separate facts from guesses, mitigate safely, and verify recovery.
Updated October 10, 2026Read article - A Beginner's Guide to SRE7 min read
Blameless SRE postmortems: turning incidents into improvements
Blameless SRE postmortems explained with a fictional booking incident. Learn what blameless means, what the report contains, and how to write actions.
Updated October 10, 2026Read article - A Beginner's Guide to SRE7 min read
SRE capacity planning: overload and cascading failures
SRE capacity planning explained with simple booking sums. Work out how many instances you need, how fast a queue grows, and why retries make overload worse.
Updated October 10, 2026Read article - A Beginner's Guide to SRE6 min read
Data integrity and disaster recovery: RPO, RTO, and restore tests
SRE disaster recovery explained: what RPO and RTO mean, a timed restore of ticket records, and why live replicas and a successful backup are not enough.
Updated October 10, 2026Read article - A Beginner's Guide to SRE6 min read
Error budgets: balancing reliability and release velocity
SRE error budget explained with booking failures and worked percentages. Learn to calculate budget used, set a policy, and spot what the budget cannot excuse.
Updated October 10, 2026Read article - A Beginner's Guide to SRE6 min read
SRE incident management: roles, communication, and recovery
SRE incident management explained with a booking outage. Learn when to declare an incident, who leads, how to write updates, and how to hand over.
Updated October 10, 2026Read article - A Beginner's Guide to SRE6 min read
SRE monitoring: golden signals and useful observability
SRE golden signals explained with booking examples: latency, traffic, errors, saturation. Learn how metrics, logs, and traces reveal problems an average hides.
Updated October 10, 2026Read article - A Beginner's Guide to SRE7 min read
SRE on-call: troubleshooting with evidence and safe mitigation
SRE on-call troubleshooting explained with a booking failure. Learn to test a hypothesis, pick a safe mitigation, ask for help, and prove recovery.
Updated October 10, 2026Read article - A Beginner's Guide to SRE6 min read
Safe releases in SRE: canaries, rollback, and launch readiness
SRE release engineering explained with a booking app: run a canary, compare it with a control, decide when to stop, and plan recovery when stored records have changed.
Updated October 10, 2026Read article - A Beginner's Guide to SRE6 min read
Reliability testing in SRE: failure drills and production readiness
SRE reliability testing explained with a booking failure drill. Learn to write a testable claim, run a controlled instance loss, and judge launch readiness.
Updated October 10, 2026Read article - A Beginner's Guide to SRE7 min read
SLI vs SLO vs SLA: choosing meaningful service levels
SLI vs SLO vs SLA explained with a ticket-booking example. Learn to define what you count, write a clear objective, and work out success rates yourself.
Updated October 10, 2026Read article - A Beginner's Guide to SRE6 min read
SLO alerting: error-budget burn rates without pager noise
SLO burn rate alerting explained step by step. Calculate burn rate, combine long and short windows, and decide which booking problems deserve an urgent page.
Updated October 10, 2026Read article - A Beginner's Guide to SRE7 min read
SRE interview practice: a reliability design capstone
SRE interview preparation capstone: plan a busy booking launch, work the error budget and capacity sums, judge a canary, and handle a growing queue.
Updated October 10, 2026Read article - A Beginner's Guide to SRE5 min read
Toil in SRE: identifying work worth automating
Toil in SRE explained with repeated booking repairs. Learn to recognise toil, estimate automation payback in weeks, and keep an automated repair safe.
Updated October 10, 2026Read article - A Beginner's Guide to SRE8 min read
What is SRE? The role and engineering mindset
What Site Reliability Engineering is, what an SRE actually does all day, and how the role differs from software engineering, explained with a ticket-booking example.
Updated October 10, 2026Read article - Company interview guides9 min readEditorial preview
Stripe SRE interview preparation: debugging and reliability
How to prepare for a Stripe SRE interview: practise reading unfamiliar code, debug a payments retry bug, rehearse reliability design, and follow a four-week study plan.
Updated October 10, 2026Read preview