SRE stands for Site Reliability Engineering. It is the engineering work of keeping websites and apps working correctly, so that people can finish what they came to do: stream a show, book a flight, pay a bill, or practise questions for an upcoming interview. A person who does this work is a site reliability engineer, usually shortened to SRE. Google, which coined the term, describes SRE as what you get when you treat operations as a software problem.

Think about buying a train ticket on your phone. You choose a journey, pay, and expect a confirmed ticket. If the app takes your money but never shows you a ticket, something has gone wrong. An SRE investigates problems like this, helps fix them, and then changes the system so the same failure is less likely to happen again.

If you are new to all of this, you are in the right place. This lesson assumes no technical background, and every term is explained the first time it appears.

Tux comic illustrating the sre role.
  1. 01Book a ticket
  2. 02Spot the problem
  3. 03Fix it and prevent repeats
A customer pays but receives no ticket. SREs help restore the booking experience safely and prevent the same problem from recurring.

What is an online service?

An online service is any website or app that lets you do something over the internet. A ticket-booking site, a messaging app, and an online shop are all online services.

Behind every link you click and every screen you tap are computers following instructions written by people. Those instructions are called software. The people who write and improve software are software developers, also called software engineers.

When you tap “Book,” the software has several jobs to do: check whether a seat is available, take the payment, save your reservation, and show you the result. You never see those steps. You just see your ticket.

Any one of those steps can fail, and when it does, someone has to own the problem. That is where reliability engineering begins.

What does “reliable” mean for an app?

A reliable app does what you reasonably expect it to do, often enough and quickly enough to be useful.

For a booking app, opening the home page is only the start. You also need to find a journey, book an available seat, and receive the correct confirmation without waiting forever.

Suppose the page opens instantly but every payment fails. The app is reachable, yet you cannot buy a ticket. That is a reliability problem. So is an app that works fine for a few customers but stops responding when thousands arrive at once for a big sale, the way shopping sites get hammered during Flipkart's Big Billion Days or Amazon's Black Friday.

The simplest way to think about reliability is to ask one question: what are our users trying to do, and can they actually finish it?

How is an SRE different from a software engineer?

Less than you might think. Both write code and both debug failures. The difference is what each role is primarily responsible for. Software engineers build the features of the app. SREs build and run the systems that keep the app working as it grows, as hardware fails, and as new code keeps arriving.

Throughout this track we use TicketDesk, a fictional ticket-booking service, to make the difference concrete.

A product software engineer at TicketDesk might build seat selection and cancellation. Their main responsibility is making those features do what customers need, correctly. An SRE's main responsibility is helping the whole service keep doing those things reliably as more people use it, as servers fail, and as developers ship new features and bug fixes.

Take the missing-confirmation problem. A developer might find a mistake in the booking code and fix it. An SRE might work alongside that developer on the fix, while also setting up a warning for when paid bookings go unconfirmed and building a safe way to recover stuck confirmations without charging anyone twice. Afterwards, they would measure whether customers are now receiving their tickets and whether staff still have to repair bookings by hand.

Developers can do all of that reliability work too. The reason to have an SRE is to give that work dedicated engineering time and sustained attention, so that reliability, availability, and scalability do not get squeezed out by feature deadlines. The two roles overlap; what differs is each role's primary focus.

Does every app need a separate SRE?

No. A small team can build the app and look after its reliability themselves. The work still has to be done, even when nobody has “SRE” in their job title.

A dedicated SRE role starts to make sense when keeping the service dependable needs substantial, ongoing effort. For TicketDesk, repeated booking repairs or failures during every busy sale could justify that focus. Hiring an SRE would not take away the developers' responsibility for making their own software work correctly.

What does an SRE do when something breaks?

Imagine TicketDesk customers have paid but cannot see their tickets. An SRE first needs to understand who is affected and what they are experiencing. Is every booking failing, or only some? Did the payments go through? Were the reservations saved anywhere?

The people investigating can read records of what the software did. They might discover that reservations were saved correctly, but the step that sends confirmations stopped working. That finding points to a safer repair: resend the confirmations without charging customers again.

Then they check whether customers can actually see their tickets now. Watching one computer come back to life is not enough to prove the booking experience has recovered.

The exact repair always depends on the evidence. SRE work is investigation and judgment. There is no single button that fixes every failure.

How does an SRE prevent the same problem next time?

Now imagine someone has to repair those confirmations every single morning. That person helps today's customers, but tomorrow's customers will hit the same wall.

An SRE digs into why confirmations keep getting stuck. They might change the software so that it can safely retry after a temporary failure. “Safely” is the important word: a retry must never create a second booking or charge the customer twice.

After the change, the team checks how many confirmations still fail and how often someone has to step in by hand. If both numbers fall, there is evidence that the change worked.

This is what the engineering in Site Reliability Engineering means: building lasting improvements to how a service works. Responding to a failure is part of the job. Making the next failure less likely, or easier to recover from, is the other part.

How do SREs notice problems?

SREs rely on monitoring, which means continuously collecting information about how a service is behaving. For TicketDesk, useful information includes how many bookings succeed and how long people wait for confirmation.

An alert is a notification that asks someone to investigate a problem. If many customers suddenly cannot book tickets, the service should notify the person who is responsible for helping it recover.

SREs also help prepare for busy periods and check that changes to the app work properly before more customers receive them. Later modules cover each of those practices. The common purpose is always the same: spot trouble early and protect the customer's experience.

Does an SRE work alone?

Rarely. Several people contribute to reliability. Developers understand the software they built. Product managers understand what customers need. SREs bring the focus on keeping the service dependable and recovering from failures quickly.

Before anyone becomes responsible for responding to problems, everyone should agree on who will help, who can change the software, and whether the service is ready for that arrangement. Handing someone responsibility without the access or support they need just makes problems harder to solve.

You may also hear the word DevOps. It describes an approach where the people building software and the people running it work closely together. SRE is one specific, well-defined way to manage reliability within that cooperation. Using a particular tool, or putting “SRE” in a job title, does not by itself make a service reliable.

Why not make everything work perfectly, all the time?

Because failures happen even with careful engineering, and preventing more of them costs time, money, and effort. Teams have to choose where that effort matters most.

On TicketDesk, losing a confirmed reservation is far more serious than briefly hiding a list of suggested events. The people who understand customers' needs and the engineers who understand the service should decide together what level of performance to aim for.

They can then measure whether the service meets that goal. A target for successful bookings guides improvement work; it does not make a duplicate charge acceptable. Keeping records and payments correct always needs care.

The service-level lesson introduces the names for these measurements and goals. You do not need to memorise them yet.

Exercise: explain SRE in your own words

A friend asks: “If TicketDesk's developers can fix missing tickets, why hire an SRE?” Work out your own answer first, then compare your reasoning with this one:

TicketDesk's developers can fix missing tickets. In a small team, they may handle all the reliability work themselves. But if confirmations fail every morning, helping affected customers keeps pulling them away from planned work, such as adding features or fixing other bugs.

An SRE has dedicated time to investigate why confirmations fail and to work with developers on a lasting fix. They also help detect failed bookings sooner and build safe recovery, so customers get their tickets without being charged twice. Then they check whether fewer bookings fail or need manual repair. TicketDesk might hire an SRE when this ongoing work needs dedicated attention alongside product development. Developers still share responsibility for reliability.

Quick answers

What does SRE stand for?

SRE stands for Site Reliability Engineering. It is also used as a job title: a site reliability engineer keeps online services working reliably and uses software engineering to reduce manual repair work.

Is SRE the same as DevOps?

No. DevOps is a broad approach to how development and operations teams collaborate. SRE is a specific engineering discipline, with its own practices such as service level objectives and error budgets, for delivering reliability within that collaboration.

Do I need to be a programmer to become an SRE?

Yes, eventually. SREs write code to automate repairs, build monitoring, and improve systems. You can start learning the ideas in this track without any programming experience, and pick up the coding skills alongside them.