Toil in SRE is recurring operational work that could be automated and leaves no lasting improvement behind. It usually means repeating a manual repair just to keep a service running. Reducing toil gives people more time to improve the software, so that serving more customers does not always require more manual repairs.
Picture someone fixing missing ticket confirmations every morning. Each repair helps a customer, yet the following morning brings another pile of confirmations. That recurring effort is a reason to investigate how the service could work better, and this lesson shows you how to decide what to do about it.

- 01Repair again tomorrow
- 02Remove the cause
- 03Check time saved
How do you recognise toil?
You recognise toil by asking what remains after the task is finished. If someone clears today's stuck bookings but the same problem returns tomorrow, the work has not left a lasting improvement, and that is the defining sign.
Toil is often manual, repetitive, and done in reaction to a problem. It tends to grow as the service grows. A person pressing a button to run a repair can still be doing toil, even if the computer performs most of the steps. Someone still has to notice the problem and start the repair.
Record how often the work happens and how much hands-on time it takes. Also record interruptions and the consequences of a mistake. This gives the team a clearer basis for choosing improvements than a general complaint that everyone is busy.
Is every task that runs a service toil?
No. Operational work that teaches the team something or leaves the service better is not toil. Investigating a new failure can teach the team why it happened. Building a safer way to recover can improve future operation, including the protections for correct booking records. Those tasks are different from repeating a repair whose steps and outcome are already familiar.
A repetitive step may also contain an important safeguard. Suppose a person checks whether a ticket already exists before repairing a confirmation. Removing that check could create duplicate bookings. Understand what a task protects before deciding to eliminate it.
Automation means making software perform work that previously needed a person. It is useful when the work can be performed safely, but it also needs maintenance. A repair tool that breaks every week creates more work of its own.
How do you estimate whether automation is worth it?
You estimate whether automation is worth it by comparing the weekly time it saves against the time it costs to build and maintain. In this fictional TicketDesk example, someone checks and repairs waiting bookings 12 times a week. Each intervention takes ten minutes. An automated process would take 20 engineering hours to build and half an hour a week to maintain.
Current weekly effort = 12 × 10 minutes = 2 hours
Net weekly saving = 2 - 0.5 = 1.5 hours
Simple payback = 20 / 1.5 ≈ 13.3 weeks
First convert the current effort to hours. Subtract the expected maintenance from the time saved each week. Then divide the build time by that weekly saving.
On these assumptions, the saved effort pays back the build effort after roughly 14 weeks. Extra testing, training, or handling unusual cases could extend that period, so include them when they matter.
Time is only part of the decision. Two hours of work spread across interrupted nights can be harder to sustain than two planned daytime hours. A short repair that occasionally creates a wrong booking may need attention before a longer but safer task.
Should you automate the repair or remove its cause?
Where you can, remove the cause; automate the repair when the cause cannot be fixed quickly or completely. Suppose a booking passes through several steps: payment, reservation storage, then confirmation. A defect between those steps can leave the software unsure whether to continue. Engineers call the move from one step or condition to another a state transition.
Fixing that transition might prevent most stuck bookings. Automating the existing repair could provide faster relief while the lasting fix is developed. Detection and safe recovery may still be needed for failures that cannot be prevented completely.
Choose the approach using the problem's cause and consequences. Saying “I would write a script” in an interview leaves out that decision. A script is a small program; the useful explanation is what it should change and how you would know the change helped.
What makes an automated repair safe?
An automated repair is safe when it can be interrupted, repeated, and inspected without creating new damage. Imagine the repair updates half the bookings and then stops. Can it resume without creating another ticket for the bookings already repaired? Can someone see which records still need attention?
The design should answer those questions. Check the input, record the result of each operation, and limit how much work runs at once. Give the tool only the access it needs and a way to stop if results look wrong.
A dry run shows intended changes without applying them. It helps a person review the plan, but it does not prove that recovery works after an interrupted live repair. Test interruptions, repeat attempts, old input, and unusual cases as well as the normal path. The reliability-testing lesson explains how to make those checks controlled experiments.
After the repair runs, check booking correctness alongside the number of manual interventions. Fewer button presses would be a poor improvement if more customers received duplicate tickets.
Exercise: which problem comes first?
Task A takes four hours a week and rarely causes harm. Task B takes one hour and sometimes duplicates reservations. You have one week to improve the service.
Explain your reasoning using customer risk as well as time saved. Task B may need an immediate measure that stops further duplicates while a lasting fix is built. Task A could be a good next project because it frees more time each week.
Describe the result you would check afterward: fewer repeated repairs, correct bookings, fewer affected customers, and manageable maintenance. If you can say all four, you have answered the question the way an experienced engineer would.
Quick answers
What is an example of toil?
Manually resending stuck booking confirmations every morning is toil: it is repetitive, reactive, grows with traffic, and tomorrow's pile is just as big as today's.
How do you measure toil?
Record how often the task happens, how much hands-on time each occurrence takes, how often it interrupts other work, and what a mistake would cost. Those figures make the case for automation concrete.
Should you automate everything?
No. Automate work that is repetitive and safe to perform by software, and keep any step that protects customers, such as a check for existing tickets. Where the cause can be removed, fixing it beats automating the repair.