SRE disaster recovery is the plan for restoring usable service after a major failure, with two limits attached: the recovery point objective (RPO), which caps how much recent data may be lost, and the recovery time objective (RTO), which caps how long recovery may take. Data integrity, the companion idea, means keeping a service's records correct in the first place. Together they answer the question customers care about most after a failure: is my booking still there, and is it right?
For TicketDesk, our fictional booking app, a page that loads is not enough. Customers need their confirmed reservations to remain correct after a failure and after the repairs that follow.
If RPO and RTO sound like the same thing with different letters, you are not alone. The worked example below shows how a recovery can meet one and miss the other.

- 01Protect correct bookings
- 02Restore a saved point
- 03Check records and reopen
How can a running app have incorrect data?
A running app can have incorrect data because the software can show “booking confirmed” while saving the wrong reservation, saving it twice, or losing it later. Data means the stored information: the customer, the seat, the payment status, and the confirmation record. A green response on a screen and a correct record in storage are two different things.
Define what must always remain true. Engineers call those rules invariants. One confirmed reservation should refer to one valid booking. Repeating the same accepted operation should not create another booking. Accepted work should eventually have a known result.
These rules give the team concrete checks during recovery. Compare the work that was accepted with the records that are still saved after the initial response. Matching the confirmation the customer saw to an actual reservation matters more than whether any one piece of software is still running.
What are RPO and RTO?
RPO is the maximum acceptable data loss expressed as a time interval, and RTO is the maximum acceptable time to restore usable service. A five-minute recovery point objective means the plan aims to lose no more than five minutes of recent work when restoring from an earlier saved point.
For the recovery time objective, define when the clock starts and what counts as restored service. Copying files may finish long before customers can safely book again, and the objective should cover the whole journey back.
Choose both objectives according to the consequences. Losing suggested-event updates is different from losing reservations that customers believe are confirmed, and the objectives should reflect that difference.
A backup is a saved copy used for recovery. Its schedule alone does not prove an RPO can be met. A copy might fail, be unreadable, or become inaccessible. What the team needs is a usable recovery point and evidence that restoring from it works.
Can recovery meet one objective and miss the other?
Yes, recovery can meet the RPO and still miss the RTO, and this example shows how. TicketDesk's failure begins at 12:00. The latest usable saved point is 11:56. Restoring takes 12 minutes, checking the results takes eight, and gradually reopening the service takes five.
Assume a five-minute RPO and a 20-minute RTO:
Potential lost interval = 12:00 - 11:56 = 4 minutes
Recovery duration = 12 + 8 + 5 = 25 minutes
RPO: met in this exercise
RTO: missed by 5 minutes
The four-minute gap fits inside the data-loss objective. The full 25-minute recovery misses the time objective. Reporting only the 12-minute copy would leave out the checking and reopening that customers depend on, and would make the recovery look faster than it was.
This result assumes the saved point is valid and the service is safe afterwards. There is a further complication: a payment service outside TicketDesk may have completed actions after 11:56. Restoring older internal records does not reverse those payments. The team must identify and reconcile them, meaning compare the two sets of records and resolve the differences, without charging customers again.
Why do you need backups if there are other live copies?
You need backups because live copies faithfully replicate mistakes as well as data. A replica is another live copy of the data. It helps when one copy becomes unreachable. It also receives a mistaken deletion or corrupt update within moments, so every live copy can end up sharing the same error.
A historical backup gives you an earlier point to return to. Check what it actually survives. A separately stored copy may still be deletable using the same account that caused the problem. Recovery may depend on an encryption key, which unlocks protected data, or on compatible software that may no longer be installed.
A reported successful backup is limited evidence. A restore exercise checks access, readable information, suitable software, correct records, and the full time to return to service. Practise with the permissions responders will actually have on the day, not with an administrator's.
What should happen while records are being corrupted?
While records are being corrupted, the first priority is to stop the damage spreading, which may mean temporarily stopping new writes. A write is an operation that adds or changes stored information. Coordinate that decision with the people responsible for the product and the data through the incident response; it affects customers immediately.
Preserve evidence and identify the affected operations. Choose a trustworthy recovery point and work out how to reconcile the activity around it. Some records may need targeted repairs rather than restoring everything from the saved point.
Check the invariants before reopening. Equal record counts do not prove equal correctness: the right number of bookings could still belong to the wrong customers. Reopen gradually and watch for recurring problems or delayed work.
Exercise: is “we have a backup” enough?
Someone proposes a risky migration, meaning a change to the way stored information is organised, because a backup exists. The release lesson explains why returning to old software may not undo the new records. What evidence would make recovery credible?
Ask for a usable saved point, a timed restore test, working access and keys, compatible software, and checks for correct bookings. Ask what happens to new writes and external payments during restoration. Identify who decides that the service can reopen.
Explain your reasoning across the whole path back to usable service. The four-minute gap and the 25-minute recovery show why limiting lost data and restoring promptly are separate goals. In an interview, make both measurable and include the work customers depend on after the files have been copied.
Quick answers
What is the difference between RPO and RTO?
RPO limits how much recent data you may lose, measured as a time interval back to the last usable saved point. RTO limits how long the whole recovery may take, from failure to safely restored service. A recovery can meet one and miss the other.
Is a database replica the same as a backup?
No. A replica is a live copy that receives the same changes almost immediately, including mistaken deletions and corrupt updates. A backup is a historical saved point you can return to after such a mistake.
How often should you test backup restores?
Often enough that you have recent evidence a restore works with the access, keys, software, and time limits you would face in a real failure. A backup that has never been restored is an assumption, not a recovery plan.