Disaster Recovery Planning: What Actually Works
The plan for restoring service after a major failure, including targets for how quickly and how much data can be lost.
The key figures
- Recovery time objective
- how long restoration may take
- Recovery point objective
- how much data loss is acceptable
- Dependencies
- DNS, certificates and third parties affect recovery
- Exercises
- plans should be rehearsed
Why this is worth getting right
Recovery targets drive architecture and cost, and without them teams discover their real limits during an outage.
Do this, not that
Do
- Agree recovery targets with the business
- Document dependencies and access
- Keep contact and escalation lists current
- Rehearse recovery at least annually
- Store the plan somewhere reachable during an outage
Don’t
- Plans stored only on the affected systems
- Targets nobody has validated
- Single points of failure
- Plans that assume key people are available
When to bring in help
Our advice Bring in help when downtime would seriously damage the business.
Where this comes from
- National Institute of Standards and Technology — Contingency planning guide
- Microsoft Learn — Reliability design principles
The figures and practices above come from the sources listed.
Working on something like this?
We take on Web Design & Development work for teams who want it done once, properly. Tell us what you are building and we will tell you honestly whether we are the right studio for it. Start a project.
Where to go next
Spotted something wrong? Report an error on this page. We correct on the page and say what changed.