Date
August 25, 2026
Topic
Disaster Recovery
IT
Disaster
Recovery
Planning:
How
to
Protect
Operations
When
Disruption
Occurs
Most organisations discover the limits of their recovery plan on the worst possible day. Here is how to find those limits first.
IT Disaster Recovery Planning: How to Protect Operations When Disruption Occurs

Almost nobody tests their disaster recovery plan on an ordinary Tuesday. The plan gets examined during the outage, which is the one moment when there is no time to fix what is missing.

A recovery plan is not a document. It is a set of capabilities that either work under pressure or do not, and the only way to know which is to exercise them deliberately before you need them.

Start with what the business cannot do without

Recovery planning goes wrong when it starts from the server list rather than from the business. The first question is not what you are running, it is what has to be working for the organisation to operate at all.

Two numbers follow from that. How long can a given system be unavailable before the impact becomes serious, and how much data can you afford to lose if you have to fall back to your last good copy. Those two answers determine the design and the cost of everything else. Set them with the people who run the business, not with the people who run the servers.

A backup you have never restored is not a backup

The most common failure we see is a backup that has been running faithfully for years and has never once been restored. Backups fail quietly. Jobs stop covering systems that were added later, retention gets shortened to save space, and the copy sitting on the same network as the original gets encrypted alongside it during a ransomware event.

  • Restore something real, on a schedule, and time how long it takes
  • Keep at least one copy that an attacker on your network cannot reach or delete
  • Confirm coverage after every significant change to the environment
  • Check that what is restored is actually usable, not just that the job reported success

Plan for the people, not only the systems

Technical recovery is only part of it. During a real incident, someone has to decide to declare it, someone has to tell staff and customers what is happening, and someone has to make the call about what gets restored first. If those roles are not agreed in advance, the first hour is spent working out who is in charge.

Write down who decides, who communicates, and how people will reach each other if the usual systems are the ones that are down.

Exercise it before you rely on it

A tabletop walkthrough once or twice a year will surface more gaps than another round of documentation. Talk through a realistic scenario with the people who would actually be involved and note every point where the answer is uncertain. Those uncertainties are the plan.

Disruption is not a rare event any more. The organisations that come through it well are not the ones with the longest plan; they are the ones that have already practised.