Skip to content
Guide

How to Run a Disaster Recovery Failover Test That Proves Something

Attacked organizations take 24.6 days on average to recover. A workable DR test procedure, what to measure, and how often to repeat it.

Oğuzhan Gerçek··4 min read
How to Run a Disaster Recovery Failover Test That Proves Something

Short answer: For Tier 1 systems, run a non-disruptive failover test to an isolated network at least twice a year, and have business users verify the recovered application rather than taking IT's word that "the virtual machines came up." Measure the time from the moment of decision to a working service; that number, not the one in the plan, is your real RTO.

Why untested plans fail

In Veeam's 2025 ransomware research, only 10% of attacked organizations recovered more than 90% of their data, 57% recovered less than half, and recovery took 24.6 days on average. Almost all of them had backups. What they lacked was proof that those backups could actually be restored.

A recovery plan that has never been run is an assumption. The failures that show up in real incidents are consistently ordinary:

    Procedures reference servers, IP ranges or people that no longer exist.Application dependencies are incompletely mapped; the database is recovered, the service depending on it is not.Nobody has authority to declare a disaster after hours, and the first hour is spent chasing approval.Backups were reported as successful and never restored.DNS and certificate changes are not part of the plan.

Each of these is cheap to find in a test and expensive to find in an incident. At the PagerDuty 2026 figure of $4,537 per minute of downtime, a test that costs you a morning costs less than twenty minutes lost during a real event.

The procedure

1. Define the scenario. Skip "the data center was destroyed" and pick something concrete and plausible: a storage array failure, ransomware at 2am on a Sunday, the loss of one critical application.

2. Write the success criteria in advance. Define what "recovered" means for this scenario, in business language. For example: a user can place an order and the order appears correctly in the ERP.

3. Isolate the test environment. Both Zerto and Veeam support recovery to an isolated network. Verify the isolation before starting; a test that leaks into production is worse than no test at all.

4. Start the clock at the moment of decision, not when the first command is typed. Approval time is part of the RTO, and usually the largest part.

5. Work from the procedure, not from memory. A procedure that cannot be followed by someone who did not write it is not a procedure. Deliberately run the test without the person who knows the system best; in a real incident there is no guarantee that person picks up the phone.

6. Let the business unit do the verification. IT confirming that machines are running proves very little. Ask a real user to complete a real transaction.

7. Record everything that went wrong, then fix the procedure the same week, while the details are fresh. A finding that is not fixed is found again in the second test, and at that point it invalidates both tests.

What to measure

    Measured RTO. From decision to working service. We covered how to set the target in how to calculate RPO and RTO.Measured RPO. How much data was actually lost at the recovery point.Manual steps. Every manual step is a failure point at 3am.Dependency gaps. Anything needed that sat outside the protected group.Decision time. How much of the total was technical work and how much was waiting for approval. If those two are not separated, the wrong problem gets optimized.

Frequency

    Tier 1: at least twice a year, full failover.Tier 2: annually.Tier 3: an annual sample restore check is usually enough.Plus after every significant architectural change.

Tabletop exercises are useful for practicing decisions but do not verify the technology. They are not a substitute for a real failover.

Do not forget failback

Teams rehearse failover and neglect failback, then discover during a real incident that returning to the primary site is harder than leaving it. The reason is simple: the data produced while running at the recovery site has to be carried back to the primary, and that requires synchronization in the reverse direction. Practice both.

Frequently asked questions

Does the test affect production? Not if it runs on an isolated network; both major platforms support this. Verify the isolation first.

How long does a test take? Plan half a day for a first Tier 1 test, including verification and review. Subsequent ones get faster.

Who should take part? Infrastructure, application owners, at least one business verifier, and someone with authority to declare an incident. You can test whether that authority is written down using the seven lines in the incident escalation matrix.

What if the test fails? That means the test worked. A failed test costs you a morning; the same failure during an incident costs a great deal more.

Sources