Skip to content
Modernization & Resilience Engineering

Disaster Recovery & Continuity Engineering (DRaaS)

Disaster recovery design and managed DRaaS operation. The goal is being able to recover, which takes more than backups: RPO and RTO targets are set against what the business can tolerate, proven in drills, and documented.

What we deliver

  1. RPO is how much data you can afford to lose; RTO is how long you can afford to be down. Both are business decisions: an RPO in minutes means synchronous replication and the cost that comes with it, while an RPO in hours is far cheaper. Systems are tiered against these two numbers.

  2. Standby site, replication direction, network routing and DNS failover are designed together. If the servers are up but the application is unreachable, recovery has not happened, so the architecture covers the whole chain end to end.

  3. An untested recovery plan is an assumption. A scheduled drill performs a real failover, measures the elapsed time and records what did not go to plan. When you go into an audit, the evidence that recovery works is already prepared.

  4. Applications do not come back alone: authentication, DNS, message queues and integrations have to return in a particular order. The failover sequence comes out of that map.

  5. Immutable backups and logical separation stop encrypted production data from propagating into the backup set. In a ransomware scenario the recovery point is chosen based on when the infection began, not when it was detected.

  6. The plan, drill records and role definitions required for ISO 22301 business continuity and the relevant ISO 27001 clauses are kept ready.

<15mTypical RTO
<5mTypical RPO
100%Drill Success
24/7Recovery Ops
Operational architecture

How it works

Every engagement follows the same five steps: baseline the current state, design the target model, roll out in stages, operate it, and improve against measurements.

01

Assess

Baseline the current state, name the gaps and put the success criteria in writing.

02

Design

Architect the target operating model and the toolchain it needs.

03

Deploy

Implement, configure and validate in a staged rollout.

04

Operate

24/7 management with contracted response times and proactive monitoring.

05

Improve

Continuous improvement driven by metrics, incidents and changes in the business.

Contracted service levels

Every engagement runs under a written SLA: a commitment, not a best-effort promise.

Run by engineers

Dedicated engineers who know your stack. No generalist help-desk tier in between.

Continuous improvement

Service reviews every two weeks, roadmap updates every quarter.

The technologies we run this on

IN PRODUCTIONIN TRIALUNDER ASSESSMENTON HOLDManaged DR platformZerto
The technologies below are taken from the Eclit technology radar. The ring a technology sits in does not rate how good it is: it says how far we have taken it in our own operation.
The full technology radar →

The concepts behind this service

Disaster recovery
Bringing systems and data back after a major outage, data loss or disaster.
Disaster recovery as a service (DRaaS)
Disaster recovery delivered as a service.
Disaster recovery site
The secondary facility that takes over workloads when the primary data center becomes unusable.
Business continuity
Keeping critical business processes running through a disruption or crisis.
RTO (recovery time objective)
How quickly a system must be back after an outage.

This section explains the technical terms used on this page. The definitions come from Eclit's own technology glossary, and each term links through to its full entry there.

The full technology glossary →
01If we have backups, do we still need a disaster recovery plan?

Yes. A backup brings the data back; disaster recovery brings the service back. The difference lies in where and how the servers, the network and the identity layer come up, and in what order.

The question that makes this concrete: with the backup you hold today, starting from an empty environment, how many hours until your application is open to users? In organizations that have backups and nothing else, that figure is usually measured in days, because restoring means working through server provisioning, the IP plan, DNS records, certificates, license keys and authentication one at a time, and none of that is written inside the backup.

A disaster recovery plan writes those steps down in advance: which system comes up in what order, where the network points, when DNS cuts over, and who makes which call. The backup is the material this plan works with, not the plan itself.

02Who sets RPO and RTO?

The business sets them; we show what each costs and what it takes technically.

RPO is the amount of data you are willing to lose: an RPO of 15 minutes means the last 15 minutes of data may be gone after an incident. RTO is how long you can be down. Both are commercial decisions rather than technical ones, because both translate directly into cost.

Roughly: an RPO measured in minutes needs synchronous or continuous replication and standing capacity on both sides. An RPO measured in hours is met with periodic replication and costs markedly less. An RPO measured in days can usually be built on the backup arrangement you already have.

In practice, the right approach is to tier the systems instead of picking one number: the platform that takes payments and the internal reporting tool do not need the same target, and forcing one on both spends most of the budget in the wrong place. Once decided, the targets are written down; an RPO that stays verbal becomes an argument during the incident.

03Do drills affect production?

Set up correctly, no: recovery runs on an isolated network with no path back to production.

What happens technically is that replica copies are brought up in a network segment fenced off from production. The application genuinely starts, a login is attempted and data integrity is checked, but that environment does not reach production DNS, the production database or external integrations. If it did, you would create a real risk of two systems writing at once.

The real cost of a drill falls on the calendar rather than the systems: preparation, execution and reporting take the team's time. Refuse that cost, and the plan gets tried for the first time during a real disaster, the worst moment of the year to try anything.

04How often should drills run?

At least twice a year for critical systems, and again after any significant change to the architecture.

That second condition is the one most organizations skip, and it is where the risk actually sits. When a new application goes live, when the identity layer changes, when a database is upgraded or the network topology changes, the plan may still describe the old environment. A failover order that worked six months ago stops working with one dependency inserted in the middle.

Every drill records the measured recovery time. The most useful figure is the trend across drills: if the time is climbing, the environment is drifting away from the plan, and measuring is the only way to see that.

If you are audited under ISO 22301 the drill record is required anyway; the same exercise covers both.

05Do we need to build a secondary site ourselves?

Not necessarily. A cloud-based recovery site removes the cost of a second data center sitting idle: resources run only during drills and real incidents.

The comparison works like this. Your own secondary site means paying continuously for hardware, floor space, power and maintenance; in return, control is entirely yours and the data never leaves the organization. On a cloud-based site, storage and replication are continuous but compute is billed only when you need it.

Regulation usually weighs more than cost: if rules limit the country or facility where the data may be held, the site choice starts from that constraint. Where there is no such constraint, the decision comes down to the balance between RTO and budget.

06Is disaster recovery the same thing as high availability?

No, and confusing the two leads to an expensive mistake.

High availability (HA) absorbs the failure of a single component inside the organization, usually within seconds: a server drops, the cluster takes over. It runs in the same data center, on the same storage.

Disaster recovery absorbs the loss of that entire data center: fire, flood, an extended power failure, or ransomware that encrypts the whole estate. An HA cluster helps in none of those, because both of its nodes are hit by the same event.

The two are complementary layers: HA makes everyday failures invisible, and disaster recovery keeps the company standing.

07What if ransomware encrypts the backups too?

This is the weakest point of classic backup and it has to be designed for at the start.

Ransomware today does not encrypt the moment it lands; it usually sits quiet for weeks looking for the backup infrastructure. A backup repository reachable over the network is the attacker's priority target: once the backup is unusable, the attacker holds all the leverage.

The countermeasure has two layers. First, immutable backups: copies that, once written, no account, including the administrator's, can delete or alter for a set period. Second, logical separation: the backup infrastructure authenticating independently of the production directory service, so a compromised credential in production does not carry over.

The recovery point is chosen differently too: it is set by when the infection began, which is earlier than when the attack was detected. The gap between those is usually days, which means enough history has to be retained to reach a clean point.

08Who decides to fail over, and how quickly?

This is the most frequently skipped clause in a plan and the one that costs the most time during an incident.

Even with the technical preparation done, moving to the secondary site is a commercial decision: failing back also carries a cost, and an early failover can turn a temporary outage into lasting data inconsistency.

The plan therefore states three things in writing: the name of the role that decides (the role, not the person; that person may be on leave), who the authority passes to if they cannot be reached, and the threshold at which the decision is treated as automatic. A threshold like "a full outage lasting over two hours with no identified cause" is defined in advance so that nobody hesitates.

These clauses are exercised in a drill as well: the decision chain needs rehearsing as much as the technical steps do.

09How does failback work?

Most recovery plans describe the move to the secondary site in detail and cover the return in a single sentence. The return is the harder part.

During the time spent on the secondary site, new data has been produced there. Failback means moving that data to the primary site without loss and without conflict: the replication direction is reversed, both sides are synchronized, and only once parity is verified does a short planned outage bring you back.

Unlike the disaster, failback is a planned operation and is not rushed. Continuing to run on the secondary site until the primary is genuinely ready is often the right call.

The plan covers the failback steps and their verification criteria; if it does not, the plan is half-written.

10Can you review our existing disaster recovery plan?

Yes, and most engagements start there, by testing the existing plan against reality instead of writing a new one from scratch.

A review asks four questions. Do the RPO and RTO targets in the plan still match what the business units expect? Does the dependency map describe the current architecture, or the one from two years ago? When was the last drill, and where did the measured time land against the target? Are the backups logically separated from an attacker who has taken production?

The output is a report with the findings ranked by severity. Some of them can usually be closed by your own team; not all of them have to become a service, and we do not present them that way.

Let's work out where to start

Within two weeks you get it in writing: what works, what carries risk, and a prioritized roadmap.

Request a conversation