Replication & Recovery Readiness
Continuous measurement of replication health, recovery exercises, and evidence that the plan actually works.
What we deliver
Replication lag monitored continuously with thresholds set against the RPO target. What matters is not that the replica is up but how current it is.
Planned failover exercises with measured duration and failing steps reported. Having a plan is not enough; you need evidence that it works.
Actual recovery time measured in exercises and compared with the target. In most organizations the RTO on paper and the RTO measured are different numbers.
The order in which systems come back, based on application dependencies. The wrong order is the most common reason a recovery runs long.
Every exercise documented with date, participants, measured values and findings. That record is the evidence an audit asks for.
The recovery plan updated as the environment changes. Last year's plan does not recover this year's infrastructure.

How it works
Every engagement follows the same five steps: baseline the current state, design the target model, roll out in stages, operate it, and improve against measurements.

Assess
Baseline the current state, name the gaps and put the success criteria in writing.
Design
Architect the target operating model and the toolchain it needs.
Deploy
Implement, configure and validate in a staged rollout.
Operate
24/7 management with contracted response times and proactive monitoring.
Improve
Continuous improvement driven by metrics, incidents and changes in the business.
Every engagement runs under a written SLA: a commitment, not a best-effort promise.
Dedicated engineers who know your stack. No generalist help-desk tier in between.
Service reviews every two weeks, roadmap updates every quarter.
Who uses this
The industries we run Replication & Recovery Readiness for.
The concepts behind this service
- Replication
- Continuously copying data from a source to one or more targets.
- RPO (recovery point objective)
- The maximum acceptable data loss in an outage, measured in time.
This section explains the technical terms used on this page. The definitions come from Eclit's own technology glossary, and each term links through to its full entry there.
The full technology glossary →Knowledge Hub
What we have written about running and managing technology, collected in one place.
01If replication is working, does that mean we can recover?
It does not, and confusing the two produces the most expensive surprise during an incident.
Replication produces a second copy of the data. Recoverability is whether a working system can be brought up from that copy. The difference is network routing, authentication, dependency order, DNS, certificates and license keys.
A concrete example: your database may be replicating to the secondary site without a problem, but if the application servers are not there, or are there but still trying to reach the production directory service, what you have is a copy that does not run.
You only learn this by trying. A "replication green" dashboard says nothing about recoverability; it only says the data is flowing.
02How do you measure readiness?
We measure four things, and unless all four are green, the readiness claim is incomplete.
Replication lag: how far behind the primary the secondary copy is. That number is your actual RPO: if your target is 15 minutes and the lag is hours, the target exists only on paper.
Scope completeness: are all the systems on the recovery list genuinely being replicated? A newly added server is usually the one nobody remembers to add to scope, and that gets discovered during the incident.
Date of the last drill: a drill older than six months loses value depending on what has changed in the architecture since.
Recovery time measured in the drill: where it sits against the target, and whether it is lengthening or shortening over time.
All four are reported together; none of them alone means readiness.
03What should replication lag be?
You set the target; the measurement takes its meaning from that target.
In synchronous replication, lag is theoretically zero: a write is not acknowledged until it completes on the second side as well. The price is that network latency between the two sites is added directly to every write the application makes, which is why synchronous replication is built between geographically close sites.
In asynchronous replication, lag can be minutes or hours and application performance is unaffected. In return, a disaster costs you that much data.
What actually needs watching is how the lag behaves over time, more than its absolute value: if it stays within a steady band, the system is healthy. If it climbs and falls regularly, a batch job or backup window is usually saturating the bandwidth, and a disaster during that window would cost far more than the average suggests.
04Do drills put production at risk?
A properly designed drill does not touch production.
Recovery runs in a network segment fenced off from production. Copies are brought up there, the application starts, data integrity is checked, but that environment does not reach production DNS, the production database or external integrations.
Isolation is therefore the most critical technical part of a drill. Leave connectivity in place and you create real risks: two systems writing at once, duplicate emails going out, a test transaction landing at the payment provider.
The real cost of a drill is team time: preparation, execution and reporting. Skip it, and the plan gets its first real test during an actual disaster.
05What is the report good for?
It serves three purposes, and each is worth as much as the drill itself.
Audit. The relevant clauses of ISO 22301 and ISO 27001 ask for evidence that recovery has been tested. With the report on file, nobody has to hunt for evidence on audit day.
Improvement. The report lists the steps that faltered, in order: which system came up later than expected, which dependency had been missed, which permission was absent. That list becomes the agenda for the next drill.
Decision. If the measured recovery time is above target, either the target or the architecture has to change. That is a budget conversation and it cannot be had without a number. The report provides the number.
As reports accumulate, a trend also emerges: if recovery time is lengthening, the environment is drifting away from the plan.
06Which systems fall into replication scope?
Not all of them, and putting everything in scope is usually the wrong call.
Replication means continuous bandwidth and standing capacity on the second side. Replicating every system at the same level spends most of the budget on low-priority workloads.
The right method is tiering. First tier: systems whose failure stops the business (payments, orders, authentication). These earn the most aggressive replication. Second tier: systems that can be down for hours but not a day. Third tier: those where restoring from backup within a few days is acceptable (internal reporting, archives, test environments).
Tiering is done with the business and written down. When it is not written down, everyone's system becomes first tier during the incident.
07Do you replicate cloud workloads as well?
Yes, but the mechanism differs from on-premises and that difference matters.
On-premises, replication is usually done at the storage or hypervisor layer. In the cloud, the provider's own replication and snapshot services come into play, and these offer different guarantees within a region and across regions.
The critical distinction: a cloud provider's "high durability" commitment protects against data loss, not against an account-level mistake. A resource deleted by accident, or by a compromised account, cannot be recovered from redundancy inside that same account.
This is why a separate copy is kept for cloud workloads too: a different account, a different region, and a location that production credentials cannot reach.
08How often do we receive the readiness report?
Measurement is continuous; the report is periodic.
Replication lag and scope completeness are monitored continuously; when a system falls out of scope or lag crosses a threshold, it is handled as an incident right away instead of waiting for the report.
The periodic report brings all four measures together and makes the trend visible. For critical systems, a six-month cycle strikes the right balance in most organizations: it lines up with the drill calendar and is frequent enough to put in front of a board.
An interim report is issued after a significant architectural change. A new application, a change in the identity layer or a data-center move all bear directly on whether the previous report still holds.
09What if recovery time comes out above our target?
There are three options and all three are legitimate; the wrong move is not choosing.
Change the target. If the stated RTO is not realistic (often it is a number set without consulting the technical team), it is revisited with the business. An unreachable target is worse than no target, because it creates false confidence.
Change the architecture. The things that shorten recovery are well known and all cost money: standing capacity at the secondary site, more frequent replication, pre-built network configuration, automated failover steps.
Narrow the scope. Bring only the first tier to target rather than every system. In most organizations this yields the best cost-to-benefit balance.
Whichever is chosen, the decision is written down; an acceptance that stays verbal turns into a finding at the next audit.
Let's work out where to start
Within two weeks you get it in writing: what works, what carries risk, and a prioritized roadmap.
Request a conversation