Skip to content
Article

What Is Replication? Synchronous vs. Asynchronous Replication

Replication keeps copying data to a second target. Synchronous vs. asynchronous, distance limits, database replication, and why a replica is not a backup.

Oğuzhan Gerçek··7 min read
What Is Replication? Synchronous vs. Asynchronous Replication

Short answer: Replication is the continuous copying of data from a source to one or more targets. Its purpose is to keep an up-to-date second copy ready to take over when the primary system stops. With synchronous replication, a write is not acknowledged until both sides have committed it: no data is lost, but distance is limited. With asynchronous replication, the target lags slightly behind: there is no distance limit, but whatever was in that lag is lost when a failure hits. Replication is not a backup: it copies deleted or corrupted data just as fast.

What does replication mean?

In IT, replication means sending the changes in a database, a virtual machine or a storage volume to another server, storage system or data center. The copy at the target is called a replica.

The question replication answers is not "how do I get yesterday's data back?" but "if the primary stops now, where do I carry on, and how many minutes will it take?" That is why it is a core part of high-availability and disaster recovery designs. The cost of downtime is not small either: in Uptime Institute's 2025 survey, 57% of respondents said their most recent major outage cost more than $100,000.

Synchronous vs. asynchronous replication

In SNIA's definition, synchronous replication means data must be committed to stable storage at both the primary and the secondary site before the write is acknowledged. With asynchronous replication, only the primary site has to commit the write; data is forwarded to the secondary site as the network allows.

That difference sets your RPO directly. As Microsoft puts it, synchronous replication can guarantee consistency, so it can support an RPO of zero. With asynchronous replication, the replica might not have the latest data at failover, so the RPO must be greater than zero. The PostgreSQL documentation says the same thing in one sentence: the amount of data lost is proportional to the replication delay at the time of failover.

  • Synchronous: RPO is zero. Every write waits for the other side to confirm, so write latency goes up and distance is limited.
  • Asynchronous: RPO ranges from seconds to minutes. It does not slow the application down and has no distance limit, but data that has not reached the target when a failure hits is lost.

We explain how to set an acceptable RPO from business impact in How to calculate RPO and RTO.

The distance limit of synchronous replication

With synchronous replication, every write takes longer by the round-trip time between the two sites. That is why vendors state their limits in milliseconds or kilometers:

  • VMware vSAN stretched cluster: At most 5 ms of round-trip latency between the two main sites.
  • NetApp SnapMirror active sync: Inter-cluster round-trip latency must be under 10 ms.
  • Dell PowerMax SRDF/S: A recommended maximum of 200 km between arrays. Asynchronous SRDF/A has no distance limit, with an RPO measured in seconds.

As a rule of thumb from NetApp, every 100 km adds about 1 ms of round-trip latency. Synchronous replication between two data centers in the same city is usually feasible; over intercity distances, latency approaches these limits and asynchronous replication comes into play. A common design combines the two: synchronous within the metro area, asynchronous to a distant site.

At which layer does replication happen?

  • Storage layer: The storage array copies blocks to an array on the other side. It works independently of the servers, but usually needs arrays from the same vendor on both sides.
  • Hypervisor layer: Changes to virtual machines are shipped through the hypervisor. In vSphere Replication, the RPO can be set between 5 minutes and 24 hours. Zerto ships changes continuously and keeps them in a journal, so you can roll back to checkpoints a few seconds apart. We explain its architecture in How Zerto works.
  • Database layer: The database engine's own mechanism ships its transaction log to another server.
  • File and application layer: File sync tools or the application's own copy features do the work.

How does database replication work?

Most common databases replicate asynchronously by default and have to be switched to synchronous mode explicitly:

  • PostgreSQL: Streaming replication is asynchronous by default. To make it synchronous, synchronous_standby_names must be set.
  • MySQL: Replication is asynchronous by default; if the source crashes, transactions it committed may never have reached a replica. In semi-synchronous mode, the source does not complete a transaction until at least one replica confirms it has received the events.
  • SQL Server Always On: Synchronous-commit mode is for high availability; asynchronous-commit mode is for a disaster recovery copy over long distances. In asynchronous mode only forced failover is possible, with possible data loss. Since SQL Server 2019, up to five synchronous replicas are supported.
  • Oracle Data Guard: There are three protection modes. The default is Maximum Performance. Maximum Availability prevents data loss as long as at least one standby is synchronized. Maximum Protection shuts the primary down rather than let it process unprotected transactions.

A replica being up is not enough; you need to measure how far behind it is. We show how to do that for PostgreSQL in Reliability Friday 01.

Replication is not a backup

Microsoft's resiliency guidance says it plainly: replication does not protect you against risks that result in data loss or corruption. If you accidentally delete data, the deletion is replicated to every replica and the data is gone everywhere. Ransomware encryption is a write too, and it travels to the target the same way.

Replication and backup are therefore not alternatives but complements: replication gives you speed, a backup gives you a history you can return to. We explain the difference in What is backup?

Continuous data protection (CDP) sits between the two. It copies changes continuously but also lets you go back to a point in the past; in NIST's words, unlike replication, CDP allows playback of the copied data to previous points in time. We cover how that difference is used in a ransomware scenario in our ransomware recovery playbook.

Split-brain and the witness

When the link between two sites drops, both sides may decide they are the primary and start writing to the same data separately. This is called split-brain. According to Microsoft's clustering documentation, quorum exists precisely to prevent it. In vSAN, a witness host at a third site does the same job: if the connection between the two main sites is lost, vSAN carries on with the preferred site.

In active-active designs, where both sides write, the witness or quorum design matters as much as the replication itself.

Checklist

  • Is the accepted RPO written down for each system, and does the replication mode match it?
  • For synchronous replication, has the latency between the two sites been measured, and is it under the vendor's limit?
  • Is replication lag monitored, and does crossing the threshold raise an alert?
  • Is the split-brain witness or quorum at a third site?
  • Alongside the replica, is there an immutable backup that lets you go back in time?
  • When were failover and failback last tested? The method is in our disaster recovery test guide.
  • If the replica is abroad, are KVKK's cross-border transfer conditions met? Details in KVKK and cross-border data transfer.

Frequently asked questions

What does replication mean? Continuously copying data from a source to one or more targets. The copy at the target is called a replica.

Is replication a backup? No. Replication copies deleted or corrupted data too. To go back to a point in the past, you also need a backup.

What is the distance limit for synchronous replication? It depends on the vendor; limits are usually given as 5 to 10 ms of round-trip latency or a few hundred kilometers. The rule of thumb is about 1 ms per 100 km.

How much data is lost with asynchronous replication? Whatever had not yet reached the target when the failure hit. That amount depends on the replication lag at that moment, which is why lag has to be monitored continuously.

What is the difference between database replication and storage replication? Database replication ships the engine's own transaction log and understands the database's consistency. Storage replication copies blocks; it is independent of the application but needs extra measures to produce a consistent copy.

Sources