Skip to content
Guide

Getting Started with Zerto: Your First Virtual Protection Group

A first-steps guide for engineers new to Zerto: installing the VRA, building your first VPG, testing it without disrupting production, and failing back.

Oğuzhan Gerçek··5 min read
Getting Started with Zerto: Your First Virtual Protection Group

Short answer: Start with a non-critical multi-VM application, not your most critical system. Install Zerto Virtual Manager and a VRA on every host, create a single Virtual Protection Group containing all the virtual machines in that application, wait for it to synchronize, and run a non-disruptive failover test. You will learn more from that one cycle than from any document.

We explained the concepts in our article on how Zerto works; this one covers the order in which to build them in your own environment for the first time.

Step 1: understand what you are protecting

Before touching the console, list the virtual machines that make up the application and how they depend on each other. Zerto recovers groups, not individual machines; a wrong dependency map produces a recovery that is technically successful but functionally broken.

Write down which machine holds state, which is stateless, and what talks to what. Three dependencies are most often missed when drawing up that list: the authentication server, the license server, and the file share. All three tend to be classed as "infrastructure" and left outside the VPG, and all three break the recovery.

Step 2: install the components

Two things get installed:

    Zerto Virtual Manager (ZVM), one per site, the management layer.Virtual Replication Appliance (VRA), one per hypervisor host, which does the replication work.

Nothing goes inside the guest operating system. Give the VRAs enough resources; an undersized VRA is the most common reason an RPO drifts above target.

Step 3: plan bandwidth honestly

Continuous replication wants continuous bandwidth. Estimate the daily change rate of the protected virtual machines and compare it with the link between sites. If the link cannot carry the change rate, you have three options: reduce what you protect, increase bandwidth, or accept a higher RPO. Throttling does not remove that constraint; it only hides it.

When estimating, work from the peak rather than the average. Overnight batch jobs and database maintenance windows generate several times the daily average in writes; your RPO breaks during those hours and nobody notices.

Step 4: create the first VPG

Group all the virtual machines belonging to the application. Choose journal retention deliberately: retention multiplied by change rate determines the storage you need at the recovery site.

Starting at 24 hours while you learn is reasonable. When you move to production, your target should be the vendor recommendation of eight days, a window that covers the ransomware scenario, because the attacker is usually inside for days before encryption starts.

Do the network mapping in this step too: which port group to connect to on the recovery side, and a new IP address if one is needed. Leaving those fields blank is the most common reason a first test fails.

Then wait for the initial synchronization to complete before drawing conclusions about your RPO.

Step 5: test before you trust

Zerto supports a non-disruptive failover test that brings the VPG up on an isolated network without affecting production. Run one. Then do what most people skip: do not stop at confirming the virtual machines powered on; log into the application and verify that it works.

Write down what broke. Something always breaks the first time: usually DNS, IP addressing, or a dependency left outside the VPG. We cover what makes a test meaningful in the disaster recovery failover test guide.

Step 6: know what a real failover does

There is a difference between a test and a real failover, and it is one you do not want to learn under pressure: after a real failover, Zerto puts you inside a commit window. Within that window you verify the recovered environment and make one of two decisions: commit and the recovery becomes permanent, or roll back and you return to the source side.

Decide the length of that window and its default behavior in advance. If nobody decides within the window, the default behavior decides for you.

Failback is a separate job and requires replication in the reverse direction. Teams tend to neglect this step, and simply continuing to run at the recovery site amounts to having no plan.

A realistic 30-day learning plan

Week 1. Install in a lab, protect a single test virtual machine, watch the journal grow, and understand the RPO graph. Week 2. Build a multi-machine VPG, run a test failover, deliberately break something, and observe the result. Week 3. Practice failback; this is the step teams neglect and struggle with under pressure. Week 4. Protect a real, non-critical application and write down the procedure.

If you want to plan the certification side, we mapped out the paths in our guide on how to learn Veeam and Zerto.

Frequently asked questions

Do I need to install anything inside the virtual machines? No. Zerto is agentless and works at the hypervisor layer.

How long does the initial synchronization take? It depends on data volume and bandwidth. For your first sizable VPG, plan for anywhere from hours to days.

Does a failover test affect production? Not if it is configured correctly: the test runs on an isolated network. Verify network isolation before starting.

What should the initial journal retention be? Start around 24 hours in a lab, measure real consumption, and move to eight days in production.

Does Zerto replace backup? No. The journal is a recovery window, not an archive, and it is not immutable. We covered the difference in our article on the recovery gap.

Sources