Everyone Uses Terraform, but Who Is Catching the Drift?
90% of cloud users run IaC, but 64% cannot find skilled staff. A method that makes configuration drift measurable.
Short answer: Defining infrastructure as code has become table stakes. According to CNCF data, 90% of cloud users manage infrastructure as code in some form, and Terraform holds 76% of that space. The hard part now is operating it: in the same source, 64% of organizations say they cannot find skilled cloud and automation staff, and 45% name security and compliance risk as a major IaC challenge.
This article covers the work that starts after Terraform is installed: measuring the gap between the infrastructure your code describes and the infrastructure actually running.
What drift is, and why it is silent
Drift is the divergence between the state written in code and the state that actually exists in the cloud. It appears when an engineer adds a security group rule from the console, when an autoscaling event changes a tag, or when a provider updates a default.
Drift is dangerous because it is silent, whatever its size. It raises no alert, appears on no dashboard, and usually surfaces at one of two moments: during a disaster recovery drill, or during an audit. Both are too late.
In a mature practice, drift detection runs continuously and every divergence found is treated as a bug, not as an exception.
Terraform and Ansible do not do the same job
The two are often treated as alternatives, but they sit at opposite ends of the pipeline.
Terraform is the provisioning layer. It works declaratively: you describe the target state and it calculates the difference. It is strongest at defining which resources exist and how they relate: networks, virtual machines, managed services, access policies.
Ansible is the configuration layer. It works procedurally, configuring what runs on a system once it is up: package installation, service settings, security baseline.
Drawing that boundary clearly is one of the highest-return decisions in an IaC project. When the boundary blurs, the same setting gets defined in two places and which one wins is left to runtime.
From mutable to immutable infrastructure
The approach that kills drift at the source is to replace a running server rather than update it. That turns the patching window into a deployment job and removes the question of which version is on which server.
Not every workload suits this. Stateful systems, legacy applications and software with hardware-bound licenses stay mutable. What matters is documenting which system follows which model; in a mixed fleet without that record, patch management becomes guesswork.
Four measurable checks
The health of an IaC pipeline comes down to these four:
- How often does drift detection run? If the answer is "before a deployment," drift is being discovered at deployment time, which is the most expensive moment.How many divergences were found in the last thirty days, and how many were written back into code? A divergence found and not fixed is no different from one never found.Are manual console changes recorded? If not, root cause analysis starts from scratch with every incident.Is the disaster recovery environment built from the same code as production? If not, the drill is testing a different system, not production.
The fourth one matters most: any claim that the recovery environment mirrors production is only an assumption until you run a disaster recovery test that proves something.
The skills gap is a design problem, not a training problem
The 64% staffing gap shows the distance between Terraform's ubiquity and the expertise needed to run it well at scale. The way to close that gap is to offer fewer choices: standard modules, reviewed defaults, and policy checks.
The same discipline applies at the application layer. We covered how to declare that a service is genuinely ready in our article on Kubernetes readiness probes, and how to establish a hardening baseline on the system hardening page.
You can see how Eclit builds its automation pipeline on the DevOps page.