Infrastructure as Code / DevOps
September 28, 20268 min read
Infrastructure Drift Never Left. Now There Are Agents Chasing It.
Terraform and OpenTofu promised to eliminate infrastructure drift. They didn't. What changed in 2026 is who's chasing it: autonomous agents that detect, document, and escalate — without replacing the governance enterprises still have to build.

Infrastructure as code (IaC) made a simple promise: what's in the repository is what runs in production. A decade later, that promise still isn't fully kept. Drift — when the real state of infrastructure diverges from the code that's supposed to describe it — remains, according to HashiCorp, one of the most persistent operational risks in cloud environments.
Why drift never went away
The cause isn't technical, it's operational. Drift happens when someone changes infrastructure outside the IaC flow — almost always for a good reason. During a break-glass incident, response teams sometimes decide to bypass standard procedures to fix the problem as quickly as possible, and those kinds of shortcuts cause changes to resources that are tough to track and resolve in code.
The problem isn't just code hygiene. Unrecognized drift creates multiple security risks that need to be addressed before they become real problems, dramatically increasing the probability of critical data exposures from systems accidentally left open to public access. And it's not just security — it's money, too. Temporary infrastructure changes can go unnoticed, costing thousands of dollars a month in unnecessary provisioning.
Info
For a business with meaningful infrastructure spend, drift isn't a technical footnote. It's the difference between knowing what you're paying for and finding out at the end of the month.
What's new: agents that don't just alert, they act
For years, the answer to drift was the same pattern: a scheduled job checks state, generates a report, and someone on the team has to read it, prioritize it, and decide what to do. That changed with the arrival of operational agents capable of running full tasks without constant supervision. AWS DevOps Agent, launched in preview in late 2025 and generally available since March 2026, is the most visible example.
What's interesting isn't that it detects drift — that already existed. It's how it does it. The prompt for the typical use case is literally: "Create an agent that reviews production AWS resources for drift from our infrastructure standards (encryption, network isolation, tagging, credential handling, and observability) and creates recommendations for anything out of compliance." The team describes the business rule in plain language; the agent translates it into a recurring check.
- Detects: compares real infrastructure against defined policies, not just against a static state file.
- Documents: for each finding, it creates a recommendation with the affected resource, current vs. expected state, and a suggested remediation.
- Avoids duplicate noise: updating an existing recommendation instead of creating a duplicate keeps the backlog clean and shows whether an issue is new, ongoing, or resolved.
- Schedules like any other job: supports EventBridge-compatible cron expressions, running for example every Monday ahead of the weekly operational review.
The result, per AWS, is that from there, the same finding becomes action — it lands as a ticket, notification, or handoff for the owning team rather than becoming another report to read. That's the difference between a dashboard nobody checks and a flow that actually closes the loop.
The other side of the coin: more surface area, more governance required
While the detection layer gets automated, the change layer keeps getting more complex, not less. The Terraform provider for AWS — the bridge between code and the actual AWS API — keeps growing at a notable pace. Version v6.62.0, released in August 2026, added support for new resources across services including Amazon DSQL, ECS, ECR, SES, and Pinpoint, alongside further enhancements to Bedrock AgentCore, CloudFront, ElastiCache, Resilience Hub, and Secrets Manager.
That growth rate isn't free. Provider upgrades can affect schemas, defaults, state, and resource behavior, and the withdrawal of version v6.57.0 earlier in 2026 demonstrated that even mature infrastructure tooling can introduce upgrade risk. In that specific incident, a bug forced HashiCorp to pull the version from the registry and temporarily recommend pinning to the prior release while a patch was published.
For enterprise teams, provider versioning should be treated much like application dependency management: pin versions, test upgrades, review changelogs, and promote changes through controlled environments rather than automatically consuming every new release.
Terraform and OpenTofu: competition doesn't solve governance
It's worth putting this in context. Since OpenTofu split from Terraform in 2023 following HashiCorp's license change, the two projects have competed closely on features — and in 2026 that race is still active. Terraform 1.15 closed the gap with capabilities like dynamic module sources, but as InfoQ notes, some of its headline features had already been available in the open-source fork OpenTofu for nearly two years. OpenTofu 1.12, for its part, resolved a nearly decade-old community request by wiring prevent_destroy to a variable — something Terraform still doesn't offer natively.
The practical takeaway for an enterprise isn't "which tool wins." We already covered that ground comparing AWS CDK against its alternatives. The takeaway here is different: more options and more features don't reduce the governance load — they shift it. Every new capability (native state encryption, policy as code, detection agents) is one more tool someone on the team has to configure, maintain, and audit with judgment.
| Approach | What it does well | What it still demands from the team |
|---|---|---|
| Manual `terraform plan` review | Full control, zero external dependencies | Doesn't scale; depends on someone running it on time |
| Native drift detection (HCP Terraform) | Continuous checks, centralized alerts | Still generates reports a human has to prioritize |
| Autonomous agent (AWS DevOps Agent and similar) | Detects, documents, and routes remediation without constant intervention | Requires well-defined business rules and periodic review of its recommendations |
What this means for a business running production infrastructure
Drift-detection agents don't eliminate the need for a change process. They make it more visible. An agent that checks encryption, network isolation, and tagging every Monday is only useful if someone has already defined what "correct" means for that specific infrastructure, and if there's a real flow that turns each finding into a closed action — not another report piling up without an owner.
That's exactly where the difference between having tools and having a system shows up. A team that already manages infrastructure as code on AWS with clear governance criteria — disciplined provider versioning, explicit policies, and continuous ownership of the pipeline — can use these agents as an added layer of watchfulness. A team without those criteria just automates the noise faster.
Quick Tip
The question worth asking isn't "do we have drift detection?" It's "who reviews what the agent finds, how often, and what happens if nobody does this week?"
Solving that operationally — not as a one-time project, but as a function someone owns continuously — is exactly what separates infrastructure that maintains itself on paper from infrastructure that actually maintains itself in production. Our work with AWS on platforms like Multimedios starts from that same premise: infrastructure as code is only worth as much as the operational discipline around it.
Need help with this?
This is exactly what we do in Build (Development). Our team can implement it in your company.
Want to implement this in your company?
We can help you take your business to the next level with cutting-edge technology.