Terraform Applied. Reality Diverged. Your Infrastructure Code Isn't the Source of Truth Anymore.
Infrastructure as Code was supposed to solve the 'it works on my machine' problem at the infrastructure layer. Define everything in version-controlled files, apply it consistently, and never again argue about what's actually running in prod versus what someone thought they configured last March.
Great idea. Genuinely transformative when it works. But in a lot of organizations, IaC has quietly become something else: a snapshot of what infrastructure looked like at some point in the past, dressed up with enough automation to feel trustworthy.
The code is there. The pipelines run. The README says 'fully reproducible.' And none of it reflects reality anymore.
The Drift Nobody Talks About in the Demo
Drift is what happens between the last [terraform](https://en.wikipedia.org/wiki/Terraform_(software)) apply and right now. It's the security group rule someone added manually through the AWS console because the incident was happening and there wasn't time for a PR. It's the instance type that got resized by a well-meaning SRE who forgot — or didn't know — that it needed to go through the IaC workflow. It's the S3 bucket policy that got tweaked directly because the Terraform module didn't expose that parameter and nobody wanted to refactor it at 11pm.
Each of these is a small, understandable decision. Collectively, they add up to infrastructure that's running in a configuration your code has never seen.
Run terraform plan on a mature production environment that hasn't had strict drift controls and you'll often see a wall of changes — resources to update, settings to revert, configurations that diverged months ago. Most teams look at that output, feel a little queasy, and then decide not to apply it because they don't know what those changes will break.
So the drift stays. And the code stays. And now you have two separate systems: the one that's checked in, and the one that's actually running.
IaC as Documentation That Happens to Have IAM Permissions
Here's the uncomfortable framing: if your IaC isn't being continuously reconciled with real infrastructure, it's basically documentation. Useful documentation, maybe — it describes intent, it captures decisions that were made at some point — but documentation nonetheless. It just also happens to have the ability to make real changes if someone runs it.
That's a weird thing to have. Documentation that's mostly accurate but diverges in unknown ways, sitting in a repo next to your application code, giving everyone the feeling that infrastructure is under control.
The feeling is the problem. Teams that believe their IaC reflects reality make decisions based on that belief. They estimate migration complexity by reading Terraform files. They audit security posture by reviewing the code. They plan capacity based on what the modules say. And when the actual infrastructure doesn't match, those decisions are wrong in ways that are hard to trace.
Why 'Set It and Forget It' Doesn't Work at the Infrastructure Layer
Application code gets touched constantly. PRs, reviews, CI runs, deployments — there are a dozen forcing functions that keep application code connected to reality. Infrastructure code doesn't have the same forcing functions.
You might apply a Terraform module once when you spin up a service, and then not touch it for eight months. In that eight months, the underlying cloud provider has released new features, your team has made manual adjustments, and the module itself may have drifted out of sync with the actual resource configuration. Nobody noticed because nothing was checking.
This is compounded by the organizational pressure to treat infrastructure as stable. 'If it's not broken, don't touch it' is reasonable advice for a lot of things. Applied to IaC, it means the code gets staler and staler while the infrastructure underneath keeps evolving.
The Versioning Problem Makes It Worse
Module versioning in Terraform — and the equivalent in Pulumi, CDK, or whatever your team prefers — is supposed to help. Pin a module version, know what you're getting. But in practice, version pinning creates its own problems.
Old pinned versions don't get security patches. They don't get support for new resource types. Teams end up running ancient module versions not because they're the right choice but because upgrading them is a project nobody has time for. The code that's checked in becomes a liability, not an asset.
And when someone does try to upgrade, the diff between the old module behavior and the new one, combined with the existing drift, creates a plan output that's genuinely difficult to reason about. So the upgrade gets deferred. The drift compounds. The cycle continues.
What Treating IaC Like Real Code Actually Looks Like
If you want your infrastructure code to mean something, it needs the same rigor you apply to application code — and then some.
Drift detection needs to be automated and visible. Tools like Atlantis, Spacelift, or even a simple scheduled terraform plan that posts output to a Slack channel will surface divergence before it becomes a crisis. The goal is to make drift uncomfortable to ignore, not something you only discover when you're trying to reproduce an environment.
Manual console changes need to be a process violation, not a cultural norm. That means having an actual runbook for 'I need to make an urgent infrastructure change,' and that runbook needs to include a follow-up step for codifying the change. Urgency is real; the answer is a faster IaC workflow, not an escape hatch that bypasses the workflow entirely.
Treat infrastructure modules like libraries. They need changelogs, semantic versioning, and deprecation policies. If a module is too rigid to expose the parameters your team needs, fix the module — don't work around it in the console.
And run regular reconciliation reviews. Not when something breaks — quarterly, as a habit. Pull the plan output, look at the drift, understand it, and make a conscious decision about what to do with it. That's the only way to know whether your code and your infrastructure are still in the same conversation.
The Source of Truth Has to Actually Be True
IaC is a genuinely powerful idea. The ability to version, review, and reproduce infrastructure is worth the investment. But the investment doesn't stop at writing the initial Terraform.
The source of truth only works if it's maintained as one. Otherwise you're carrying the cognitive overhead of two systems — the code and the reality — while telling yourself you only have one. That gap is where incidents are born, where compliance audits get awkward, and where 'just spin up a new environment' turns into a two-week project.
The code in your repo should describe the infrastructure that's running. If you're not sure whether it does, that uncertainty is the problem worth solving first.