Terraform State Drift: How to Detect and Fix Manual AWS Changes

Terraform state drift happens when the real infrastructure managed by Terraform no longer matches the values Terraform expects from its configuration and state. Manual AWS console changes are a common cause, but drift can also come from automation, autoscaling, external controllers, provider-side defaults, scripts, or changes made by another team.

The important part is not to “force Terraform back into sync” immediately. First determine what changed, why it changed, and whether the live change should be accepted or reverted. A rushed state update can make an accidental production change look intentional, while a rushed apply can undo a legitimate emergency fix.

This guide uses the safer workflow recommended by HashiCorp: inspect drift first, decide the desired state, then update either the Terraform configuration, the state, or the remote infrastructure as appropriate.

What Causes Terraform State Drift?

Terraform tracks managed resources in state, while your HCL configuration describes the desired infrastructure. During terraform plan and terraform apply, Terraform refreshes its view of remote objects and compares the real infrastructure with the configuration.

Drift appears when those values no longer line up. Typical causes include:

  • manual changes in the AWS Management Console
  • AWS CLI or SDK changes made outside the Terraform workflow
  • autoscaling or controllers changing values such as desired capacity
  • emergency production fixes that were never added back to HCL
  • multiple pipelines or engineers operating on the same infrastructure
  • provider or service behavior changing attributes after deployment

Manual change does not automatically mean Terraform will destroy a resource. The next plan shows the actions Terraform believes are required to make remote infrastructure match the configuration. Depending on the resource and attribute, that may be an in-place update, no action, or a replacement.

Safe Terraform State Drift Workflow

Step 1: Pause Competing Applies

If production is already in an uncertain state, temporarily stop automated Terraform applies and make sure another engineer or pipeline is not changing the same state at the same time. State locking helps prevent concurrent writes, but operationally you still want one controlled recovery path.

Step 2: Run a Normal Plan First

terraform plan

A normal plan refreshes Terraform’s view of managed objects and shows how the real environment differs from the configuration. Read the proposed actions carefully, especially any line marked for replacement or destruction.

Do not assume every difference should be accepted into state. Ask whether the out-of-band change was intentional and whether it should become the new desired configuration.

Step 3: Use Refresh-Only Mode to Review State Changes

If you want to inspect how Terraform would update state to reflect the current remote objects without changing the infrastructure, use:

terraform plan -refresh-only

HashiCorp recommends reviewing the refresh-only plan before committing it. If the proposed state changes are correct, you can then run:

terraform apply -refresh-only

A refresh-only apply updates Terraform state and outputs to match the current remote objects; it does not make further changes to the remote infrastructure. That does not mean it is always the right fix. If the manual AWS change was accidental, you may instead want to keep the original HCL and use a normal plan/apply to restore the intended configuration.

Step 4: Choose One Source of Desired Truth

After you identify the drift, choose one of these paths:

  • Accept the live change: update the Terraform configuration so HCL describes the new desired value, then plan again.
  • Reject the live change: keep the Terraform configuration as-is and use a reviewed normal apply to restore the managed resource.
  • Bring an unmanaged resource under Terraform: write the matching resource configuration and use an import block or terraform import as appropriate.

The goal is to end with configuration, state, and real infrastructure describing the same intended environment.

Do Not Use Refresh-Only as an Automatic Fix

One of the easiest mistakes is to run terraform apply -refresh-only immediately and assume drift is fixed. Refresh-only records remote changes in state, but it does not decide whether those changes were correct.

HashiCorp specifically recommends reviewing refresh-only changes first. A bad provider configuration or wrong credentials can make Terraform believe managed objects are missing, so blindly updating state can remove objects from Terraform’s tracking even though the infrastructure still exists elsewhere.

Modern AWS Backend Setup for Terraform State

Remote state does not eliminate drift, but it makes team workflows safer by centralizing state, controlling access, and preventing concurrent writes. For AWS, Amazon S3 is a common backend.

ControlRecommended ApproachWhy It Matters
State storageAmazon S3 remote backendCentralizes state for CI/CD and teams.
State lockingS3 native lockfile with use_lockfile = trueReduces concurrent state-write risk.
RecoveryS3 bucket versioningProvides previous state object versions after accidental changes.
AccessLeast-privilege IAMLimits who can read or modify sensitive state.
Manual changesControlled break-glass accessKeeps emergency access possible without making console edits the normal workflow.

Current HashiCorp documentation supports native S3 state locking through use_lockfile. DynamoDB-based locking is now deprecated and is retained mainly for compatibility with older Terraform versions. AWS Prescriptive Guidance likewise recommends S3 native locking for current deployments.

terraform {
  backend "s3" {
    bucket       = "myorg-terraform-states"
    key          = "prod/network/terraform.tfstate"
    region       = "us-east-1"
    use_lockfile = true
  }
}

Also enable S3 versioning and tightly restrict access to the state bucket. Terraform state can contain sensitive values, so it should be treated as security-sensitive data.

When Should You Use ignore_changes?

The lifecycle.ignore_changes argument can be useful when another system intentionally manages a specific attribute. A common example is an autoscaling system changing desired capacity.

lifecycle {
  ignore_changes = [desired_capacity]
}

Use it narrowly. ignore_changes tells Terraform not to plan updates for selected attributes after creation; it should not be used to hide unexplained drift across security groups, IAM policies, encryption settings, or other security-sensitive configuration.

How to Prevent Terraform Drift in AWS

  • Run reviewed plans in CI/CD: require plan review before production applies.
  • Use remote state and locking: centralize state and prevent simultaneous writers.
  • Restrict direct write access: use least-privilege IAM and a documented break-glass process for emergencies.
  • Feed emergency fixes back into code: any approved console change should be represented in Terraform afterward.
  • Schedule drift detection: run terraform plan or appropriate HCP Terraform drift checks on a regular cadence.
  • Protect credentials: prefer short-lived identities for CI/CD instead of permanent AWS keys. See our analysis of leaked AWS keys and long-lived credential risk.

Drift prevention is also a process problem, not only a Terraform setting. Clear ownership, peer review, and controlled deployment paths are part of the same DevOps discipline discussed in our guide to how Agile and DevOps interrelate.

Official Documentation

Frequently Asked Questions

Does terraform plan detect state drift?

Yes. A normal Terraform plan refreshes its view of managed remote objects before calculating proposed changes, so it can reveal differences between the current infrastructure and the configuration.

Should I always run terraform apply -refresh-only when drift appears?

No. First review the drift and decide whether the remote change should be accepted or reverted. Refresh-only is appropriate when you intentionally want state to record the current remote values without changing the infrastructure.

Is DynamoDB still required for Terraform S3 state locking?

No. Current Terraform supports native S3 state locking with use_lockfile = true. DynamoDB-based locking is deprecated and mainly remains for compatibility with older Terraform versions.

Is the terraform refresh command deprecated?

Yes. HashiCorp recommends using terraform plan -refresh-only to review proposed state updates and terraform apply -refresh-only when you intentionally want to commit those updates.

About the author

Kiran Sonawane

Kiran Sonawane is a DevOps engineer working with AWS, Azure, Google Cloud, Terraform and Kubernetes. His experience includes infrastructure automation, CI/CD pipelines, cloud security and deploying AI applications. At TechUpdate24, he writes about cloud engineering, DevOps, security and AI tooling.