Quick answer: Troubleshooting Terraform means working through a fixed sequence: read the error carefully, run terraform validate and terraform plan to isolate the failing resource, turn on TF_LOG for detail, inspect state with terraform state list and show, and fix root causes rather than forcing the apply. Monitoring Terraform means watching for drift, failed pipeline runs and stale locks, and shipping the infrastructure you create with its own observability built in.
Most Terraform problems fall into a small number of categories β syntax, provider/API errors, state mismatches, locking, and dependency ordering. This guide shows you how to recognise each one quickly, which commands to reach for, and how to set up monitoring so the next problem is caught before it reaches production.
Understand where Terraform can fail
Every Terraform run has four stages, and each stage fails in a characteristic way:
| Stage | Typical failure | First command to run |
|---|---|---|
| init | Provider download, backend auth, lock file mismatch | terraform init -upgrade or check backend credentials |
| validate | Syntax errors, unknown arguments, type mismatches | terraform validate |
| plan | Missing variables, unreadable data sources, auth errors, drift | terraform plan -var-file=... |
| apply | API quota limits, IAM denials, name conflicts, timeouts | Re-run with TF_LOG=DEBUG |
Knowing which stage failed narrows the search immediately. A problem during init is almost never about your resources; a problem during apply is almost never about HCL syntax.
Step 1: Validate and format before anything else
Run these two commands before every commit. They are fast, free and catch a surprising share of issues:
terraform fmt -check -recursive # consistent formatting
terraform validate # syntax, types, required arguments
# Typical validate output you want to see:
# Success! The configuration is valid.
validate does not contact any cloud API, so it cannot detect wrong AMI IDs or missing permissions β but it will tell you exactly which line has a typo in an argument name or references a variable that does not exist.
Step 2: Read the plan like a reviewer
terraform plan is your most important diagnostic tool. Save the plan so you can inspect it and apply exactly what you reviewed:
terraform plan -out=tfplan
terraform show tfplan # human-readable
terraform show -json tfplan | jq '.resource_changes[] | {address, actions: .change.actions}'
# Apply precisely the plan you reviewed
terraform apply tfplan
Watch for three warning signs in the plan output. Unexpected replacements (-/+) mean a change forces a new resource β common with name changes or AMI updates. Changes you did not make indicate drift: someone edited the resource in the console. “Known after apply” everywhere can signal a dependency problem, which is explained in Understanding Resource Dependencies and Ordering in Terraform.
Step 3: Turn on detailed logging
When the error message is vague β “unexpected EOF”, “error reading resource”, a bare HTTP 403 β enable provider-level logging. Terraform writes logs based on the TF_LOG environment variable, and TF_LOG_PATH sends them to a file instead of flooding the terminal:
export TF_LOG=DEBUG # TRACE, DEBUG, INFO, WARN, ERROR
export TF_LOG_PATH=./terraform-debug.log
terraform apply
# Only the provider plugin, not Terraform core:
export TF_LOG_CORE=ERROR
export TF_LOG_PROVIDER=DEBUG
# Search the log for the actual API response
grep -n "HTTP Response" terraform-debug.log | head
Debug logs include every request and response the provider sends to the cloud API, which usually reveals the real cause β a quota error, a malformed ARN, a region mismatch. Remember to unset TF_LOG afterwards, and never commit log files; they can contain sensitive values.
Step 4: Inspect and repair state
Many “mysterious” errors are state problems: a resource exists in the cloud but not in state (or the reverse), or an address changed after a refactor. The state subcommands are your toolkit:
terraform state listβ see every address Terraform tracks.terraform state show aws_instance.webβ view the recorded attributes of one resource.terraform import aws_s3_bucket.logs tkh-logs-prodβ adopt an existing resource into state. In Terraform 1.5+ prefer animportblock so the operation is reviewable in a plan.movedblocks β rename or move resources without destroying them when you refactor into modules.terraform state rmβ stop managing a resource without deleting it. Use carefully.terraform force-unlock LOCK_IDβ release a stale lock after a crashed run, only once you have confirmed no other run is active.
Always back up state before manual surgery. With a versioned S3 or GCS backend you already have history; with local state, copy the file first.
Monitoring Terraform in production
Troubleshooting is reactive. Monitoring makes problems visible before anyone files a ticket. Four practices cover most teams:
- Drift detection. Schedule
terraform plan -detailed-exitcodenightly in CI. Exit code 2 means changes are pending; alert on it. HCP Terraform and Spacelift offer this as a built-in health check. - Pipeline observability. Treat every plan and apply as an event: record duration, success or failure, number of resources changed, and who triggered it. Export these to Prometheus, Datadog or CloudWatch and dashboard them.
- Cloud audit logs. Alert when resources tagged
ManagedBy = terraformare modified by a human principal in CloudTrail or Azure Activity Log β that is drift being created in real time. - Monitor what you build. Terraform should provision alarms alongside resources. The snippet below creates a CloudWatch alarm with the same code that creates the instance, so no server ships without monitoring.
resource "aws_cloudwatch_metric_alarm" "cpu_high" {
alarm_name = "${aws_instance.web.tags["Name"]}-cpu-high"
comparison_operator = "GreaterThanThreshold"
evaluation_periods = 3
metric_name = "CPUUtilization"
namespace = "AWS/EC2"
period = 300
statistic = "Average"
threshold = 80
alarm_actions = [aws_sns_topic.alerts.arn]
dimensions = {
InstanceId = aws_instance.web.id
}
}
Eight troubleshooting mistakes that make things worse
- Re-running
applyrepeatedly hoping it works. Read the error, fix the cause, then re-run. - Editing
terraform.tfstateby hand. Usestate mv,movedblocks andimportinstead; manual edits corrupt serials and lineage. - Using
force-unlockwhile a colleague’s apply is still running. Two writers to one state file means a corrupted state. - Testing in production first. Keep a dev or staging workspace with the same code and smaller sizes.
- Applying without a saved plan. The plan you reviewed and the plan that applies can differ if the world changed in between.
- Ignoring provider warnings about deprecations. They become hard errors on the next major upgrade.
- Leaving
TF_LOG=TRACEon. It slows runs dramatically and fills disks with logs that may include secrets. - Fixing drift in the console. Fix it in code, or import the manual change; otherwise it will reappear.
Frequently asked questions
What does “Error acquiring the state lock” mean?
Another process holds the lock, or a previous run crashed before releasing it. Check whether a pipeline or colleague is mid-apply. If nobody is, use terraform force-unlock with the lock ID shown in the error.
Why does plan show changes when I changed nothing?
Either someone modified the resource outside Terraform (drift), a provider upgrade changed default values, or a computed attribute is unstable. Run terraform plan -refresh-only to see what changed in the cloud without proposing any fix.
Can I see the exact API calls Terraform makes?
Yes. Set TF_LOG_PROVIDER=TRACE; the log includes full HTTP requests and responses between the provider and the cloud API.
Key takeaways
- Identify which stage failed β init, validate, plan or apply β and the cause is usually obvious.
- Use
TF_LOGandTF_LOG_PATHto see real API responses behind vague errors. - Repair state with
import,movedandstatesubcommands, never by hand-editing JSON. - Monitor drift, pipeline runs and cloud audit logs; provision alarms in the same code as the resources.
- Save plans, apply the saved plan, and test in a non-production workspace first.
Ready to debug and operate real Terraform pipelines with confidence? Our DevOps course teaches Terraform, CI/CD and cloud monitoring through live projects, mentor support and placement assistance. Prefer learning by watching? Subscribe to our YouTube channel.


