Terraform

Best Practices for Using Terraform in Production Environments

AGAnurag Gupta04 Apr 2025 Β· Updated 04 Oct 2026 Β· 8 min read
Best Practices for Using Terraform in Production Environments

Quick answer: Production Terraform comes down to a handful of non-negotiables: code in Git with mandatory review, remote state with locking and encryption, pinned Terraform, provider and module versions, strict separation between environments, applies run only by CI with a saved plan, no secrets in code, and guard rails such as prevent_destroy, policy checks and drift detection. Teams that follow these rarely have Terraform incidents; teams that skip them have one eventually.

Getting Terraform to work on a laptop is easy. Running it safely against infrastructure that serves customers is a different discipline. This guide turns the familiar checklist into concrete, current practices with working examples: repository layout, version pinning, the plan-and-apply pipeline, secrets handling, safety features built into Terraform 1.x, and the production mistakes that still cause outages.

Treat infrastructure code like application code

Everything starts with version control. Terraform code lives in Git, changes arrive as pull requests, and nobody applies from a personal branch. Three habits make this work:

  • Automated checks on every PR: terraform fmt -check, terraform validate, a linter such as tflint, and a security scanner such as trivy or checkov.
  • A speculative plan posted to the PR so reviewers see exactly which resources will change, not just which lines of HCL did.
  • Branch protection requiring at least one approval and green checks before merge to main.

Organise the repository so blast radius is obvious. A common, battle-tested layout:

infra/
β”œβ”€β”€ modules/                 # reusable building blocks, versioned by tag
β”‚   β”œβ”€β”€ vpc/
β”‚   └── eks-cluster/
└── live/                    # one root module per environment and component
    β”œβ”€β”€ dev/
    β”‚   β”œβ”€β”€ network/         # own backend key, own state
    β”‚   └── platform/
    β”œβ”€β”€ staging/
    β”‚   β”œβ”€β”€ network/
    β”‚   └── platform/
    └── prod/
        β”œβ”€β”€ network/
        └── platform/

Each leaf directory under live/ has its own state file. A mistake in prod/platform cannot touch prod/network, and dev never shares state with prod. Workspaces are an alternative for small projects, but separate directories make environment boundaries visible in the file tree and in IAM.

Pin every version

Unpinned versions are the most common cause of “it worked yesterday”. Pin Terraform itself, every provider and every module, and commit the lock file.

# versions.tf in every root module
terraform {
  required_version = "~> 1.9.0"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.70"
    }
  }

  backend "s3" {
    bucket       = "tkh-terraform-state-123456789012"
    key          = "prod/platform/terraform.tfstate"
    region       = "ap-south-1"
    encrypt      = true
    use_lockfile = true
  }
}

module "vpc" {
  source  = "terraform-aws-modules/vpc/aws"
  version = "~> 5.13"
  # ...
}

Commit .terraform.lock.hcl; it pins exact provider builds and checksums across every machine and CI runner. Upgrade deliberately with terraform init -upgrade in its own pull request, so a provider bump is reviewed separately from feature changes. Variables keep this configuration flexible without editing files per environment; Using Variables and Expressions in Terraform covers the patterns.

Remote state, locking and encryption

Local state has no place in production. Use a remote backend in the cloud you already run on, with encryption at rest, object versioning for recovery and locking so concurrent applies cannot corrupt state. The S3 block above shows the AWS version; Azure Blob and GCS lock automatically. Restrict read access to the state bucket as tightly as to a secrets store, because state contains every sensitive attribute in plain text. Full configuration details are in Storing Terraform State Remotely.

Apply only through a pipeline, and only a saved plan

Humans review; machines apply. Running terraform apply from a laptop against production bypasses every control you built. A minimal GitHub Actions workflow that plans on pull requests and applies the exact reviewed plan after merge:

name: terraform-prod-platform
on:
  pull_request:
    paths: ["live/prod/platform/**"]
  push:
    branches: [main]
    paths: ["live/prod/platform/**"]

permissions:
  id-token: write    # OIDC to AWS, no long-lived keys
  contents: read
  pull-requests: write

jobs:
  plan:
    if: github.event_name == 'pull_request'
    runs-on: ubuntu-latest
    defaults: { run: { working-directory: live/prod/platform } }
    steps:
      - uses: actions/checkout@v4
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::123456789012:role/tf-plan-readonly
          aws-region: ap-south-1
      - uses: hashicorp/setup-terraform@v3
        with: { terraform_version: "1.9.8" }
      - run: terraform init -input=false
      - run: terraform plan -input=false -lock-timeout=5m -out=tfplan
      - run: terraform show -no-color tfplan > plan.txt
      # post plan.txt as a PR comment with your preferred action

  apply:
    if: github.event_name == 'push'
    runs-on: ubuntu-latest
    environment: production        # requires manual approval in GitHub
    defaults: { run: { working-directory: live/prod/platform } }
    steps:
      - uses: actions/checkout@v4
      - uses: aws-actions/configure-aws-credentials@v4
        with:
          role-to-assume: arn:aws:iam::123456789012:role/tf-apply
          aws-region: ap-south-1
      - uses: hashicorp/setup-terraform@v3
        with: { terraform_version: "1.9.8" }
      - run: terraform init -input=false
      - run: terraform plan -input=false -lock-timeout=10m -out=tfplan
      - run: terraform apply -input=false tfplan

Key points: the plan role is read-only, the apply role is separate and protected by a GitHub environment with required reviewers, credentials come from OIDC rather than stored keys, and apply consumes a saved plan file so what runs is exactly what was approved. GitLab CI, Jenkins, Atlantis, HCP Terraform and Spacelift all support the same pattern.

Keep secrets out of code and state where possible

Approach How Trade-off
Never in .tf or committed .tfvars Pass via TF_VAR_* environment variables or CI secret stores Values still land in state
Mark variables and outputs sensitive = true Hides values in CLI output and logs Does not encrypt state
Generate secrets in the cloud aws_secretsmanager_secret with random_password, or managed passwords such as RDS manage_master_user_password Application reads the secret at runtime, Terraform never sees the value
Ephemeral values (Terraform 1.10+) ephemeral variables and write-only arguments keep secrets out of plan and state entirely Requires provider support for the specific argument

Combine these with .gitignore entries for *.tfstate*, *.tfvars (except non-sensitive example files) and .terraform/, and a pre-commit hook such as gitleaks to catch accidents.

Use Terraform’s built-in guard rails

Terraform 1.x ships several safety features that cost nothing to enable:

resource "aws_db_instance" "main" {
  identifier     = "payments-prod"
  engine         = "postgres"
  instance_class = "db.r6g.large"
  # ...

  lifecycle {
    prevent_destroy = true                  # refuse any plan that deletes this
    ignore_changes  = [password]            # managed outside Terraform

    precondition {
      condition     = var.environment == "prod" ? var.multi_az : true
      error_message = "Production databases must be Multi-AZ."
    }
  }
}

check "alb_is_healthy" {
  data "http" "health" {
    url = "https://${aws_lb.app.dns_name}/healthz"
  }
  assert {
    condition     = data.http.health.status_code == 200
    error_message = "Load balancer health check failed after apply."
  }
}

Add policy as code on top: Sentinel on HCP Terraform, or Open Policy Agent via conftest in any pipeline, to enforce rules such as “all S3 buckets must block public access” before apply. Finally, schedule a nightly terraform plan -detailed-exitcode for every production state and alert on exit code 2, which means drift has been detected.

Eight production mistakes that still happen

  1. Applying from a laptop. Every control above is bypassed. Restrict apply credentials to the CI role.
  2. One giant state file. Plans take twenty minutes and one typo threatens everything. Split by environment and component.
  3. Running apply -auto-approve without a saved plan. What applies may differ from what was reviewed. Always plan -out then apply tfplan.
  4. Unpinned providers or modules. A major version lands unannounced and proposes to replace resources.
  5. Using -lock=false or force-unlock casually. This reintroduces the concurrency bug locking prevents.
  6. Long-lived cloud keys in CI secrets. Use OIDC federation; keys leak, short-lived tokens expire.
  7. No prevent_destroy on stateful resources. Databases, storage buckets and KMS keys deserve it.
  8. No drift detection. Console changes accumulate silently until an apply reverts a critical hotfix.

Frequently asked questions

Workspaces or separate directories for environments?

Separate directories (or separate repositories) for production. Workspaces share one backend configuration and one set of credentials, which makes it too easy to apply the wrong environment. Workspaces are fine for short-lived feature environments.

Do I need Terragrunt or HCP Terraform?

No. Plain Terraform with a good CI pipeline is enough for most teams. Terragrunt reduces duplication across many root modules; HCP Terraform adds a hosted runner, policy and run history. Adopt them when you feel the specific pain they solve.

How should I test Terraform before production?

Validate and lint on every PR, run terraform test for modules, apply to dev and staging through the same pipeline, and use check blocks for post-apply assertions. Promotion to production should be the same code with different variables.

How do I roll back a bad apply?

Terraform has no rollback command. Revert the commit and apply again; the plan will show the reverse change. For destroyed data you need backups, which is why prevent_destroy and backend versioning matter.

Key takeaways

  • Git, pull requests and automated checks are the foundation; nobody applies from a laptop.
  • Pin Terraform, provider and module versions, commit the lock file, and upgrade in dedicated PRs.
  • Remote, encrypted, locked state with one state file per environment and component.
  • CI applies a saved plan with OIDC credentials; secrets stay out of code; prevent_destroy, policy checks and drift detection catch what reviews miss.

If you want to build this entire production workflow yourself, from repository layout to a gated CI/CD pipeline deploying to AWS and Azure, our DevOps course covers Terraform, GitHub Actions, Jenkins and monitoring with mentor support and placement assistance. For free tutorials, subscribe to our YouTube channel.

AG
Written byAnurag Gupta

Part of the Techknowledgehub team of industry mentors, writing practical guides to help you build a job-ready tech career.

More articles by Anurag Gupta β†’
Keep reading

Related articles

Leave a Reply