DevOps

Terraform for DevOps Engineers: Infrastructure as Code Mastery

AGAnurag Gupta18 Mar 2023 Β· Updated 04 Oct 2026 Β· 9 min read
Terraform for DevOps Engineers: Infrastructure as Code Mastery

Quick answer: Terraform is the Infrastructure as Code (IaC) tool DevOps engineers use to create, change and version cloud infrastructure from declarative HCL files instead of clicking through consoles. You write what you want (a VPC, a cluster, a database), run terraform plan to preview the changes and terraform apply to make them, and Terraform tracks everything in a state file so the same code can be applied repeatedly, safely, across AWS, Azure, GCP and Kubernetes.

This article is a complete, structured path to Terraform mastery for DevOps engineers. It starts with why IaC matters, walks through the core workflow with real HCL, then covers variables and modules, state management, providers, testing, CI/CD integration, Kubernetes, multi-cloud and security, before finishing with a learning roadmap, the mistakes that hurt production teams, and answers to common questions.

Why Terraform belongs in every DevOps toolchain

Before IaC, infrastructure was built by hand or with ad-hoc shell scripts. Environments drifted apart, nobody knew who had changed what, and rebuilding a region after an outage took days. Terraform fixes this by treating infrastructure like application code: it lives in Git, goes through pull requests, is reviewed, tested and deployed through a pipeline. The result is infrastructure that is repeatable (dev, staging and prod built from the same modules), auditable (every change is a commit) and scalable (one engineer can manage thousands of resources).

Compared with cloud-specific tools such as AWS CloudFormation or Azure Bicep, Terraform is provider-agnostic: the same language and workflow manage AWS, Azure, GCP, Kubernetes, GitHub, Datadog and hundreds of other platforms. Since 2023 Terraform ships under the BUSL licence, and the community fork OpenTofu is a drop-in alternative; everything in this guide applies to both.

Core concepts and the Terraform workflow

Four ideas carry you through almost everything:

  • Providers – plugins that talk to a platform’s API (hashicorp/aws, hashicorp/azurerm, hashicorp/google, hashicorp/kubernetes).
  • Resources – the things you create, such as aws_instance or azurerm_storage_account. Data sources read existing things.
  • State – a JSON record of what Terraform created, used to compute the diff on the next run.
  • Dependency graph – Terraform works out the order of operations automatically from references between resources.

The lifecycle is always terraform init (download providers, configure the backend) β†’ terraform fmt and terraform validate (format and sanity-check) β†’ terraform plan (preview) β†’ terraform apply (execute) β†’ terraform destroy (tear down). The plan step is the single biggest reason DevOps teams trust Terraform: nothing changes until a human or pipeline approves the diff.

Hands-on: from a single resource to reusable modules

A minimal, production-shaped configuration that launches a web server in Mumbai with a remote backend and pinned versions:

terraform {
  required_version = ">= 1.6.0"

  required_providers {
    aws = {
      source  = "hashicorp/aws"
      version = "~> 5.0"
    }
  }

  backend "s3" {
    bucket         = "tkh-terraform-state"
    key            = "web/prod/terraform.tfstate"
    region         = "ap-south-1"
    dynamodb_table = "tkh-terraform-locks"
    encrypt        = true
  }
}

provider "aws" {
  region = var.region
}

variable "region" {
  type    = string
  default = "ap-south-1"
}

variable "instance_type" {
  type        = string
  description = "EC2 size for the web tier"
  default     = "t3.micro"
}

data "aws_ami" "ubuntu" {
  most_recent = true
  owners      = ["099720109477"] # Canonical
  filter {
    name   = "name"
    values = ["ubuntu/images/hvm-ssd/ubuntu-jammy-22.04-amd64-server-*"]
  }
}

resource "aws_instance" "web" {
  ami           = data.aws_ami.ubuntu.id
  instance_type = var.instance_type

  tags = {
    Name        = "tkh-web"
    Environment = "prod"
    ManagedBy   = "terraform"
  }
}

output "public_ip" {
  value = aws_instance.web.public_ip
}

Notice that credentials are absent (they come from environment variables or an IAM role), the provider version is pinned, and the state is stored in S3 with DynamoDB locking so two engineers cannot apply at the same time. For a deeper walkthrough of this file structure see How to Write Terraform Code: A Beginner’s Guide.

Once a pattern repeats, wrap it in a module. A module is just a folder of .tf files with inputs and outputs; the root configuration calls it like a function:

# modules/vpc/main.tf
variable "name"       { type = string }
variable "cidr_block" { type = string }
variable "azs"        { type = list(string) }

resource "aws_vpc" "this" {
  cidr_block           = var.cidr_block
  enable_dns_hostnames = true
  tags = { Name = var.name }
}

resource "aws_subnet" "public" {
  for_each          = toset(var.azs)
  vpc_id            = aws_vpc.this.id
  availability_zone = each.value
  cidr_block        = cidrsubnet(var.cidr_block, 8, index(var.azs, each.value))
  tags = { Name = "${var.name}-public-${each.value}" }
}

output "vpc_id"     { value = aws_vpc.this.id }
output "subnet_ids" { value = [for s in aws_subnet.public : s.id] }
# envs/prod/main.tf
module "vpc" {
  source     = "../../modules/vpc"
  name       = "tkh-prod"
  cidr_block = "10.0.0.0/16"
  azs        = ["ap-south-1a", "ap-south-1b"]
}

module "vpc_staging" {
  source     = "git::https://github.com/techknowledgehub/terraform-modules.git//vpc?ref=v1.2.0"
  name       = "tkh-staging"
  cidr_block = "10.1.0.0/16"
  azs        = ["ap-south-1a"]
}

Version your modules with Git tags or a private registry, and never edit a published version in place. Variables, locals and expressions such as for_each, cidrsubnet() and conditionals are what make modules flexible; we cover them in detail in Using Variables and Expressions in Terraform.

State management: the part that breaks production

State maps your code to real resource IDs. Lose it and Terraform wants to recreate everything; corrupt it and you get phantom diffs. Production rules:

  • Store state remotely (S3 + DynamoDB, Azure Blob, GCS, or HCP Terraform) with encryption and versioning enabled.
  • Enable locking so concurrent applies are rejected.
  • Split state by environment and by blast radius: networking, data and application tiers in separate state files.
  • Use import blocks (Terraform 1.5+) to bring hand-made resources under management, and moved blocks when refactoring so resources are renamed instead of destroyed.
  • Treat terraform state rm and manual state edits as emergency surgery, never routine.

Testing, CI/CD and integrations

Terraform fits naturally into a pipeline. A typical GitHub Actions, GitLab CI or Jenkins flow runs fmt -check, validate, a linter such as tflint, a security scanner such as tfsec, checkov or trivy, then posts the plan output on the pull request. Merging to main triggers apply with an approval gate. Terraform 1.6+ also has a native terraform test command for module unit tests written in HCL, and tools such as Terratest let you write integration tests in Go.

Terraform provisions; configuration managers configure. A common pattern is Terraform creating the servers and then handing off to Ansible for OS-level setup, or, better, baking images with Packer so servers are immutable. Terraform also pairs with AWS CodePipeline, Argo CD and Helm for application delivery. Avoid provisioner blocks except as a last resort; they are hard to test and hide state outside Terraform.

Kubernetes, cloud-native and multi-cloud

Terraform commonly creates the cluster (EKS, AKS or GKE) and foundational add-ons via the kubernetes and helm providers, while application manifests flow through GitOps tools. The same configuration can declare a serverless API on AWS Lambda, a container on Azure Container Apps and DNS on Cloudflare, which is why Terraform is the default choice for multi-cloud teams. Use workspaces sparingly for lightweight environment switching; for real multi-environment setups prefer separate directories and state files per environment.

Security and migration

Lock down state buckets (they contain secrets in plain text), use OIDC federation instead of long-lived cloud keys in CI, pull secrets from Vault or a cloud secrets manager at apply time, and enforce guardrails with policy-as-code (Sentinel or Open Policy Agent). When migrating existing infrastructure, decide between a lift-and-shift import of what exists and a cloud-native rebuild from clean modules; most teams import critical stateful resources (databases, DNS) and rebuild everything stateless.

Terraform learning roadmap for DevOps engineers

Stage Topics Outcome
1. Foundations IaC vs manual ops, HCL syntax, providers, resources, data sources, the plan/apply lifecycle Deploy a VM and a bucket on one cloud
2. Reusability Variables, outputs, locals, expressions, count/for_each, modules, module versioning Build a reusable VPC module
3. State and teams Remote backends, locking, import/moved blocks, workspaces, environment layout Multi-environment repo with shared state
4. Production Testing (terraform test, Terratest), linting, security scanning, CI/CD pipelines, HCP Terraform Pull-request driven infrastructure delivery
5. Advanced Kubernetes and Helm providers, multi-cloud networking, policy-as-code, custom providers with the Plugin Framework Platform-engineering level ownership

Eight Terraform mistakes that hurt real teams

  1. Local state on a laptop. One lost disk and the environment is unmanaged. Use a remote backend from day one.
  2. Unpinned provider versions. A major-version bump can rewrite your plan. Pin with ~> and commit .terraform.lock.hcl.
  3. One giant state file. Every plan takes minutes and every mistake has a huge blast radius. Split by layer and environment.
  4. Secrets in .tfvars committed to Git. Use environment variables, Vault or a secrets manager.
  5. Applying without reading the plan. The word destroy in a plan should always stop you.
  6. Copy-pasting instead of modules. Duplication guarantees drift between environments.
  7. Abusing workspaces for prod/staging. Shared backend configuration and easy mistakes; use directories.
  8. Clicking in the console after Terraform runs. Manual changes create drift; fix it in code or import it.

Frequently asked questions

Do I need to know a cloud before learning Terraform?

Basic familiarity with one provider (AWS, Azure or GCP), Linux command-line comfort and an understanding of the software delivery lifecycle are enough. Terraform will teach you the cloud faster, because every resource argument is documented.

Terraform or Ansible, which should a DevOps engineer learn?

Both, but for different jobs. Terraform provisions infrastructure declaratively; Ansible configures machines and deploys software. Most job descriptions list both.

What is the difference between Terraform, HCP Terraform and Terraform Enterprise?

The CLI is free and open. HCP Terraform (formerly Terraform Cloud) is HashiCorp’s SaaS for remote state, runs, policy and team access. Terraform Enterprise is the self-hosted version for organisations with strict compliance needs.

Is Terraform still worth learning with OpenTofu around?

Yes. The language, providers and workflow are the same; OpenTofu is a compatible open-source fork. Learning one means you know both.

Key takeaways

  • Terraform turns infrastructure into reviewed, versioned, repeatable code, which is the heart of modern DevOps.
  • Master the lifecycle, then variables and modules, then state; that order mirrors how teams actually grow.
  • Remote state with locking, pinned versions and CI-driven plans are non-negotiable in production.
  • Pair Terraform with Ansible, Packer, Kubernetes and policy-as-code for a complete platform.

Ready to go from reading about Terraform to deploying real infrastructure on AWS and Azure through a CI/CD pipeline? Our DevOps course covers Terraform, Docker, Kubernetes, Jenkins and cloud end to end with live projects, mentor support and placement assistance. Prefer video? Follow along on our YouTube channel.

AG
Written byAnurag Gupta

Part of the Techknowledgehub team of industry mentors, writing practical guides to help you build a job-ready tech career.

More articles by Anurag Gupta β†’
Keep reading

Related articles

Leave a Reply