A basic Terraform demo proves infrastructure as code works. Real infrastructure needs more than one config and a manual terraform apply.

I built this project to look more like a real workflow. It has reusable Terraform modules and separate dev and prod environments. It uses remote state, dependency handling, and automated CI/CD. It authenticates to AWS with short-lived credentials. It also runs security checks, policy checks, and cost estimates.

The stack:

  • Terraform for AWS resource definitions
  • Terragrunt for environment orchestration and dependency handling
  • GitHub Actions for CI/CD
  • AWS OIDC for short-lived authentication
  • Checkov for static infrastructure security checks
  • OPA/Conftest for policy-as-code validation
  • Infracost for cost estimation

1. Reusable Terraform modules

I split the infrastructure into three modules:

modules/
├── vpc/
├── ec2/
└── s3/

Each one has its own job:

modules/vpc
  VPC
  public subnet
  internet gateway
  route table

modules/ec2
  EC2 instance
  security group

modules/s3
  S3 bucket
  versioning
  encryption
  public-access blocking

The main rule is that modules hold no environment-specific values. Things like dev, prod, CIDR ranges, instance names, and environment tags come in as variables from each environment.

That way one module serves every environment. No duplicated code.

Keep variable names consistent

One of my first bugs came from naming. I used one variable name in the module and a different one in the calling config.

I had:

environment

in one place, and:

env

in another.

Pick your naming rules before the config spreads across environments. It saves a lot of pain.

2. Organizing environments with Terragrunt

Terraform can handle multiple environments by itself. I used Terragrunt to cut down on repeated config.

The big win is the folder layout. Dev and prod stay isolated but share the same backend and provider setup.

Here’s the layout:

live/
├── terragrunt.hcl
├── dev/
│   ├── vpc/
│   │   └── terragrunt.hcl
│   ├── ec2/
│   │   └── terragrunt.hcl
│   └── s3/
│       └── terragrunt.hcl
└── prod/
    ├── vpc/
    │   └── terragrunt.hcl
    ├── ec2/
    │   └── terragrunt.hcl
    └── s3/
        └── terragrunt.hcl

The root config lives in:

live/terragrunt.hcl

It sets the shared behavior for every environment.

3. One place for backend and provider config

The root Terragrunt file defines three shared things:

S3 backend
DynamoDB state locking
Generated AWS provider configuration

Every environment-level terragrunt.hcl includes the root file.

So I write the backend and provider logic once. I don’t copy it into every dev and prod folder.

That’s the main reason I like Terragrunt for multiple environments. The environments stay separate and the repeated config stays in one spot.

4. Linking modules with dependencies

The EC2 module can’t deploy by itself. It needs things the VPC module creates.

At minimum it needs:

vpc_id
subnet_id

Hardcoding IDs like:

vpc-0a1b2c3d

would break every time the VPC gets recreated with a new ID.

Terragrunt fixes this with a dependency block.

Here’s what I used:

dependency "vpc" {
  config_path = "../vpc"

  mock_outputs = {
    vpc_id           = "vpc-00000000"
    public_subnet_id = "subnet-00000000"
  }

  mock_outputs_allowed_terraform_commands = [
    "validate",
    "plan"
  ]
}

inputs = {
  vpc_id    = dependency.vpc.outputs.vpc_id
  subnet_id = dependency.vpc.outputs.public_subnet_id
}

It points the EC2 config at the VPC config and reads its outputs.

Once the VPC exists, Terragrunt passes the real values into EC2 on its own.

Why mock outputs help

Before the VPC is ever deployed, there are no outputs to read.

mock_outputs gives placeholder values so validate and plan still work:

mock_outputs = {
  vpc_id           = "vpc-00000000"
  public_subnet_id = "subnet-00000000"
}

I limited the placeholders to:

validate
plan

so they never get used in a real deployment.

To plan the whole environment I ran:

terragrunt run --all plan

Terragrunt reads the dependency graph and knows the VPC has to go before EC2.

5. Bootstrapping remote state

Remote state has a chicken-and-egg problem.

Terraform needs an S3 bucket to store state. But the same config can’t create that bucket before the backend exists.

I solved it with a one-time bootstrap.

I created two resources by hand first:

S3 bucket
DynamoDB table

The bucket stores state and has versioning on.

The DynamoDB table handles state locking.

Once those existed, Terragrunt could set up the remote backend.

6. CI/CD with GitHub Actions

Next I moved from local commands to a repeatable pipeline.

I split the pipeline so changes get checked before they deploy.

The flow:

Pull request
  │
  ├── Terraform/Terragrunt plan
  ├── Checkov
  ├── OPA/Conftest
  └── Infracost

Main/manual workflow
  │
  └── Terragrunt apply

Additional workflow
  │
  └── Drift detection

Now the repo controls infrastructure changes. Nothing depends on commands from one person’s laptop.

7. GitHub Actions to AWS with OIDC

I didn’t want long-lived AWS keys in GitHub secrets.

So GitHub Actions authenticates with OpenID Connect.

On the AWS side there are two pieces:

IAM OIDC identity provider
IAM role trusted by that provider

The role’s trust policy is scoped to my repo with a subject like:

repo:username/reponame:*

GitHub Actions gets a short-lived token and uses it to assume the role.

No permanent keys stored anywhere in the repo.

Fixing the missing OIDC provider

My first pipeline run failed with:

Could not assume role with OIDC:
No OpenIDConnect provider found in your account

The role existed and the trust policy was set. But I never created the OIDC provider itself.

I confirmed it with:

aws iam list-open-id-connect-providers

Once I registered the GitHub OIDC provider in IAM, the role assumption worked.

The role and the OIDC provider are separate pieces. You need both.

8. Terragrunt CLI changes

Another failure came from a Terragrunt syntax change.

The old command was:

terragrunt run-all plan

I needed the new form:

terragrunt run --all plan

The old syntax gave me:

ERROR  unknown command: "run-all"

The destroy command is:

terragrunt run --all destroy

When a command from an older example fails for no clear reason, check your installed CLI:

terragrunt --help

9. Adding Infracost

I wanted cost estimates in CI too.

My first try used:

infracost/actions/comment@v3

and failed with:

Error: Can't find 'action.yml', 'action.yaml' or 'Dockerfile'
for action 'infracost/actions/comment@v3'

The action’s structure had changed.

The new pattern:

1. Install the Infracost CLI
2. Run infracost breakdown
3. Run infracost comment github

The setup action is:

infracost/actions/setup@v3

The rest runs as shell commands instead of separate actions.

This lesson goes beyond Infracost. When a third-party action breaks, read its action.yml. It shows what the action expects and what you can redo with normal run: steps.

10. OPA and Conftest policy checks

I also added policy-as-code checks.

I started with:

--parser hcl

and Conftest returned:

Error: running test: parse configurations:
new parser: unknown parser: hcl

The fix was changing the parser to:

--parser hcl2

Small change. But it blocked the whole policy stage until I figured it out.

11. The first full deployment

With the Terragrunt, OIDC, pipeline, and tool issues fixed, a push to main ran the apply workflow successfully.

It created:

VPC
public subnet
internet gateway
route table
security group
EC2 instance
S3 bucket

The instance type was:

t2.micro

The whole deploy ran through the pipeline with OIDC. I never ran terraform apply from my laptop.

I kept it up for about six hours and then tore it down with:

terragrunt run --all destroy

The cost for that short run was tiny. That number was specific to my account and pricing at the time, so don’t treat it as a current estimate.

12. Handling Checkov results

Checkov reported:

34 checks total
20 passed
14 failed

I didn’t want every failure to block the pipeline. So I set:

soft_fail: true

and wrote down each failed check.

Some were things I’d fix in a production design. Others I accepted for a portfolio project because fixing them would add cost or complexity.

Two examples:

S3 cross-region replication
VPC flow logs

For each finding I recorded whether it was:

accepted temporarily
planned for production remediation
or deliberately excluded

A written decision beats quietly ignoring a failed check.

13. What I’d tell someone building this

Here’s the order I’d follow:

  1. Get plain Terraform working before adding Terragrunt.
  2. Learn Terragrunt dependencies and outputs before building the full environment tree.
  3. Do a real deployment early. Auth, AMI, backend, and provider issues don’t always show up in static checks.
  4. Use terragrunt run --all destroy to get a test environment back to clean.
  5. Write down design and security trade-offs when you make them. Don’t try to remember them later.

I documented choices like using a public subnet instead of NAT, broad IAM permissions for testing, and letting Checkov soft-fail. Those are project trade-offs. They are not production defaults.

14. Final architecture

Here’s the full design:

GitHub repository
       │
       ▼
GitHub Actions
       │
       ├── Plan
       ├── Checkov
       ├── OPA / Conftest
       ├── Infracost
       └── Apply / drift workflows
       │
       ▼
AWS OIDC
       │
       ▼
IAM Role
       │
       ▼
Terragrunt
       │
       ├── dev
       │   ├── VPC
       │   ├── EC2
       │   └── S3
       │
       └── prod
           ├── VPC
           ├── EC2
           └── S3

Terraform provides the reusable modules.

Terragrunt handles environment separation, remote state, shared config, and dependencies.

GitHub Actions runs validation and deployment.

OIDC gives temporary AWS credentials instead of stored keys.

The repo is here:

https://github.com/dcepeda31415/multi-env-aws-infra

15. Proof it deployed

Here’s what the pipeline created.

The S3 bucket holding Terraform remote state:

Terraform remote state stored in the S3 backend

Terraform remote state stored in the S3 backend

The EC2 instance from the automated deploy:

EC2 instance created by the infrastructure pipeline

EC2 instance created by the infrastructure pipeline

And the VPC:

VPC created by the Terraform and Terragrunt deployment

VPC created by the Terraform and Terragrunt deployment

Wrapping up

This goes past creating one resource with Terraform. The value is in how the pieces work together:

reusable Terraform modules
+ environment isolation
+ dependency-aware Terragrunt orchestration
+ remote state
+ GitHub Actions
+ OIDC
+ security checks
+ policy checks
+ cost visibility

Most of what I learned came from wiring things together, not from writing Terraform resources. Remote state bootstrapping, dependencies, OIDC setup, CLI changes, and action updates are the exact problems you hit when infrastructure code moves from a local demo into an automated pipeline.