What is FinOps? Cloud Cost Management for DevOps Engineers

FinOps -- short for Financial Operations -- is the practice of bringing financial accountability to cloud infrastructure spending. The discipline emerged as cloud adoption matured and organisations discovered that the flexibility of pay-as-you-go pricing is also a mechanism for runaway costs when not actively managed. An unchecked Terraform module, an oversized RDS instance left running after a project ends, or an auto-scaling group without a maximum limit can cost tens of thousands of dollars a month without anyone noticing until the bill arrives.

For DevOps engineers, FinOps matters because you are the person provisioning the infrastructure. You make decisions about instance types, storage tiers, data transfer paths, and retention policies. Understanding the cost implications of those decisions is part of engineering competence, not just a finance team concern.

FinOps is relevant whenever you are working with Terraform, AWS, or any infrastructure that bills on consumption. It connects directly to the infrastructure-as-code and cloud modules in the DevOps tools guide.

The FinOps Lifecycle

The FinOps Foundation describes three phases that organisations cycle through continuously:

Inform: Make cost visible. Tag resources so you can attribute spend to teams, services, and environments. Use AWS Cost Explorer, Azure Cost Management, or Google Cloud Billing to understand where money is going. Build dashboards that show cost trends alongside utilisation metrics.

Optimize: Identify and act on savings opportunities. Rightsize underutilised instances. Buy Reserved Instances for stable baseline workloads. Move infrequently accessed S3 objects to cheaper storage tiers. Implement auto-scaling to stop paying for capacity you are not using at night.

Operate: Embed cost awareness into engineering processes. Require cost estimates in architecture reviews. Integrate cost diff tools into pull request pipelines. Set budget alerts. Review cloud spend in regular engineering meetings alongside reliability and performance metrics.

Tagging Strategies

Tags are the foundation of cost attribution. Without consistent tagging, you cannot answer "how much does the payments service cost?" or "what is staging infrastructure costing us per month?" A tagging strategy should be enforced at the Terraform level, not the console level -- human beings clicking in AWS consoles will tag inconsistently.

A minimal useful tag set:

# terraform/modules/base-tags/variables.tf
variable "tags" {
  type = map(string)
  default = {}
}

# Applied to every resource in the module
locals {
  common_tags = merge(
    {
      Environment = var.environment          # prod, staging, dev
      Service     = var.service_name         # payments, auth, api
      Team        = var.team                 # platform, backend, data
      ManagedBy   = "terraform"
      CostCenter  = var.cost_center          # finance reference
    },
    var.tags
  )
}

resource "aws_instance" "app" {
  ami           = var.ami_id
  instance_type = var.instance_type
  tags          = local.common_tags
}

AWS Service Control Policies (SCPs) can enforce mandatory tags at the organisation level, rejecting resource creation that does not include required tag keys.

Reserved Instances, Savings Plans, and Spot

On-demand pricing is what you pay with no commitment. It is the baseline and the most expensive option for sustained workloads.

Reserved Instances (RIs) and Savings Plans commit to 1 or 3 years of usage in exchange for 30--72% discounts. Savings Plans are more flexible than RIs -- they apply to any instance in a region rather than a specific instance type. For production workloads that run continuously, buying Savings Plans is one of the highest-impact FinOps actions available.

Spot instances use AWS spare capacity and can be interrupted with a 2-minute notice. Discounts of 60--90% are common. The trade-off is interruption risk. Spot works well for: CI/CD build runners, batch data processing, ML training jobs, and stateless worker pools that can be interrupted and restarted without data loss. It is a poor fit for databases, stateful services, or anything where a 2-minute interruption would cause user-visible failures.

Infracost in CI/CD

Infracost estimates the monthly cost of a Terraform plan before terraform apply is run. This integrates cleanly into pull request workflows:

# .github/workflows/infracost.yml
name: Infracost

on:
  pull_request:

jobs:
  infracost:
    name: Cost Estimate
    runs-on: ubuntu-latest

    steps:
      - uses: actions/checkout@v4

      - name: Setup Infracost
        uses: infracost/actions/setup@v3
        with:
          api-key: ${{ secrets.INFRACOST_API_KEY }}

      - name: Generate Infracost diff
        run: |
          infracost diff \
            --path=terraform/ \
            --format=json \
            --out-file=/tmp/infracost.json

      - name: Post Infracost comment
        uses: infracost/actions/comment@v1
        with:
          path: /tmp/infracost.json
          behavior: update

This posts a comment on the pull request showing the cost difference between the base branch and the proposed change -- for example, "This change will increase monthly costs by $1,240 -- adding 3 x m5.xlarge instances." Engineers see the cost impact before the code merges, not after the bill arrives.

OpenCost and Unit Economics

OpenCost is an open-source project (CNCF sandbox) that provides Kubernetes-native cost monitoring. It allocates cloud costs at the pod, namespace, and deployment level -- answering questions like "what does running the recommendations service cost per month?" rather than just "what does the cluster cost?"

Unit economics -- cost per user, cost per API request, cost per transaction -- is the end goal of FinOps work. When you know that processing a payment costs $0.003 in infrastructure, you can make informed decisions about caching, batch processing, and service architecture that would otherwise be purely technical.

Frequently Asked Questions