What is Containerization? Docker and Containers Explained

Before containers, deploying software meant fighting with environment differences. Works on my machine, fails in staging, crashes in production -- the root cause was almost always a mismatch between the environment the code was developed in and the environment it ran in. Containers solved this by making the environment part of the deployable artefact.

How Containers Work

A container is a process running in an isolated environment on the host operating system. The isolation comes from two Linux kernel features:

Namespaces control what the process can see: its own filesystem view, network interfaces, process tree, and hostname. A container's process cannot see the host filesystem or other containers' processes unless explicitly permitted.

Control groups (cgroups) control what the process can use: CPU time, memory, network bandwidth, and disk I/O. You can limit a container to 512MB of RAM and 0.5 CPU cores and the kernel enforces it.

This is fundamentally different from a virtual machine. A VM emulates hardware and runs a full guest OS kernel. A container shares the host kernel and just wraps the process in isolation. The practical difference: containers start in milliseconds, use tens of megabytes, and can run hundreds per host. VMs start in seconds to minutes, use gigabytes, and run tens per host.

Docker: The Container Toolchain

Docker popularised containers in 2013 by making them accessible. Before Docker, using Linux namespaces and cgroups directly required low-level system programming. Docker wrapped them in a usable CLI and introduced the container image format.

A Dockerfile defines how to build an image -- a read-only template for containers. Here is a minimal example for a Python API:

FROM python:3.12-slim

WORKDIR /app

COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY . .

EXPOSE 8000
CMD ["uvicorn", "main:app", "--host", "0.0.0.0", "--port", "8000"]

Each instruction creates a layer. Docker caches layers -- if requirements.txt has not changed, the pip install step does not rerun. This makes builds fast in practice.

# Build the image
docker build -t payments-api:v1.2.0 .

# Run a container from the image
docker run -p 8000:8000 payments-api:v1.2.0

# Push to a registry
docker push registry.example.com/payments-api:v1.2.0

The image pushed to the registry is exactly what runs in production. No installation scripts, no "make sure Python 3.12 is installed" -- the environment is encoded in the image.

Container Images and Registries

A container image is a layered, read-only filesystem snapshot. Images are stored in registries -- Docker Hub for public images, AWS ECR, Google Artifact Registry, or GitHub Container Registry for private ones.

The CI/CD guide covers how container images fit into a delivery pipeline: on every merge to main, CI builds the image, tags it with the commit SHA, runs tests inside the container, and pushes it to the registry. Deployment is then a matter of telling your orchestrator to run the new image tag.

Containers vs Virtual Machines: When to Use Each

Containers replaced VMs for many workloads but not all. The choice depends on your isolation requirements:

ContainersVirtual Machines
Startup timeMillisecondsSeconds to minutes
SizeMegabytesGigabytes
Kernel sharingYes (shared host kernel)No (each VM has own kernel)
Security isolationProcess-levelHardware-level
Best forMicroservices, apps, CI jobsMulti-tenant environments, different OS needs

Multi-tenant cloud infrastructure, where you are running workloads from different customers on shared hardware, typically still uses VMs for the stronger isolation boundary. Application workloads within a trusted environment are almost always better served by containers.

Container Orchestration

A single container on a single host is simple. Hundreds of containers across dozens of hosts, with rolling deployments, autoscaling, health checks, and service discovery, is not. That problem is container orchestration.

Kubernetes is the dominant orchestration platform. You declare what you want -- three replicas of the payments service, using image tag v1.2.0, with 256MB memory limit -- and Kubernetes figures out which nodes to schedule them on, restarts them when they crash, and routes traffic to healthy instances.

Kubernetes pods are the smallest deployable unit in Kubernetes. Understanding how pods work is the foundation for understanding the rest of the platform.

Containerization is also the foundation of infrastructure as code practices -- when your runtime environment is defined in code (a Dockerfile), it gets the same version control, code review, and automated testing as your application logic.

Frequently Asked Questions