A production-ready Kubernetes cluster on EKS with Terraform
What "production-ready" means for an EKS cluster, how I lay out the Terraform, and the checklist I go through before the first workload goes live.
Creating an EKS cluster takes one command. Running one that a team can trust with production traffic is a different job: networking, identity, add-ons, autoscaling, observability and an upgrade path all have to be decided up front, and all of it has to live in code. This is how I approach it.
What "production-ready" means
Before writing any Terraform I agree on a short definition with the team. A cluster is ready for production when:
- Everything is in code. The VPC, the cluster, node groups, add-ons and access rules are created by Terraform from a reviewed merge request. Nobody clicks in the console.
- It survives a zone failure. Nodes and load balancers span at least three availability zones, and workloads run more than one replica.
- Access is identity-based. People and pipelines get Kubernetes permissions through IAM, and pods get AWS permissions through their own roles, never through the node role.
- It scales on its own. Pods scale on load, nodes scale on pending pods, and there are limits so a bug cannot scale the bill.
- You can see what it does. Metrics, logs and alerts exist before the first release, not after the first incident.
- It can be upgraded. There is a tested procedure to move to the next Kubernetes version every few months.
Network first
The cluster lives in its own VPC with private subnets for nodes and public subnets only for load balancers. A few decisions here are expensive to change later:
- CIDR size. The VPC CNI gives every pod a VPC IP address. A /16 with /19 private subnets leaves room to grow; a /24 does not. If addresses are tight, prefix delegation or a secondary CIDR for pods are the usual fixes.
- Three availability zones, one private and one public subnet in each.
- NAT gateways. One per zone for resilience, or a single one to save cost in non-production environments. VPC endpoints for ECR, S3 and STS cut NAT traffic for image pulls and API calls.
- Subnet tags so the load balancer controller knows where to place internal and internet-facing load balancers.
Terraform layout
I keep the cluster in its own Terraform root module with remote state and locking, separate from the applications that run on it. Most of the heavy lifting is done by the well-maintained community modules for the VPC and EKS, pinned to a major version:
module "eks" {
source = "terraform-aws-modules/eks/aws"
version = "~> 21.0"
name = "platform-prod"
kubernetes_version = var.kubernetes_version
vpc_id = module.vpc.vpc_id
subnet_ids = module.vpc.private_subnets
endpoint_public_access = false # API reachable from the VPC / VPN only
enable_irsa = true
eks_managed_node_groups = {
system = {
instance_types = ["m7g.large"]
ami_type = "AL2023_ARM_64_STANDARD"
min_size = 3
max_size = 6
desired_size = 3
}
}
}
A few habits that save trouble later:
- Pin everything: provider versions, module versions and the Kubernetes version. Upgrades become deliberate merge requests instead of surprises.
- One state per environment, with the same code and different variables, so staging really rehearses production.
- Plan in CI, apply after review. The pipeline posts the plan on the merge request; a person reads it before anything changes.
Identity and access
Two separate questions: who can talk to the Kubernetes API, and what AWS resources a pod can reach.
- People and pipelines get access through EKS access entries mapped to IAM roles, with Kubernetes RBAC on top. Engineers assume a role through SSO; the deploy pipeline has its own narrow role.
- Pods get AWS permissions through IAM roles for service accounts (or EKS Pod Identity). Each service gets only what it needs, such as reading one S3 bucket or one secret.
- Nodes keep a minimal role, and IMDSv2 with a hop limit of 1 stops pods from borrowing the node's credentials.
Add-ons that every cluster needs
I manage these with Terraform or Helm, versioned like everything else:
- Core EKS add-ons: VPC CNI, CoreDNS, kube-proxy and the EBS CSI driver for persistent volumes.
- AWS Load Balancer Controller to expose services through ALBs and NLBs from Ingress and Service objects.
- External DNS and cert-manager, if the cluster owns DNS records and certificates.
- External Secrets to sync values from Secrets Manager or Parameter Store into Kubernetes, so secrets never sit in Git.
- metrics-server, which the horizontal pod autoscaler depends on.
Autoscaling and cost
- Pods: horizontal pod autoscalers on CPU, memory or request rate, with sensible minimum replicas and pod disruption budgets.
- Nodes: Karpenter (or the Cluster Autoscaler) adds capacity when pods are pending and removes empty nodes. Graviton instances and Spot capacity for stateless workloads usually cut compute cost noticeably.
- Guardrails: resource requests and limits on every workload, namespace quotas and a maximum node count, so a runaway deployment cannot grow without bounds.
Observability from day one
- Metrics: Prometheus scraping the cluster and the applications, Grafana dashboards per service, and alerts on symptoms users feel (error rate, latency, saturation) rather than on every CPU spike.
- Logs: container logs shipped to a central store with a retention policy.
- Control plane logs: API server and audit logs enabled in EKS, so you can answer "who changed this?".
Security baseline
- Private API endpoint, or a public one restricted to known CIDRs.
- Pod Security Standards enforced per namespace, with
restrictedas the default for application namespaces. - Network policies so namespaces cannot talk to each other unless they need to.
- Image scanning in the pipeline, and images pulled only from your own registry.
- Encryption of Kubernetes secrets with a KMS key.
Upgrades
EKS versions leave standard support after about 14 months, so upgrades are routine work, not a project. My usual order is: read the upgrade notes and check for deprecated APIs, upgrade staging, then the control plane in production, then add-ons, then node groups with a rolling replacement that respects pod disruption budgets. Because everything is in Terraform, the upgrade is a short merge request plus a checklist.
Checklist before the first workload
- VPC across three zones, private nodes, endpoints for ECR, S3 and STS.
- Cluster, node groups and add-ons created from Terraform in CI, state remote and locked.
- Access entries and RBAC for people and pipelines; IRSA or Pod Identity for workloads; IMDSv2 enforced.
- Load balancer controller, External Secrets and metrics-server installed and versioned.
- Pod and node autoscaling with limits; requests and limits on every workload.
- Prometheus, Grafana, logs and alerts working, with a test alert delivered.
- Pod Security Standards, network policies and secret encryption enabled.
- Upgrade procedure written down and rehearsed on staging.
If your team needs a cluster like this, or a review of one that already runs in production, I can help: I design and build it with you, and hand it over with documentation your team can maintain.