AWS Builder Center

Upgrading EKS with Terraform: Control Plane, Nodes, Then Add-ons

How I move EKS clusters from 1.32 to 1.34 one minor version at a time through a GitLab pipeline: a new worker group on each AMI, old nodes cordoned and drained by hand, and the compatibility checks that mattered more than the control plane.

cloud engineer, indie hacker, devops

Why We Had to Do This

Two of the EKS clusters I work on were running Kubernetes 1.32.
EKS 1.32 left standard support on March 23, 2026. After that date a cluster moves into extended support, which costs $0.60 per cluster per hour instead of $0.10. The extra $0.50 an hour is about $365 a month per cluster. Extended support for 1.32 also ends on March 23, 2027, and then EKS upgrades the control plane for you whether you are ready or not.
So the target was 1.34, which has standard support until December 2, 2026. That is not far off, so this is not the last upgrade for these clusters. The same steps apply to 1.35, and writing them down once was part of the point.
The control plane part of an EKS upgrade is well documented. What I wanted to get right was everything around it: the order of the changes, how the worker nodes move without surprises, and how all of it goes through a pipeline instead of someone clicking in the console.
You cannot jump two versions on EKS. You go one minor version at a time, so the path was:
1
1.32 → 1.33 → 1.34
Each step is the same process, done twice:
  1. Control plane
  2. Worker nodes
  3. Add-ons
All of it goes through Terraform and a GitLab pipeline. Nothing is changed by hand in the console. The one manual part is cordoning and draining old nodes, and that is on purpose.

The Setup

A bit of context, because the setup decides how the upgrade works.
PieceHow it is set up
Control planeaws_eks_cluster in our own Terraform module, version set by cluster_version in tfvars
AWS add-onsaws_eks_addon resources driven by a cluster_addons map in the same tfvars, with pinned versions
WorkersSelf-managed EC2 Auto Scaling groups with launch templates, AL2023 EKS-optimized AMIs pinned by ID
Other add-onsHelm jobs in the pipeline (node termination handler, External Secrets, monitoring and security agents)
API endpointPrivate only, so kubectl runs from a bastion in the same VPC
PipelineGitLab, dev branch deploys non-production, main deploys production
The worker setup matters most. These are not EKS managed node groups, so there is no "update node group version" button. A node gets its Kubernetes version from the AMI in its launch template. Changing the version means changing the AMI.
The cluster and its add-ons live in the same Terraform state, and the workers are in a separate one. That split ends up shaping the order of the MRs.

The Whole Flow

This is one version hop in one environment.
One hop, start to finish. Non-production first, then the same steps in production.
The order is control plane, then nodes, then add-ons. That matches the order in the EKS user guide, and it fits how our add-ons are split. Some of them, like CoreDNS, the EBS CSI driver and metrics-server, run as pods on the workers, so I want the new workers in place before I change them.

Before Touching Anything

Run the upgrade insights for the target version

EKS checks the cluster for things that will break on the next version, like deprecated API calls. I filter to the version I am moving to:
1
2
3
aws eks list-insights \
--cluster-name my-cluster \
--filter categories=UPGRADE_READINESS,kubernetesVersions=1.33
Everything should be PASSING before the control plane MR.
One catch: you can only check 1.34 readiness once the cluster is on 1.33. So the pre-checks are not a one-time thing. I ran them again at the start of the second hop.

Check every add-on against both versions

1
2
3
aws eks describe-addon-versions \
--kubernetes-version 1.34 \
--addon-name aws-efs-csi-driver
I did this for every add-on, for both 1.33 and 1.34, and put the results in a table in the MOP. Most add-ons had a version that worked across both. One did not: the EFS CSI driver on the second cluster was on v2.0.7, which is not listed for 1.34. The plan for that cluster is to move it during the 1.33 step, to a version that is listed for both 1.33 and 1.34, so it is not blocking the second hop.

Check the tools, not just the cluster

The pipeline's Helm jobs used an alpine/helm:3.14.0 image. Helm supports a limited range of Kubernetes versions, and 3.14 is too old for 1.34. That needs its own MR, pinned to one exact tested 3.19 or 3.20 patch version, before the Helm jobs run against an upgraded cluster.
The same goes for kubectl on the bastion. It should be within one minor version of the cluster.

Look at PodDisruptionBudgets and unhealthy pods

1
2
kubectl get pdb -A
kubectl get pods -A | grep -v Running | grep -v Completed
During the review, several application PDBs allowed zero disruptions, and a few workloads were already unhealthy before anything changed. Neither is something I can fix myself, because the applications belong to the vendor team. So both went on a watchlist and were reviewed with them before the worker phase.
I am glad this happened before the drain and not during it. A PDB that allows zero disruptions makes kubectl drain wait until it times out.

Save a baseline

From the bastion:
1
2
3
4
5
6
aws eks update-kubeconfig --name my-cluster
kubectl get nodes -o wide
kubectl get deploy,statefulset,daemonset -A
kubectl get pdb -A
aws eks list-addons --cluster-name my-cluster
helm list -A
After every phase I run the same commands and compare. It is much easier to spot "this was already broken" versus "this broke just now" when you have the before picture.

How Changes Go Through the Pipeline

Every phase takes the same route.
Non-production on dev, production on main. Nothing applies without someone clicking it.
A few details that matter here:
  • Pipelines only run on dev and main. A feature branch does not deploy anything. The plan shows up after the MR is merged into dev.
  • Apply is manual and uses the saved plan. The plan job saves tfplan as an artifact, and the apply job runs terraform apply tfplan. What gets applied is exactly what I reviewed.
  • The plan expires after one hour. If I take longer than that, or the branch changes, I run a new plan and read it again.
  • Jobs are gated by variables. Each group of jobs only appears when a flag like RUN_EKS_JOBS is true.
That last one caught me. The worker jobs for non-production were gated by a different flag, which was set to false, so after merging the worker MR the plan job did not appear at all. It took a small CI MR to put the worker jobs behind the same flag as the cluster jobs.
The plan review itself is simple. For the control plane MR I expect to see one change: the cluster version. For the worker MR I expect new resources only. If the plan shows a delete or a replacement I did not expect, I stop.

Phase 1: Control Plane

The change is one line in the environment tfvars:
1
cluster_version = "1.33" # was 1.32
In my first version of this MR, I changed the non-production and production tfvars together. Then I took the production change back out, so production stays on 1.32 until non-production has finished the full hop. Since main deploys production, anything sitting in the production tfvars goes out with the next dev → main promotion, even if that promotion was meant for something else.
After merging, the dev pipeline plans the cluster deployment. The plan shows the version change on aws_eks_cluster and nothing else. Then the manual apply.
During the update, EKS replaces the API server instances in a rolling way, so the API stays reachable. You cannot pause or cancel the update once it starts. Workloads keep running, because the workers have not changed yet.
After the apply:
1
2
3
4
aws eks describe-cluster --name my-cluster \
--query 'cluster.{status:status,version:version,health:health.issues}'

kubectl get nodes
The cluster should be ACTIVE on 1.33, and the nodes should still show 1.32. That mixed state is allowed. The kubelet can be up to three minor versions behind the API server, although you do not want to stay that way.

If something goes wrong

EKS can roll the control plane back one minor version within 7 days of the upgrade. That is the first step of the rollback plan in the MOP:
  1. Stop. Do not continue to workers, add-ons, production, or the next version.
  2. Roll the control plane back to the previous minor version, then reconcile Terraform straight away.
  3. For workers, stop the migration and keep the old group until workloads pass on the new one.
  4. For add-ons, only go back to versions AWS lists for the active control plane version.
We did not need it. The 7-day window is still a good reason to let each step sit before moving to the next one.

Phase 2: Worker Nodes

This is where most of the time goes.

Add a new group instead of changing the old one

Each worker group in tfvars has its own AMI. To move workers to 1.33, I added a new group next to the old one, with the same settings and the 1.33 AMI:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
self_managed_node_groups = {
worker-az1 = {
ami_id = "ami-xxxxxxxx" # AL2023 EKS 1.32, left unchanged
instance_type = "c6i.4xlarge"
min_size = 6
max_size = 7
desired_size = 6
subnet_ids = ["subnet-xxxxxxxx"]
# ...
}

worker-az1-133 = {
ami_id = "ami-yyyyyyyy" # AL2023 EKS 1.33
instance_type = "c6i.4xlarge"
min_size = 6
max_size = 7
desired_size = 6
subnet_ids = ["subnet-xxxxxxxx"] # same subnet as the old group
# ...
}
}
The new group copies the old one: instance type, size, subnet, labels, and volume settings. Only the AMI changes. Keeping the same subnet also keeps the new nodes in the same Availability Zone, which matters for anything with an EBS volume.
To get the AMI ID for the target version, I query the public SSM parameter:
1
2
3
aws ssm get-parameter \
--name "/aws/service/eks/optimized-ami/1.33/amazon-linux-2023/x86_64/standard/recommended" \
--query Parameter.Value --output text | jq -r '.image_id'

Why not just change the AMI on the old group

This is the part I would point out to anyone with a similar setup.
Our Auto Scaling groups have an instance_refresh block with a rolling strategy and min_healthy_percentage = 50. If I change the AMI on the existing group, the launch template changes, and the Auto Scaling group starts replacing instances on its own, up to half of them at a time. It does not wait for me, and it does not check the vendor's applications between nodes.
Adding a new group avoids that completely. Terraform only creates things. The old nodes stay exactly as they are until I move the pods off them.
The plan for this MR should show only new resources: a launch template and an Auto Scaling group. Nothing on the old group.

Cordon and drain, one node at a time

Once the new nodes are Ready, I move workloads over from the bastion.
One old node at a time. Every drain waits for a check before the next one.
See what is running on the node first:
1
kubectl get pods -A -o wide --field-selector spec.nodeName=<old-node>
Cordon it, so nothing new gets scheduled there:
1
kubectl cordon <old-node>
Drain it:
1
2
3
4
kubectl drain <old-node> \
--ignore-daemonsets \
--delete-emptydir-data \
--timeout=10m
Drain uses the eviction API, so it respects PDBs. If a PDB blocks an eviction, drain keeps retrying until the timeout. That is what you want. It is better for the drain to stop than for it to remove the last replica of something.
Then check before touching the next node:
1
2
kubectl get pods -A -o wide | grep -v Running | grep -v Completed
kubectl get nodes
And the vendor team runs their application checks.
The MOP has a rule for this: show one controlled drain in non-production before anyone does it in production. I like that rule. The first drain is where you find out about the PDB you missed.

Stateful pods

Some of the workloads are stateful, like databases with EBS volumes. These need more care when a node drains, because the pod has to stop, the volume has to detach, and the pod can only start again on a node in the same Availability Zone as the volume.
Things to check for these:
  • A new node in the same AZ. The new group uses the same subnet as the old one, so this was covered. If your groups span different AZs, check before draining.
  • PDBs that allow zero disruptions. A single-replica StatefulSet with minAvailable: 1 can never be evicted.
  • Shutdown time. The default terminationGracePeriodSeconds is 30 seconds. A database that needs longer gets killed.
  • Single instances. If there is only one pod, it is unavailable while it moves. Plan that with the application team rather than calling it zero downtime.

Remove the old group later, in its own MR

After the drains, the old group is still there with cordoned, empty nodes. I leave it while workloads run on the new nodes, in case I need to move things back. Removing it is a separate MR that I raise only after the environment checks pass.
Keep in mind that until that MR goes in, you are paying for both groups.

Phase 3: Add-ons

The AWS add-ons are a map in the cluster tfvars. Each one has a pinned version:
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
cluster_addons = {
vpc-cni = {
addon_version = "v1.22.4-eksbuild.3"
resolve_conflicts = "OVERWRITE"
# configuration_values stays exactly as it was
}
kube-proxy = {
addon_version = "v1.33.10-eksbuild.29"
resolve_conflicts = "OVERWRITE"
}
coredns = {
addon_version = "v1.12.4-eksbuild.38"
resolve_conflicts = "OVERWRITE"
}
# ...
}
The MR changes only the addon_version values. I do not rewrite the map. Some add-ons carry configuration_values (our VPC CNI has custom networking and prefix delegation settings in there), and some have IAM roles. If any of that changes by accident, the plan shows it, and that is a reason to stop.
I pin exact versions instead of using "latest". With pinned versions, an add-on only changes when an MR says so.
These are the versions I moved through in non-production:
Add-on1.321.33 step1.34 step
VPC CNIv1.20.5-eksbuild.1v1.22.4-eksbuild.3no change
kube-proxyv1.32.13-eksbuild.5v1.33.10-eksbuild.29v1.34.6-eksbuild.29
CoreDNSv1.11.4-eksbuild.33v1.12.4-eksbuild.38no change
EBS CSI driverv1.59.0-eksbuild.1v1.66.0-eksbuild.1no change
metrics-serverv0.8.0-eksbuild.1v0.8.1-eksbuild.20v0.9.0-eksbuild.11
Pod Identity Agentv1.3.10-eksbuild.3no changeno change
Do not copy these. Query the versions for your own cluster when you do this. kube-proxy is the one that follows the Kubernetes minor version closely, so it changes on every hop.
After the apply, every add-on should be ACTIVE:
1
2
3
aws eks list-addons --cluster-name my-cluster
aws eks describe-addon --cluster-name my-cluster --addon-name coredns \
--query 'addon.{version:addonVersion,status:status}'
The Helm-based add-ons are separate pipeline jobs. Node termination handler, External Secrets, and the monitoring and security agents each have their own compatibility to check. They are not covered by describe-addon-versions.

Validating the Environment

After all three phases, from the bastion:
1
2
3
4
5
6
7
kubectl get nodes -o wide
kubectl get pods -A -o wide
kubectl get deploy,statefulset,daemonset -A
kubectl get pdb -A
kubectl get events -A --sort-by='.lastTimestamp'
aws eks list-addons --cluster-name my-cluster
helm list -A
What I check against the baseline:
  • every node Ready, all on the new version
  • no new failed or pending pods
  • storage, load balancers, DNS, external secrets, security agents and monitoring working
  • the vendor's application tests passing
Then one more step: run the cluster and worker plan jobs again with no changes. Both should come back with nothing to do. If a plan shows changes nobody expected, something drifted, and I want to know before production.

The Second Hop

The second hop, 1.33 → 1.34, follows the same process: control plane, a new worker-az1-134 group on the 1.34 AMI, the drains, then the add-ons. The insights for 1.34 can only be checked at this point, now that the cluster is on 1.33.
The add-on diff was smaller this time. Only kube-proxy and metrics-server needed new versions.
Once the workloads were running on the 1.34 workers, a separate MR removed the old worker groups.
If your workers span more than one Availability Zone, migrate one AZ at a time and check between them. It keeps the blast radius to one zone if something goes wrong.

What I Would Do Differently

Keep production out of the first MR. With a dev → main flow, production values in the tfvars ride along with the next promotion. I would only touch the production tfvars once non-production has finished the whole hop.
Check that the jobs you need actually appear. Before merging, I would confirm the plan and apply jobs show up for the target branch with the current flags. Finding out after the merge costs an extra CI MR.
Update the tools in CI first. The Helm image, kubectl on the bastion, and anything else that talks to the cluster. These are easy to forget because they are not part of the cluster.
Treat the add-on review as a two-version check. Checking add-ons against both the next version and the final target is what catches something like the EFS CSI driver early, instead of in the middle of the second hop.

My Take

The control plane is the shortest part of an EKS upgrade. It is one line of Terraform, and EKS does the work.
The workers take the most time, and how long they take depends on how your worker groups are built. With self-managed Auto Scaling groups, adding a new group per version and draining the old nodes by hand is slower than letting an instance refresh roll everything. I still prefer it. Each node moves when I decide, and I can stop after any one of them.
The other thing I would repeat is checking compatibility well beyond the AWS add-ons. The control plane was fine. The things that needed attention were a storage driver version, a Helm image in CI, a pipeline flag, and PDBs that allowed zero disruptions.

References

  • EKS Kubernetes version lifecycle and support dates: https://docs.aws.amazon.com/eks/latest/userguide/kubernetes-versions.html
  • Update an existing cluster to a new Kubernetes version: https://docs.aws.amazon.com/eks/latest/userguide/update-cluster.html
  • Roll back a cluster to a previous Kubernetes version: https://docs.aws.amazon.com/eks/latest/userguide/rollback-cluster.html
  • Update self-managed nodes: https://docs.aws.amazon.com/eks/latest/userguide/update-workers.html
  • Migrate to a new node group: https://docs.aws.amazon.com/eks/latest/userguide/migrate-stack.html
  • Add-on compatibility: https://docs.aws.amazon.com/eks/latest/userguide/addon-compat.html
  • Cluster insights: https://docs.aws.amazon.com/eks/latest/userguide/cluster-insights.html
  • Retrieve recommended Amazon Linux AMI IDs: https://docs.aws.amazon.com/eks/latest/userguide/retrieve-ami-id.html
  • EKS pricing: https://aws.amazon.com/eks/pricing/
  • Kubernetes version skew policy: https://kubernetes.io/releases/version-skew-policy/
  • Safely drain a node: https://kubernetes.io/docs/tasks/administer-cluster/safely-drain-node/
  • Helm version support policy: https://helm.sh/docs/v3/topics/version_skew/
Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article