Amazon EKS Now Supports Kubernetes Version Rollback
Amazon EKS now lets you roll back a minor version upgrade to the previous version within 7 days, backed by automated Rollback Readiness Insights that check safety beforehand. This article covers what's in and out of scope, how the mechanism works under the hood, how to decide between rolling back and fixing forward, and a step-by-step walkthrough of the process.
Introduction
On July 1, 2026, AWS announced Kubernetes version rollback for Amazon EKS. The headline capability is straightforward: after performing a minor version upgrade, you can now revert your cluster's control plane back to the previous version within a limited window, instead of the upgrade being a one-way operation.
If you operate EKS clusters, you may have run into a situation where, after an upgrade, something unexpected broke — an application incompatibility, a deprecated API still in use, or a third-party tool that didn't behave as expected — and there was no way to go back. This update addresses that gap directly.
This article walks through what changed, when this feature is useful, how it works under the hood, and how to decide between rolling back and fixing forward, followed by a hands-on walkthrough. It's aimed at infrastructure and platform engineers running EKS, as well as anyone shaping their organization's upgrade strategy.
What Changed
Until now, an Amazon EKS minor version upgrade was a one-way operation — once you upgraded, there was no way back. With this update, you can now roll back the control plane to the previous version within 7 days of completing an upgrade. The version you roll back to isn't an emulated or reconstructed state — it's the exact, previously validated version that was running in production.
Rollback Scope
The table below shows what is, and isn't, reverted during a rollback.
| Category | Scope |
|---|---|
| Rolled back | The Kubernetes API server, control plane components, the EKS platform version, and (for EKS Auto Mode clusters) worker nodes |
| Not rolled back | etcd data, workloads (Pods, Deployments, etc.), persistent volumes, EKS add-ons, Managed Node Groups, and self-managed or hybrid nodes |
etcd data and running workloads are preserved throughout. Rollback only affects the Kubernetes control plane — it doesn't touch your cluster state or application data.
Rollback Constraints
A few conditions apply before you can roll back:
| Constraint | Details |
|---|---|
| Time window | The rollback must be initiated within 7 days of the upgrade completing |
| Target version | You can only roll back one minor version (N to N-1) — not two or more versions back, and not to the version the cluster was originally created at |
| Cluster status | The cluster must be in ACTIVE status, with no other update in progress |
| Nodes | Managed Node Groups, self-managed nodes, and Fargate are out of scope and require separate handling |
Since Managed Node Group and self-managed node versions aren't reverted automatically, you'll need to handle the worker node side separately, either before or after rolling back the control plane.
Cost and Region Availability
This feature is available at no additional cost, in every AWS Region where Amazon EKS is offered. It works with existing clusters with no extra setup required.
Use Cases
Until now, there was no way to undo an EKS minor version upgrade. Problems like deprecated API usage, incompatibilities with custom controllers, or friction with third-party tooling often don't surface until a cluster is running real production traffic — something staging environments can't always catch. Version rollback acts as a safety net for exactly this kind of "you don't know until you try it" risk.
Before This Feature Existed
Without a native rollback path, teams typically relied on one of a few workarounds when something went wrong after an upgrade:
| Approach | Trade-offs |
|---|---|
| Blue/green cluster migration | Running old and new clusters side by side so you can switch back. Reliable, but requires standing up and maintaining a second cluster |
| Extensive staging validation | Testing thoroughly before touching production. Still can't always catch issues that only appear under production-scale load or traffic patterns |
| Rebuilding the cluster | A last resort: recreating the cluster on the older version. Costly in both downtime and effort |
Each of these approaches came with real infrastructure and operational overhead. With version rollback, a single API call is enough to revert the control plane to the previous version on your existing cluster.
Scenarios Where This Feature Helps
A few situations where this capability is particularly useful:
| Scenario | How it helps |
|---|---|
| Troubleshooting after a production upgrade | If something breaks right after an upgrade, you can revert to the previous version to limit impact while you investigate the root cause |
| Validating deprecated API or custom controller compatibility | You can try an upgrade in production knowing you have up to 7 days to roll back if something doesn't work |
| Simplifying operations for EKS Auto Mode users | On Auto Mode clusters, worker node rollback is handled automatically by EKS, so operators don't need to manage node reverts separately |
Beyond helping with production incidents, this also lowers the cost of delaying upgrades in the first place. Knowing you can roll back makes it easier to spend less time running versions with known CVEs, which in turn helps meet compliance requirements that call for running actively patched software.
The next section covers how this safety net actually works.
How It Works
Amazon EKS surfaces whether a rollback can be performed safely through cluster insights, specifically a category called
ROLLBACK_READINESS. These insights are generated automatically after an upgrade and remain available throughout the 7-day rollback window, in both the console and the CLI.Rollback Readiness Insights
These checks cover API compatibility (including field-level changes introduced between versions), cluster health, kubelet and kube-proxy version skew, and EKS add-on version compatibility. For EKS Auto Mode clusters, EKS also evaluates NodePool disruption budgets, do-not-disrupt annotations, and PodDisruptionBudget configuration.
Each check reports one of four statuses:
| Status | Meaning | Effect on rollback |
|---|---|---|
| PASSING | No issues detected | Rollback allowed |
| WARNING | A potential issue was detected | Rollback allowed (advisory only) |
| ERROR | A blocking issue was detected | Rollback blocked unless resolved or bypassed with --force |
| UNKNOWN | Status couldn't be determined | Rollback blocked unless resolved or bypassed with --force |
Insights are also re-evaluated at the moment a rollback is triggered. If the existing insight data is stale, EKS refreshes it automatically before proceeding.
--force Flag Behavior
If ERROR-status insights are present, you can pass the
--force flag to bypass all insight checks (ERROR, WARNING, and UNKNOWN) and proceed with the rollback anyway.That said,
--force only bypasses insight checks. It has no effect on other prerequisites — the 7-day window, the check for whether the cluster was created at its current version, or the check that blocks rolling back more than one minor version. On EKS Auto Mode clusters, --force also doesn't override disruption controls: NodePool disruption budgets, PodDisruptionBudgets, and do-not-disrupt annotations are still honored.EKS Auto Mode Rollback Flow
On clusters running EKS Auto Mode, worker node rollback is handled automatically through a Karpenter-based mechanism. The overall sequence looks like this:
| Phase | Cluster status | What's happening |
|---|---|---|
| Node rollback in progress | ACTIVE | Karpenter replaces existing nodes with ones running the previous version's AMI, while the control plane keeps serving traffic on its current version |
| Control plane rollback | UPDATING | The API server and control plane components are reverted to the previous version |
| Rollback complete | ACTIVE | The entire cluster is now on the previous version |
Here's what that sequence looks like as a flow:

Node rollback happens first, and the control plane rollback only proceeds once every node falls within the version skew policy for the target version. Kubernetes' version skew policy allows worker nodes to run up to three minor versions behind the kube-apiserver, so this intermediate state is fully supported.
Node rollback has a default timeout of 720 minutes (12 hours), configurable between 120 minutes (2 hours) and 10,080 minutes (7 days) via the
timeoutMinutes field in rollbackConfig. If the timeout is reached, nodes drift back to the current version, the control plane rollback never starts, and the update is marked as failed.Before node rollback finishes, you can also cancel the operation with the
CancelUpdate API. Cancelling reverts nodes to the current version and skips the control plane rollback entirely. As long as you're still within the 7-day window, you can retry the rollback afterward.Decision Framework
Before this feature existed, the only real option when something broke after an upgrade was to fix forward. Now that rolling back is on the table too, the right call depends on the situation.
Key Considerations
| Factor | Favors rollback | Favors fixing forward |
|---|---|---|
| Scope of impact | Cluster-wide or affecting multiple workloads | Limited to a specific Pod or feature |
| Time to root cause | Root cause is unclear and likely to take time to isolate | Root cause is clear and the fix is quick |
| Remaining window | Close to the 7-day deadline, where waiting risks losing the option entirely | Plenty of time left to keep investigating |
| Add-on compatibility | Add-on downgrades are already planned for | Many add-ons are incompatible with the previous version, making downgrades costly |
Data Is Not Rolled Back
Rollback only reverts the control plane version — etcd data and workload state remain exactly as they were. If new resources were created using new APIs or fields between the upgrade and the rollback, those resources persist after the rollback completes. Resources created while bypassing insight errors with
--force aren't automatically cleaned up either. When deciding whether to roll back, it's worth considering how much has been written to the cluster in that window.It's also worth remembering that insight evaluations are a point-in-time snapshot. If you make changes to the cluster after checking insights but before the rollback runs, those changes won't be reflected until insights are re-evaluated at rollback time.
Shared Responsibility
EKS is responsible for safely reverting the control plane. Verifying that your applications, custom controllers, and third-party tools work correctly on the previous version is on you. If you choose to roll back, that decision should be made on the assumption that your workloads are compatible with the version you're rolling back to.
One more thing worth factoring in: rolling back from a version under Standard Support to one under Extended Support restarts Extended Support charges from that point forward.
Hands-On: Performing a Rollback
This section walks through the actual rollback process with CLI examples. It breaks down into four steps: reviewing insights, preparing worker nodes and add-ons, rolling back the control plane, and monitoring progress.
Reviewing Insights
Start by checking for any issues under the
ROLLBACK_READINESS category. In the console, this is available on the cluster's Upgrade insights tab. From the CLI:1
2
3
4
aws eks list-insights \
--cluster-name <cluster-name> \
--region <region> \
--filter '{"categories": ["ROLLBACK_READINESS"]}'
To get details on a specific insight:
1
2
3
4
aws eks describe-insight \
--cluster-name <cluster-name> \
--region <region> \
--id <insight-id>
If there are ERROR or UNKNOWN insights, review and address them. If that's not practical, you can proceed with
--force, but keep in mind that EKS can't guarantee the rollback is safe once insight checks are bypassed.Preparing Worker Nodes
Before rolling back the control plane, check what needs to happen on the worker node side:
| Node type | Required action |
|---|---|
| EKS Auto Mode | None — EKS rolls back nodes automatically as part of the rollback |
| Managed Node Groups | Update to the previous version ahead of time with update-nodegroup-version |
| Self-managed or hybrid nodes | Update AMIs or configuration to the previous version yourself |
| Fargate | Not supported for rollback. Fargate Pods still running the current version will trigger an ERROR-status kubelet version skew insight |
To update a managed node group:
1
2
3
4
5
aws eks update-nodegroup-version \
--cluster-name <cluster-name> \
--nodegroup-name <nodegroup-name> \
--kubernetes-version <kubernetes-version> \
--region <region>
If you're using Fargate, Pods still running on the same version as the control plane will trigger an ERROR insight. The standard approach is to delete those Pods before the rollback and let them relaunch on the previous version afterward. You can bypass this check with
--force, but doing so means those Fargate workloads keep running with a kubelet version skew violation until they're replaced, which can lead to unexpected behavior in the meantime.Checking Add-ons
EKS doesn't automatically roll back add-on versions along with the control plane. Use insights to check add-on compatibility with the target version, and downgrade any that aren't compatible beforehand:
1
2
3
4
5
aws eks update-addon \
--cluster-name <cluster-name> \
--addon-name <addon-name> \
--addon-version <addon-version> \
--region <region>
Note that insights only check EKS-managed add-on versions. If you're running self-managed add-ons, or have overridden a managed add-on's version outside the normal lifecycle, you'll need to verify compatibility with the previous version yourself.
Rolling Back the Control Plane
Once everything's ready, initiate the rollback. In the console, go to the cluster's Actions menu and choose Rollback cluster version, then review and confirm.
From the CLI, use the existing
update-cluster-version command with the previous (N-1) version:1
2
3
4
aws eks update-cluster-version \
--name <cluster-name> \
--kubernetes-version <kubernetes-version> \
--region <region>
A control plane rollback typically takes about as long as a standard upgrade — around 20 minutes.
Monitoring Progress
Track progress with
describe-update:1
2
3
4
aws eks describe-update \
--name <cluster-name> \
--region <region> \
--update-id <update-id>
In the console, check the status of the relevant update ID on the cluster's Update history tab. Status transitions from
InProgress to either Successful or Failed. For Auto Mode clusters, the cluster remains ACTIVE while nodes are rolling back, and only switches to UPDATING once the control plane rollback begins.Summary
With this update, Amazon EKS minor version upgrades can now be reverted to the previous version within 7 days of completing the upgrade. What you roll back to isn't a reconstructed approximation — it's the exact version that was previously running in production. The feature is available at no additional cost in every AWS Region, and for EKS Auto Mode clusters, worker node rollback is handled automatically.
That said, only the control plane is rolled back automatically. Managed Node Groups, self-managed nodes, Fargate, and add-ons all require separate handling. This isn't a universal undo button — it's a safety net with a defined scope and a defined time limit.
Adding a "roll back" option alongside "fix forward" changes how production upgrades can be approached. Checking what insights show on an existing cluster is the natural first step toward putting this feature to use.
References
- Amazon EKS now supports Kubernetes version rollback (official announcement)
- Announcing Amazon EKS Rollback for safe and reliable management of cluster upgrades (AWS Containers Blog)
- Upgrade Amazon EKS clusters with confidence using Kubernetes version rollbacks (AWS News Blog)
- Rollback cluster to previous Kubernetes version
- Rollback EKS Auto Mode clusters
- Prepare for Kubernetes version upgrades and troubleshoot misconfigurations with cluster insights
- Update existing cluster to new Kubernetes version
- Best Practices for Cluster Upgrades
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article