AWS Builder Center

Amazon EKS Now Supports Kubernetes Version Rollback

Amazon EKS now lets you roll back a minor version upgrade to the previous version within 7 days, backed by automated Rollback Readiness Insights that check safety beforehand. This article covers what's in and out of scope, how the mechanism works under the hood, how to decide between rolling back and fixing forward, and a step-by-step walkthrough of the process.

Introduction

On July 1, 2026, AWS announced Kubernetes version rollback for Amazon EKS. The headline capability is straightforward: after performing a minor version upgrade, you can now revert your cluster's control plane back to the previous version within a limited window, instead of the upgrade being a one-way operation.
If you operate EKS clusters, you may have run into a situation where, after an upgrade, something unexpected broke — an application incompatibility, a deprecated API still in use, or a third-party tool that didn't behave as expected — and there was no way to go back. This update addresses that gap directly.
This article walks through what changed, when this feature is useful, how it works under the hood, and how to decide between rolling back and fixing forward, followed by a hands-on walkthrough. It's aimed at infrastructure and platform engineers running EKS, as well as anyone shaping their organization's upgrade strategy.

What Changed

Until now, an Amazon EKS minor version upgrade was a one-way operation — once you upgraded, there was no way back. With this update, you can now roll back the control plane to the previous version within 7 days of completing an upgrade. The version you roll back to isn't an emulated or reconstructed state — it's the exact, previously validated version that was running in production.

Rollback Scope

The table below shows what is, and isn't, reverted during a rollback.
CategoryScope
Rolled backThe Kubernetes API server, control plane components, the EKS platform version, and (for EKS Auto Mode clusters) worker nodes
Not rolled backetcd data, workloads (Pods, Deployments, etc.), persistent volumes, EKS add-ons, Managed Node Groups, and self-managed or hybrid nodes
etcd data and running workloads are preserved throughout. Rollback only affects the Kubernetes control plane — it doesn't touch your cluster state or application data.

Rollback Constraints

A few conditions apply before you can roll back:
ConstraintDetails
Time windowThe rollback must be initiated within 7 days of the upgrade completing
Target versionYou can only roll back one minor version (N to N-1) — not two or more versions back, and not to the version the cluster was originally created at
Cluster statusThe cluster must be in ACTIVE status, with no other update in progress
NodesManaged Node Groups, self-managed nodes, and Fargate are out of scope and require separate handling
Since Managed Node Group and self-managed node versions aren't reverted automatically, you'll need to handle the worker node side separately, either before or after rolling back the control plane.

Cost and Region Availability

This feature is available at no additional cost, in every AWS Region where Amazon EKS is offered. It works with existing clusters with no extra setup required.

Use Cases

Until now, there was no way to undo an EKS minor version upgrade. Problems like deprecated API usage, incompatibilities with custom controllers, or friction with third-party tooling often don't surface until a cluster is running real production traffic — something staging environments can't always catch. Version rollback acts as a safety net for exactly this kind of "you don't know until you try it" risk.

Before This Feature Existed

Without a native rollback path, teams typically relied on one of a few workarounds when something went wrong after an upgrade:
ApproachTrade-offs
Blue/green cluster migrationRunning old and new clusters side by side so you can switch back. Reliable, but requires standing up and maintaining a second cluster
Extensive staging validationTesting thoroughly before touching production. Still can't always catch issues that only appear under production-scale load or traffic patterns
Rebuilding the clusterA last resort: recreating the cluster on the older version. Costly in both downtime and effort
Each of these approaches came with real infrastructure and operational overhead. With version rollback, a single API call is enough to revert the control plane to the previous version on your existing cluster.

Scenarios Where This Feature Helps

A few situations where this capability is particularly useful:
ScenarioHow it helps
Troubleshooting after a production upgradeIf something breaks right after an upgrade, you can revert to the previous version to limit impact while you investigate the root cause
Validating deprecated API or custom controller compatibilityYou can try an upgrade in production knowing you have up to 7 days to roll back if something doesn't work
Simplifying operations for EKS Auto Mode usersOn Auto Mode clusters, worker node rollback is handled automatically by EKS, so operators don't need to manage node reverts separately
Beyond helping with production incidents, this also lowers the cost of delaying upgrades in the first place. Knowing you can roll back makes it easier to spend less time running versions with known CVEs, which in turn helps meet compliance requirements that call for running actively patched software.
The next section covers how this safety net actually works.

How It Works

Amazon EKS surfaces whether a rollback can be performed safely through cluster insights, specifically a category called ROLLBACK_READINESS. These insights are generated automatically after an upgrade and remain available throughout the 7-day rollback window, in both the console and the CLI.

Rollback Readiness Insights

These checks cover API compatibility (including field-level changes introduced between versions), cluster health, kubelet and kube-proxy version skew, and EKS add-on version compatibility. For EKS Auto Mode clusters, EKS also evaluates NodePool disruption budgets, do-not-disrupt annotations, and PodDisruptionBudget configuration.
Each check reports one of four statuses:
StatusMeaningEffect on rollback
PASSINGNo issues detectedRollback allowed
WARNINGA potential issue was detectedRollback allowed (advisory only)
ERRORA blocking issue was detectedRollback blocked unless resolved or bypassed with --force
UNKNOWNStatus couldn't be determinedRollback blocked unless resolved or bypassed with --force
Insights are also re-evaluated at the moment a rollback is triggered. If the existing insight data is stale, EKS refreshes it automatically before proceeding.

--force Flag Behavior

If ERROR-status insights are present, you can pass the --force flag to bypass all insight checks (ERROR, WARNING, and UNKNOWN) and proceed with the rollback anyway.
That said, --force only bypasses insight checks. It has no effect on other prerequisites — the 7-day window, the check for whether the cluster was created at its current version, or the check that blocks rolling back more than one minor version. On EKS Auto Mode clusters, --force also doesn't override disruption controls: NodePool disruption budgets, PodDisruptionBudgets, and do-not-disrupt annotations are still honored.

EKS Auto Mode Rollback Flow

On clusters running EKS Auto Mode, worker node rollback is handled automatically through a Karpenter-based mechanism. The overall sequence looks like this:
PhaseCluster statusWhat's happening
Node rollback in progressACTIVEKarpenter replaces existing nodes with ones running the previous version's AMI, while the control plane keeps serving traffic on its current version
Control plane rollbackUPDATINGThe API server and control plane components are reverted to the previous version
Rollback completeACTIVEThe entire cluster is now on the previous version
Here's what that sequence looks like as a flow:
Node rollback happens first, and the control plane rollback only proceeds once every node falls within the version skew policy for the target version. Kubernetes' version skew policy allows worker nodes to run up to three minor versions behind the kube-apiserver, so this intermediate state is fully supported.
Node rollback has a default timeout of 720 minutes (12 hours), configurable between 120 minutes (2 hours) and 10,080 minutes (7 days) via the timeoutMinutes field in rollbackConfig. If the timeout is reached, nodes drift back to the current version, the control plane rollback never starts, and the update is marked as failed.
Before node rollback finishes, you can also cancel the operation with the CancelUpdate API. Cancelling reverts nodes to the current version and skips the control plane rollback entirely. As long as you're still within the 7-day window, you can retry the rollback afterward.

Decision Framework

Before this feature existed, the only real option when something broke after an upgrade was to fix forward. Now that rolling back is on the table too, the right call depends on the situation.

Key Considerations

FactorFavors rollbackFavors fixing forward
Scope of impactCluster-wide or affecting multiple workloadsLimited to a specific Pod or feature
Time to root causeRoot cause is unclear and likely to take time to isolateRoot cause is clear and the fix is quick
Remaining windowClose to the 7-day deadline, where waiting risks losing the option entirelyPlenty of time left to keep investigating
Add-on compatibilityAdd-on downgrades are already planned forMany add-ons are incompatible with the previous version, making downgrades costly

Data Is Not Rolled Back

Rollback only reverts the control plane version — etcd data and workload state remain exactly as they were. If new resources were created using new APIs or fields between the upgrade and the rollback, those resources persist after the rollback completes. Resources created while bypassing insight errors with --force aren't automatically cleaned up either. When deciding whether to roll back, it's worth considering how much has been written to the cluster in that window.
It's also worth remembering that insight evaluations are a point-in-time snapshot. If you make changes to the cluster after checking insights but before the rollback runs, those changes won't be reflected until insights are re-evaluated at rollback time.

Shared Responsibility

EKS is responsible for safely reverting the control plane. Verifying that your applications, custom controllers, and third-party tools work correctly on the previous version is on you. If you choose to roll back, that decision should be made on the assumption that your workloads are compatible with the version you're rolling back to.
One more thing worth factoring in: rolling back from a version under Standard Support to one under Extended Support restarts Extended Support charges from that point forward.

Hands-On: Performing a Rollback

This section walks through the actual rollback process with CLI examples. It breaks down into four steps: reviewing insights, preparing worker nodes and add-ons, rolling back the control plane, and monitoring progress.

Reviewing Insights

Start by checking for any issues under the ROLLBACK_READINESS category. In the console, this is available on the cluster's Upgrade insights tab. From the CLI:
1
2
3
4
aws eks list-insights \
--cluster-name <cluster-name> \
--region <region> \
--filter '{"categories": ["ROLLBACK_READINESS"]}'
To get details on a specific insight:
1
2
3
4
aws eks describe-insight \
--cluster-name <cluster-name> \
--region <region> \
--id <insight-id>
If there are ERROR or UNKNOWN insights, review and address them. If that's not practical, you can proceed with --force, but keep in mind that EKS can't guarantee the rollback is safe once insight checks are bypassed.

Preparing Worker Nodes

Before rolling back the control plane, check what needs to happen on the worker node side:
Node typeRequired action
EKS Auto ModeNone — EKS rolls back nodes automatically as part of the rollback
Managed Node GroupsUpdate to the previous version ahead of time with update-nodegroup-version
Self-managed or hybrid nodesUpdate AMIs or configuration to the previous version yourself
FargateNot supported for rollback. Fargate Pods still running the current version will trigger an ERROR-status kubelet version skew insight
To update a managed node group:
1
2
3
4
5
aws eks update-nodegroup-version \
--cluster-name <cluster-name> \
--nodegroup-name <nodegroup-name> \
--kubernetes-version <kubernetes-version> \
--region <region>
If you're using Fargate, Pods still running on the same version as the control plane will trigger an ERROR insight. The standard approach is to delete those Pods before the rollback and let them relaunch on the previous version afterward. You can bypass this check with --force, but doing so means those Fargate workloads keep running with a kubelet version skew violation until they're replaced, which can lead to unexpected behavior in the meantime.

Checking Add-ons

EKS doesn't automatically roll back add-on versions along with the control plane. Use insights to check add-on compatibility with the target version, and downgrade any that aren't compatible beforehand:
1
2
3
4
5
aws eks update-addon \
--cluster-name <cluster-name> \
--addon-name <addon-name> \
--addon-version <addon-version> \
--region <region>
Note that insights only check EKS-managed add-on versions. If you're running self-managed add-ons, or have overridden a managed add-on's version outside the normal lifecycle, you'll need to verify compatibility with the previous version yourself.

Rolling Back the Control Plane

Once everything's ready, initiate the rollback. In the console, go to the cluster's Actions menu and choose Rollback cluster version, then review and confirm.
From the CLI, use the existing update-cluster-version command with the previous (N-1) version:
1
2
3
4
aws eks update-cluster-version \
--name <cluster-name> \
--kubernetes-version <kubernetes-version> \
--region <region>
A control plane rollback typically takes about as long as a standard upgrade — around 20 minutes.

Monitoring Progress

Track progress with describe-update:
1
2
3
4
aws eks describe-update \
--name <cluster-name> \
--region <region> \
--update-id <update-id>
In the console, check the status of the relevant update ID on the cluster's Update history tab. Status transitions from InProgress to either Successful or Failed. For Auto Mode clusters, the cluster remains ACTIVE while nodes are rolling back, and only switches to UPDATING once the control plane rollback begins.

Summary

With this update, Amazon EKS minor version upgrades can now be reverted to the previous version within 7 days of completing the upgrade. What you roll back to isn't a reconstructed approximation — it's the exact version that was previously running in production. The feature is available at no additional cost in every AWS Region, and for EKS Auto Mode clusters, worker node rollback is handled automatically.
That said, only the control plane is rolled back automatically. Managed Node Groups, self-managed nodes, Fargate, and add-ons all require separate handling. This isn't a universal undo button — it's a safety net with a defined scope and a defined time limit.
Adding a "roll back" option alongside "fix forward" changes how production upgrades can be approached. Checking what insights show on an existing cluster is the natural first step toward putting this feature to use.

References

Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article