Kubernetes Version Upgrades on EKS Without Downtime: A Production Runbook

Most EKS outages I’ve seen during upgrades weren’t caused by Kubernetes. They were caused by a PodDisruptionBudget nobody remembered writing, a CNI add-on upgraded in the wrong order, or a Helm chart still using an API that was removed two versions ago. The upgrade itself is boring when the preparation is done. This post is the preparation, plus the upgrade, plus what to check afterward.

It assumes you have at least one production EKS cluster, manage it with Terraform, and would like the next upgrade to be an event nobody notices.

Why You Can’t Put This Off

EKS follows upstream Kubernetes closely. Each minor version gets roughly 14 months of standard support from its EKS release date, followed by up to 12 months of extended support. Extended support is not free: the control plane is billed at six times the standard price, enrollment is automatic because the default upgrade policy is EXTENDED, and when extended support ends AWS auto-upgrades your control plane to the oldest supported version at a time of its choosing.

That last part is what should worry you. A forced upgrade you didn’t schedule, on a cluster whose workloads you haven’t tested against the new version, is the worst-case scenario this post is trying to prevent.

There’s a second reason not to fall behind: you can only upgrade one minor version at a time. Three versions behind means three full upgrade cycles, back to back, with the clock running.

Check where you stand right now:

aws eks describe-cluster --name prod-eu --query 'cluster.version'
aws eks describe-cluster-versions \
  --query 'clusterVersions[].{v:clusterVersion,status:status,eos:endOfStandardSupportDate}' \
  --output table

What Actually Happens During an EKS Upgrade

An EKS upgrade is three separate upgrades that need to happen in a specific order:

  1. Control plane — AWS manages this. You request the new version, AWS rolls the API servers and etcd behind the scenes. It takes 20–40 minutes. The API stays available, though long-lived connections (watches, kubectl logs -f) may drop and reconnect. Your running pods are not touched.
  2. Add-ons — kube-proxy, CoreDNS, VPC CNI, EBS CSI driver, and anything else installed as an EKS add-on. Each has a compatibility matrix against the Kubernetes version.
  3. Data plane — the worker nodes. This is where downtime comes from, because every node has to be replaced with one running the new kubelet, and every pod on it has to move.

“Zero downtime” is almost entirely a data-plane concern. If your workloads can survive a node being drained, they can survive an upgrade. If they can’t, no amount of care in the control plane step will save you.

The Pre-Upgrade Checklist

Do all of this at least a week before you touch the cluster. Most of it is one-time work you’ll reuse every quarter.

1. Read the changelog

The EKS version documentation has a per-version page listing what changed, what was removed, and what AWS did differently. Read the target version’s page end to end. For example, the 1.34 notes flag that AWS is not releasing an Amazon Linux 2 AMI for 1.34, AppArmor is deprecated, and VolumeAttributesClass moved from the v1beta1 to the v1 storage API. Any one of those could break a cluster that isn’t ready.

2. Scan for removed APIs

Two tools, both worth having in your CI:

# Scan live cluster for deprecated/removed APIs in the target version
kubent --target-version 1.35.0

# Scan Helm releases and raw manifests
pluto detect-helm -o wide --target-versions k8s=v1.35.0
pluto detect-files -d ./manifests --target-versions k8s=v1.35.0

kubent inspects what’s actually running; pluto inspects what you’re about to deploy. Anything flagged as REMOVED in the target version must be fixed first. Anything flagged DEPRECATED should be fixed now while you have time.

3. Check third-party compatibility

For every operator and Helm chart in the cluster (cert-manager, ingress controller, Kyverno, Istio, Prometheus operator, external-secrets), find the chart version’s supported Kubernetes range. Chart maintainers publish this in Chart.yaml under kubeVersion or in their release notes. An operator with a validating webhook that panics on a new API version can block every pod creation in the cluster.

4. Check add-on target versions

for addon in vpc-cni coredns kube-proxy aws-ebs-csi-driver; do
  echo "== $addon"
  aws eks describe-addon-versions --kubernetes-version 1.35 --addon-name $addon \
    --query 'addons[].addonVersions[].{v:addonVersion,default:compatibilities[0].defaultVersion}' \
    --output table | head -8
done

Write the target versions down. You’ll pin them in Terraform in the next step.

5. Audit PodDisruptionBudgets

This is the one that hangs upgrades:

kubectl get pdb -A -o custom-columns='NS:.metadata.namespace,NAME:.metadata.name,MINAVAIL:.spec.minAvailable,MAXUNAVAIL:.spec.maxUnavailable,ALLOWED:.status.disruptionsAllowed'

Any PDB with ALLOWED = 0 will block a drain indefinitely. Common causes: maxUnavailable: 0, minAvailable equal to the replica count, or a single-replica Deployment with a PDB attached. Fix these before the upgrade, not during.

Every workload that matters should have: at least 2 replicas, a readiness probe that actually reflects readiness, a topologySpreadConstraints or anti-affinity rule across zones, and a PDB that allows at least one disruption.

6. Back up what you own

etcd is managed by AWS and you can’t snapshot it. What you can back up is your declared state. If you’re doing GitOps, git is the backup. If not, at minimum:

kubectl get all,cm,secret,ing,pvc,pdb,sa,role,rolebinding -A -o yaml > pre-upgrade-$(date +%F).yaml

For stateful workloads, take EBS snapshots or run Velero before you start.

Pin Versions in Terraform

The single most important configuration decision is that upgrades happen when you choose, not when AWS’s rollout schedule decides. That means pinning the cluster version, every node group version, and every add-on version explicitly.

With the terraform-aws-eks module:

module "eks" {
  source  = "terraform-aws-eks/aws"
  version = "~> 21.0"

  name               = "prod-eu"
  kubernetes_version = "1.34"   # <- the only line you change to upgrade the control plane

  upgrade_policy = {
    support_type = "STANDARD"   # fail loudly instead of drifting into paid extended support
  }

  addons = {
    kube-proxy = {
      addon_version = "v1.34.0-eksbuild.2"   # pin; use describe-addon-versions output
    }
    coredns = {
      addon_version = "v1.12.1-eksbuild.2"
    }
    vpc-cni = {
      addon_version  = "v1.20.1-eksbuild.1"
      before_compute = true
    }
    aws-ebs-csi-driver = {
      addon_version = "v1.45.0-eksbuild.1"
    }
  }

  eks_managed_node_groups = {
    general = {
      ami_type       = "AL2023_x86_64_STANDARD"
      instance_types = ["m6i.xlarge"]
      min_size       = 3
      max_size       = 12
      desired_size   = 6

      # Pin the node group independently of the control plane
      kubernetes_version = "1.34"

      update_config = {
        max_unavailable_percentage = 25
      }
    }
  }
}

Three things to note:

  • support_type = "STANDARD" makes Terraform refuse to leave the cluster on a version that has fallen out of standard support. You’ll get an error at plan time instead of a surprise on the bill.
  • Node groups pin their own kubernetes_version. Without it, the module tracks the cluster version and Terraform will try to roll all your nodes in the same apply as the control plane. You don’t want that.
  • Add-on versions are pinned as literal strings. The version strings above are illustrative; use the output from describe-addon-versions for your target.

terraform plan is now your upgrade preview. If the plan shows more than you expect to change, stop.

The Upgrade, Step by Step

Step 1: Control plane

Change one line:

kubernetes_version = "1.35"

Run terraform plan and confirm the only change is the cluster version. Then apply. Wait for the update to finish:

aws eks describe-update --name prod-eu --update-id <id-from-terraform-output>
# or just poll
watch aws eks describe-cluster --name prod-eu --query 'cluster.{v:version,status:status}'

What you should see: status: ACTIVE, version: 1.35, and kubectl get nodes still showing every node on v1.34.x and Ready. Nothing about your workloads has changed yet. Kubernetes supports kubelets up to three minor versions behind the API server, so this state is safe to sit in for a while.

Step 2: Add-ons, in order

Update the pinned add-on versions in Terraform and apply. The order matters:

  1. kube-proxy — must match the control plane minor version.
  2. CoreDNS — verify with kubectl -n kube-system rollout status deploy/coredns.
  3. VPC CNI — this is the one that bites. Check its release notes for the exact supported upgrade path; it does not always support skipping versions. Watch kubectl -n kube-system rollout status ds/aws-node and confirm no pods are stuck ContainerCreating afterward (that’s the symptom of a broken CNI).
  4. CSI drivers — EBS, EFS, whatever you use.

After each one, run a quick DNS and networking check from a fresh pod:

kubectl run nettest --rm -it --image=busybox:1.36 --restart=Never -- \
  sh -c 'nslookup kubernetes.default && wget -qO- --timeout=3 http://my-service.my-namespace:8080/healthz'

Step 3: Nodes

You have two choices for managed node groups.

Option A — rolling update in place. Bump kubernetes_version = "1.35" on the node group and apply. EKS launches surge nodes on the new version, cordons and drains old nodes respecting your PDBs, and terminates them. max_unavailable_percentage controls how aggressive it is. This is simple and works well for non-critical clusters.

The downside: if a drain hangs on a bad PDB, the whole node group update stalls, and if something on the new version is broken, your old nodes are already gone.

Option B — blue/green node groups. This is what I recommend for production. Add a second node group on the new version alongside the old one:

eks_managed_node_groups = {
  general-134 = { kubernetes_version = "1.34", desired_size = 6, ... }
  general-135 = { kubernetes_version = "1.35", desired_size = 6, ... }
}

Apply, wait for the new nodes to be Ready, then move workloads over at your own pace:

# Stop scheduling onto old nodes
kubectl cordon -l eks.amazonaws.com/nodegroup=general-134

# Drain them one at a time, respecting PDBs
for node in $(kubectl get nodes -l eks.amazonaws.com/nodegroup=general-134 -o name); do
  kubectl drain $node --ignore-daemonsets --delete-emptydir-data --timeout=600s
  sleep 60   # let the scheduler and any autoscaler settle
done

Watch your dashboards between drains. If anything looks wrong, kubectl uncordon the old nodes and you have an instant rollback — the old node group is still there, still on 1.34. Once you’re confident (I’d say 24 hours minimum for a production cluster), remove general-134 from Terraform and apply.

It costs you double node capacity for a day. That’s cheap compared to the alternative.

Step 3b: If you run Karpenter

Karpenter replaces nodes when they drift from their EC2NodeClass. Update the amiSelectorTerms to the new version’s alias and Karpenter will roll the fleet:

apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
spec:
  amiSelectorTerms:
    - alias: al2023@v20260812   # pin a specific AMI release, not "latest"

Throttle the rollout so it doesn’t replace everything at once:

apiVersion: karpenter.sh/v1
kind: NodePool
spec:
  disruption:
    budgets:
      - nodes: "10%"
      - nodes: "0"
        schedule: "0 9 * * mon-fri"   # freeze during business hours
        duration: 8h

Karpenter respects PDBs, so the same audit from the checklist applies.

Rolling Out Across a Fleet

If you run more than a couple of clusters, never upgrade them all at once. The order I use:

dev → staging → one production canary → remaining production, in batches.

Each promotion is gated on health signals from the previous environment, not on a calendar. What “healthy” means depends on your stack, but at minimum: all nodes Ready on the new version, no pods in CrashLoopBackOff or Pending, error rate and p99 latency within their normal bands for at least a few hours, and — if you’re running Flux or Argo — every HelmRelease or Application reporting Ready.

Running a control plane that manages dozens of customer clusters has taught me that this is the only sane approach. Version pinning plus staged canaries means an upgrade that breaks something breaks it in one place, where you’re watching, instead of everywhere at 3 a.m.

If your version is a single string in Terraform, the promotion is a one-line pull request per environment. That’s the goal.

Verification

After the last node is on the new version:

# Every node on the target version and Ready
kubectl get nodes -o custom-columns='NAME:.metadata.name,VERSION:.status.nodeInfo.kubeletVersion,STATUS:.status.conditions[-1].type'

# No sad pods anywhere
kubectl get pods -A --field-selector=status.phase!=Running,status.phase!=Succeeded

# Add-ons healthy
aws eks list-addons --cluster-name prod-eu
aws eks describe-addon --cluster-name prod-eu --addon-name vpc-cni \
  --query 'addon.{v:addonVersion,status:status,health:health}'

# Deprecated API scan is clean
kubent --target-version 1.35.0

Then run whatever end-to-end smoke test you trust — a synthetic login, a checkout, a message through the queue. Don’t declare victory on kubectl get nodes alone.

The Rollback Reality

There is no rollback for an EKS control plane. Once it’s on 1.35, it stays on 1.35. Plan accordingly:

  • The API-deprecation work happens before the control plane upgrade, because after it you can’t go back to a version that still serves the old APIs.
  • Your real rollback is the old node group. That’s why blue/green is worth the cost — the 1.34 nodes are your escape hatch until you delete them.
  • If a specific workload breaks on the new version and you can’t fix it quickly, pin it to the old node group with a nodeSelector on eks.amazonaws.com/nodegroup and keep that group alive while you sort it out.

Common Failure Modes

Drain hangs forever. A PDB with zero allowed disruptions, or a pod using emptyDir without --delete-emptydir-data. Run kubectl get pdb -A and look for ALLOWED DISRUPTIONS: 0.

New pods stuck in ContainerCreating. Almost always the VPC CNI — either the upgrade broke it or the new nodes can’t allocate IPs. Check kubectl -n kube-system logs ds/aws-node and your subnet’s free IP count.

CoreDNS in CrashLoopBackOff. Usually upgraded before kube-proxy or before the control plane. Order matters.

Workloads fail on new nodes only. A removed API, or an AMI change (AL2 → AL2023 changed cgroups, the init system, and the default container runtime version). Compare a failing pod’s events on a new node to a healthy one on an old node.

Nothing schedules at all. A validating or mutating webhook (Kyverno, OPA Gatekeeper, cert-manager, Istio) is rejecting requests because it doesn’t understand the new API version. Check kubectl get validatingwebhookconfigurations and each webhook’s failurePolicy; anything set to Fail can take down the cluster when the backing pod is unhealthy.

The One-Screen Checklist

Before

  • ☐ Read the target version’s EKS release notes
  • ☐ kubent and pluto scans clean for the target version
  • ☐ Every operator and chart confirmed compatible
  • ☐ Add-on target versions recorded
  • ☐ Every PDB allows ≥1 disruption; critical workloads have ≥2 replicas and readiness probes
  • ☐ Backup taken (git state, Velero, EBS snapshots)
  • ☐ Cluster, node group, and add-on versions all pinned in Terraform

During

  • ☐ Control plane upgraded; ACTIVE on the new version
  • ☐ Add-ons upgraded in order: kube-proxy → CoreDNS → VPC CNI → CSI
  • ☐ DNS and network smoke test passes
  • ☐ New node group created; old one cordoned and drained node by node
  • ☐ Dashboards watched between drains

After

  • ☐ All nodes Ready on the new version
  • ☐ No pods Pending or CrashLoopBackOff
  • ☐ Add-ons ACTIVE
  • ☐ End-to-end smoke test passes
  • ☐ Old node group removed after 24h+ of stability
  • ☐ Next upgrade date on the calendar

Cadence

Upgrade every quarter, and never sit more than one minor version behind the latest in standard support. Done regularly, an EKS upgrade is a one-line pull request per environment and an afternoon of watching dashboards. Done once every two years under an end-of-support deadline, it’s a project. You get to choose which one it is.

Atiqur Rahman

I am MD. Atiqur Rahman graduated from BUET and is an AWS-certified solutions architect. I have successfully achieved 6 certifications from AWS including Cloud Practitioner, Solutions Architect, SysOps Administrator, and Developer Associate. I have more than 8 years of working experience as a DevOps engineer designing complex SAAS applications.

Leave a Reply