Amazon EKS 1.37 Upgrade Guide: SELinuxMount, Kubelet Changes and the Auto Mode Default That Flips

Amazon EKS 1.37 Upgrade Guide: SELinuxMount, Kubelet Changes and the Auto Mode Default That Flips

Amazon EKS and EKS Distro added Kubernetes 1.37 on October 2, 2026, in every Region where EKS runs, including AWS GovCloud (US). The headline features are pleasant: the HPA can scale to zero by default, HPA tolerance is configurable per autoscaler, and DRA device taints went stable. The part worth reading before you click "Update now" is shorter and less pleasant. Four changes in 1.37 can leave a Pod stuck in ContainerCreating, stop a kubelet from starting on a custom AMI, blank out a dashboard, or quietly change how often EKS Auto Mode evicts your Pods. None of them shows up as an API deprecation, so a clean "no deprecated APIs" check will not catch them.

What actually changes in EKS 1.37

This table is built from the EKS release notes for Kubernetes 1.37 and the EKS Auto Mode release notes dated October 1, 2026.

ChangeWho is affectedSymptom if you ignore it
SELinuxMount enabled by defaultPods that set seLinuxOptions and share a volume on SELinux-enabled nodesPod stays in ContainerCreating
Kubelet refuses deprecated cAdvisor flagsCustom AMIs or extra kubelet argumentsKubelet does not start, node never joins
eventRecordQPS: 0 now means unlimitedAnyone who set it to 0 explicitlyEvent volume is no longer capped at 5 per second
Three cAdvisor metrics removedAnyone scraping them, on any AMIEmpty panels, alerts that never fire
Auto Mode NodePools default to BalancedNew NodePools without consolidationPolicy, plus the built-in poolsDifferent eviction behaviour, GitOps drift
Kubelet logs its full config at startupEveryoneNode logs now contain the effective kubelet configuration

Why SELinuxMount is the one to take seriously

Before 1.37, when a Pod with an SELinux label mounted a volume, the container runtime walked the volume and relabeled every file. That is slow on large volumes, and it is why the Kubernetes project has been moving to a different mechanism: mount the volume once with mount -o context=<label> and skip the relabel entirely.

Kubernetes 1.36 started that move. In 1.37 the SELinuxMount feature is on by default. When a Pod sets seLinuxOptions and the volume's CSI driver declares seLinuxMount: true, the kubelet uses the mount option instead of relabeling.

The catch is physical: one mount can carry only one SELinux label. So if a second Pod on the same node uses the same volume with a different SELinux label, a different seLinuxChangePolicy, or a different privilege level, the kubelet cannot mount it and the second Pod sits in ContainerCreating. Before 1.37 the same pair of Pods started fine, because the files were just relabeled again.

Nodes without SELinux enabled are not affected. On EKS that makes Bottlerocket the first place to look, since it runs SELinux in enforcing mode. Do not assume the rest of your fleet is out of scope: check what your own nodes report (Step 2 below), especially if you build custom AMIs.

Step 1: Run the EKS upgrade insights first

EKS cluster insights still do the baseline work (deprecated API usage, add-on compatibility, kubelet version skew), so start there.

1aws eks list-insights \
2  --cluster-name prod-cluster \
3  --filter 'categories=UPGRADE_READINESS' \
4  --query 'insights[].{name:name,status:insightStatus.status,version:kubernetesVersion}' \
5  --output table
6
7# Details for anything not PASSING
8aws eks describe-insight --cluster-name prod-cluster --id <insight-id>

Treat a clean result as necessary, not sufficient. The four checks below are the ones insights are not designed to make for you.

Step 2: Audit SELinux volume sharing

Find the CSI drivers that opt in to mount-based labeling. This is the command AWS gives in the release notes:

1kubectl get csidriver \
2  -o custom-columns=NAME:.metadata.name,SELINUXMOUNT:.spec.seLinuxMount

If nothing shows true, this change cannot bite you and you can move on. If a driver does, list the Pods that set SELinux options, along with the node they run on:

 1kubectl get pods -A -o json | jq -r '
 2  .items[]
 3  | select(
 4      (.spec.securityContext.seLinuxOptions != null)
 5      or any(.spec.containers[]; .securityContext.seLinuxOptions != null)
 6    )
 7  | [.metadata.namespace, .metadata.name, .spec.nodeName,
 8     (.spec.securityContext.seLinuxOptions // {} | tostring),
 9     (.spec.securityContext.seLinuxChangePolicy // "unset")]
10  | @tsv'

Then confirm whether SELinux is on for a given node. From a debug Pod or SSM session on the host:

1getenforce    # Enforcing, Permissive or Disabled

What you are looking for is two or more Pods that mount the same PersistentVolumeClaim, land on the same node, and differ in seLinuxOptions.level, in seLinuxChangePolicy, or in privilege (one privileged, one not). The classic case is an application Pod plus a privileged backup or log-shipping sidecar Pod reading the same volume.

The fix is to make the Pods agree. Either give them the same seLinuxOptions.level, or opt the shared volume back into the old relabeling behaviour with seLinuxChangePolicy: Recursive. Apply it to every Pod that shares the volume, not just one of them:

 1apiVersion: v1
 2kind: Pod
 3metadata:
 4  name: app
 5spec:
 6  securityContext:
 7    seLinuxChangePolicy: Recursive
 8    seLinuxOptions:
 9      level: "s0:c123,c456"
10  containers:
11    - name: app
12      image: public.ecr.aws/docker/library/busybox:stable
13      command: ["sleep", "infinity"]
14      volumeMounts:
15        - name: data
16          mountPath: /data
17  volumes:
18    - name: data
19      persistentVolumeClaim:
20        claimName: shared-data

Recursive costs you the faster mount on that volume. That is the trade: correctness now, speed later once the Pods share a label.

Step 3: Check kubelet flags and eventRecordQPS

In 1.37 the kubelet will not start if a deprecated cAdvisor flag is set. The single exception is --housekeeping-interval. The EKS-optimized AMIs do not set these flags, so managed node groups on stock AMIs are fine. You are exposed if you bake custom AMIs or pass extra kubelet arguments through user data or a launch template.

On a running node, print what the kubelet was actually started with:

1ps -o args= -C kubelet | tr ' ' '\n' | grep '^--'

Compare the output against the deprecated flags in the kubelet reference. Flags from the old cAdvisor set, such as --enable-load-reader, --global-housekeeping-interval, --event-storage-age-limit or --storage-driver-*, are the kind to remove. Also grep your bootstrap scripts, since a flag there only fails when the next node launches:

1grep -rnE "kubelet-extra-args|kubelet:|flags:" ./ami ./terraform ./userdata

The second kubelet change is easy to miss. eventRecordQPS: 0 used to apply a limit of 5 events per second. In 1.37 it means no limit. If you set 0 deliberately, you now have to write the old limit down explicitly. With AL2023 and nodeadm:

1apiVersion: node.eks.aws/v1alpha1
2kind: NodeConfig
3spec:
4  kubelet:
5    config:
6      eventRecordQPS: 5

Step 4: Find dashboards and alerts using the removed metrics

The kubelet drops three cAdvisor metrics in 1.37, and this one applies on every AMI:

  • container_cpu_load_average_10s
  • container_cpu_load_d_average_10s
  • container_tasks_state

Ask Prometheus whether you ingest any of them today:

1count by (__name__) (
2  {__name__=~"container_cpu_load_average_10s|container_cpu_load_d_average_10s|container_tasks_state"}
3)

Then search the places that would break silently:

1grep -rnE "container_cpu_load_(d_)?average_10s|container_tasks_state" \
2  ./dashboards ./alert-rules ./recording-rules

An alert built on a metric that no longer exists does not fire and does not error. It just goes quiet, which is the worst way for an alert to fail.

Step 5: Pin consolidationPolicy on EKS Auto Mode

Starting with EKS 1.37, a new Auto Mode NodePool that does not set consolidationPolicy gets Balanced instead of WhenEmptyOrUnderutilized. AWS describes the difference this way: WhenEmptyOrUnderutilized disrupts a node for any saving, as little as USD 0.01 per hour, while Balanced only approves a disruption when the hourly saving is worth the cost of moving the Pods. Expect fewer evictions at roughly the same cost.

Three details matter in practice:

  1. The default is applied only when the NodePool is created. Existing NodePools keep their behaviour through the upgrade.
  2. A GitOps sync that deletes and recreates a NodePool with no value set will come back as Balanced. Argo CD or Flux will also show a live value your Git source does not have.
  3. The built-in general-purpose and system NodePools move to Balanced and are reconciled hourly, so you cannot override the field on them.

If you want a specific behaviour, say so in the manifest:

1apiVersion: karpenter.sh/v1
2kind: NodePool
3metadata:
4  name: batch
5spec:
6  disruption:
7    consolidationPolicy: WhenEmptyOrUnderutilized
8  # ...rest of your NodePool spec

A workload that needs the old aggressive consolidation should target a NodePool of your own that sets it, not the built-in pools.

Step 6: Upgrade the control plane, add-ons, then nodes

EKS upgrades one minor version at a time, so the cluster must already be on 1.36.

 1# Control plane
 2aws eks update-cluster-version \
 3  --name prod-cluster \
 4  --kubernetes-version 1.37
 5
 6aws eks describe-update --name prod-cluster --update-id <update-id>
 7
 8# Add-ons: find a version built for 1.37, then update
 9aws eks describe-addon-versions \
10  --kubernetes-version 1.37 --addon-name vpc-cni \
11  --query 'addons[].addonVersions[].addonVersion' --output text
12
13aws eks update-addon --cluster-name prod-cluster \
14  --addon-name vpc-cni --addon-version <version>
15
16# Managed node groups
17aws eks update-nodegroup-version \
18  --cluster-name prod-cluster --nodegroup-name workers

With eksctl the control plane step is eksctl upgrade cluster --name prod-cluster --version 1.37 --approve.

Repeat the add-on step for CoreDNS, kube-proxy, the EBS CSI driver and anything else you run as an EKS add-on. Upgrade a non-production cluster first and leave it running for a full day of real traffic, because the SELinux problem only appears when two specific Pods land on the same node.

Best practices

  • Run the Step 2 audit on the 1.36 cluster. After the upgrade you are debugging in production instead of reading a list.
  • Set consolidationPolicy explicitly on every NodePool you own, whatever value you choose. Explicit beats default when defaults move.
  • Remove the deprecated kubelet flags while still on 1.36. They are already deprecated there, so the change is safe to ship early.
  • The kubelet now logs its full effective configuration at startup. Review who can read node logs in CloudWatch or your log pipeline.
  • Know your exit. EKS has supported version rollback since July 2026, and it has limits worth reading before you need it.

Common mistakes to avoid

  • Treating "no deprecated APIs" as "safe to upgrade". None of the four changes here is an API removal.
  • Adding seLinuxChangePolicy: Recursive to one Pod only. Pods sharing the volume must match, so one Pod with Recursive and one without is still a conflict.
  • Assuming existing Auto Mode NodePools flip to Balanced. They do not. Recreated ones do, and that is the surprise.
  • Testing on an empty staging cluster. A staging cluster without the shared-volume Pods proves nothing about the shared-volume Pods.
  • Changing eventRecordQPS after noticing the event flood. If you had 0, change it to 5 before the nodes roll.

Troubleshooting

Pod stuck in ContainerCreating after the node upgrade. Run kubectl describe pod and read the events for a volume mount failure that mentions the SELinux context. Find the other Pod using the same PVC on that node, compare seLinuxOptions and seLinuxChangePolicy, and align them. As a stopgap, set seLinuxChangePolicy: Recursive on all Pods sharing the volume.

New nodes never reach Ready on a custom AMI. Check the kubelet unit on the instance with journalctl -u kubelet --no-pager | tail -50. A rejected flag is reported at startup. Remove the flag from the AMI or the launch template user data and roll the node group again.

Grafana panels empty after the upgrade. Compare the panel query against the three removed metric names above. There is no drop-in replacement from the kubelet for load average, so decide whether node-level load from node_exporter answers the same question.

Argo CD reports a NodePool as out of sync. The live object shows consolidationPolicy: Balanced and Git has no value. Add the field to Git.

FAQ

Can I go from EKS 1.35 straight to 1.37? No. Upgrade to 1.36 first, and read the 1.36 notes on the way, since the gitRepo volume removal and strict IP/CIDR validation land there.

Do I need to act on SELinuxMount if my nodes do not use SELinux? No. The EKS release notes say nodes without SELinux enabled are not affected. Verify with getenforce instead of assuming.

Does HPA scale-to-zero just work now? The HPAScaleToZero feature gate is on by default, but it only applies to an HPA that scales on Object or External metrics, and those need a metrics adapter that EKS does not install for you.

Is DRA usable on EKS Auto Mode in 1.37? Not with the NVIDIA DRA driver. AWS supports it with Karpenter static capacity, managed node groups and self-managed nodes. On Auto Mode, use the NVIDIA device plugin.

Do G7 GPU instances need anything special? They need NVIDIA driver 595. The AL2023 NVIDIA AMIs include it on all supported versions. The Bottlerocket NVIDIA AMIs include it from Kubernetes 1.37.

Key takeaways

CheckCommand or actionDo it when
Upgrade insightsaws eks list-insights --filter 'categories=UPGRADE_READINESS'Before anything else
SELinux mount conflictskubectl get csidriver plus the Pod auditOn 1.36, before upgrade
Deprecated kubelet flagsps -o args= -C kubelet and grep bootstrap scriptsOn 1.36, before upgrade
eventRecordQPS: 0Change to 5 to keep the old limitBefore nodes roll
Removed metricsGrep dashboards and alert rulesBefore upgrade
Auto Mode consolidationSet consolidationPolicy explicitlyBefore any NodePool is recreated

Further Reading