The Death of cgroup v1 and the Rise of Node Swap: How Kubernetes v1.35 Rewrites Node Memory Economics
With Kubernetes v1.35 blocking cgroup v1 by default and v1.34 promoting Node Swap to GA, platform teams must overhaul Linux kernel configs to survive bursty agentic AI workloads.
Published: 2026.10.10
The Sudden End of Dual-Stack Kernel Support in Cloud-Native Infrastructure
For nearly a decade, enterprise platform engineering teams treated Linux control groups as invisible plumbing. Fleet managers ran a patchwork of operating system distributions across bare metal, virtual machines, and managed cloud environments. Some nodes operated on the legacy cgroup v1 hierarchy designed in 2007. Others operated on cgroup v2, introduced into the Linux kernel in 2016. Because Kubernetes maintained backward compatibility, few teams felt an immediate operational urgency to clean up their underlying host kernels.
That grace period is now officially over. Starting in Kubernetes v1.35, the kubelet agent refuses to boot on Linux nodes running cgroup v1 by default. What was once a gentle deprecation notice in release documentation has converted into a hard operational blocker. If a worker node boots with the legacy hierarchy, the kubelet fails its preflight checks, marks the node as unready, and ejects it from scheduling.
The Kubelet v1.35 Control Group Cutover
From silent kernel drift to hard initialization failure
cgroup v1 Multi-Hierarchy
Fragmented controller trees hide true memory usage and prevent unified out-of-memory container handling.
Kubelet v1.35 Hard Exit
The node agent terminates immediately during initialization if cgroup v1 controllers are detected.
cgroup v2 & NVMe Node Swap
Single unified hierarchy unlocks memory QoS, rootless pods, and burst shock-absorption on fast NVMe drives.
This structural shift collides directly with a radical change in modern compute consumption: the explosion of bursty, agentic AI workloads. Enterprise developers are deploying autonomous reasoning agents, retrieval-augmented generation (RAG) pipelines, and local inference workers directly into Kubernetes clusters. Unlike traditional stateless microservices that consume memory in predictable, flat lines, agentic workflows generate sharp, volatile memory spikes.
When a multi-step reasoning model spins up vector retrieval or builds dynamic context windows, its container demands a massive wave of immediate RAM. Minutes later, the workload finishes, and that memory sits idle. Under older Kubernetes configurations, cluster architects had to size node capacity around those extreme peak allocations. Because nodes ran out of RAM long before their CPUs approached full saturation, enterprises ended up paying for massive fleets of half-empty cloud instances.
The confluence of two major milestones—the hard phase-out of cgroup v1 in v1.35 and the General Availability of Node Swap in v1.34—rewrites the financial and operational baseline of container orchestration. Engineers can no longer rely on unmaintained Linux images, but in exchange, they gain the ability to run worker nodes with unprecedented hardware density.
Benchmarking the Real Hardware Cost: cgroup v1 Bottlenecks Versus cgroup v2 with NVMe Swap
To understand why Kubernetes maintainers pulled the plug on cgroup v1, infrastructure leaders must look at how the Linux kernel handles memory pressure. Under cgroup v1, every resource—CPU, memory, blkio, and networking—lives in a separate controller hierarchy. If a container exhausts its memory limit, the kernel out-of-memory (OOM) killer fires blindly, often terminating critical helper processes inside a pod while leaving orphan threads running.
Furthermore, cgroup v1 makes it impossible to implement writeback throttling accurately. When a container writes data to disk, the kernel buffers those pages in system memory. Under cgroup v1, the memory controller and the block I/O controller cannot communicate, meaning disk writes are charged to arbitrary system pools rather than the pod that triggered them.
Kubernetes Memory Architecture Shift
Performance and density metrics across the cgroup v2 transition
Pod Density Multiplier
Workload packing gain on nodes backed by fast NVMe SSD swap partitions
Kubelet v1.35 Tolerance
Zero-boot policy for nodes detected with legacy cgroup v1 controllers
Unified Tree Hierarchy
Eliminates split accounting across CPU, memory, and blkio subsystems
The introduction of cgroup v2 introduces a single, unified hierarchy where every process belongs to exactly one control group. This structural change enables granular Memory Quality of Service (QoS), container-aware OOM handling (which kills the entire container cleanly instead of terminating random worker threads), and native support for rootless containers. Crucially, cgroup v2 provides the accounting foundation that allowed Kubernetes maintainers to finally stabilize Node Swap in v1.34.
The operational and financial differences between these architectures become immediately apparent when evaluated across standard cloud instances.
| Operating Parameter | Legacy cgroup v1 (Kubelet ≤ v1.34) | Standard cgroup v2 (No Swap) | Modern cgroup v2 + NVMe Swap (Kubelet v1.34+) |
|---|---|---|---|
| Kubelet v1.35+ Compatibility | Fails startup (Refused) | Supported natively | Supported natively |
| Memory Accounting Tree | Split per controller | Unified single hierarchy | Unified single hierarchy |
| Out-of-Memory (OOM) Handling | Arbitrary per-thread kill | Container-level atomic kill | Container-level atomic kill |
| Swap Support Safety | Disabled / High risk of thrashing | Optional (Limited visibility) | Full cgroup v2 swap accounting |
| Agentic Burst Handling | Immediate Pod OOMKill | Immediate Pod OOMKill | Absorbed into NVMe swap space |
| Average Node RAM Utilization | 40% – 55% (Capped by burst sizing) | 50% – 60% | 85% – 92% |
| Simulated Monthly Cost per 100 Nodes | $38,400 (Overprovisioned instances) | $32,000 (Standard sizing) | $12,800 – $15,360 (High-density packing) |
Simulation Context: Monthly cost benchmark based on 100 cloud compute nodes (equivalent to 16 vCPU / 64 GB RAM instances at an industry baseline of $0.533 per node/hour). The modern cgroup v2 configuration with Node Swap uses high-density worker nodes with attached fast NVMe storage, reducing the total node count required to support bursty enterprise workloads by roughly 60%.
When nodes run without swap, Kubernetes treats RAM as a rigid wall. If an agentic workflow briefly requests 12 GB of memory to run a contextual search across internal documentation, the scheduler must reserve that capacity indefinitely. If multiple pods burst simultaneously, the host runs out of physical memory, triggering a cascade of pod evictions and restarts.
By contrast, benchmark tests by Kubernetes upstream contributors show that pairing cgroup v2 with fast enterprise NVMe SSD swap partitions allows nodes to absorb temporary memory spikes seamlessly. In bursty inference and reasoning scenarios, clusters achieve up to three times higher pod density. Cold pages—such as initialization code and inactive service workers—are automatically paged out to disk, keeping expensive physical RAM open for active token generation.
Operational Friction: Three Critical Points Where Clusters Break
Migrating an enterprise cluster fleet from legacy kernel flags to modern control groups is not a painless toggle. For large platforms supporting multi-tenant environments, the transition exposes real architectural friction across three distinct operational areas.
Infrastructure Sizing: Fixed Allocations vs. Dynamic Buffering
How node swap transforms memory provisioning for AI services
Traditional Rigid RAM
High Capital Waste- • Clusters sized for maximum burst peaks
- • Average node memory sits 50% idle
- • Sudden memory spikes trigger pod eviction
- • Fleet requires 2.5x more physical nodes
cgroup v2 + NVMe Swap
High Node Density- • Physical RAM dedicated to active processing
- • Idle container pages move to fast NVMe
- • Spikes buffered without OOM kills
- • Node count drops by up to 60%
Escalating Infrastructure Bills from Artificially Capped Density
Before the stabilization of Node Swap, platform engineers had no choice but to over-provision hardware. In production Kubernetes environments, memory is almost always the first ceiling that clusters hit. A 32-core node frequently exhausts its 128 GB of RAM while its processor load hovers around 25%.
Because cluster autoscale rules trigger based on requested resource thresholds, cloud providers spin up additional worker nodes solely to satisfy unallocated memory reservations. Enterprise finance teams end up paying thousands of dollars every month for physical DRAM that does nothing. Without cgroup v2 and swap accounting, teams cannot safely enable disk-backed memory without risking severe system thrashing, locking them into an expensive, low-density deployment model.
Cascading Out-of-Memory Evictions During Agentic Traffic Surges
Agentic AI workloads do not behave like monolithic web applications. When multiple autonomous agents execute complex plans in parallel, their memory profiles resemble a serrated saw. An agent running document analysis might consume 2 GB for ten seconds, jump to 16 GB while parsing multi-modal image tokens, and drop back down to baseline.
Under cgroup v1, when a node’s physical memory fills up, the Linux kernel cannot distinguish between mission-critical orchestrator pods and temporary worker processes. The kernel executes an uncontrolled OOM kill, often knocking out network mesh sidecars, logging daemons, or ingress controllers. Platforms experience sudden latency spikes, broken client connections, and pod crash-loop cycles that require manual engineering intervention.
Deployment Gridlock in Air-Gapped and Edge Environments
The difficulty deepens when enterprises deploy Kubernetes outside hyperscaler cloud environments. According to engineering analysis from organizations running customer-hosted and edge clusters, installation scripts frequently stall due to operating system discrepancies.
When vendors deliver self-hosted software packages to customer-owned private infrastructure, they often discover that host base images run outdated Linux kernels pinned to cgroup v1. When the vendor installer attempts to deploy a modern Kubernetes v1.35 control plane, the installation grinds to an absolute halt. Site reliability teams face delayed go-live schedules, complicated cross-team security audits, and friction with corporate infrastructure teams who refuse to update their long-term support (LTS) kernel images without months of formal review.
Buffer Technologies: Edge Appliances, Advanced Telemetry, and Purpose-Built Compute
To counter the instability of memory-heavy workloads, enterprise engineering organizations are turning away from generic virtual machines in favor of purpose-built architectures designed specifically for AI inference.
At the physical infrastructure layer, hardware manufacturers like Hewlett Packard Enterprise are addressing this bottleneck by pairing edge-optimized servers, such as ProLiant platforms, directly with specialized cloud-native runtimes. Instead of backhauling massive data streams to central public clouds—which triggers steep network egress charges and introduces unacceptable network latency—enterprises run inference locally on hardened hardware. These edge appliances come pre-configured with Linux distributions that support cgroup v2 out of the box, ensuring that local inference workers can leverage high-speed NVMe storage to absorb context spikes without crashing edge control planes.
The Agentic Memory Shock-Absorption Pipeline
How cgroup v2 and NVMe swap prevent OOM kills during reasoning tasks
Traffic Surge
Agentic pod receives sudden request requiring large context retrieval.
Physical RAM Exhaustion
Active working set exceeds available node DRAM capacity.
cgroup v2 Page Migration
Linux kernel cleanly evicts cold background pages to NVMe swap space.
Zero Eviction Completion
Inference finishes, freeing DRAM; pod completes without OOM termination.
At the networking and telemetry level, modern cloud-native environments rely heavily on extended Berkeley Packet Filter (eBPF) technology. Using tools like Cilium alongside comprehensive observability platforms such as Datadog, site reliability engineers can monitor memory pressure metrics in real time. Because cgroup v2 exposes standardized memory pressure stall information (PSI) via the /proc/pressure interface, monitoring agents can detect when a container is experiencing I/O bottlenecks long before it triggers a failure.
Furthermore, cloud-native frameworks like KubeEdge allow distributed enterprises to manage these edge and on-premise nodes from a centralized Kubernetes control plane. By combining modern kernel primitives with edge-aware schedulers, platforms can route compute-heavy inference jobs exclusively to nodes equipped with verified NVMe swap buffers, ensuring reliable execution even when operating under constrained hardware footprints.
Operational Roadmap: The Two-Stage Migration from Legacy Kernels to Node Swap
Migrating an enterprise infrastructure fleet to Kubernetes v1.35 and enabling Node Swap requires a disciplined execution plan. Platform teams must separate the mandatory operating system upgrade from the performance-tuning phase to prevent cluster instability.
Enterprise cgroup v2 and Node Swap Cutover Roadmap
Two-stage implementation for production Kubernetes clusters
Kernel Audit & Host Conversion
Verify distribution flags, update systemd configurations, and enforce cgroup v2 across base AMI images.
Control Plane Validation
Deploy Kubernetes v1.35 on upgraded nodes and confirm clean kubelet initialization without legacy fallbacks.
NVMe Swap Partitioning
Provision dedicated high-speed swap files and configure NodeSwap flags inside kubelet configuration profiles.
Agentic AI Stress Testing
Run bursty multi-agent benchmarks, measure PSI memory metrics, and optimize pod density allocations.
Short-Term Milestones: Pilot Kernel Audits and cgroup v2 Cutover
The immediate priority for every infrastructure team is ensuring that existing worker nodes can run modern kubelet versions without boot failures.
- Audit Host Kernel Parameters: Connect to every production node pool and run
stat -fc %T /sys/fs/cgroup/. If the command outputscgroup2fs, the node is operating correctly on cgroup v2. If the command outputstmpfs, the host is still running cgroup v1 and will fail during a Kubernetes v1.35 upgrade. - Update Operating System Boot Flags: For distributions running older systemd versions, update the GRUB bootloader configuration by adding
systemd.unified_cgroup_hierarchy=1toGRUB_CMDLINE_LINUX. Rebuild the bootloader configuration and reboot the host. - Refresh Base Machine Images: If your organization provisions infrastructure using immutable machine images (AMIs, golden images, or Packer templates), ensure that legacy operating system versions (such as Ubuntu 18.04, CentOS 7, or outdated Debian releases) are replaced with modern distributions (such as Ubuntu 22.04+, RHEL 9+, or Amazon Linux 2023) that enable cgroup v2 by default.
- Validate Container Runtime Engines: Ensure that underlying container runtimes—such as containerd (v1.6+) or CRI-O (v1.20+)—are explicitly configured to use the systemd cgroup driver rather than the legacy cgroupfs driver.
Long-Term Milestones: Node Swap Tuning and Production Agentic Sizing
Once the fleet runs safely on cgroup v2, engineering teams can turn their attention to unlocking hardware density gains by activating Node Swap on nodes designated for bursty workloads.
- Provision Enterprise NVMe Storage: Never configure swap space on standard network-attached cloud storage (such as default AWS EBS volumes or GCP Persistent Disks). High I/O latency will stall the Linux kernel during memory paging. Provision swap strictly on locally attached, low-latency NVMe solid-state drives.
- Configure Kubelet Swap Flags: In the
KubeletConfigurationresource, setfailSwapOntofalseand configure thememorySwapblock. SetswapBehaviortoLimitedSwapto ensure that regular pods do not exhaust the host’s swap partition and disrupt critical system daemons. - Establish Memory Pressure Alerts: Instrument worker nodes with eBPF agents to monitor Pressure Stall Information (PSI). Configure operational alerts to trigger if
somememory pressure exceeds 10% for more than 30 consecutive seconds, which indicates that disk paging is beginning to degrade workload execution speed. - Recalibrate Pod Requests and Cluster Density: Re-evaluate sizing guidelines for engineering teams deploying agentic AI services. Reduce static memory requests down to the container’s true baseline idle consumption, allowing the local NVMe swap buffer to absorb temporary reasoning spikes. Gradually scale down total worker node counts as average cluster density increases toward 85%.