Live
OpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendAccess Cloudflare Skills Directly Through the API MCP ServerCodeQL 2.27.2 expands language models and tightens macOS build support – what engineers need to knowTangible Certification: Turning a Kubernetes Badge into a Gold NecklaceGoogle Data Cloud GA updates: agent‑centric tooling, hybrid Spanner, and expanded Lakehouse catalogCloudflare folds Deno runtime into Workers: practical impact on serverless deploymentsManaging Copilot Code Review Costs and License Scope with New Org‑Level ControlsOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendAccess Cloudflare Skills Directly Through the API MCP ServerCodeQL 2.27.2 expands language models and tightens macOS build support – what engineers need to knowTangible Certification: Turning a Kubernetes Badge into a Gold NecklaceGoogle Data Cloud GA updates: agent‑centric tooling, hybrid Spanner, and expanded Lakehouse catalogCloudflare folds Deno runtime into Workers: practical impact on serverless deploymentsManaging Copilot Code Review Costs and License Scope with New Org‑Level Controls
Kubernetes

Kubernetes shifts to cgroup v2: practical steps for engineers

AI SummaryPowered by AI

Kubernetes 1.35 will refuse to start on nodes using Linux cgroup v1, and node swap reached GA in 1.34. The changes affect resource isolation, workload density, and edge AI deployments, requiring concrete operational adjustments.

Kubernetes 1.35 enforces a move to Linux cgroup v2, and the platform already introduced node swap as a general‑availability feature in 1.34. Engineers responsible for AI inference at the edge, platform reliability, or security must audit node configurations, update kubelet start‑up flags, and rethink resource‑management patterns to keep clusters functional and performant.

cgroup v2 becomes the required runtime

The older cgroup v1 hierarchy suffers from fragmented interfaces and limited isolation capabilities. cgroup v2 offers a single unified hierarchy, more consistent APIs, and enables newer features such as memory quality‑of‑service updates, container‑aware out‑of‑memory handling, and rootless operation. According to the DaoCloud open‑source lead, any Kubernetes version that relies on cgroup v2 (starting with v1.35) will refuse to start the kubelet on nodes still on v1. This makes pre‑upgrade migration of every Linux node to cgroup v2 a prerequisite, otherwise the control plane will reject the node.

Node swap as a density lever for edge AI

Agentic AI workloads often allocate large memory blocks that sit idle after a spike, limiting cluster density. The Kubernetes blog highlighted that node swap, now GA in v1.34, can act as a “shock absorber” for such spikes. Benchmarks from Ocean Xie and Yuan Wang show up to three‑fold density improvements when swap resides on fast NVMe SSDs. Practitioners should evaluate swap size, SSD performance, and monitoring of swap activity to avoid unintended latency in latency‑sensitive edge inference.

Edge compute hardware and operational context

HPE positions its ProLiant edge servers as purpose‑built for resource‑constrained, high‑security environments. The hardware is marketed as optimized for AI inference at the edge, reducing data egress and compliance risk. The upcoming Kubernetes on Edge Day at KubeCon will surface observability and security patterns specific to distributed edge clusters, reinforcing the need for tight integration between hardware capabilities and Kubernetes resource controls.

Self‑hosted AI platform deployment checklist

Fairwinds warns that a working installer is only the first milestone for self‑hosted AI platforms. Teams must also define IAM responsibilities, establish ongoing maintenance processes, verify compatibility with private clouds, and configure DNS and networking according to the customer’s constraints. Skipping these preparatory steps can cause installation delays or project stalls, especially when combined with the new cgroup v2 requirement.

Related CloudNinjas coverage: hands-on guides.

What This Means For Practitioners

  • Audit every node’s kernel configuration; upgrade to cgroup v2 before upgrading to Kubernetes 1.35.
  • Enable and size node swap on clusters that run bursty AI inference; pair swap with low‑latency NVMe storage.
  • Align edge hardware choices (e.g., HPE ProLiant) with Kubernetes resource‑isolation features to meet security and compliance goals.
  • Incorporate a pre‑install checklist covering IAM, DNS, and maintenance when delivering self‑hosted AI platforms.
  • Monitor upcoming Edge Day sessions for emerging observability and security best practices that complement the cgroup v2 transition.
Originally published atThe New Stack