Live
AI‑Generated Code Halves Manual Effort – Redesigning CI/CD and Governance for the New Development PaceGemini 4 Argon expands token limits and tops knowledge‑work benchmarks – what engineers need to knowData Agent Kit GA unlocks direct agent access to BigQuery Graph, Bigtable, and Spark for AI‑driven pipelinesRunning gcloud and bq via the Google Cloud CLI remote MCP server: practical implications for AI and platform engineersS3 Tables add full Iceberg V3 type support, deletion vectors, and row lineageAI‑First CI/CD Pivot at CloudBees Redefines Enterprise Pipeline PracticesMetadata Pre‑Filtering in Amazon S3 Vectors Improves Filtered Search RecallCommand Injection via Branch Name Exposes GitHub Token in AI Coding AgentsAI‑Generated Code Halves Manual Effort – Redesigning CI/CD and Governance for the New Development PaceGemini 4 Argon expands token limits and tops knowledge‑work benchmarks – what engineers need to knowData Agent Kit GA unlocks direct agent access to BigQuery Graph, Bigtable, and Spark for AI‑driven pipelinesRunning gcloud and bq via the Google Cloud CLI remote MCP server: practical implications for AI and platform engineersS3 Tables add full Iceberg V3 type support, deletion vectors, and row lineageAI‑First CI/CD Pivot at CloudBees Redefines Enterprise Pipeline PracticesMetadata Pre‑Filtering in Amazon S3 Vectors Improves Filtered Search RecallCommand Injection via Branch Name Exposes GitHub Token in AI Coding Agents
Kubernetes

Kubernetes DRA vs HAMi GPU Scheduling

AI SummaryPowered by AI

The introduction of Dynamic Resource Allocation (DRA) in Kubernetes v1.35 changes how fractional GPUs are requested, but it does not render the popular project HAMi obsolete for container-level enforcement.

For cloud engineers and AI practitioners managing high-performance computing clusters on Kubernetes, resource scheduling has long been a complex balancing act between efficiency and isolation. Historically, sharing expensive GPU hardware required workarounds around standard APIs because device plugins could only count whole cards as available units: one card or none. This binary limitation made it impossible to express nuanced requirements like allocating 80% of memory for an inference workload while reserving the remainder.

Evolution from Binary Allocation

  • The original Device Plugin API treated GPUs as atomic resources, forcing users to accept entire cards regardless of actual utilization needs.
This rigidity drove projects like HAMi (Hardware Accelerated Memory Isolation) into existence. Accepted by the CNCF Technical Oversight Committee in 2026, HAMI built a sophisticated pipeline involving mutating webhooks and scheduler extenders specifically designed to express fractional requests that standard Kubernetes vocabulary could not support.

Dynamic Resource Allocation Integration

Kubernetes DRA (Dynamically Requested Allocations) reached general availability in version 1.34, with native enablement by default starting at v1.35 via the consumable capacity feature. This shift allows pods to request slices of device memory directly from the scheduler without relying on annotations or external plugins for basic allocation logic.

Persistence and Enforcement Layers

While DRA handles fractional requests natively, it lacks a specific design focus: enforcing those fractions inside containers at CUDA-call granularity. This distinction is critical because standard scheduling does not prevent processes from consuming more memory than requested once the container starts running without additional constraints.

HAMI's response to this architectural shift has been strategic rather than reactive. The project split its responsibilities accordingly, retaining enforcement capabilities while rebuilding encoding logic on top of DRA across three distinct repositories for better integration with modern Kubernetes versions.

Certification Relevance and Exam Scenarios

Kubernetes certifications (CKA or CKS), understanding the difference between scheduler-level resource requests and runtime enforcement is vital. While DRA simplifies manifest definitions, candidates must still understand how to configure device plugins alongside native features.

What This Means For You

Kubernetes DRA (Dynamically Requested Allocations)-to focusing on the enforcement layer. Engineers should leverage native scheduling for initial resource assignment but rely on specialized tools like updated versions of HAMi to ensure strict memory limits within containers.

Originally published atCNCF