GPU allocation for AI workloads has shifted from isolated, per‑team clusters to a shared pool that multiple teams can draw from simultaneously. Kubernetes now offers primitives such as Dynamic Resource Allocation (DRA) and a growing ecosystem of CNCF‑aligned projects that make fine‑grained sharing possible while preserving isolation, which directly impacts utilization, cost, and security for engineers.
What Changed in GPU Allocation
Historically, the device‑plugin model required a pod to request nvidia.com/gpu: 1, which pinned an entire accelerator even if the workload used only a fraction of its capacity. The introduction of DRA in Kubernetes 1.34 allows the scheduler to treat GPUs as rich devices with attributes like memory and topology. DRA alone does not split a GPU, but it enables the underlying layers—NVIDIA MIG, HAMi, or time‑slicing solutions—to expose fractional resources to the scheduler. This change removes the primary bottleneck of low utilization caused by whole‑GPU pinning.
Key Building Blocks for a Shared AI Factory
Implementing a shared GPU pool involves several layers, each backed by open‑source or CNCF projects:
- Hardware lifecycle: Provision bare‑metal servers with
Metal3,Ironic, orTinkerbell, and record inventory inNetBoxusing Redfish for remote management. - Cluster lifecycle: Create versioned clusters via
Cluster APIand apply GitOps withArgo CDorFlux. NVIDIA AICR can be used for GPU‑specific extensions. - Node inventory: Detect GPU capabilities, MIG profiles, and network topology with
Node Feature Discoveryand the NVIDIA/AMD GPU & Network Operators. - Tenant isolation: Deploy per‑team virtual clusters using
vClusteror sandboxed runtimes to keep workloads separate on the same hardware. - GPU allocation: Combine DRA with MIG, HAMi, or time‑slicing, and schedule with topology‑aware schedulers such as
KAI Scheduler,Volcano, orKueue. - Inference serving: Run models behind APIs with
vLLM,KServe, orllm‑d. Batch and HPC jobs can be submitted viaSLURMon Kubernetes usingSlinky. - VM support: Offer full VMs per tenant with
KubeVirtwhen isolation at the VM level is required. - Networking: Use
Cilium,Multus, and SR‑IOV or RDMA to provide high‑performance, tenant‑aware networking. - Storage: Persist datasets and checkpoints via CSI drivers backed by
Rook/Ceph, parallel filesystems, or object storage. - Observability: Track utilization and health with
Prometheus,OpenTelemetry, and theDCGM exporter. - Identity & policy: Enforce authentication and authorization with
Keycloak,ResourceQuota/Kueue, and policy engines likeKyvernoorOPA. - Secrets & supply chain: Secure runtime secrets via
OpenBaoandExternal Secrets, and scan images withFalcoandTrivy. - Reliability: Detect node failures using
DCGM health checksandNode Problem Detector, and automate drain/cordon actions. - Self‑service & billing: Expose provisioning APIs through
OpenTofuor GitOps, and measure GPU‑seconds withOpenCostandDCGM.
Operational and Security Implications
Sharing GPUs forces a rethink of isolation boundaries. While dedicated clusters guarantee physical separation, they waste capacity. Virtual clusters (vCluster) and sandboxed runtimes provide logical isolation, but they rely on correct policy enforcement (e.g., ResourceQuota, Kyverno) and accurate inventory data. Mis‑configured quotas can lead to noisy‑neighbor effects, and insufficient network segmentation may expose traffic between tenants. Observability becomes critical: without fine‑grained metrics from DCGM exporter, it is hard to detect over‑commit or contention. The hardware lifecycle also gains importance; automated burn‑in tests and inventory validation (via NetBox) reduce the risk of deploying faulty GPUs that could corrupt workloads.
Related CloudNinjas coverage: hands-on guides.
What This Means For Practitioners
Adopt DRA as soon as your cluster version supports it and pair it with a fractional allocation mechanism such as MIG. Deploy virtual clusters or sandboxed runtimes to keep teams isolated while maximizing utilization. Integrate inventory tools (Node Feature Discovery, NetBox) early to ensure accurate scheduling decisions. Harden policy enforcement with Kyverno or OPA and monitor GPU health and usage continuously. Finally, evaluate billing pipelines (OpenCost) to translate improved utilization into cost visibility for stakeholders.

