AI21 migrated from ad‑hoc Slack‑based GPU capacity requests to Google Cloud AI Hypercomputer, running its training workloads on a shared GKE cluster of thousands of A3 and A3 Ultra GPU nodes. The change eliminated manual scheduling steps and cut high‑priority job wait times from 72 hours to 12 hours.
GPU cluster scheduling challenges
When utilization approaches 100 %, two problems emerge. First, contention: multiple teams compete for the same GPUs, requiring human arbitration. Second, fragmentation: free GPUs are scattered across nodes in small chunks that cannot satisfy a large, multi‑node job, creating a bin‑packing deadlock.
AI Hypercomputer solution
AI Hypercomputer provides a managed orchestration layer on top of GKE that aggregates the entire GPU fleet into a single scheduling domain. By evaluating open‑source batch schedulers such as YuniKorn, Volcano, and Kueue, the team selected a solution that integrates with the managed service and can admit jobs only when sufficient contiguous capacity exists, thereby removing the need for manual negotiation.
Operational implications
- Capacity is pooled across the whole cluster, allowing any team to draw on the full fleet rather than a fixed slice.
- Job admission is automated, reducing weekly manual interventions from ~20 to zero.
- Utilization remains high because the scheduler can pack jobs tightly and release fragmented resources.
- Monitoring must shift from individual request queues to cluster‑wide metrics such as GPU fragmentation, pending job count, and scheduler latency.
Security considerations
Running many teams on a shared GPU cluster introduces a multi‑tenant surface. Practitioners should verify that namespace isolation, RBAC, and GKE node security policies are correctly configured to prevent cross‑team data leakage. The managed service inherits Google Cloud’s underlying security controls, but teams remain responsible for securing their workloads, secrets, and access to the GPU nodes.
Related CloudNinjas coverage: Google Cloud.
What This Means For Practitioners
- Adopt a cluster‑wide scheduler that can enforce all‑or‑nothing admission for large GPU jobs to avoid fragmentation deadlocks.
- Instrument GPU fragmentation metrics and set alerts for prolonged pending jobs.
- Review GKE namespace and RBAC configurations when moving to a shared GPU pool.
- Evaluate the trade‑off between dedicated slices and a pooled model based on your workload’s predictability and latency requirements.


