CoreWeave has introduced NVIDIA Vera Rubin NVL72 GPU systems with Spectrum‑X 102.4 T Ethernet networking to its cloud platform and announced the forthcoming availability of NVIDIA Vera CPUs, a processor built specifically for agentic AI workloads. The change gives AI engineers, platform teams, and security operators a newer hardware tier for production inference, a high‑density CPU option for thousands of isolated agent sandboxes, and a unified environment—CoreWeave Forge—to keep training, evaluation, and observability tightly coupled.
New GPU Generation for Agentic Inference
The Vera Rubin NVL72 GPUs replace earlier Volta‑based V100 hardware that CoreWeave has been running for nearly a decade. Cognition, the lab behind the Devin AI software engineer, is the first production customer on the new GPUs. In benchmark tests against a GB200 NVL72 baseline, Cognition measured up to 4.8× higher token throughput for a software‑engineering inference workload (SWE‑2). The higher throughput translates directly into faster real‑time code generation and lower per‑token cost for multi‑step reasoning tasks.
CPU‑Scale Sandboxes for Agent Isolation
NVIDIA Vera CPUs are designed to host large numbers of concurrent, hardware‑isolated agent environments. A single rack houses 128 CPUs and 11,264 cores, supporting more than 11,000 concurrent environments at one core each. CoreWeave reports more than 3× faster sandbox startup times compared with prior CPU generations, and a 1.7× performance gain on the Terminal‑Bench suite across all passing tasks. Sandboxes run alongside training jobs, leveraging Spectrum‑X Ethernet switches and BlueField‑4 DPUs to provide low‑latency, secure communication between agents and the underlying infrastructure.
CoreWeave Forge: Closing the Training‑Inference Loop
CoreWeave Forge bundles several tools—Weights & Biases, OpenPipe post‑training expertise, and the open‑source marimo notebook—into a single environment that spans model training, evaluation, and production monitoring. New services include:
- CoreWeave ARIA (generally available) that analyzes runs, proposes experiments, recommends code changes, and pushes updates to GitHub.
- CoreWeave Agent Lens (new) that converts production agent traces into actionable insights, improving failure detection by 20 % and reducing fix cost by 50 %.
- CoreWeave Sandboxes (generally available) that let users run agents, tool calls, reinforcement‑learning loops, and evaluations in isolated CPU or GPU environments.
All services are accessible through CoreWeave Kubernetes Service, SUNK, Mission Control, and the Sandboxes UI, keeping the stack consistent from token serving to model iteration.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Engineers should consider moving high‑throughput inference workloads—especially those with long contexts and high concurrency—to Vera Rubin NVL72 to capture the reported token‑throughput gains. Teams building large fleets of autonomous agents can evaluate Vera CPUs for sandbox density, using the documented 3× faster startup and 11 k+ concurrent environment capacity to reduce provisioning latency. The integrated observability provided by Agent Lens offers a concrete path to cut failure‑detection cycles and operational spend, while ARIA’s experiment‑suggestion loop can shorten the feedback cycle between production behavior and next‑generation training runs.
Next steps include testing the Vera Rubin GPUs in a staging Kubernetes cluster, profiling sandbox startup latency on Vera CPUs, and enabling CoreWeave Forge services to align training pipelines with production metrics. Keep an eye on the GA timeline for the Vera CPU offering and on any updates to Spectrum‑X networking configurations that may affect bandwidth planning for token‑heavy workloads.


