At a joint Microsoft‑NVIDIA event, the companies announced that Windows PCs will now host full‑stack NVIDIA AI hardware—branded RTX Spark—and an OS‑level container runtime called Microsoft Execution Containers (MXC) for running AI agents locally. The change means developers can move large language models and multi‑agent workloads onto laptops, compact desktops, or a new DGX Station for Windows without relying on cloud clusters.
Hardware and software stack for on‑device agents
RTX Spark integrates a Blackwell RTX GPU with up to 6,144 cores and a Grace CPU with up to 20 cores, linked by a 600 GB/s interconnect. The platform offers up to 128 GB of unified memory and a peak of one petaflop of FP4 AI performance. It runs the standard NVIDIA CUDA stack, so existing CUDA‑based code, libraries, and toolchains can be used unchanged. The hardware is being shipped in Surface Laptop Ultra devices and in a range of OEM laptops and compact desktops, with pre‑orders opening immediately.
For enterprise‑grade workloads, NVIDIA unveiled DGX Station for Windows, built on the GB300 Grace‑Blackwell Ultra Superchip. It provides 748 GB of coherent memory and up to 20 petaFLOPS of FP4 AI compute, enough to run trillion‑parameter models locally. This is the first deskside AI supercomputer that runs natively on Windows, eliminating the need for a separate Linux environment for heavy AI workloads.
Operational impact for development and deployment
Because RTX Spark and DGX Station expose the same CUDA APIs as NVIDIA’s data‑center GPUs, developers can keep a single code base from laptop to server. CI/CD pipelines can be extended to include on‑device testing, but they now need to provision machines with the specific GPU/CPU configuration and ensure driver and CUDA version alignment. Model developers gain the ability to iterate on large models locally, reducing latency and cost associated with cloud‑based training or inference.
Always‑on agents, a core use case highlighted by Microsoft, will run continuously in the background. Teams should plan for power, thermal, and firmware management on devices that may be used for both productivity and high‑throughput AI workloads. Monitoring tools must be extended to capture GPU utilization, memory pressure, and container health in real time.
Security considerations with MXC
MXC introduces an OS‑level container abstraction that isolates agents from the rest of the system while allowing the OS to observe and govern their behavior. Practitioners should evaluate MXC’s isolation guarantees, integrate its lifecycle hooks into existing security tooling, and define policies for data access, persistence, and network egress. Because agents can run persistently, audit logging and runtime integrity checks become essential to detect drift or malicious behavior.
Windows’ existing security stack (e.g., Windows Security, Agent 365) is being extended to manage MXC workloads, but organizations will need to map their compliance requirements onto the new primitives. This may involve configuring Windows policies to restrict hardware access, defining container resource quotas, and ensuring that any secrets used by agents are stored in approved vaults rather than embedded in the container image.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
- Assess whether your AI workloads fit within the RTX Spark memory and compute envelope; if not, consider DGX Station for Windows.
- Update build pipelines to target the CUDA version shipped with RTX Spark and verify driver compatibility on Windows.
- Incorporate MXC container images into your deployment workflow, treating them as a distinct security boundary that requires policy definition and monitoring.
- Plan for continuous‑operation considerations: power budgeting, thermal throttling, and firmware updates on devices that will host always‑on agents.
- Leverage Windows’ extended security controls to audit agent activity, enforce data‑handling policies, and integrate with existing SIEM solutions.


