OpenAI’s new custom inference chip, Jalapeño, moved from design to silicon in nine months and now shows measurable gains in throughput, latency, and power efficiency on large language models. The chip’s architecture and the use of AI‑generated kernel code directly affect how engineers design, tune, and operate inference workloads.
Performance and Power Gains
OpenAI released benchmark data for Jalapeño using the InferenceX suite with GPT‑OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. Across these models the chip delivered 1.5–1.9× more work per watt while cutting end‑to‑end latency by 1.7–3.6×. For highly interactive workloads the speedup ranged from 2.1–4.1× compared with the reference systems. When measured at the previous best time‑between‑tokens point, Jalapeño achieved 8.6–104.3× more work per watt, depending on the model. The accelerator is rated at 700 W but never exceeded 550 W during testing, indicating a headroom between specification and actual draw.
AI‑Generated Kernel Optimizations
During development OpenAI ran its own models to explore design alternatives, shortening the tape‑out cycle to nine months. Selected attention and mixture‑of‑experts blocks that were generated by the company’s AI models ran 1.5–1.8× faster than hand‑written equivalents. While the full model speedup is not claimed, the result demonstrates that the chip’s programming model is simple enough for AI to produce performant implementations. Engineers should note that each new model family still requires dedicated kernels and optimizations, even if AI can accelerate the creation of those kernels.
Architectural Shifts for Agent Workloads
OpenAI emphasizes that agents invoke models repeatedly, so small per‑inference delays compound over a task. Jalapeño addresses this by keeping the KV cache local and by providing networking that keeps more of the workload within a single connected system. This reduces data movement between cores and chips, which traditionally adds waiting time during both the compute‑heavy prompt phase and the memory‑bandwidth‑heavy token‑generation phase. The design aims to improve throughput without the latency penalty that batch‑oriented accelerators often incur.
Operational Considerations
From an ops perspective the chip’s power envelope (rated 700 W, observed ≤550 W) and its latency characteristics affect capacity planning and cooling design. Because Jalapeño is intended to coexist with existing Nvidia accelerators, teams will need to evaluate scheduling policies that balance batch‑oriented and interactive workloads. The reliance on AI‑generated kernels introduces a verification step: while the generated code is faster, practitioners must still validate correctness and security before deployment.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
- Benchmark your own models against the reported work‑per‑watt and latency figures to decide if Jalapeño offers a net benefit over current accelerators.
- Plan for model‑specific kernel development; consider integrating AI‑assisted code generation to shorten the optimization cycle, but allocate time for testing and validation.
- Review data‑movement patterns in your inference pipelines; local KV caching and reduced inter‑chip traffic can lower tail latency for agent‑style workloads.
- Update capacity and power budgeting to reflect the chip’s observed draw (≤550 W) rather than its rated maximum.
- Monitor OpenAI’s upcoming generations for further architectural refinements and for guidance on mixed‑accelerator deployments.


