DoorDash has replaced its vendor‑first approach with an internal open‑weight GenAI platform, moving from external LLM services to a mix of open‑weight models and custom agent gateways. The shift impacts anyone who builds, runs, or secures AI workloads because it changes where latency, cost, and accuracy are managed and introduces new operational surfaces.
Architectural Shifts
The new platform centralises LLM access behind a gateway that can route requests to either vendor‑hosted models or internally hosted open‑weight models. This dual‑path design lets teams choose the most appropriate model for a given workload while keeping a consistent API surface for downstream services.
Implementation Considerations
Adopting open‑weight models requires teams to provision compute, storage, and networking resources that were previously abstracted away by vendors. Engineers must therefore plan for model loading times, hardware acceleration, and the lifecycle of model artefacts. The gateway also becomes a point where request shaping (e.g., throttling, payload validation) can be applied before reaching the model.
Operational Implications
Balancing accuracy, latency, and cost across 5,000 internal users introduces monitoring challenges. Metrics around inference latency, error rates, and compute spend need to be collected per model and per gateway route. Alerting must differentiate between vendor‑side outages and internal resource saturation.
Security and Governance Outlook
Running open‑weight models internally expands the attack surface: model artefacts, inference pipelines, and the gateway must be protected from tampering and data leakage. While the source does not detail specific controls, practitioners should consider isolation of model workloads, audit logging of gateway traffic, and validation of model provenance.
Related CloudNinjas coverage: DevOps.
What This Means For Practitioners
Teams should evaluate their current reliance on vendor LLMs, identify workloads that could benefit from open‑weight alternatives, and plan for the required infrastructure and monitoring changes. Early focus on gateway design, cost tracking, and security hygiene will smooth the transition and keep the platform responsive for a large internal user base.
