Microsoft has added a new decision‑making service, Microsoft‑Decision‑1, to its Foundry platform. The model is a post‑trained version of Alibaba’s Qwen3.5‑9B, priced at $0.042 per million input tokens with free output, and is already being used internally for incident handling, experiment scoring, and feedback analysis. Engineers should treat the launch as a low‑cost, high‑throughput option for structured decision workloads, while keeping an eye on its integration constraints, calibration behavior, and the announced plan to rebase the model on Microsoft’s own MAI and OpenAI models.
What Changed: Microsoft‑Decision‑1 on Foundry
Three days after OpenAI opened its Decisions API to public beta, Microsoft released Decision‑1 on the Foundry marketplace. Unlike the OpenAI‑based offerings, Decision‑1 is built on the Qwen3.5‑9B checkpoint from Alibaba. The service accepts up to 32,768 tokens of plain text and returns a JSON payload; it does not support image inputs, a capability present in OpenAI’s Decisions API and Cloudflare’s Clef model. Pricing matches TypeSafe’s Jev at $0.042 per million input tokens, with no charge for output tokens. Microsoft’s internal rollout includes Xbox Research, the Copilot team, on‑call engineers, and the Discovery group, each reporting latency improvements (14× faster than GPT‑6 Sol on Xbox) and higher consistency (46× on Discovery).
Implications for Architecture and Operations
Decision‑1’s token limit and JSON‑only response shape how it can be wired into existing pipelines. Engineers must design request‑building logic to stay within the 32 k token ceiling and parse the JSON result without expecting multimodal data. The model is offered through Foundry, so integration follows the platform’s authentication and billing mechanisms; there is no indication that the service is open‑source or that weights are publicly available.
- Cost model: At $0.042 per million input tokens, decision calls are inexpensive, but the revenue per call remains low, suggesting Microsoft’s primary goal is to keep downstream generative traffic within Azure.
- API compatibility: The Foundry sample invokes a
/systemoneendpoint, but Microsoft has not confirmed full compliance with the System One API adopted by competitors such as AWS and Upstage. - Performance expectations: Internal benchmarks claim significant latency and consistency gains, but external validation is pending.
Security and Reliability Considerations
Microsoft advertises calibrated probabilities, meaning a 90 % confidence score should be correct nine times out of ten on representative data. However, a pre‑print study of a similar model (Jev) showed that short, natural‑sounding context changes can flip correct decisions in over 60 % of cases. Microsoft’s own documentation advises customers to validate calibration on their own datasets, indicating that the model’s confidence may drift under adversarial or noisy inputs.
Because the service runs inside the Foundry environment, any data sent to Decision‑1 is subject to the platform’s security posture, but the lack of open weights or a public licensing model limits external auditability. Practitioners should treat the model as a trusted component within Azure but still enforce input sanitization and monitor for unexpected probability shifts.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Engineers evaluating Decision‑1 should:
- Prototype the model against existing decision workloads to verify latency and cost claims within the 32 k token limit.
- Implement calibration checks on representative data sets, especially if the model influences automated routing or on‑device vs. cloud execution decisions.
- Confirm whether the
/systemoneendpoint meets the System One API contract required by existing tooling. - Plan for the announced rebasing to MAI and OpenAI models, which may affect performance, pricing, or compatibility.
- Monitor Microsoft communications for any changes to weight availability, licensing, or security certifications that could impact compliance requirements.



