Over the past three weeks a de‑facto standard for decision‑model services has taken shape. TypeSafe’s System One schema – a JSON contract that defines three question types (choice, score, and noul) and returns a probability distribution plus a confidence metric – is now implemented by AWS, Upstage, Ollama, Strands Decider, and a growing set of open‑weight models. Practitioners can point a client at any endpoint that respects the schema (for example /v1/systemone) and receive interchangeable, structured answers without rewriting business logic.
Standardized Schema, Divergent Endpoints
The core contract is identical across implementations, but the network address varies. TypeSafe introduced the contract with its Jev service; AWS’s Strands Decider 2B, built on Alibaba’s Qwen3.5, serves it at the same path. Upstage, Ollama (v0.35), and community projects such as Open‑Jev expose identical request shapes while hosting the API under their own domains. A second wave of adopters keeps the semantics but publishes a different URL – OpenRouter’s native Decisions API, Venice’s beta endpoint, and Perplexity’s Decisions API all accept the same three question types but use distinct paths. The result is code‑level portability: swapping the base URL in TypeSafe’s SDK points the application at a new provider, though providers differ in option limits, confidence calculations, and answer quality, so functional testing remains required.
Architectural Patterns Enabled by Decision Models
Three usage patterns have surfaced quickly:
- Routing to the appropriate model. OpenRouter’s Jev Router decides, per request, whether a cheap decision model can answer or whether a full‑blown LLM is needed, preserving expensive compute for high‑value work.
- Gating tool calls. Strands Decider provides a
before_tool_callhook that runs two yes/no questions before invoking an external tool (e.g., a weather API), ensuring the agent asks for missing context instead of guessing. - Fallback to LLMs. Maxim AI’s Bifrost gateway detects an unavailable decision service and proxies the request to an LLM via the provider’s Responses API, reshaping the LLM output to match the System One schema.
These patterns let engineers treat decision models as lightweight conditional branches – essentially an if statement with a calibrated probability – while keeping the LLM for generative tasks.
Operational and Security Considerations
Deploying a decision model introduces new operational knobs. Confidence values reflect distribution concentration, not correctness; teams must define thresholds based on observed calibration (AWS reports 0.9 confidence yields ~95 % accuracy on unseen short‑classification tasks). Monitoring should capture confidence trends and fallback rates to detect drift.
Because the same contract can be served by multiple third‑party endpoints, supply‑chain trust becomes a factor. Engineers should verify the provenance of hosted models (e.g., open‑weight models on Hugging Face) and consider network isolation or mutual TLS when calling external decision services. Switching providers is straightforward at the SDK level, but differences in rate limits, option caps, and confidence algorithms mean performance testing is essential before production rollout.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
Adopt the System One SDK and treat the endpoint URL as a configurable variable. Define confidence thresholds that align with your risk tolerance, and instrument fallback paths to an LLM for resilience. Validate each provider’s limits and confidence behavior in a staging environment before committing to a production endpoint. Finally, keep an eye on the emerging “endpoint war”: OpenAI’s preview Decisions API, Open Responses spec, and any new vendor paths may shift the balance of features, pricing, and support, influencing long‑term architecture decisions.


