Live
GitHub scheduled code scanning now waits for code changes before running weekly scansNew Cloudflare WAF Rule Blocks Citrix NetScaler ADC/Gateway Input Validation Flaw (CVE‑2026‑88771)Workers OAuth split API reaches v1: separate auth and resource Workers with Service BindingRethinking AI Factory Design: Productivity, Durability, and Fungibility for EngineersBridging the Kubernetes Ownership Gap After Day 2Edge Decision Models on Workers AI: Clef and Clef‑Flash Enable Fast Structured InferenceEvent‑Driven Ambient Agents on Amazon Bedrock AgentCore: A Serverless PatternIntegrating Amazon S3 Vectors as a Persistent Memory Backend for NVIDIA NeMo Agent ToolkitGitHub scheduled code scanning now waits for code changes before running weekly scansNew Cloudflare WAF Rule Blocks Citrix NetScaler ADC/Gateway Input Validation Flaw (CVE‑2026‑88771)Workers OAuth split API reaches v1: separate auth and resource Workers with Service BindingRethinking AI Factory Design: Productivity, Durability, and Fungibility for EngineersBridging the Kubernetes Ownership Gap After Day 2Edge Decision Models on Workers AI: Clef and Clef‑Flash Enable Fast Structured InferenceEvent‑Driven Ambient Agents on Amazon Bedrock AgentCore: A Serverless PatternIntegrating Amazon S3 Vectors as a Persistent Memory Backend for NVIDIA NeMo Agent Toolkit
Cloudflare

Edge Decision Models on Workers AI: Clef and Clef‑Flash Enable Fast Structured Inference

AI SummaryPowered by AI

Cloudflare has added two open‑source decision models, Clef and Clef‑Flash, to the Workers AI platform, providing structured probability outputs instead of free‑form text. The models run on edge GPUs, deliver sub‑200 ms latency, and can be fine‑tuned via a new RL service, giving engineers a low‑latency alternative for routing, blocking, or escalation logic.

Cloudflare has added two open‑source decision models, @cf/cloudflare/clef and @cf/cloudflare/clef-flash, to the Workers AI platform, delivering structured probability outputs that can be consumed directly by edge services. Engineers care because the models run on Cloudflare’s edge GPUs, return results in tens of milliseconds, and expose a reinforcement‑learning fine‑tuning workflow, offering a low‑latency alternative to text‑based LLM calls for routing, blocking, or escalation decisions.

Clef and Clef‑Flash: Structured Decision at the Edge

Both models belong to the decision‑model family: they accept a state description and a set of typed questions, then emit a probability for each permitted answer. The output is a deterministic data structure, eliminating the need for downstream parsing or token‑by‑token waiting. The models are released under the Apache 2.0 license on Hugging Face, and the code paths are compatible with Typesafe’s Jev model, allowing a drop‑in replacement where Jev is already used.

Performance Compared to Existing Decision Models

Benchmark runs show a clear latency advantage. Median response times are 209 ms for clef, 38.8 ms for clef‑flash, and 524 ms for Jev. At the 95th percentile, the numbers rise to 238.6 ms, 122.4 ms, and 536 ms respectively. In a set of ten decision benchmarks, Clef achieved the top score on seven, surpassing Jev and other open decision models. Highlights include:

  • BFCL (case exact): Clef 98.47, Clef‑Flash 98.76, Jev 95.75
  • BANKING77 (macro‑F1): Clef 94.20, Clef‑Flash 90.93, Jev 79.74
  • CLINC150+OOS (macro‑F1): Clef 97.43, Clef‑Flash 66.77, Jev 89.27
  • Home appliances (case exact): Clef 82.95, Clef‑Flash 97.73, Jev 52.27

On Typesafe’s workflow evaluations, Clef outperformed Jev in three of four domains—invoice processing, customer service, and security incidents—demonstrating practical relevance across common enterprise use cases.

Operational Implications on Workers AI

Deploying the models on Workers AI means inference runs on GPUs distributed across Cloudflare’s edge network, keeping round‑trip latency low for end‑users. Practitioners can place a decision call directly in the request path of a Cloudflare Worker, then optionally forward to a generative LLM for follow‑up actions. The open‑source weights enable offline experimentation or custom deployment, while a newly announced reinforcement‑learning fine‑tuning service lets teams adapt the models to proprietary data sets. Signing up as a design partner is the only required step to access that service.

Related CloudNinjas coverage: AI engineering.

What This Means For Practitioners

Consider replacing text‑based LLM calls with clef or clef‑flash when you need deterministic, sub‑200 ms decisions at the edge. Evaluate latency and accuracy against your existing decision pipeline, and plan for a potential RL fine‑tuning phase if your workload deviates from the benchmark datasets. Keep an eye on the upcoming design‑partner program for early access to the fine‑tuning API, and monitor the Hugging Face model cards for any updates to performance or licensing.

Originally published atCloudflare Developer Platform