Workers AI now hosts @cf/cloudflare/clef-omni, a decision-model that accepts text, images, audio (WAV/MP3) and video (MP4/WebM) in a single call. At the same time Cloudflare announced a price cut for clef-flash and a speed improvement for the original clef model.
One request for every modality
Previously, a workflow that needed to reason over a voice recording or a video clip required chaining a transcription model, an audio-splitting step, and a text-decision model. clef-omni removes that choreography: the model ingests base64-encoded media in the images, audio, and videos fields and returns a scored answer for each supplied question.
The model is built on a 30 B-parameter mixture-of-experts backbone with 3 B active parameters. It does not generate free-form text; instead it scores a predefined set of answers in a single pass, which keeps inference latency low.
- Text-only request: ~20 ms
- Image or audio input: < 100 ms
- 21-second video with sound: ~300 ms
Pricing and speed updates for the Clef family
Cloudflare reduced the cost of clef-flash to a level described as “less than Jev”. The original clef model also received a performance boost, though the exact magnitude is not quantified in the announcement. For teams that already use these models, the lower price and faster response time translate directly into lower operational spend and tighter service-level targets.
Operational considerations for Workers AI deployments
Because media are sent as base64 data URLs, request payloads can grow quickly. Practitioners should verify that their Workers AI edge functions respect the platform’s request-size limits and that any upstream API gateway or client can handle the increased bandwidth.
The shift to a single-call decision model reduces the number of secrets, IAM tokens, and network hops required to stitch together multiple AI services. Fewer moving parts generally mean a smaller attack surface, but the larger payloads still need proper validation to avoid malformed data causing runtime errors.
Running a 30 B-parameter MoE model with 3 B active parameters at the edge may increase memory consumption per invocation. Teams should monitor Workers AI instance metrics and adjust concurrency limits if they encounter out-of-memory throttling.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
- Re‑architect pipelines that previously chained transcription, audio, and text models into a single
clef-omnicall to cut latency and simplify code. - Update budgeting models to reflect the lower price of
clef-flashand the faster execution ofclef. - Validate request size handling and memory usage in your edge functions before scaling production traffic.
- Continue to monitor Cloudflare announcements for any further pricing or capability changes that could affect cost‑performance trade‑offs.
const response = await env.AI.run("@cf/cloudflare/clef-omni", {
model: "clef-omni",
state: "Review the installation: a photo, an audio clip, and a video of the fan.",
images: ["data:image/png;base64,"],
audio: ["data:audio/mpeg;base64,"],
videos: ["data:video/mp4;base64,"],
questions: {
label_visible: { type: "noul", instructions: "Is the label visible?" },
sounds_normal: { type: "noul", instructions: "Is the unit running smoothly?" },
fan_running: { type: "noul", instructions: "Is the fan running?" }
}
});


