Live
Microsoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendAccess Cloudflare Skills Directly Through the API MCP ServerCodeQL 2.27.2 expands language models and tightens macOS build support – what engineers need to knowTangible Certification: Turning a Kubernetes Badge into a Gold NecklaceGoogle Data Cloud GA updates: agent‑centric tooling, hybrid Spanner, and expanded Lakehouse catalogMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendAccess Cloudflare Skills Directly Through the API MCP ServerCodeQL 2.27.2 expands language models and tightens macOS build support – what engineers need to knowTangible Certification: Turning a Kubernetes Badge into a Gold NecklaceGoogle Data Cloud GA updates: agent‑centric tooling, hybrid Spanner, and expanded Lakehouse catalog
Cloudflare

Edge-native multimodal decisions: Clef-omni adds audio/video to Workers AI, Clef-flash gets cheaper, and Clef inference speeds up

AI SummaryPowered by AI

Clef-omni is now available on Workers AI, allowing a single request to evaluate text, images, audio, and video, while Clef-flash pricing has been reduced and the original Clef model now runs faster. The changes simplify pipelines, lower costs, and tighten latency budgets for AI-driven services at the edge.

Workers AI now hosts @cf/cloudflare/clef-omni, a decision-model that accepts text, images, audio (WAV/MP3) and video (MP4/WebM) in a single call. At the same time Cloudflare announced a price cut for clef-flash and a speed improvement for the original clef model.

One request for every modality

Previously, a workflow that needed to reason over a voice recording or a video clip required chaining a transcription model, an audio-splitting step, and a text-decision model. clef-omni removes that choreography: the model ingests base64-encoded media in the images, audio, and videos fields and returns a scored answer for each supplied question.

The model is built on a 30 B-parameter mixture-of-experts backbone with 3 B active parameters. It does not generate free-form text; instead it scores a predefined set of answers in a single pass, which keeps inference latency low.

  • Text-only request: ~20 ms
  • Image or audio input: < 100 ms
  • 21-second video with sound: ~300 ms

Pricing and speed updates for the Clef family

Cloudflare reduced the cost of clef-flash to a level described as “less than Jev”. The original clef model also received a performance boost, though the exact magnitude is not quantified in the announcement. For teams that already use these models, the lower price and faster response time translate directly into lower operational spend and tighter service-level targets.

Operational considerations for Workers AI deployments

Because media are sent as base64 data URLs, request payloads can grow quickly. Practitioners should verify that their Workers AI edge functions respect the platform’s request-size limits and that any upstream API gateway or client can handle the increased bandwidth.

The shift to a single-call decision model reduces the number of secrets, IAM tokens, and network hops required to stitch together multiple AI services. Fewer moving parts generally mean a smaller attack surface, but the larger payloads still need proper validation to avoid malformed data causing runtime errors.

Running a 30 B-parameter MoE model with 3 B active parameters at the edge may increase memory consumption per invocation. Teams should monitor Workers AI instance metrics and adjust concurrency limits if they encounter out-of-memory throttling.

Related CloudNinjas coverage: AI engineering.

What This Means For Practitioners

  • Re‑architect pipelines that previously chained transcription, audio, and text models into a single clef-omni call to cut latency and simplify code.
  • Update budgeting models to reflect the lower price of clef-flash and the faster execution of clef.
  • Validate request size handling and memory usage in your edge functions before scaling production traffic.
  • Continue to monitor Cloudflare announcements for any further pricing or capability changes that could affect cost‑performance trade‑offs.
const response = await env.AI.run("@cf/cloudflare/clef-omni", {
  model: "clef-omni",
  state: "Review the installation: a photo, an audio clip, and a video of the fan.",
  images: ["data:image/png;base64,"],
  audio: ["data:audio/mpeg;base64,"],
  videos: ["data:video/mp4;base64,"],
  questions: {
    label_visible: { type: "noul", instructions: "Is the label visible?" },
    sounds_normal: { type: "noul", instructions: "Is the unit running smoothly?" },
    fan_running: { type: "noul", instructions: "Is the fan running?" }
  }
});
Originally published atCloudflare Developer Platform