Live
Self‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendSelf‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model Backend
Cloudflare

AI Search adds six Workers AI models with expanded context windows

AI SummaryPowered by AI

AI Search now includes six new Workers AI text‑generation models with context windows up to 1 M tokens. This gives engineers more flexibility for large‑context workloads while removing the need to manage external provider credentials.

AI Search now supports six additional Workers AI text‑generation models, each offering context windows from 128 k up to 1 M tokens. The change expands the toolbox for engineers building search‑enhanced applications while eliminating the need to manage separate provider credentials.

New model roster and context capabilities

The added models are:

  • @cf/deepseek-ai/deepseek-v4-flash-0731 – 1,048,576 token window
  • @cf/deepseek-ai/deepseek-v4-pro-0813 – 1,048,576 token window
  • @cf/openai/gpt-oss-120b – 128,000 token window
  • @cf/openai/gpt-oss-20b – 128,000 token window
  • @cf/qwen/qwen3.8-27b – 262,144 token window
  • @cf/moonshotai/kimi-k2.7-code – 262,144 token window

All models execute on the Workers AI runtime, meaning they are hosted at the edge and share the same deployment model as other Workers AI workloads.

Operational impact

Selection of a model occurs when creating or updating an AI Search instance, either through the Cloudflare dashboard or via the API. Because the models run on Workers AI, no external provider key is required, simplifying secret management and deployment pipelines. Teams should verify that the chosen model’s context size aligns with their query or document length requirements and test latency characteristics in their target region.

Security considerations

Removing the need for an additional provider key reduces the surface area for credential leakage. However, the data sent to any text‑generation model remains under the same data‑handling policies that apply to Workers AI, so practitioners must continue to treat model output as untrusted and apply appropriate validation before downstream use.

Related CloudNinjas coverage: AI engineering.

What This Means For Practitioners

  • Update AI Search configurations to include the new models where larger context windows are beneficial.
  • Audit deployment scripts to drop any now‑unnecessary provider key handling.
  • Run functional tests to confirm that model responses meet quality and latency expectations for your workload.
  • Monitor the official supported‑models list for future additions or deprecations that could affect your architecture.
Originally published atCloudflare Developer Platform