AI Search now supports six additional Workers AI text‑generation models, each offering context windows from 128 k up to 1 M tokens. The change expands the toolbox for engineers building search‑enhanced applications while eliminating the need to manage separate provider credentials.
New model roster and context capabilities
The added models are:
@cf/deepseek-ai/deepseek-v4-flash-0731–1,048,576token window@cf/deepseek-ai/deepseek-v4-pro-0813–1,048,576token window@cf/openai/gpt-oss-120b–128,000token window@cf/openai/gpt-oss-20b–128,000token window@cf/qwen/qwen3.8-27b–262,144token window@cf/moonshotai/kimi-k2.7-code–262,144token window
All models execute on the Workers AI runtime, meaning they are hosted at the edge and share the same deployment model as other Workers AI workloads.
Operational impact
Selection of a model occurs when creating or updating an AI Search instance, either through the Cloudflare dashboard or via the API. Because the models run on Workers AI, no external provider key is required, simplifying secret management and deployment pipelines. Teams should verify that the chosen model’s context size aligns with their query or document length requirements and test latency characteristics in their target region.
Security considerations
Removing the need for an additional provider key reduces the surface area for credential leakage. However, the data sent to any text‑generation model remains under the same data‑handling policies that apply to Workers AI, so practitioners must continue to treat model output as untrusted and apply appropriate validation before downstream use.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
- Update AI Search configurations to include the new models where larger context windows are beneficial.
- Audit deployment scripts to drop any now‑unnecessary provider key handling.
- Run functional tests to confirm that model responses meet quality and latency expectations for your workload.
- Monitor the official supported‑models list for future additions or deprecations that could affect your architecture.

