Cloudflare’s AI Search service is now generally available, and the platform begins usage‑based billing on 1 November 2026. The rollout introduces hybrid search as the default indexing mode, bundles Workers AI embedding and reranking costs into the AI Search price, adds official support for multimodal embedding models, and makes OCR available on all accounts with revised file‑size limits. Engineers and operators need to adjust cost monitoring, index configuration, data ingestion pipelines, and observability practices to accommodate these shifts.
Billing and Included Usage
From the start of November, AI Search will be charged on a usage‑based model. The pricing tier includes a monthly allowance for data ingestion, storage, semantic vector queries, and traditional full‑text queries. Cloudflare will send a reminder email a week before the first bill is generated, giving teams a chance to verify that the included usage aligns with expected traffic.
Hybrid Search as the Default Index
New AI Search instances now run hybrid search out of the box. Hybrid search blends semantic vector retrieval with conventional full‑text matching, offering broader relevance without additional configuration. When creating an instance, teams can still select an alternative indexing method, but the default change means existing automation that assumes a pure‑text index may need to be reviewed.
Workers AI Embeddings, Reranking, and Multimodal Models
Calls to Workers AI for embedding generation and result reranking that originate from AI Search are now covered by the AI Search pricing tier. Consequently, these calls no longer appear on the separate Workers AI bill or in AI Gateway logs, simplifying cost attribution but also removing a visibility point for those monitoring Workers AI usage.
AI Search also officially supports two multimodal embedding models: @cf/qwen/qwen3-vl-embedding-2b and google-ai-studio/gemini-embedding-2. The REST API and public endpoint accept image data alongside text, enabling image‑augmented search and chat scenarios without additional service integration.
OCR Availability and Revised File Limits
Optical character recognition is enabled for every account, allowing scanned PDFs to be indexed. File‑size limits have been adjusted: plain‑text files, code files, and PDFs with OCR can be up to 10 MiB, while PDFs without OCR and other supported formats remain capped at 4 MiB. This change may affect batch ingestion jobs and storage planning.
Related CloudNinjas coverage: AI engineering.
What This Means For Practitioners
- Cost tracking: Update budgeting dashboards to reflect the November billing start and the bundled usage allowances; watch for any residual Workers AI calls that remain outside AI Search.
- Index strategy: Review automated instance creation scripts to confirm they either accept the hybrid default or explicitly set a different index type if required.
- Observability: Adjust logging and alerting to account for the disappearance of embedding and reranking entries from Workers AI logs; consider instrumenting AI Search‑specific metrics instead.
- Data pipeline: Validate that ingestion processes respect the new 10 MiB limit for OCR‑enabled PDFs and that OCR is enabled where needed.
- Feature adoption: Evaluate whether the supported multimodal models meet your use‑case requirements and test image‑based queries before rolling them into production.


