What to know
- AI Search has moved to general availability.
- The release adds native visual retrieval and PDF OCR.
- Cloudflare says billing begins November 1.
Retrieval extends beyond extracted text
Cloudflare made AI Search generally available on October 1 and expanded its handling of images and scanned documents. The release adds native image embeddings, optical character recognition for PDFs and support for larger files. Cloudflare says billing starts on November 1, with a free tier continuing across Workers plans.
The company’s earlier image workflow relied on generating a caption and embedding its text. The newer approach can also represent the pixels themselves. That change matters when a useful detail is visible in an image but absent from a caption, such as a diagram’s spatial arrangement or a small difference between product photographs.
Better retrieval is not the same as a verified answer
The practical test is whether the system retrieves the evidence a user needs. A visually similar image can still show a different product revision. A scanned document can contain an OCR error in a date or amount. If an application uses the result to generate an answer, that next step needs to preserve the distinction between a possible match and a confirmed fact.
A representative evaluation should therefore include difficult examples rather than only obvious matches. Teams could test similar-looking diagrams with different labels, scans containing tables and records in which the relevant fact appears in a footnote. The expected answer should identify both the correct source and the evidence inside it.
Access controls require a separate check. A successful search should return only material the requesting user is allowed to see. Restricting the final answer is insufficient if an unauthorized document has already entered a model’s context or an application log. Permission changes and document deletion also need to propagate to the searchable index.
Operational costs arrive with the lifecycle
Byte Watchr’s analysis is that managed retrieval can reduce integration work, but it cannot decide which documents deserve to be trusted. Organizations still need ownership of source collections, rules for versioning and a process for replacing obsolete files. Without that work, a search service can deliver outdated evidence more efficiently.
Cost measurement should cover the lifecycle of a collection: initial ingestion, updates, storage and queries. A dataset that changes frequently may behave differently from a mostly static archive. Visual material and large scanned files should be included in a trial if they are part of the intended workload, rather than added after a text-only cost estimate has been accepted.
The release broadens the types of material developers can search through one service. Its value will depend on retrieval quality and the controls surrounding that service. Teams preparing for the billing change have a useful opportunity to establish both before a pilot becomes a permanent dependency.
Sources & further reading
Factual statements are grounded in the linked material. Interpretation and illustrative examples are Byte Watchr analysis. Vendor claims are identified as claims, rather than independent testing.
The event date records the source announcement or documented operation. The coverage edition groups recent developments and is separate from the publication date. Actual publication is recorded above.
Corrections policy · About this byline

