What to know

  • RAG supplies information at answer time rather than retraining a model.
  • Retrieval quality, permissions and document freshness affect the final answer.
  • A citation needs to support the claim, not merely resemble the question.

A library beside the model

Retrieval-augmented generation, usually called RAG, adds a search step before a model produces an answer. The application finds relevant material, places it in the model’s context and asks the model to respond using that material. The approach is especially useful for documents that change more often than a model is trained.

A company handbook illustrates the distinction. The model may know generally how parental leave works, but it cannot infer the current policy of a particular employer. Retrieval can provide that policy. It cannot make an outdated handbook accurate or resolve a contradiction between two approved versions.

The retrieval pipeline is the product

Cloudflare’s technical explanation describes ingestion, embeddings, storage, retrieval and generation as separate stages. Treating them separately helps teams diagnose failures. A poor answer may come from missing documents, an unsuitable split between passages, an overly broad search or the model’s interpretation of otherwise correct evidence.

For a support service, keep product version, publication date and document owner alongside each passage. A technically relevant answer for an older version may be actively unhelpful for a current customer. Metadata filters can be more valuable than a larger collection of semantically similar documents.

Permissions belong before the answer

A retrieval system should not fetch a restricted payroll file and then rely on the model to conceal it. Access control needs to constrain what a particular user can retrieve. This becomes more complicated when a search index combines material from multiple repositories with different permission models.

Test the same question under several identities. A manager and a contractor may have access to different documents, even when the wording of their requests is identical. Permission changes should also reach the index quickly enough that revoked access does not persist through an old cached representation.

Citations are evidence to check

A generated citation can create confidence without proving anything. Check whether the cited passage actually establishes the particular statement. A document about refunds may be relevant to a question while offering no support for the amount or deadline in the answer.

Evaluate retrieval and generation independently. First measure whether the correct evidence appears in the candidate results. Then measure whether the answer stays within that evidence, handles conflicting sources and admits when information is missing. Include questions that genuinely have no answer in the corpus.

When a simpler solution wins

RAG is not required for every application. A deterministic lookup may be better for a shipping status, tax code or account balance. Structured data often deserves a structured query, with a language model helping interpret the request rather than inventing the result.

Start with a small, maintained corpus and a set of real questions. Measure the time needed to correct the data and repair retrieval errors. The promise of grounded AI depends on this ordinary information work as much as it depends on the model.

Sources & further reading

  1. Cloudflare: Building retrieval-augmented generation
  2. Liu et al.: Lost in the Middle

Factual statements are grounded in the linked material. Interpretation and illustrative examples are Byte Watchr analysis. Vendor claims are identified as claims, rather than independent testing.

This article belongs to Byte Watchr’s launch collection. The edition date organizes evergreen coverage and does not imply historical publication. Actual publication is recorded above.

Corrections policy · About this byline