DTarform
Menu

Home / Glossary / Retrieval-Augmented Generation

AI

Retrieval-Augmented Generation (RAG)

In one sentence

RAG fetches the passages of your own documents that relate to a question, hands them to a language model, and asks it to answer using only those — so the answer can cite where it came from.

Also called grounded generation. If you have ever asked a model a question about your own business and got a confident, entirely invented answer, RAG is the standard fix.

Why it exists

A language model knows what was in its training data. It does not know your pricing, your protocols, your contracts or anything that happened after training finished. Asked anyway, it will usually produce something plausible rather than admit the gap — the failure mode known as hallucination.

Retraining the model on your documents is expensive, slow, and has to be redone every time a document changes. RAG takes the other route: leave the model alone, and give it the right pages at the moment of the question.

How it works

Four steps, and only the first happens ahead of time.

  1. Index. Your documents are split into passages — see chunking — and each passage is converted into an embedding, a numeric fingerprint of its meaning. These are stored in a vector database.
  2. Retrieve. A question is converted the same way, and the passages closest in meaning are pulled back. Good systems combine this with keyword search, because exact terms like a product code still matter.
  3. Augment. Those passages are placed into the model's context window alongside the question and an instruction to answer only from what is supplied.
  4. Generate. The model answers, and the system attaches the source passages as citations, so a reader can verify the claim rather than trust it.

When to use it

RAG is the right default when the answer lives in documents that change, and when being able to check the source matters.

  • Support and internal help desks answering from current policy.
  • Clinical, legal or regulatory questions where a citation is non-negotiable.
  • Engineering and design archives where the answer exists but nobody can find it.
  • Anything where a wrong answer has a real cost and a human needs to verify quickly.

RAG or fine-tuning?

They solve different problems and are frequently confused in vendor pitches.

Use RAG when the model needs to know facts that change — documents, records, prices, policies. Updating means re-indexing a file, not retraining anything.

Use fine-tuning when the model needs to adopt a behaviour — a house tone of voice, a strict output format, a narrow classification task. Fine-tuning teaches style and shape, not current facts.

Plenty of production systems use both: fine-tuned for format, RAG for content.

What it costs

Three cost lines, and only one of them is the model.

  • Indexing — a one-off per document, plus a small recurring cost as documents change.
  • Storage — the vector database. Usually the smallest line, and often over-specified.
  • Inference — charged per token, and this is the one that scales with usage. Retrieved passages count toward it, so retrieving ten passages when three would do makes every question permanently more expensive.

The single biggest cost mistake we see is retrieving too much context to compensate for weak retrieval. Better chunking is cheaper than a bigger context window.

Where it goes wrong

  • Bad chunking. Passages split mid-table or mid-clause retrieve well and read as nonsense.
  • Stale index. The document was updated in March; the index was not. The system now cites an obsolete policy with full confidence.
  • No permission model. Retrieval that ignores access control will happily surface an HR file to the wrong person. Permissions belong in the retrieval step, not in a disclaimer.
  • No “I don’t know”. If retrieval finds nothing relevant, the system must say so. Otherwise you have reintroduced the problem you bought RAG to fix.
  • Prompt injection. A retrieved document can itself contain instructions aimed at the model. Treat retrieved text as data, never as instructions.

How we build it

Retrieval quality is measured before anything is shown to a user — we test whether the correct passage is returned for a set of real questions your team writes. If retrieval is wrong, no amount of model quality will save the answer.

Deployments run inside your environment, permissions are enforced at retrieval, and every answer is logged with the passages it saw, so a decision can be reconstructed months later.

Thinking about RAG on your own documents?

Two weeks of discovery tells you whether your content is ready, what retrieval quality to expect, and what it will cost to run.

Book an AI Discovery