Yann Torrent
Yann Torrent · CTO, co-founder View Yann Torrent's profile

RAG: preventing LLM hallucinations on business data

RAG: preventing LLM hallucinations on business data

RAG (Retrieval-Augmented Generation) has become the standard way to make an AI answer on your data: your procedures, your contracts, your product documentation, your FAQs. The principle fits in one sentence: you search your documents, pass the relevant passages to the LLM as context, and the model answers based on them.

It is the right architecture. But it does not eliminate hallucinations — it moves them, and makes them more insidious. A hallucinated answer about your data looks like an answer from your AI: your tone, your vocabulary, sometimes a reference to a passage that does not exist.

One error that circulates — a parameter applied wrongly, an invented clause, a nonexistent procedure — is enough for everyone to doubt the whole feature.

In this article: exactly where RAG hallucinations come from, how to measure them before going to production, and the concrete levers to limit them. If you first want a broader picture of LLM hallucination causes and consequences in general, read our article on LLM hallucinations.

RAG in 30 seconds

Two steps, two components:

  1. Retrieval: the user’s question is turned into a search in your index — your documents, chunked into passages, represented as semantic vectors. The closest passages are selected.
  2. Generation: the LLM receives the question and the retrieved passages, and produces an answer.

The essential point: the LLM does not “consult” your database. It reads an excerpt that someone selected for it. Anything that is not in that excerpt, it does not know. That is where everything happens.

Two failure points, two different fixes

The classic mistake is to treat “hallucination” as a single cause. In practice there are two, and they require different answers.

1. Retrieval failed: the answer is not in the context

The relevant passage was not retrieved: ambiguous question, wrong passage selected, document never ingested. Or more simply: the answer does not exist in your data — a question about a recent, undocumented decision, for example.

The LLM has no default “I don’t know”. Its objective is to generate the most plausible continuation — so it fills in the gap. This is precisely the most dangerous case: the fabricated answer is written in the style of your documents, so it looks like yours.

2. Retrieval succeeded: but the generation distorts it

The information is in the context. The model paraphrases it, mixes up two passages, or invents a detail: an article number, a date, a threshold, a name. On dense texts — legal, finance, HR — it is frequent.

Why? The objective function of an LLM is not faithfulness, it is the plausibility of the next token. Faithfulness to your context is not a model objective: it is something you impose, and something you must verify.

The reflex to adopt

Faced with a hallucination observed in production, the first question is not “change the model” or “improve the prompt”. It is: was the right information in the context?

  • If no, it is a retrieval problem.
  • If yes, it is a generation problem.

Mixing the two is the number one reason RAGs stall: you change the LLM while the search is broken, or the opposite.

How to measure hallucinations before go-live

Without measurement, you discover hallucinations through your users. And at that point, it is too late to quantify the problem, compare solutions, or commit to contractual guarantees.

1. An evaluation set first

Build 50 to 200 real business questions, with the expected answer known. Sources: your real FAQs, your recurring support tickets, your procedures.

And the point everyone forgets: include questions whose answer is not in your data. They are the ones that test refusal. A RAG tested only on questions it knows how to answer teaches you nothing about its real behavior.

2. The metrics to track

Four metrics are enough to start — they are the RAGAS standard:

  • Context relevance: did the search retrieve the right passage? It measures retrieval.
  • Faithfulness (groundedness): is every claim in the answer supported by the retrieved context? It measures generation.
  • Answer relevance: does the answer address the question asked?
  • Refusal rate: on questions with no answer in the data, did the AI refuse or invent? This is the most important one.

3. How to compute them

  • LLM-as-judge: an LLM — often a stronger model — evaluates each answer against the context, automatically and reproducibly. This is what makes the evaluation set scalable.
  • Human sample: manually re-read 10 to 20% of the cases, refusals and edge cases first. This is your ground truth.
  • On every change: re-run the set when you change the chunking, the index, the model or the prompt. The before/after comparison is what makes the work measurable.

4. In production

  • The rate of answers without a cited source: if it is not zero, your pipeline is broken by definition.
  • User feedback “wrong answer”.
  • Monthly audits on a sample of real conversations.

The metric that ultimately matters is not a percentage: it is zero critical error — an error that leads to a wrong action (a payment, a contract, an HR decision). That is not a measurement target, it is a design target.

The four levers to limit hallucinations

1. Retrieval first — it is most of the problem

  • Hybrid search: semantic alone fails on keyword questions (procedure numbers, clause references). You need lexical and semantic, together.
  • Structure-respecting chunking: a clause must not be cut in two, a table must stay whole. You split according to the document structure, not the size.
  • Business metadata: document type, version, effective date, scope. The search can then filter on them.
  • Reranking: the initial search is noisy. A rerank selects the 3 to 5 passages that matter.

2. Constrain the generation

Here, the prompt is a contract, not a suggestion:

  • “Answer only from the provided context.”
  • “If the answer is not in the context, say so. Do not invent.”
  • Cite the passages used.

And the point of view that changes everything: refusing is a feature. An AI that says “I do not have this information in our documents, here is where to look” is an AI you can use. An AI that always answers is an AI you must re-read — so an useless AI.

3. Cite the sources — verifiability

Every answer must point to the exact passage that supports it. This is not a UX detail:

  • The user verifies in five seconds. Trust is built on the verifiable, not on the guarantee.
  • Your support team identifies the origin of an answer in one minute, instead of re-reading it with a magnifying glass.
  • An answer without a source is not an answer: it is the model talking.

That is the very architecture of our Chercher feature: every answer is sourced, and the chain validates faithfulness before returning it.

4. Control the critical perimeter

  • On sensitive actions (HR, contracts, payments): human validation before the action. The model never acts alone.
  • Full traceability: question, retrieved context, answer, model and version. Auditable, reproducible.
  • Monitoring: a drop in faithfulness is almost always a documents problem — poorly updated, poorly ingested — not an AI problem. The alert must tell you “check your data”.

And a chatbot without your data, what is that?

An LLM without RAG knows neither your rules, nor your versions, nor your history. On your questions, it answers with the average of what it saw during training. And that average, for you, is wrong by construction.

RAG changes the error profile: you move from a plausible generic answer to a sourced, verifiable one. That is the difference between an AI feature to show in a demo, and a feature you can put in a contract.

Key takeaways

  • Two distinct causes: retrieval does not bring the right information, or generation distorts the information it brought. Identify which one before touching anything.
  • Measure before go-live: evaluation set with no-answer questions, faithfulness, refusal rate. Re-run on every change.
  • Limit via retrieval first: hybrid, structural chunking, metadata, reranking.
  • Refusing is a feature, citing sources is an architecture, and the critical perimeter stays validated by a human.

Retrieval is the most fragile link of a RAG — and the one that depends most on your business context. On the Chercher page, see how Agora builds a document search where every answer is sourced and verified.

Bring AI into your software with Agora Software.

Let's talk