Evaluating RAG Pipelines: Retrieval, Generation, and Release Testing

How to evaluate RAG in two steps (retrieval then generation), why happy-path evals miss failures, attack packs for drift and injection, and Release Gate decisions after corpus changes.

RAG is two jobs pretending to be one: fetch the right context, then answer without overstepping it. Agnostics evaluates both hops on your configured target, records findings when retrieval or generation fails, and gives you a ship or fix call before users trust a wrong citation.

RAG evaluation is two steps, not one score

Most RAG failures are misdiagnosed because teams collapse retrieval and generation into a single pass or fail metric. That hides whether the break happened in ranking, chunk boundaries, citation logic, or the model filling gaps.

Step one is retrieval evaluation: did the system fetch chunks that actually support an answer to this question? Step two is generation evaluation: given what was fetched, did the model stay inside those bounds, refuse when grounding is thin, and cite honestly?

Split testing makes fixes faster. Reranker tweaks address retrieval findings. Prompt and refusal rules address generation findings. Mixing them produces prompt changes that mask ranking bugs and vice versa.

Agnostics attack packs pressure both hops on your live target so findings name the failing stage with evidence, not guesswork.

Retrieval metrics that matter for launch

Academic retrieval metrics help research comparisons. Launch teams need metrics tied to harm: wrong policy retrieved, stale doc wins over current policy, hostile chunk ranks first, empty retrieval still produces an answer.

Track whether cited passages support the final claim line by line. Track behavior when two sources disagree. Track what happens when the user phrasing pulls a header match without supporting body text.

Precision on a golden set is not enough if adversarial phrasing never appears in that set. Pressure ranking with near-duplicate topics, time-sensitive policies, and questions designed to pull misleading chunks.

Generation metrics that matter for launch

Generation evaluation asks whether the model respects retrieved boundaries. Paraphrase is fine until it invents policy, blends incompatible chunks, or cites a source that was never retrieved.

Tone matters for launch bars. A hesitant wrong answer is still a finding if your UX presents it as authoritative. Customer-facing doc search needs stricter generation rules than internal exploratory search.

Refusal behavior is a generation metric too. When retrieval is empty or weak, does the app say it does not know, ask a clarifying question, or fabricate? Your release policy should encode the expected behavior.

Would you ship if the model answered confidently with no supporting chunk? If not, measure refusal under pressure, not only on easy questions.

A practical RAG metrics table for release reviewers

Release reviewers need a short table they can scan without reading fifty eval notebooks. Group metrics by hop, severity, and whether a failure blocks launch under your policy.

Retrieval rows cover ranking mistakes, stale wins, poisoned fetches, and empty context. Generation rows cover overreach, fabrication, bad citations, and missing refusals. Security rows cover instruction conflicts from retrieved text and injection through user messages.

The table is not a substitute for findings with reproduction evidence. It is how you communicate coverage to product and compliance stakeholders who will not read raw scan logs.

Why happy-path RAG evals miss real failures

Golden question sets are curated to be answerable from your corpus. Attackers and confused users are not curated. They ask with wrong keywords, mixed languages, fake urgency, and instructions embedded in uploaded files.

Offline evals often score final text against reference answers. They skip retrieval ranking under pressure, instruction layering at runtime, and the model synthesizing policy that appears nowhere in sources.

Indexing QA proves documents are reachable. It does not prove hostile documents stay inert or that the model refuses when chunks do not support a reply.

Agnostics attack packs for RAG pipelines

Retrieval drift pack pressures answers that sound grounded but misread, overextend, or mis-cite retrieved text. Use it on launch-critical knowledge domains first: refunds, security, pricing, medical boundaries, legal disclaimers.

Hallucination pressure pack stresses behavior when retrieval is thin or empty. It finds confident fabrication, fake section references, and answers that ignore missing context.

Prompt injection pack covers direct overrides and indirect injection through retrieved content. RAG apps fail indirect paths constantly because retrieval imports text the model treats as instructions.

Add boundary bypass when strict topic scope is a promise. Add sensitive context exposure when the corpus includes HR, account, or internal runbook material.

End-to-end scans vs split hop testing

End-to-end scanning attacks the full pipeline as users experience it. That is what you need for Release Gate decisions because failures emerge from the interaction of retrieval, prompts, tools, and models.

Split hop testing helps debugging after you have a finding. Inspect what ranked, what chunks arrived, what the model saw, and what it said. Fix the hop that failed, then rerun end-to-end coverage to catch regressions.

Do not optimize split metrics in isolation while ignoring end-to-end behavior. A perfect reranker still ships wrong if generation invents when chunks disagree.

Ship decisions come from end-to-end findings on your configured target. Split views are for engineering diagnosis.

How to read RAG evaluation findings

Skim summaries and you will ship the wrong fix. Read evidence: the question, retrieved snippets, citations shown to the user, and final model output.

Group findings by pattern. Three drift findings on refund wording are one chunking or citation fix, not three unrelated tickets.

Severity should reflect launch impact and user trust. Wrong refund policy on a customer widget is not the same bar as a awkward paraphrase on an internal wiki search.

Hand engineering reproduction steps that include retrieval state, not just the final answer text.

Release Gate decisions for RAG pipelines

The Release Gate summarizes whether to ship, monitor, fix, or block this target for this scan based on your release policy.

RAG apps often land in Monitor or Fix when launch-critical docs show drift or injection even if casual questions look fine. Align policy with product promises about citations, scope, and refusal.

Document accepted risks: corpus version tested, packs run, findings waived, and retest dates. Future you will forget why a green light was granted.

Retest after corpus changes, not after hope

Corpus drift is continuous. New uploads, scraped pages, connector syncs, and chunking strategy changes all move retrieval behavior without a code deploy.

Retest with the same attack packs after material corpus or prompt changes. Launch retests from the original finding or scan so improved means verified under the same coverage.

Treat retest outcomes honestly: resolved, still open, improved, regressed, or could not verify. Only then move the Release Gate from Fix to Ready.

Hallucinations in RAG: evaluation overlap

Hallucination in RAG often masquerades as retrieval success. The model cites a real document but claims a clause that is not there, or blends two chunks into a rule neither supports.

Evaluate hallucination as both a generation failure and a retrieval ranking failure. Empty retrieval that still produces an authoritative answer is hallucination pressure. Populated retrieval with overreach is drift.

Pair retrieval drift and hallucination pressure packs on launch-critical targets. Read the dedicated hallucination guide when refusal behavior is part of your product promise.

Run your first RAG pipeline evaluation

Pick one launch-critical RAG target. Name three never events: wrong policy, leaked internal doc, obeying instructions from an uploaded PDF.

Run retrieval drift, prompt injection, and hallucination pressure packs. Read findings like a release reviewer. Fix patterns. Retest. Then widen coverage.

That loop turns RAG evaluation from a research exercise into a release habit your team can repeat every sprint.

Questions

How do you evaluate a RAG pipeline before launch?

Evaluate retrieval and generation separately, then confirm end-to-end behavior on a configured target. Run attack packs for retrieval drift, prompt injection, and hallucination pressure, read findings with evidence, fix patterns, retest, and record a Release Gate decision.

What is the difference between retrieval and generation evaluation?

Retrieval evaluation asks whether the right context was fetched. Generation evaluation asks whether the model stayed inside that context, cited honestly, and refused when grounding was thin. Failures in each hop need different fixes.

Which RAG metrics matter most for release decisions?

Focus on launch impact: citation fidelity, conflict handling between sources, behavior on empty retrieval, adversarial ranking to wrong chunks, and instruction conflicts from retrieved text. Academic averages help research; these metrics help ship or fix calls.

Why do offline RAG evals miss security failures?

Golden sets rarely include poisoned documents, adversarial phrasing, or instruction overrides in retrieved text. Offline scores measure expected answers, not failure behavior under attack on your live target path.

When should we retest a RAG app after corpus changes?

Retest after new connectors, uploads, scrape updates, chunking changes, or prompt edits that touch grounding rules. Also retest before major launches even if application code did not change, because corpus drift is continuous.

Does Agnostics replace RAG quality evals?

No. Quality evals help regression on known questions. Agnostics adds adversarial runtime testing tied to findings and a Release Gate. Use both: evals for expected behavior, scans for failure behavior before launch.

What attack packs should RAG teams run first?

Start with retrieval drift, prompt injection, and hallucination pressure on your launch-critical target. Add boundary bypass for strict scope and sensitive context exposure when the corpus includes privileged material.