Evaluating RAG Pipelines: Retrieval, Generation, and Release Testing
How to evaluate RAG in two steps (retrieval then generation), why happy-path evals miss failures, attack packs for drift and injection, and Release Gate decisions after corpus changes.
RAG is two jobs pretending to be one: fetch the right context, then answer without overstepping it. Agnostics evaluates both hops on your configured target, records findings when retrieval or generation fails, and gives you a ship or fix call before users trust a wrong citation.
RAG evaluation is two steps, not one score
Most RAG failures are misdiagnosed because teams collapse retrieval and generation into a single pass or fail metric. That hides whether the break happened in ranking, chunk boundaries, citation logic, or the model filling gaps.
Step one is retrieval evaluation: did the system fetch chunks that actually support an answer to this question? Step two is generation evaluation: given what was fetched, did the model stay inside those bounds, refuse when grounding is thin, and cite honestly?
Split testing makes fixes faster. Reranker tweaks address retrieval findings. Prompt and refusal rules address generation findings. Mixing them produces prompt changes that mask ranking bugs and vice versa.
Agnostics attack packs pressure both hops on your live target so findings name the failing stage with evidence, not guesswork.
Retrieval metrics that matter for launch
Academic retrieval metrics help research comparisons. Launch teams need metrics tied to harm: wrong policy retrieved, stale doc wins over current policy, hostile chunk ranks first, empty retrieval still produces an answer.
Track whether cited passages support the final claim line by line. Track behavior when two sources disagree. Track what happens when the user phrasing pulls a header match without supporting body text.
Precision on a golden set is not enough if adversarial phrasing never appears in that set. Pressure ranking with near-duplicate topics, time-sensitive policies, and questions designed to pull misleading chunks.
- Citation fidelity: does the cited text support the answer?
- Conflict handling: which source wins when docs disagree?
- Empty or low-confidence retrieval: does the app refuse or bluff?
- Adversarial ranking: can phrasing pull the wrong chunk on purpose?
Generation metrics that matter for launch
Generation evaluation asks whether the model respects retrieved boundaries. Paraphrase is fine until it invents policy, blends incompatible chunks, or cites a source that was never retrieved.
Tone matters for launch bars. A hesitant wrong answer is still a finding if your UX presents it as authoritative. Customer-facing doc search needs stricter generation rules than internal exploratory search.
Refusal behavior is a generation metric too. When retrieval is empty or weak, does the app say it does not know, ask a clarifying question, or fabricate? Your release policy should encode the expected behavior.
Would you ship if the model answered confidently with no supporting chunk? If not, measure refusal under pressure, not only on easy questions.
A practical RAG metrics table for release reviewers
Release reviewers need a short table they can scan without reading fifty eval notebooks. Group metrics by hop, severity, and whether a failure blocks launch under your policy.
Retrieval rows cover ranking mistakes, stale wins, poisoned fetches, and empty context. Generation rows cover overreach, fabrication, bad citations, and missing refusals. Security rows cover instruction conflicts from retrieved text and injection through user messages.
The table is not a substitute for findings with reproduction evidence. It is how you communicate coverage to product and compliance stakeholders who will not read raw scan logs.
Why happy-path RAG evals miss real failures
Golden question sets are curated to be answerable from your corpus. Attackers and confused users are not curated. They ask with wrong keywords, mixed languages, fake urgency, and instructions embedded in uploaded files.
Offline evals often score final text against reference answers. They skip retrieval ranking under pressure, instruction layering at runtime, and the model synthesizing policy that appears nowhere in sources.
Indexing QA proves documents are reachable. It does not prove hostile documents stay inert or that the model refuses when chunks do not support a reply.
- Golden sets rarely include adversarial phrasing or poisoned documents
- Single-number pass rates hide which hop failed
- Staging corpora trimmed for demos hide upload and scrape risks
- Eval regression does not connect to a ship or fix Release Gate
Agnostics attack packs for RAG pipelines
Retrieval drift pack pressures answers that sound grounded but misread, overextend, or mis-cite retrieved text. Use it on launch-critical knowledge domains first: refunds, security, pricing, medical boundaries, legal disclaimers.
Hallucination pressure pack stresses behavior when retrieval is thin or empty. It finds confident fabrication, fake section references, and answers that ignore missing context.
Prompt injection pack covers direct overrides and indirect injection through retrieved content. RAG apps fail indirect paths constantly because retrieval imports text the model treats as instructions.
Add boundary bypass when strict topic scope is a promise. Add sensitive context exposure when the corpus includes HR, account, or internal runbook material.
End-to-end scans vs split hop testing
End-to-end scanning attacks the full pipeline as users experience it. That is what you need for Release Gate decisions because failures emerge from the interaction of retrieval, prompts, tools, and models.
Split hop testing helps debugging after you have a finding. Inspect what ranked, what chunks arrived, what the model saw, and what it said. Fix the hop that failed, then rerun end-to-end coverage to catch regressions.
Do not optimize split metrics in isolation while ignoring end-to-end behavior. A perfect reranker still ships wrong if generation invents when chunks disagree.
Ship decisions come from end-to-end findings on your configured target. Split views are for engineering diagnosis.
How to read RAG evaluation findings
Skim summaries and you will ship the wrong fix. Read evidence: the question, retrieved snippets, citations shown to the user, and final model output.
Group findings by pattern. Three drift findings on refund wording are one chunking or citation fix, not three unrelated tickets.
Severity should reflect launch impact and user trust. Wrong refund policy on a customer widget is not the same bar as a awkward paraphrase on an internal wiki search.
Hand engineering reproduction steps that include retrieval state, not just the final answer text.
Release Gate decisions for RAG pipelines
The Release Gate summarizes whether to ship, monitor, fix, or block this target for this scan based on your release policy.
RAG apps often land in Monitor or Fix when launch-critical docs show drift or injection even if casual questions look fine. Align policy with product promises about citations, scope, and refusal.
Document accepted risks: corpus version tested, packs run, findings waived, and retest dates. Future you will forget why a green light was granted.
Retest after corpus changes, not after hope
Corpus drift is continuous. New uploads, scraped pages, connector syncs, and chunking strategy changes all move retrieval behavior without a code deploy.
Retest with the same attack packs after material corpus or prompt changes. Launch retests from the original finding or scan so improved means verified under the same coverage.
Treat retest outcomes honestly: resolved, still open, improved, regressed, or could not verify. Only then move the Release Gate from Fix to Ready.
- Record corpus version and source list with each scan
- Fix retrieval or generation patterns on staging mirroring production
- Launch retests from original findings or full scan
- Update Release Gate and release report with retest evidence
Hallucinations in RAG: evaluation overlap
Hallucination in RAG often masquerades as retrieval success. The model cites a real document but claims a clause that is not there, or blends two chunks into a rule neither supports.
Evaluate hallucination as both a generation failure and a retrieval ranking failure. Empty retrieval that still produces an authoritative answer is hallucination pressure. Populated retrieval with overreach is drift.
Pair retrieval drift and hallucination pressure packs on launch-critical targets. Read the dedicated hallucination guide when refusal behavior is part of your product promise.
Run your first RAG pipeline evaluation
Pick one launch-critical RAG target. Name three never events: wrong policy, leaked internal doc, obeying instructions from an uploaded PDF.
Run retrieval drift, prompt injection, and hallucination pressure packs. Read findings like a release reviewer. Fix patterns. Retest. Then widen coverage.
That loop turns RAG evaluation from a research exercise into a release habit your team can repeat every sprint.
Questions
How do you evaluate a RAG pipeline before launch?
Evaluate retrieval and generation separately, then confirm end-to-end behavior on a configured target. Run attack packs for retrieval drift, prompt injection, and hallucination pressure, read findings with evidence, fix patterns, retest, and record a Release Gate decision.
What is the difference between retrieval and generation evaluation?
Retrieval evaluation asks whether the right context was fetched. Generation evaluation asks whether the model stayed inside that context, cited honestly, and refused when grounding was thin. Failures in each hop need different fixes.
Which RAG metrics matter most for release decisions?
Focus on launch impact: citation fidelity, conflict handling between sources, behavior on empty retrieval, adversarial ranking to wrong chunks, and instruction conflicts from retrieved text. Academic averages help research; these metrics help ship or fix calls.
Why do offline RAG evals miss security failures?
Golden sets rarely include poisoned documents, adversarial phrasing, or instruction overrides in retrieved text. Offline scores measure expected answers, not failure behavior under attack on your live target path.
When should we retest a RAG app after corpus changes?
Retest after new connectors, uploads, scrape updates, chunking changes, or prompt edits that touch grounding rules. Also retest before major launches even if application code did not change, because corpus drift is continuous.
Does Agnostics replace RAG quality evals?
No. Quality evals help regression on known questions. Agnostics adds adversarial runtime testing tied to findings and a Release Gate. Use both: evals for expected behavior, scans for failure behavior before launch.
What attack packs should RAG teams run first?
Start with retrieval drift, prompt injection, and hallucination pressure on your launch-critical target. Add boundary bypass for strict scope and sensitive context exposure when the corpus includes privileged material.