How to Measure and Prevent LLM Hallucinations Before You Ship
Define hallucination for launch decisions, test wrong facts and real-time claims, limits of prompt tuning, RAG grounding, refusal behavior, hallucination pressure packs, and honest Release Gate calls.
Hallucination is not a academic curiosity. It is a launch risk when your app states false facts confidently, cites sources that do not support the answer, or guesses when it should refuse. Agnostics measures that failure mode under pressure and ties it to a ship or fix Release Gate.
Define hallucination for launch decisions, not papers
Research papers debate fine-grained definitions of hallucination. Launch teams need a practical one: the model asserts something your product should not stand behind, and the UI presents it as reliable.
That includes invented policies, wrong numbers, fake citations, answers about events that did not happen, and confident replies when retrieval returned nothing useful.
Not every imprecise paraphrase is a launch blocker. Align your definition with product promises. A creative writing assistant tolerates fiction. A refund bot does not.
Write the definition in your release policy so engineering, product, and support share the same bar for Fix, Monitor, and Ready.
Wrong facts, stale facts, and real-time claims
Static wrong facts are easy to picture: inventing a warranty period, citing a section that does not exist, mixing two products into one spec.
Stale facts are harder because the model may cite a real document that is out of date. The answer is technically grounded in retrieval but still wrong for launch.
Real-time claims are the sharp edge: stock prices, outage status, election results, medical dosing, legal deadlines. If your product implies freshness, measure whether the app refuses, qualifies, or bluffs when it cannot know.
Hallucination pressure testing should include questions where the honest behavior is refusal or deferral, not a clever guess.
- Fabrication: claims with no supporting source
- Overreach: real sources stretched beyond what they say
- Staleness: grounded in old docs that are no longer true
- False freshness: authoritative tone on unknowable real-time facts
Why prompt tuning alone does not prevent hallucinations
Longer system prompts can reduce obvious fabrication on demo questions. They do not guarantee refusal when retrieval is empty or when users pressure the model to guess.
Prompt changes fix one transcript while leaving the pattern open. Attackers and confused users vary phrasing. Without adversarial reruns on the same target, you are shipping on hope.
Useful fixes combine policy prompts with architecture: citation requirements, confidence thresholds, retrieval gates, and deterministic checks on high-risk actions outside the model.
Blocking one "always answer" instruction in a prompt does not stop the model from guessing under social pressure. Test behavior, not prompt prose.
RAG grounding: when retrieval helps and when it hides risk
RAG reduces some hallucinations by giving the model text to cite. It introduces others when ranking pulls the wrong chunk, when poisoned docs read as authoritative, or when the model cites correctly but overstates.
Grounding is not binary. Evaluate whether citations support the final claim, what happens on empty retrieval, and how the app behaves when two sources conflict.
Pair RAG evaluation with retrieval drift and hallucination pressure packs. Fixes may live in chunk boundaries, rerankers, citation UI, or refusal rules, not only in the system prompt.
When to refuse vs when to answer
Products differ on silence. Internal search may prefer a best-effort guess with low confidence. Customer support bots may need hard refusals when policy chunks are missing.
Decide explicitly per target type. Encode expected behavior in tests: empty retrieval should trigger refusal, clarifying questions, or human handoff, not fabrication.
UX copy matters. If the interface looks authoritative, users treat guesses as facts. Your hallucination bar should match how answers are presented, not only model text.
- Define expected behavior when retrieval is empty or low confidence
- Define expected behavior for real-time or high-stakes factual domains
- Test adversarial pressure to guess anyway
- Record waivers in the Release Gate when stakeholders accept risk
Hallucination pressure pack: what it tests
The hallucination pressure pack sends questions designed to elicit confident wrong answers: missing context, trick phrasing, requests for unknowable real-time facts, and pressure to cite sources that were not retrieved.
Findings include reproduction steps, severity, and evidence showing what the model said versus what sources support.
Use it on any target where user trust depends on factual accuracy: doc search, support bots, compliance assistants, and APIs that downstream systems treat as authoritative.
Why happy-path evals miss hallucinations
Golden datasets often include answerable questions with reference text in the corpus. They under-sample empty retrieval, conflicting sources, and adversarial pressure to invent.
Human review of ten demo transcripts feels reassuring. It does not scale to the combinatorial space of user phrasing and corpus drift.
Hallucination is a failure mode under stress. Measure it with adversarial coverage on your configured target, not only with cooperative QA scripts.
A high average score on curated questions is not evidence your app refuses honestly when it should.
How to read hallucination findings
Read evidence, not summaries. Compare retrieved snippets, displayed citations, and final answers line by line.
Group findings by pattern: empty retrieval guesses, citation overreach, stale doc wins, real-time bluffing. One pattern, one fix priority.
Severity reflects launch impact. Invented refund rules on a customer widget outrank awkward wording on an internal brainstorm tool.
Fix patterns, not individual wrong answers
Patching single bad replies in a log does not change system behavior. Fix retrieval gates, citation requirements, refusal templates, and human escalation paths.
For RAG, tune chunk boundaries and conflict rules when overreach repeats on the same doc family. For plain chatbots, tighten scope and refusal when questions exceed product knowledge.
Ship fixes on staging with the same target configuration as production. Retest with hallucination pressure coverage before moving the Release Gate.
- Identify the failure pattern from finding evidence
- Ship architectural or policy fixes on a staging target
- Launch retests from the original finding or scan
- Update Release Gate only after verified improvement
Release Gate calls for hallucination risk
The Release Gate translates findings into Ready, Monitor, Fix, or Blocked for this target and this scan based on your policy.
Critical fabrication on launch-critical flows often lands in Fix or Blocked until retest passes. Monitor can be valid for low-severity drift when stakeholders accept documented gaps.
Do not claim zero hallucinations. Claim you measured behavior under pressure, fixed patterns, retested, and made an informed launch call with evidence.
Retest after fixes and corpus drift
Prompt edits and reranker tweaks can fix one message while leaving the pattern intact. Retests replay the same attack coverage so improved means verified.
Corpus drift reopens hallucination paths without code changes. Rescan after material document updates, connector changes, or model swaps.
Document corpus version, packs run, gate state, and accepted risks in your release report. Future launches will reuse the same target with new content.
Hallucinations and red teaming overlap
Hallucination often pairs with injection and jailbreak. A model that guesses may also follow hostile instructions from retrieved text. Run broader red team packs when scope and grounding both matter.
LLM red teaming on your configured target catches compound failures: guess plus tool call, guess plus policy override, guess plus leaked context.
What Agnostics does not claim
Agnostics does not promise zero hallucinations after launch. Models guess under pressure. Corpora drift. Users ask unknowable questions.
Agnostics does measure whether your app fails your stated bar under adversarial coverage before users encounter the same patterns.
Sample Demo Data in the interactive demo illustrates findings and Release Gate states without live scans on your endpoints.
Tell users what the product knows and when it refuses. Tests should enforce that promise, not a fantasy of perfect accuracy.
Questions
What counts as an LLM hallucination for launch?
For launch decisions, treat hallucination as confident claims your product should not stand behind: invented policies, wrong numbers, fake or overextended citations, and answers when retrieval did not support a reply. Align severity with your product promises and UI tone.
Can prompt engineering eliminate hallucinations?
Prompt tuning reduces some obvious guesses on demo questions but does not guarantee refusal under adversarial pressure or empty retrieval. Combine prompts with retrieval gates, citation rules, escalation paths, and adversarial testing on your configured target.
How does RAG affect hallucination risk?
RAG can ground answers in sources but also enables overreach, stale citations, and confident wrong synthesis from mismatched chunks. Evaluate retrieval and generation together and run retrieval drift plus hallucination pressure packs.
When should an LLM app refuse to answer?
When retrieval is empty or low confidence, when questions require unknowable real-time facts, or when sources conflict without a policy rule to choose. Define expected refusal behavior per target and test it under hallucination pressure.
What is the hallucination pressure attack pack?
It sends adversarial scenarios designed to elicit confident wrong answers, missing citations, and guesses when honest behavior is refusal. Findings include reproduction evidence and severity for Release Gate decisions.
Does a Ready Release Gate mean no hallucinations?
No. It means this scan with these packs met your release policy bar for this target at this time. Retest after material prompt, retrieval, model, or corpus changes because hallucination risk is not static.
How often should we retest for hallucinations?
Retest after fixes to grounding or refusal logic and before major launches. Rescan when the corpus or model changes materially. Treat repeated fabrication patterns in support tickets as signals to extend staging coverage.