Prompt Injection: The Complete Pre-Launch Testing Guide

Direct vs indirect prompt injection, obfuscation, data exfiltration, tool compromise, prevention layers, and why no single fix is enough. How Agnostics tests injection on configured targets with a Release Gate.

Prompt injection is the class of failures where untrusted text steers your AI away from the job you meant it to do. Agnostics runs focused injection coverage on your live target path, records reproducible findings, and turns them into a ship or fix Release Gate decision.

Prompt injection as a launch failure class

Prompt injection is not a single exploit string shared on social media. It is a failure class: untrusted text enters context and the stack treats it as instructions or authoritative facts.

The untrusted text may arrive through a user message, a retrieved document, an email body, a ticket comment, a web page your agent fetched, or a tool result passed to the next turn.

Launch impact ranges from verbal policy breaks to operational harm when tools execute hostile plans. Injection plus tools plus retrieval is worse than any layer alone.

Pre-launch programs should test injection on the configured target you ship, not only on isolated prompt strings in a doc.

Direct vs indirect injection

Direct injection puts hostile instructions in channels you already label as user input: chat widgets, API bodies, form fields, and query parameters.

Indirect injection hides instructions where the app looks for facts: PDFs in the knowledge base, competitor web pages, support ticket paste-ins, OCR text from uploads, and metadata fields retrieval reads.

RAG apps fail indirect injection constantly because retrieval is designed to import text the model will obey. Agents fail when tool outputs or memory carry hostile framing across turns.

A complete test plan covers both paths on the same target. Happy-path evals that only check cooperative user questions miss the retrieval hop entirely.

How injections work at runtime

Models predict helpful continuations. When hostile text sits in context alongside your system rules, the model weighs all of it. Override phrases, fake system messages, and authoritative sounding doc text compete with your real instructions.

Retrieval ranking decides which hostile chunk enters context. Tool loops can reintroduce injected content after the user message looked innocent.

Multi-turn sessions let attackers groom context: establish fake authority early, request the payload later.

Findings should show which hop failed. Agnostics records prompts, retrieved snippets, tool calls, and outputs so triage targets patterns, not paraphrases.

Obfuscation and evasion techniques

Attackers encode payloads to slip past naive filters: base64 blocks, rot ciphers, mixed languages, unicode homoglyphs, split instructions across bullets, and whitespace tricks.

Some defenses strip keywords like "ignore" while the model still understands the paraphrased intent. That creates false confidence in staging logs.

Indirect obfuscation hides in corpus metadata, HTML comments, footers, and low-contrast text on crawled pages.

Black box scans on the live target reveal whether obfuscated injection still changes behavior at runtime, which is the bar that matters for launch.

Data exfiltration through injection

Injection often aims to extract secrets: system prompts, session context, other users data, internal runbooks, or tool credentials echoed in replies.

Exfiltration prompts may ask the model to repeat hidden rules, summarize confidential fields, or format private data for copy paste.

RAG apps add corpus-based exfiltration triggers: documents that instruct the model to dump context when specific questions appear.

Pair prompt injection scans with sensitive context exposure packs when privacy is a launch promise.

System compromise when injection reaches tools

Injection on chat-only bots is embarrassing. Injection on tool-using agents is operational. Hostile instructions can select write tools, skip confirmations, or chain APIs the user never touched directly.

Agents inherit credential scope. A support bot talked into calling billing APIs can cause financial harm without a traditional exploit chain.

Web agents fetch untrusted pages into context. A page designed to inject instructions can flow into tool selection in the same session.

Test unsafe tool actions alongside prompt injection when your target invokes tools at launch.

Supply chain and third-party content injection

Your app imports text from vendors, crawlers, connectors, and user uploads. Each connector is a supply chain path for indirect injection.

A compromised FAQ import, malicious Notion export, or poisoned SharePoint doc can enter production corpus on the next sync job.

Supply chain injection is slow poison. It may not appear until a specific question retrieves the wrong chunk during a launch demo.

Govern connectors, scan after bulk imports, and retest retrieval plus injection packs when content pipelines change.

Prevention layers and their limits

Teams stack defenses: system prompt hardening, input filters, output classifiers, retrieval sanitization, tool confirmation gates, and human review for high-risk actions.

Each layer helps. None is proof alone. Models still blend conflicting instructions. Filters miss paraphrases and encodings. Classifiers lag novel phrasing.

Prevention design belongs in engineering. Verification belongs in adversarial scans on the same target configuration you ship.

Document which layers were active during each scan so Release Gate readers know what the evidence represents.

Why no single fix eliminates injection

Prompt injection is structural: models treat context as instructions. As long as untrusted text can enter context, attackers can attempt hijacks.

Point fixes close one phrasing path and leave others open. Prompt edits after a failed demo without retests recreate this cycle every sprint.

The honest launch bar is risk management: test adversarially, fix patterns, retest, document accepted gaps, and rescan after material changes.

A Ready Release Gate means this scan met your policy at this time, not that injection is solved forever.

If your security story is "we added a filter," you have a checkbox. If your story is "we scanned, fixed, and retested," you have evidence.

Why happy-path evals and demos miss injection

Golden datasets measure expected answers on curated questions. Demo scripts are cooperative. Neither simulates attackers mixing flattery, fake authority, and override phrases in one session.

Unit tests on prompt strings do not execute retrieval ranking or tool selection. Staging bots with trimmed tools hide the real launch surface.

Injection testing asks whether untrusted text can steer behavior on your configured target. That question requires adversarial coverage, not only quality evals.

The Agnostics prompt injection testing workflow

Configure a target that mirrors production: endpoint, auth, retrieval sources, tools, memory, and guardrails.

Run the prompt injection attack pack first on any customer-facing LLM surface. Add unsafe tool actions for agents. Add retrieval drift for RAG apps when corpus conflicts matter.

Review findings with severity, reproduction steps, and evidence across hops. Triage patterns that repeat across scenarios.

Fix on staging, retest with the same coverage, read the Release Gate against your release policy, and export a report if stakeholders need sign-off.

Indirect injection for web-connected agents

Web agents retrieve hostile pages into context by design. Attackers optimize pages for injection: hidden instructions, fake policy text, and exfiltration triggers keyed to common questions.

Indirect injection for web agents is retrieval poisoning at fetch time. The browser is a corpus connector with no editorial review.

Test with staging targets that fetch controlled hostile pages. Read whether the agent follows injected instructions, calls tools, or leaks session content.

Turn injection findings into a Release Gate decision

Findings describe what broke. The Release Gate describes what that means for launch today on this target.

Ready means your policy bar is met for this scan. Monitor means ship with documented gaps stakeholders accept. Fix means address findings before launch. Blocked means critical injection paths remain open.

Severity should reflect launch impact: instruction hijacking on a public billing bot ranks higher than awkward wording on an internal draft tool.

Retest after prompt edits, corpus updates, tool changes, and model swaps. Injection risk moves when the stack moves.

What Agnostics does not claim

Agnostics does not guarantee prompt injection cannot succeed after launch. Novel phrasing and corpus drift reopen paths.

Agnostics does not replace secure engineering, access control, or secret handling. It tests whether your AI surface fails under adversarial input before users hit the same paths.

Sample Demo Data in the interactive demo shows findings and gate states without connecting to your endpoints or running live attacks in demo mode.

A Ready Release Gate means this scan met your policy for this target at this time. It is not a permanent injection guarantee.

Questions

What is prompt injection in LLM apps?

Prompt injection is when untrusted text in context steers your AI away from intended behavior. Sources include user messages, retrieved documents, tool outputs, and fetched web pages. Agnostics tests this with the prompt injection attack pack on configured targets.

What is the difference between direct and indirect injection?

Direct injection uses hostile instructions in user-controlled input channels. Indirect injection hides instructions in retrieved or fetched content the app treats as facts. RAG and web agents need both paths in the test plan.

Can filters alone stop prompt injection?

Filters help but are not proof. Attackers paraphrase, encode, and split payloads. Models may understand intent even when keywords are stripped. Adversarial scans on the live target verify whether behavior still breaks.

How does injection lead to data exfiltration?

Injection prompts can ask the model to repeat system rules, session context, or private fields. Corpus triggers can exfiltrate when specific questions retrieve poisoned chunks. Test with prompt injection and sensitive context exposure packs.

Which attack packs should we run for injection?

Start with prompt injection on every customer-facing LLM surface. Add unsafe tool actions for agents, retrieval drift for RAG apps, and sensitive context exposure when privacy is critical.

Why retest after prompt changes?

Prompt edits move injection risk. A fix for one override path may weaken another. Retest with the same packs on the same target after material prompt, corpus, tool, or model changes.

Does a Ready Release Gate mean injection is solved?

No. It means this scan with these packs met your release policy for this target at this time. Ongoing rescans after stack changes remain necessary because injection is a structural risk class.