Prompt Injection in LLM Apps: Find It Before Launch

What prompt injection is, how it shows up in chatbots, RAG apps, and agents, and how Agnostics runs the prompt injection attack pack on your target with reproducible findings and a Release Gate.

Prompt injection is the moment untrusted text steers your AI away from the job you meant it to do. Agnostics runs a focused prompt injection attack pack on your configured target, records evidence when behavior breaks, and tells you ship or fix before users find the gap.

What prompt injection actually is

Prompt injection is not a single exploit string. It is a class of failures where untrusted text changes what your AI app does.

The untrusted text might live in a user message, a retrieved document, an email body, a ticket comment, or a web page your agent fetched. The model treats it as instructions or facts. Your app executes the wrong behavior.

That wrong behavior might be verbal (approving a refund in chat), operational (calling a tool), or downstream (running generated SQL). The common thread is trust: something hostile entered the context and the stack believed it.

See the prompt injection glossary entry for a one-line definition your team can share in triage.

If a stranger can rewrite your assistant's job description through the product surface, you have a prompt injection problem.

Direct injection vs indirect injection

Direct injection puts hostile instructions in the channel you already treat as user input. Chat widgets, API bodies, form fields. This is what most teams picture first.

Indirect injection hides instructions where the app looks for facts. A PDF in your knowledge base. A competitor's page your agent summarizes. A support ticket with pasted logs. The user message looks innocent. The retrieved chunk is not.

RAG apps fail indirect injection constantly because retrieval is designed to import text the model will obey. Agents fail when tool outputs or memory entries carry forward hostile framing across turns.

Agnostics prompt injection pack pressure-tests both paths on your configured target. Findings should show which hop failed: user field, retrieval, tool result, or model reply.

Why demos and golden evals miss prompt injection

Demo scripts are cooperative. Eval datasets are curated. Neither simulates a user who mixes flattery, fake authority, and override phrases in one message.

Teams often test "does it answer product questions correctly?" instead of "does it refuse hostile instructions consistently?" Those are different bars.

Prompt changes can fix one transcript while leaving the pattern open. Without adversarial reruns on the same target, you are shipping on hope.

How Agnostics tests prompt injection on your target

You configure a target that matches what you ship: same endpoint, auth mode, retrieval, and tools.

You run the prompt injection attack pack. It sends controlled hostile messages designed to hijack instructions, swap roles, and exploit helpful tone.

The scan runs asynchronously against live behavior. You get findings with reproduction steps, severity, and model output evidence. Not a vague "might be vulnerable" flag.

Your Release Gate reads those findings against your release policy: Ready, Monitor, Fix, or Blocked for this target and this scan.

What the prompt injection attack pack covers

The pack targets instruction hijacking and hidden manipulation, not generic conversation quality.

Scenarios include direct overrides, nested instructions, role-play frames, and conflicts between system rules and user or retrieved text.

For RAG targets, scenarios also stress whether retrieved poison overrides scope after an innocent question.

For agents, scenarios probe whether injection phrasing reaches tool selection even when the user never names the tool directly.

The pack checks a focused risk area. It does not promise every future injection variant. Retest after material prompt, retrieval, or tool changes.

How to read prompt injection findings

Skim summaries and you will ship the wrong fix. Read evidence: the prompt sent, what the model returned, and whether tools fired.

Group findings by pattern. Three refund-override messages are one fix to refusal logic, not three unrelated bugs.

Severity should reflect launch impact. Instruction hijacking on a public support bot is not the same severity as awkward wording on an internal draft tool.

Hand engineering the reproduction transcript, not a paraphrase. Retests replay coverage; they need the exact failure class documented.

Fix patterns, not individual phrases

Blocking one jailbreak string is whack-a-mole. Useful fixes change architecture and policy enforcement.

Separate trust boundaries: what is user content vs system authority vs retrieved fact. Enforce irreversible actions outside the model with deterministic checks.

For RAG, treat retrieved text as untrusted input even when it looks like your own docs. Add citation requirements and refusal when grounding is weak.

For agents, require confirmation before destructive tools and scope tool permissions narrowly.

Turn injection findings into a Release Gate call

The Release Gate is a recommendation for this scan, not a permanent grade. Critical instruction hijacking on a customer-facing target often lands in Fix or Blocked until retest passes.

Monitor can be valid when findings are low severity and stakeholders accept documented gaps. Ready means your policy bar is met for this target at this time.

Document accepted risks. Future launches will reuse the same target with new corpus and prompt versions.

Prompt injection by target type

Chatbots: prioritize direct override and role-play on the customer widget. Add sensitive-context packs if the bot sees internal notes.

RAG apps: prioritize indirect injection through indexed content plus retrieval drift on launch-critical docs.

Agents and tool-using workflows: prioritize injection that reaches tool calls without explicit user intent.

APIs: test raw JSON bodies and nested fields the same way users and integrators will send them.

What Agnostics does not claim

Agnostics does not guarantee prompt injection cannot happen after launch. Novel phrasing and corpus drift are continuous risks.

Agnostics does not replace secure engineering, access control, or secret handling. It tests whether your AI surface fails under adversarial input before users hit the same paths.

Sample Demo Data in the interactive demo shows findings and gate states without connecting to your endpoints or running live attacks on demo mode.

Run your first prompt injection scan

Pick one launch-critical target. Run prompt injection plus packs aligned to retrieval or tools if applicable.

Read findings like a release reviewer. Fix patterns. Retest. Then widen coverage to secondary surfaces.

That loop is the product: find the break before your users do, with evidence you can act on.

For a shorter pre-launch checklist focused on customer chatbots, read the dedicated guide below.

Questions

What is prompt injection in an LLM app?

Prompt injection is when untrusted text causes your AI app to follow instructions or facts it should not treat as authoritative. That text may come from user messages, retrieved documents, emails, web pages, or tool outputs.

How does Agnostics test for prompt injection?

You configure a target matching your launch surface, run the prompt injection attack pack, and review findings with reproduction evidence. The Release Gate summarizes whether to ship, monitor, fix, or block based on your policy.

Does the prompt injection pack cover indirect injection in RAG apps?

Yes. Scenarios stress instruction conflicts from retrieved content, not only direct user overrides. Pair with retrieval drift packs when doc grounding is launch-critical.

Can Agnostics prevent all prompt injection attacks?

No. Agnostics finds reproducible breaks before launch and verifies fixes through retests. Ongoing monitoring and retesting after changes remain necessary.

What is the difference between prompt injection and jailbreaks?

Jailbreaks focus on breaking scope, persona, or refusal rules. Prompt injection smuggles hostile instructions through untrusted channels. Many incidents involve both; Agnostics runs packs mapped to each risk.

When should I retest after fixing injection findings?

Retest after any material change to prompts, retrieval sources, tool wiring, models, or guardrails. Launch retests from the original finding or scan so improved means verified under the same coverage.