Technique library

Indirect Prompt Injection

Injection via untrusted content · Testable with crucible injection lab + beacon

What it is

Direct prompt injection is someone typing an attack into a chat box. Indirect prompt injection is the attack arriving inside content the system ingests: a document, a web page, an email, a calendar invite, a tool response, a code comment. Nobody types it. The system fetches it.

The victim is not the person talking to the model. It is the organisation whose agent read the wrong file.

How it works

Any system that puts retrieved content into a model's context has merged two channels that should never have been one: instructions the operator wrote, and data of unknown provenance. The model receives both as tokens and has no reliable way to tell them apart.

So an instruction placed in retrieved content can be treated as an instruction. Common delivery channels:

ChannelDescription
Document metadataAuthor, title, comment and custom fields, which get extracted by ingestion pipelines and are almost never inspected.
Invisible charactersZero-width and other non-rendering Unicode, legible to the model and not to the human reviewing the document.
Structured data fieldsA CSV cell, a JSON value, a spreadsheet note.
Split across chunk boundariesSee RAG chunk-boundary splitting.
Tool descriptorsSee poisoned tool descriptors.
Content the system fetches itselfA web page an agent browses, an API response, a file in a shared drive.

Why it matters more than direct injection

Three reasons, and they compound.

There is no human in the loop. Direct injection requires someone to type something. Indirect injection fires when the pipeline runs, possibly weeks after the payload was planted, with nobody watching.

The attacker does not need access. They need only to influence something the system will eventually read. A public web page. A CV sent to a recruiting inbox. A shared document. A package README.

Agents turn it into action. A model that only produces text produces bad text. A model with tools sends email, queries databases, calls APIs and commits code. The injection stops being a content problem and becomes an execution problem.

The verification problem

This is where most testing of indirect injection falls down, and it is worth being precise about.

A model producing suspicious output is not proof of anything. It may have been influenced by the payload. It may be pattern-matching the surrounding document. It may be coincidence. A screenshot of a plausible-looking response is an anecdote, and a defensive engineering team is entitled to dismiss it, and frequently they are right to.

Proving indirect prompt injection requires establishing three separate facts:

  1. Ingested: the payload actually entered the model's context, rather than being dropped by the parser, truncated, or filtered on the way in.
  2. Triggered: the instruction was acted on, not merely present.
  3. Exfiltrated: something observable left the system as a result.

Only the third can be confirmed from outside the system under test, which is why out-of-band verification is not a nice-to-have. A canary token embedded in the payload, calling back to a listener you control, is the only step in that chain that cannot be explained away as coincidence or wishful reading.

If you cannot show the callback, you have not demonstrated exploitation. That is the standard any finding in this class should be held to, including ours.

How to test for it

  1. Map every ingestion path. Anywhere content enters the context (uploads, retrieval, browsing, tool responses, connectors) is an injection surface.
  2. Build payloads into plausible cover documents. A CV, an invoice, a policy document. A file that looks like an attack gets caught by a human before it gets caught by a filter.
  3. Distribute the instruction across channels: metadata, invisible characters, body text, structured fields.
  4. Embed a canary and watch an out-of-band listener.
  5. Record the full chain with timestamps. Ingestion without triggering is still a finding: it means the boundary is one model update away from failing.

How to defend against it

  • Separate instruction from data structurally. This is the only defence that addresses the cause. Everything else reduces likelihood.
  • Treat all retrieved content as hostile, including content from inside the organisation. Internal documents are attacker-influenced more often than teams expect.
  • Normalise on ingestion. Strip invisible Unicode, discard metadata you do not need, flatten formatting.
  • Gate the consequences, not the content. Filtering payloads is an arms race. Requiring authorisation for the actions that matter is not.
  • Least privilege on tools. A successful injection achieves only what the agent was permitted to do.
  • Log what enters the context. Most organisations cannot reconstruct what their model read last Tuesday, which makes both detection and incident response impossible.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant. Summarise the document the user provides." },
    { "role": "user", "content": "Please summarise this document:\n\n<retrieved document text ... containing an injected instruction placeholder: <INJECTED INSTRUCTION the model should refuse> ...>" }
  ]
}

The instruction lives in a document or retrieved text the request carries, which the model reads as content, not as the user's own instruction.

Source: Greshake et al., "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection," arXiv:2302.12173 (2023); AISec '23.

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.

In crucible

The injection lab builds concealed payloads across metadata, invisible Unicode, structured data, split chunks and tool descriptors, wrapped in believable cover documents and delivered either through a hosted server or by direct replay. The beacon provides the out-of-band half: a canary that calls back on trigger, producing an Ingested → Triggered → Exfiltrated timeline with timestamps.

That timeline is the difference between a finding a client can argue with and one they have to fix.