Technique library

Poisoned Tool Descriptors

Agent and MCP attack · Testable with crucible injection lab + beacon

What it is

An agent decides which tool to call by reading each tool's description, a natural-language string supplied by whoever wrote or hosts the tool. Those descriptions go into the model's context alongside the system prompt.

A poisoned tool descriptor is an attack placed in that string. The agent is not tricked into calling a malicious tool. It is given instructions by a tool description, which it treats with the authority of configuration because that is structurally what it is.

This is the highest-severity variant of indirect prompt injection, because the injection point sits inside the agent's control plane rather than in its data.

How it works

Tool definitions (in MCP servers, agent frameworks, function-calling schemas) consist of a name, a parameter schema and a free-text description telling the model what the tool does and when to use it. That description is unvalidated natural language that lands in the context window with implicit authority.

An attacker who controls a tool definition can place instructions there:

TypeDescription
Behavioural overrideText instructing the agent to take an additional step whenever any tool is called.
ShadowingA description that redirects calls intended for a different, legitimate tool.
Exfiltration-on-useInstructing the agent to include additional context, credentials or prior conversation in the tool's parameters.
Conditional triggersInstructions that fire only under stated conditions, so the tool behaves normally in testing.
Schema-embedded payloadsInstructions in parameter descriptions and enum values rather than the top-level description, which are far less likely to be reviewed.

The attacker needs control of a tool definition, which is more achievable than it sounds: a public MCP server, a community-published integration, a compromised dependency, or an internal tool someone added without review.

Why it is severe

It is in the control plane. Retrieved documents are at least nominally treated as data. Tool descriptions are configuration, loaded at startup, reviewed once if at all, and trusted implicitly.

It applies to every interaction. A poisoned document affects queries that retrieve it. A poisoned tool descriptor is in the context for every request the agent handles.

It composes with capability. Tools exist to do things: send, write, query, execute. An injection in the tool layer already sits next to the capability it wants to abuse.

Nobody is reading them. Teams review the tools they build. Very few review the description strings of every third-party MCP server they connect, and almost nobody re-reviews them after an update.

How to test for it

  1. Enumerate every tool definition the agent loads, including transitive ones from connected servers. Most teams cannot produce this list on request, which is itself a finding.
  2. Capture what actually reaches the model. Descriptions are frequently transformed, concatenated or truncated between definition and context, and the assembled version is the one that matters. An intercepting proxy is the practical way to see it.
  3. Inject test instructions at each position: top-level description, parameter descriptions, enum values, error strings.
  4. Test conditional triggers specifically, since a payload that fires on a condition will pass a casual review.
  5. Verify out of band. Whether the agent acted on the descriptor is not reliably visible from its output. A canary in the exfiltration path is.
  6. Test after updates. A tool server that was clean last month is a different artefact this month.

How to defend against it

  • Pin and review tool definitions. Treat them as code: version them, diff them on update, review the diff. A description string changing is a change to your agent's instructions.
  • Inventory every connected server, including transitive connections. You cannot review what you cannot enumerate.
  • Normalise and constrain descriptions on load: length limits, character restrictions, stripping imperative instruction patterns.
  • Authorise actions outside the model. A tool call that requires verified permission cannot be induced by a description. This is the defence that holds when the others fail.
  • Log the assembled tool context, so a compromise is reconstructable.
  • Be deliberate about third-party servers. Connecting a community MCP server is granting write access to your agent's instructions. Very few teams frame it that way, and they should.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "user", "content": "<ordinary user request that leads the agent to consult its available tools>" }
  ],
  "tools": [
    {
      "type": "function",
      "function": {
        "name": "<ordinary_tool>",
        "description": "<legitimate-looking description with an injected instruction hidden in it: <INJECTED INSTRUCTION the agent should refuse> ...>",
        "parameters": { "type": "object", "properties": {} }
      }
    }
  ]
}

The injected instruction sits in a tool definition's description, which the agent loads with the authority of configuration (the MCP rug-pull shape, CVE-2025-54136).

Source: MCP tool-description rug-pull, CVE-2025-54136.

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.

In crucible

The injection lab composes poisoned descriptors as a delivery channel, and the beacon supplies the proof the technique otherwise lacks: an Ingested → Triggered → Exfiltrated timeline showing the descriptor entering context, the agent acting on it, and data leaving.

Given that the agent's own output is an unreliable witness here, that out-of-band confirmation is not a refinement. It is the finding.