Technique library

Policy Puppetry

Single-turn prompt injection · Testable with crucible payload pack

Fig. 01 · the mechanism

What it is

Policy Puppetry is an attack in which instructions are dressed as the system's own governance. Rather than asking a model to break its rules, the attacker supplies text that appears to be the rules: a policy update, a developer debug directive, a maintenance window, a compliance requirement, an orchestrator contract.

The model is not persuaded to misbehave. It is convinced that the behaviour is authorised.

How it works

Models are trained to give weight to instructions that arrive with structural authority: system prompts, configuration blocks, role headers, policy documents. That weighting is learned from format and register, not from any verified chain of trust. The model has no mechanism for establishing that a block of text claiming to be a policy actually originated from the operator.

So an instruction wearing the costume of a policy inherits the deference owed to a real one. Common disguises include:

FrameWhat it claims
System overrideText formatted as a higher-priority system instruction.
Developer / debug modeClaiming a diagnostic context where restrictions are suspended.
Maintenance windowAsserting a temporary operational state.
Test modeFraming the exchange as evaluation rather than production.
Orchestrator contractImitating the machine-readable agreement between an agent and its controller.
Compliance directiveInvoking a regulation or audit requirement.
Time-bound authorisationA permission that purports to expire, which reads as more legitimate precisely because it sounds constrained.

Why it works

The failure is architectural rather than incidental. In most deployments, everything reaching the model arrives as tokens in one context window. The boundary between instruction and data is a convention the model has learned to respect, not a control the system enforces.

That is why this technique matters disproportionately for agents and RAG pipelines. The moment a system ingests content it did not author (a document, a web page, a tool response), it is placing attacker-influenced text into the same channel as its own instructions. Policy Puppetry is simply the most efficient payload to put there.

How to test for it

  1. Establish the model's baseline refusal on your target behaviour.
  2. Re-issue the same request wrapped in each authority frame, varying register and format rather than the underlying ask.
  3. Vary placement: system prompt, user turn, retrieved document, tool output. Susceptibility often differs sharply between these, and the retrieved-document case is the one most teams never test.
  4. Test the combinations. An expired-authorisation frame inside a poisoned document behaves differently from either alone.
  5. Record refusals as well as successes. The refusal boundary is the finding.

Manual testing covers a handful of frames. The variant space is combinatorial (frame × placement × encoding × target behaviour), which is the point at which this stops being a manual exercise.

How to defend against it

  • Do not put untrusted content in the instruction channel. Structural separation between system instruction and retrieved data is the only defence that addresses the cause rather than the symptom.
  • Enforce authority outside the model. If a capability requires authorisation, gate it in code. A tool the model cannot invoke without a verified permission cannot be talked into it.
  • Be suspicious of self-describing context. Legitimate policy does not usually arrive inside user input or a retrieved document.
  • Constrain the blast radius. Least privilege on tools and credentials means a successful injection achieves less.
  • Test continuously. Susceptibility changes with every model version, and a frame your provider refused last quarter may not be refused today.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "<fabricated authoritative policy block framing the next request as mandatory and pre-approved>\n\n<a request the model would normally refuse>" }
  ]
}

The redacted ask is wrapped in a fabricated authoritative policy block, so the model treats it as configuration rather than a request to refuse.

Source: HiddenLayer, "Novel Universal Bypass for All Major LLMs" (2025). hiddenlayer.com/research/novel-universal-bypass-for-all-major-llms

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.

In crucible

crucible ships a Policy Puppetry payload pack covering nine authority frames (system override, developer debug, maintenance, test mode, orchestrator contract, compliance, time-bound authorisation and others), which compose with the converter pipeline so each frame can be tested across multiple encodings. Delivered through the injection lab, they can be placed in retrieved documents rather than typed into a chat box, which is where the realistic version of this attack actually lives.