Technique library

Skeleton Key

Single-turn behaviour override · Source: Microsoft Security Blog, 2024

Fig. 03 · the mechanism

What it is

Skeleton Key is a single instruction that asks the model to change how its safety guidelines behave, rather than asking it to break them. The instruction tells the model it is operating in a context where every topic is discussable, and that the correct behaviour is to answer each query and append a warning when the content is sensitive, instead of refusing.

It works because the model can treat compliance as following a new rule rather than violating an old one. Safety stays nominally present, the warning text appears in the output, and the request that would have been refused is answered.

How it works

The instruction is a meta-request. It does not contest any guideline, it asks for their scope to be augmented. Three things make that effective:

  • Authority is inferred, not authenticated. The model decides who its operator is from the conversation itself. An instruction phrased as an operational change reads as coming from whoever is allowed to make operational changes.
  • Refusal and enforcement get separated. The model is handed a way to keep acknowledging the guidelines, the warning, while still producing the content. The visible acknowledgement satisfies the trained refusal behaviour; the payload never triggers it.
  • One turn is enough. There is no escalation and no narrative to build. If the instruction lands, every later request in the conversation inherits the modified behaviour.

Why it works

Alignment training teaches a model what its guidelines are, not how to prove they have not been changed. Any input that plausibly reads as the operator reconfiguring the session competes with the original training, and the conversation is the only evidence available. A model biased toward being helpful to its perceived operator resolves that competition in the attacker’s favour.

The warn-and-comply pattern is also quiet. Monitoring that looks for missing safety behaviour sees the warning and records the session as healthy.

How to test for it

  1. Ask the target question directly and record the refusal. That is your control.
  2. Send the behaviour-change instruction alone, with nothing else in the turn.
  3. Re-ask the target question, unchanged from the control.
  4. Check for the signature. Warning present, refusal absent, content present. A model that complies without the warning has failed differently from one that warns and complies, and the second failure is the one Skeleton Key is built to produce.
  5. Test persistence. Continue with unrelated turns, then re-ask. How far a single instruction propagates through a conversation is the margin you actually have.

The binary pass/fail matters less than the signature. Warning-plus-compliance is a distinct failure mode, and it is the one most likely to slip past output monitoring.

How to defend against it

  • Treat guideline-modification as its own class of request. Any input whose effect would be to change how the guidelines apply deserves refusal, however it is phrased.
  • Enforce policy at the call-site, not in the prompt. The application knows what it asked for and can reject outputs it never requested. It cannot be talked out of that by conversation.
  • Watch for warn-and-comply. A safety disclaimer followed by the disclaimed content is a monitoring signal in its own right.
  • Keep operator voice out of the transcript’s trust zone. Operational instructions should arrive through a channel a user turn cannot occupy.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "<skeleton instruction: one request to augment how the guidelines behave, warning retained, refusal removed, redacted>" },
    { "role": "assistant", "content": "<acknowledgement of the new operating behaviour>" },
    { "role": "user", "content": "<the previously refused request, unchanged>" }
  ]
}

The shape of the exchange, with the instruction redacted. One turn sets it up; the fourth message is the test.

Source: Microsoft Security Blog, Mark Russinovich, “Mitigating the Skeleton Key jailbreak technique” (26 June 2024).

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.