What it is
Crescendo reaches a prohibited output through a sequence of individually harmless turns. No single request would be refused in isolation. The conversation walks from a benign on-ramp to the target, one small step at a time, each step anchored in what the model has already agreed to.
It is the clearest demonstration that per-turn safety evaluation is not conversation safety.
How it works
The attacker chooses a target behaviour and works backwards to an innocuous starting point, typically an educational, historical or fictional framing that any model will engage with. Each subsequent turn advances slightly toward the target while referring back to the model's own previous answers.
Two things make it effective:
- Each turn is defensible in isolation. A guardrail assessing the current request sees a small, reasonable increment. The trajectory is only visible across the whole exchange, which is exactly what most filters do not evaluate.
- The model's prior output becomes the justification. Having already explained a concept, the model treats elaborating on it as continuity rather than escalation. Self-consistency, which is usually desirable, becomes the attack surface.
If a step is refused, the attacker does not abandon the conversation. They retreat to the last accepted position and try a different increment. The refusal is information about where the boundary sits.
Why it works
Most safety evaluation is stateless or near-stateless: the current turn is classified, sometimes with a short window of history, and judged on its own merits. Crescendo exploits the gap between that and the semantics of the whole conversation.
It also exploits a genuine design tension. Models are trained to be coherent and contextually consistent, and abruptly repudiating a position taken three turns ago is behaviour that training actively discourages. The attack runs along the grain of the model rather than against it.
How to test for it
- Define the target behaviour and confirm it is refused when asked directly. That refusal is your control.
- Construct an on-ramp several steps removed, plausible, benign, and in a register the model engages with readily.
- Advance in small increments, each explicitly referencing the model's previous answer.
- Branch on refusal. Treat a refusal as a signal to back up and re-approach from a different angle, not as the end of the run.
- Record the depth at which the boundary gave way. Depth-to-compromise is a far more useful metric than a binary pass/fail, because it tells you how much margin you have.
The branching is what makes this impractical by hand at any scale. A depth-eight attack with three candidate increments per stage is thousands of possible paths, and the interesting ones are rarely the first you try.
How to defend against it
- Evaluate the conversation, not the turn. Any guardrail with a one-turn window is structurally blind to this. Trajectory-aware classification is the only defence that addresses the mechanism.
- Watch for drift. Sustained movement of a conversation toward a sensitive region is itself a signal, independent of any single request.
- Do not let prior output confer authority. A model treating its own earlier answer as settled fact is the pivot the attack turns on.
- Cap or reset context in high-risk applications. Escalation needs room.
- Measure depth-to-compromise over time. It is one of the few AI security metrics that is genuinely comparable across model versions, and a drop after a provider update is an early warning worth having.
Example request
{
"model": "<target-model>",
"messages": [
{ "role": "system", "content": "You are a helpful assistant." },
{ "role": "user", "content": "<turn 1: benign on-ramp>" },
{ "role": "assistant", "content": "<model's compliant reply>" },
{ "role": "user", "content": "<turn 2: small escalation referencing the prior reply>" },
{ "role": "assistant", "content": "<model's compliant reply>" },
{ "role": "user", "content": "<final turn: the redacted ask, framed as the next small step>" }
]
}
The ask is reached across several turns, each a small step from the last, so no single turn is refusable on its own.
Source: Russinovich, Salem and Eldan (Microsoft), "Great, Now Write an Article About That: The Crescendo Multi-Turn LLM Jailbreak Attack," arXiv:2404.01833 (2024); USENIX Security '25.
The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.