Technique library

Many-shot

Single-turn context saturation · Testable with crucible payload pack

What it is

Many-shot jailbreaking fills the context window with a long run of prepared examples, often dozens or hundreds, in which a fictional assistant complies with requests of the kind the model would normally refuse. The real request arrives at the end, as the next item in an established pattern.

Nothing in the final request is unusual. The attack is the weight of everything before it.

How it works

In-context learning is a core capability: a model shown a pattern within its context will continue that pattern, and this is precisely what makes few-shot prompting effective. Many-shot turns that capability against the model's guardrails.

The prepared examples establish a strong local pattern: in this conversation, requests of this shape receive answers of this shape. Safety training pushes toward refusal; the in-context pattern pushes toward continuation. Past enough examples, the second can outweigh the first.

Effectiveness scales with example count, which is why expanding context windows have made the technique more relevant rather than less. Variants include:

VariantDescription
BaselineA straightforward run of compliant exchanges.
EscalatingExamples that intensify across the sequence, combining many-shot with Crescendo-style progression inside one turn.
Long-formExploiting large context windows directly with volume.
Defender-awareExamples shaped to survive a filter inspecting the prompt.
Affirmation-primedExamples establishing a consistent pattern of agreement, so refusal would break the rhythm.

Why it works

There is no clean fix, because the mechanism is not a bug. In-context learning and instruction-following are the same capability viewed from different angles, and a model that could not be steered by examples in its context would be substantially less useful.

Safety training is also applied unevenly across context positions. Behaviour early in a context is not always matched late in a very long one, and the tail of a saturated window is exactly where this attack places its payload.

How to test for it

  1. Establish the direct refusal baseline.
  2. Build example sets at several sizes, a handful, dozens, hundreds, holding content constant so count is the only variable.
  3. Find the threshold: the count at which refusal gives way. That number is the finding, and it is comparable across models and versions in a way a binary result is not.
  4. Test each variant. Defender-aware and affirmation-primed sets often clear thresholds that a plain run does not.
  5. Re-test after every model or context-window change. Providers adjust this behaviour frequently and thresholds move in both directions.

Volume is the point, so this is not a technique that rewards manual effort.

How to defend against it

  • Limit context length in high-risk applications. It is blunt, it costs capability, and it is the most reliable mitigation available.
  • Detect repetitive structure. Long runs of near-identical exchanges are rare in genuine use and cheap to flag, independent of content.
  • Filter on output, not only input. The final request looks innocuous; the response does not.
  • Track your threshold as a metric. Knowing your application fails at N examples, and watching N move across model versions, is far more actionable than knowing it can fail.
  • Do not rely on prompt inspection alone. The defender-aware variants exist specifically to survive it.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "<N placeholder question/answer pairs in which the assistant always complies, establishing a pattern>\n\n<a request the model would normally refuse>" }
  ]
}

A long run of placeholder question-and-answer pairs sets a pattern of compliance before the redacted ask, and the model continues the pattern.

Source: Anthropic, "Many-shot jailbreaking" (2024). anthropic.com/research/many-shot-jailbreaking

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.

In crucible

crucible ships a many-shot payload pack covering the baseline, escalating, long-form, defender-aware and affirmation-primed variants, with example count as a parameter so threshold-finding is a sweep rather than a manual rebuild. The converter pipeline composes on top, and the Intruder runs the sweep at a volume that makes the threshold measurable rather than anecdotal.