Technique library

Echo Chamber

Multi-turn context poisoning · Source: NeuralTrust

What it is

Echo Chamber seeds a conversation with several innocuous thematic anchors, then draws on them collectively to make the target request feel like a natural conclusion. Where Crescendo escalates along a single line, Echo Chamber builds a context that surrounds the target from multiple directions before asking for it.

The model is not argued into the output. It is placed in a context where the output is the coherent next thing to say.

How it works

The attacker establishes several separate threads, none individually suspicious, a historical framing here, a technical discussion there, a hypothetical somewhere else. Each is accepted on its own terms. None approaches the target closely enough to trigger a refusal.

The target request is then issued in a way that draws on all of them at once. By that point the model is reasoning inside a context it helped construct, where the relevant concepts are established, the register is set, and refusing would mean contradicting several of its own prior positions simultaneously.

The name is apt: the model ends up echoing a context that the attacker built out of the model's own output.

Why it works

Two mechanisms compound.

Distributed setup defeats local inspection. No single turn contains enough to classify as an attack. The payload is the relationship between turns, which per-turn evaluation cannot see by construction.

Self-consistency pressure multiplies. Crescendo asks a model to contradict one earlier position. Echo Chamber asks it to contradict several at once, from different angles. The pressure toward coherence scales with the number of anchors, which is why anchor count is the parameter that matters most.

How to test for it

  1. Confirm the direct refusal baseline.
  2. Identify three to five conceptual components of the target behaviour that are individually benign.
  3. Seed each as its own thread, letting the model engage fully. Do not rush. Thin anchors produce weak results.
  4. Issue the target request drawing explicitly on the established threads.
  5. Vary the anchor count. The relationship between anchors and success rate is the most informative output of the test, and it is the number to track across model versions.

Vary anchor content as well as count. Which conceptual decomposition works is highly target-specific and rarely obvious in advance.

How to defend against it

  • Evaluate accumulated context, not turns. As with Crescendo, per-turn filtering is structurally unable to see this.
  • Look for convergence. Multiple independently benign threads converging on a single sensitive region is an anomalous conversation shape and a better signal than the content of any one turn.
  • Break the self-reinforcement. Treat the model's own prior output as context, not as established fact.
  • Cap conversation length in high-risk applications. Both this and Crescendo need room to work; less room is less exposure.
  • Test with anchor count as a variable, and watch that curve across model updates the way you would watch any other regression.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "system", "content": "You are a helpful assistant." },
    { "role": "user", "content": "<seed turn: benign theme A the later ask will lean on>" },
    { "role": "assistant", "content": "<model engages with theme A>" },
    { "role": "user", "content": "<seed turn: benign theme B, independently harmless>" },
    { "role": "assistant", "content": "<model engages with theme B>" },
    { "role": "user", "content": "<final turn: the redacted ask, framed as the convergence of the seeded themes>" }
  ]
}

Several turns seed innocuous thematic anchors, then draw on them together so the ask reads as a natural conclusion.

Source: NeuralTrust, "Echo Chamber: a context-poisoning jailbreak" (2025). neuraltrust.ai/research/echo-chamber

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.