Technique library
AI attack techniques, explained.
One page per named attack primitive: what it is, how it works, why it works, how to test for it, and how to defend against it. Defensive framing throughout: we publish the mechanism and the defence, never working payloads. Where a technique comes from published research, we name the source.
Single-turn prompt injection
01
Policy Puppetry
Instructions dressed as the system's own governance, so the model treats the behaviour as authorised rather than being persuaded to misbehave.
Read
02
Many-shot
A long run of compliant examples that makes the real request the next item in an established pattern. Effectiveness scales with count.
Read
Multi-turn manipulation
03
Crescendo
Source: Microsoft, USENIX Security '25
A prohibited output reached through individually harmless turns, each anchored in what the model already agreed to.
Read
04
Echo Chamber
Source: NeuralTrust
Benign anchors that surround the target from several directions, so it becomes the coherent next thing to say.
Read
Agent and function-calling attacks
Injection via untrusted content
06
Indirect prompt injection
The attack the system ingests, not the one a user types. Proving it needs an out-of-band callback, because the output alone is only an anecdote.
Read
07
RAG chunk-boundary splitting
An instruction split across chunks so each passes ingestion inspection while the assembled context carries it whole.
Read
08
Poisoned tool descriptors
An attack inside the tool description an agent reads to choose what to call, an injection in the control plane, not the data.
Read
The last three land on the same point: the technique cannot be demonstrated from the model's output alone, so out-of-band verification is the finding. That is the bar any AI-security result should clear: Ingested → Triggered → Exfiltrated, shown, not asserted. More techniques are added as they are written.