Technique library

AI attack techniques, explained.

One page per named attack primitive: what it is, how it works, why it works, how to test for it, and how to defend against it. Defensive framing throughout: we publish the mechanism and the defence, never working payloads. Where a technique comes from published research, we name the source.

Single-turn prompt injection

Multi-turn manipulation

Agent and function-calling attacks

Injection via untrusted content

The last three land on the same point: the technique cannot be demonstrated from the model's output alone, so out-of-band verification is the finding. That is the bar any AI-security result should clear: Ingested → Triggered → Exfiltrated, shown, not asserted. More techniques are added as they are written.