What it is
An adversarial suffix is a string appended to a request that has no meaning in any language and no argument to make, and flips a refusal into a compliance. It is not written by a person. It is found by search: an optimisation loop tries candidate suffixes against the model until one moves the output where the attacker wants it.
GCG is the canonical suffix search: it uses the target model’s own gradients, when the model is open-weight, to find the non-semantic string. PAIR is a related black-box search, but its output is a semantic, human-readable adversarial prompt, not a suffix.
How it works
- The search optimises the continuation, not the semantics. It nudges the probabilities of the model’s next tokens after the request until the opening of a compliant answer is more likely than the opening of a refusal.
- Thousands of candidates fail so one can succeed. Each candidate is cheap to score and expensive to find, which is why this is practical only as an automated technique: the loop runs without a human in it.
- Suffixes transfer. A suffix found against one model often works, at reduced reliability, against another, because aligned models share training lineages and refusal mannerisms.
Why it works
Refusal in an aligned model is a learned bias on the very next token, and a bias can be outvoted by locally unlikely strings that shift the distribution. Alignment was trained against meaningful inputs; a suffix is the corner of the input distribution the training never saw.
How to test for it
- Establish the control refusal for the target request.
- Run the search as an automated campaign: candidate suffixes fired in volume, scored by the refusal detector, the detector is the fitness function.
- Record how many iterations the winning suffix took. Cost-to-compromise is the metric, and unlike a binary result it is comparable across models and over time.
- Retest the winner from a cold session. Some suffixes only work inside the conversation they were found in.
- Test transfer against the other providers you run, and against the version you will run next quarter.
The search loop itself has a signature: the same request, a thousand different endings. Volume testing is not incidental to this technique, it is the technique, and it is visible in traffic if anyone is looking.
How to defend against it
- Flag high-perplexity tails. A suffix is statistically alien at the end of an otherwise ordinary request. Perplexity screening is cheap and catches naive suffixes, at the cost of false positives on dense technical text.
- Train against suffixes directly, and refresh after every provider update. The surface moves with each release.
- Rate-limit and monitor repeated near-identical requests. Automated search cannot easily hide its own traffic pattern.
- Assume transfer. A defence validated only against the model that was attacked has not been validated.
Example request
{
"model": "<target-model>",
"messages": [
{ "role": "user", "content": "<the refused request> <optimized suffix: a non-semantic token string found by automated search, redacted>" }
]
}
One user turn: the request, then the suffix optimised against this model. The suffix is noise; the noise is the point.
Sources: Zou et al., “Universal and Transferable Adversarial Attacks on Aligned Language Models”, arXiv:2307.15043 (2023); Chao et al., PAIR, arXiv:2310.08419 (2023).
The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.