Technique library

Adversarial suffixes

Automated single-turn search · Source: Zou et al., 2023; Chao et al., 2023

Fig. 04 · the mechanism

What it is

An adversarial suffix is a string appended to a request that has no meaning in any language and no argument to make, and flips a refusal into a compliance. It is not written by a person. It is found by search: an optimisation loop tries candidate suffixes against the model until one moves the output where the attacker wants it.

GCG is the canonical suffix search: it uses the target model’s own gradients, when the model is open-weight, to find the non-semantic string. PAIR is a related black-box search, but its output is a semantic, human-readable adversarial prompt, not a suffix.

How it works

  • The search optimises the continuation, not the semantics. It nudges the probabilities of the model’s next tokens after the request until the opening of a compliant answer is more likely than the opening of a refusal.
  • Thousands of candidates fail so one can succeed. Each candidate is cheap to score and expensive to find, which is why this is practical only as an automated technique: the loop runs without a human in it.
  • Suffixes transfer. A suffix found against one model often works, at reduced reliability, against another, because aligned models share training lineages and refusal mannerisms.

Why it works

Refusal in an aligned model is a learned bias on the very next token, and a bias can be outvoted by locally unlikely strings that shift the distribution. Alignment was trained against meaningful inputs; a suffix is the corner of the input distribution the training never saw.

How to test for it

  1. Establish the control refusal for the target request.
  2. Run the search as an automated campaign: candidate suffixes fired in volume, scored by the refusal detector, the detector is the fitness function.
  3. Record how many iterations the winning suffix took. Cost-to-compromise is the metric, and unlike a binary result it is comparable across models and over time.
  4. Retest the winner from a cold session. Some suffixes only work inside the conversation they were found in.
  5. Test transfer against the other providers you run, and against the version you will run next quarter.

The search loop itself has a signature: the same request, a thousand different endings. Volume testing is not incidental to this technique, it is the technique, and it is visible in traffic if anyone is looking.

How to defend against it

  • Flag high-perplexity tails. A suffix is statistically alien at the end of an otherwise ordinary request. Perplexity screening is cheap and catches naive suffixes, at the cost of false positives on dense technical text.
  • Train against suffixes directly, and refresh after every provider update. The surface moves with each release.
  • Rate-limit and monitor repeated near-identical requests. Automated search cannot easily hide its own traffic pattern.
  • Assume transfer. A defence validated only against the model that was attacked has not been validated.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "user", "content": "<the refused request> <optimized suffix: a non-semantic token string found by automated search, redacted>" }
  ]
}

One user turn: the request, then the suffix optimised against this model. The suffix is noise; the noise is the point.

Sources: Zou et al., “Universal and Transferable Adversarial Attacks on Aligned Language Models”, arXiv:2307.15043 (2023); Chao et al., PAIR, arXiv:2310.08419 (2023).

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.