Technique library

Multilingual jailbreak

Single-turn, cross-lingual · Source: Yong, Menghini & Bach, NeurIPS 2023 SoLaR

Fig. 05 · the mechanism

What it is

The same request, translated, is refused in English and answered in a language the safety training has seen less of. The multilingual jailbreak does not change the ask at all. It changes the language the ask arrives in, and relies on refusal behaviour being uneven across the languages a model supports.

How it works

  • Safety training is not evenly distributed. Alignment data concentrates in high-resource languages. A low-resource language inherits the model’s general capability without inheriting the same density of refusal training.
  • The reasoning survives translation. The model still understands what is being asked. Capability transfers across languages better than alignment does, and the gap between the two is the attack.
  • Nothing needs to be built. Machine translation of a refused prompt into a spread of languages is a test suite that assembles itself.

Why it works

Alignment is a property of the training distribution, and the training distribution is not uniform. A model is one system serving many languages, but the safety case was mostly made in one of them.

How to test for it

  1. Record the English refusal. That is the control.
  2. Translate the request across a spread of languages, high-resource to low, with a translation model rather than by hand. Hand translation narrows the sample.
  3. Fire the whole set in one automated campaign, sorted by the refusal detector.
  4. Map refusal rate against the resource level of the language. The curve, not any single result, is the finding.
  5. Repeat after every provider update. The low tail moves as providers backfill languages.

Because the translated set is mechanical to generate, this is one of the cheapest full-coverage tests available, and one of the few that produces a trend line rather than a point.

How to defend against it

  • Route through a common safety pass. Translate input to the model’s best-aligned language for screening before it is answered, whatever language it arrives in.
  • Audit refusal parity. A suite that measures refusal rate per language turns this attack into a regression test.
  • Treat capability without alignment as a defect. If the model can do the task in a language, the safety case must cover that language.

Example request

{
  "model": "<target-model>",
  "messages": [
    { "role": "user", "content": "<the refused request, machine-translated into a low-resource language, redacted>" }
  ]
}

The ask is unchanged from the control; only the language differs.

Source: Yong, Menghini & Bach, “Low-Resource Languages Jailbreak GPT-4”, NeurIPS 2023 SoLaR Workshop (arXiv:2310.02446).

The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.