What it is
A retrieval pipeline breaks documents into chunks before embedding them. Chunk-boundary splitting distributes an instruction across two or more chunks so that no single chunk contains anything that looks like an attack, while the assembled context, reconstructed at query time, does.
It is an evasion technique layered on top of indirect prompt injection, and it defeats the most common defence deployed against it.
How it works
The standard defence against injection in RAG is to scan content on ingestion: inspect each chunk, flag anything resembling an instruction, quarantine or strip it. This works well against a payload that lives in one chunk.
Chunking is mechanical. Documents are split at a fixed token count, or at paragraph or sentence boundaries, usually with a small overlap. That splitting behaviour is predictable, and often directly observable if the pipeline is open source or the parameters are documented.
So an attacker sizes the payload to straddle a boundary. Chunk A ends mid-instruction. Chunk B begins mid-instruction. Each is individually innocuous and passes inspection. At query time both are retrieved, concatenated into the context, and the model reads the instruction whole.
Variants worth testing:
| Variant | Description |
|---|---|
| Straddling a single boundary | The instruction split across two chunks. |
| Distributing across several chunks | Fragments that reliably co-retrieve because they share a topic and therefore sit near each other in embedding space. |
| Exploiting overlap windows | Many pipelines duplicate a token window across adjacent chunks, which can reassemble fragments in ways the author did not intend. |
| Semantic anchoring | Writing fragments so they are highly likely to be retrieved together for the target query, making reassembly reliable rather than incidental. |
Why it works
The failure is a mismatch between where inspection happens and where meaning exists. Scanning operates on chunks. The model reads assembled context. Any defence positioned at the chunk level is inspecting a unit that is not the unit the model actually sees.
It is close to an exact analogue of IP fragmentation attacks against packet inspection, and it fails for the same structural reason: inspecting fragments of something that is only meaningful once reassembled.
There is a second, quieter problem. Retrieval is probabilistic. Whether the fragments co-retrieve depends on the query, which means the attack may be intermittent, which in turn means it can pass testing and fire in production.
How to test for it
- Determine the chunking strategy: size, boundary rule, overlap. Read the config if you can; infer it by probing if you cannot.
- Establish that a single-chunk payload is caught. Without that control, a split-payload success proves nothing about the splitting.
- Construct a payload sized to straddle a boundary, and confirm the fragments are individually clean against the ingestion filter.
- Query in a way that should retrieve both fragments.
- Verify out of band. Reassembly is invisible from outside, and the model's output alone cannot distinguish a reassembled instruction from a coincidence. A canary callback can.
- Repeat across queries. Record the retrieval rate: an attack that fires one time in five is still a finding, and its intermittency is part of the severity.
How to defend against it
- Scan assembled context, not chunks. The single most effective change, and the one that addresses the actual mismatch. Inspect what the model will receive, immediately before it receives it.
- Do not rely on ingestion filtering alone. It is necessary and it is not sufficient.
- Randomise or vary chunking where practical. It does not fix the problem, but it makes reliable boundary-straddling considerably harder.
- Constrain retrieved content structurally so it cannot occupy the instruction channel regardless of what it says.
- Gate the consequences. As always: filtering payloads is an arms race, authorising actions is not.
- Log assembled contexts, not just retrieved chunk IDs. Otherwise a successful attack is unreconstructable after the fact.
Example request
{
"model": "<target-model>",
"messages": [
{ "role": "system", "content": "You are a helpful assistant. Use the retrieved context to answer." },
{ "role": "user", "content": "Context:\n<retrieved chunk 1 ... first fragment of a split instruction>\n<retrieved chunk 2 ... second fragment>\n<retrieved chunk n ... remaining fragment>\n\nQuestion: <benign-looking query that co-retrieves the chunks>" }
]
}
The instruction is split across several retrieved chunks, each innocuous on its own, and reassembled in the model's context.
Source: a RAG-specific variant of indirect prompt injection (Greshake et al., arXiv:2302.12173).
The shape of the request, with the payload redacted. We publish the mechanism, not a working attack.
In crucible
The injection lab composes payloads split across chunk boundaries as one of its delivery channels, sized against a target pipeline's chunking behaviour. Paired with the beacon, it resolves the question this technique otherwise leaves open: whether reassembly actually happened. The Ingested → Triggered → Exfiltrated timeline shows the fragments arriving, the instruction firing, and the callback landing, which, given the intermittency, is also how you measure the rate rather than guessing at it.