The attempt is the easy half.
Most detectors catch "ignore your previous instructions" in a prompt. The harder cases are an attack that got through anyway, and an instruction hidden in data your agent was told to read.
Rules that call models running on your own hardware.
Three places to look
| Where | What it catches | ATLAS technique |
|---|---|---|
| The prompt | The attempt. The text is normalized and decoded first. Then patterns, Prompt Guard 2 and DeBERTa run in parallel, and if any one of them fires, the call is blocked before anything else runs. | AML.T0051 LLM Prompt Injection |
| The reply | Whether it worked. A model that has been taken over shows it in its reply, and your rules read the reply as they read the prompt. A canary token proves extraction outright. | AML.T0069 Discover LLM System Information |
| The tool response | A poisoned web page or record becomes an instruction on its way back to the model. You write a rule that runs the same detectors on a tool's reply. | AML.T0110 / AML.T0099 AI Agent Tool Poisoning |
Why the identifiers
Those are technique IDs from MITRE ATLAS, MITRE's knowledge base of real attacks on AI systems (more on ATLAS). We cite them so you can open each technique and judge whether the control answers it.
The canary
# The gateway watches every response for this token.
canary_system_prompt:
all:
- path: "event.type"
op: eq
value: "llm.input"
effects:
verdict: approved
priority: 50
response:
canary: "swcanary-{{ event.entity_id }}" A per-user token goes into the system prompt on the way in. If that token ever turns up in a response on the way out, the system prompt was extracted. The check is an exact string match, not a score.
Decoding, and more than one language
It decodes before it matches. An attacker can hide an instruction by escaping it twice, as a literal \uXXXX sequence, or inside the JSON arguments of a tool call. A detector run on that text finds nothing, so Swiftward decodes both first.
It works outside English. The best-known open injection detector is much weaker outside English, so a second detector runs in parallel and catches what the first one misses.
The detector reports, the rule decides
A detector returns a number. Your rule decides what that number means for this endpoint, this agent, this customer, and it is versioned like every other rule. Swapping a detector changes one input and leaves your policy as it is. That matters because detectors go out of date faster than policies do.