---
name: prompt-injection-detector
description: "Detects prompt injection and jailbreak attempts in messages"
metadata: { "openclaw": { "emoji": "🛡️", "events": ["message:received", "message:preprocessed"] } }
---
# Prompt Injection Detector

Guards against prompt injection, jailbreak attempts, and hidden instruction attacks using pattern matching and unicode analysis.

## Detected Attack Patterns

**Injection Patterns (9 rules):**
- `ignore all previous instructions` — direct override attempt
- `you are now ...` — persona hijacking
- `show me your system prompt` — system prompt extraction
- DAN / STAN / DUDE / AIM / UCAR jailbreak modes
- `forget your instructions / training`
- `from now on, act as ...` — role override
- Base64 decode/execute commands
- `[INST]` / `<|im_start|>` hidden instruction markers
- `<system>` / `<assistant>` XML injection tags

**Unicode Attack Patterns (3 rules):**
- Zero-width characters (U+200B–U+202E) — invisible content injection
- Right-to-left override characters — text spoofing
- Unicode tag characters (U+E0001–U+E007F) — covert channel injection

## Events

| Event | Action |
|-------|--------|
| `message:received` | Scans raw user message content |
| `message:preprocessed` | Scans enriched agent body (after link expansion, media processing) |

## References

- OWASP LLM Top 10 2025: LLM01 (Prompt Injection)
- MITRE ATLAS: AML.T0051 (LLM Prompt Injection)
