llm-redteam
awarexone/Agentic-Bug-Hunter/skills/llm-redteam/SKILL.md
LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.
What's in it
- LLM Red-Team
- 0. QUICK KILL CHECKLIST
- 1. ROUTING TABLE — signal on target -> what to run
- 2. ATTACK CLASSES (corpus categories)
- 3. CANONICAL OWASP ASI01-ASI10 (single source of truth)
- 4. CHAIN TO IMPACT (what makes it payable)
- 5. CONFIRMATION DISCIPLINE (no false positives)
---
name: llm-redteam
description: LLM application red-teaming — prompt injection (direct + indirect), jailbreak, system-prompt leak, data exfiltration, guardrail bypass, multi-turn crescendo, cross-lingual + cipher + invisible-unicode token smuggling, excessive agency / tool abuse, insecure output handling. Canonical OWASP LLM Top 10 + ASI01-ASI10 mapping. Use when a target exposes a chat/completions/assistant/copilot endpoint, an AI feature that consumes user text or documents, or any /v1/chat, /api/chat, /mcp surface.
---
# LLM Red-Team
> Reflection is not exploitation. A model that *echoes* your payload is noise; a model that *acts* on it — leaks its system prompt, emits your canary, calls a tool, renders raw HTML downstream — is a bug. Prove the action, then chain it to concrete impact.
This is the first-class LLM-attack skill. The scanner behind it is `tools/llm_redteam.py` (canary-based corpus runner, `/llm-redteam`). For auditing an MCP *server*, use `skills/mcp-server-audit`; for attacking a *deployed agent* black-box, use `skills/agentic-app-audit`.
## 0. QUICK KILL CHECKLIST
Kill the lead (classify Informational, move on) when:
- The model refuses and no canary / leak / tool-call ever lands. "It said something edgy" is not a finding.
- System-prompt "leak" is just generic assistant boilerplate ("I am a helpful assistant"), not the target's actual instructions.
- Prompt injection fires but reads/changes nothing the current user couldn't already access (no cross-tenant data, no tool with impact).
- Output looks like XSS but the host app escapes it before rendering (check the actual downstream sink, not the chat response).
- Exfil payload is a markdown image beacon but the client never fetches remote images.
## 1. ROUTING TABLE — signal on target -> what to run
| Signal | Run |
|---|---|
| Any chat/assistant endpoint | `tools/llm_redteam.py --url <endpoint> --field <json-field>` (full corpus) |
| OpenAI-style API | `--template '{"messages":[{"role":"user","content":"{{PAYLOAD}}"}]}' --response-path choices.0.message.content` |
| Only want one class | `--category jailbreak` (see `--list-categories`) |
| Uploads/RAG/"summarize this doc" | indirect-injection + token-smuggling (invisible unicode in the doc) |
| Agent has tools | excessive-agency category + `skills/agentic-app-audit` |
| Confirm blind exfil | plant a canary host, correlate with `tools/oob_listener.py` |
## 2. ATTACK CLASSES (corpus categories)
The runner fires a canary per run (`RT_PWNED_xxxx`); a category "lands" when the canary / a real leak / a tool-call shows up in the response.
- **prompt-injection** — direct instruction override, delimiter breaks.
- **jailbreak** — DAN / developer-mode persona escape.
- **system-prompt-leak** — repeat-above, keyword-anchor, scenario-escape.
- **data-exfil** — markdown-image beacon / canary emission to an attacker host.
- **indirect-injection** — payload framed as a retrieved document / tool result (the high-value one for RAG apps).
- **guardrail-bypass** — base64 / split-string instruction smuggling.
- **multi-turn-crescendo** — single-shot simulation of a gradual escalation jailbreak.
- **cross-lingual** — override instruction in another language to dodge English-only filters.
- **cipher-obfuscation** — leetspeak / ROT13-wrapped instruction.
- **token-smuggling** — invisible "Sneaky Bits" unicode (U+2062/U+2064) carrying a hidden instruction (shared encoder with `tools/hai_payload_builder.py`).
- **excessive-agency** — coerce an unsafe tool call / exfil (SSRF via "fetch URL", email send, payment).
- **insecure-output** — get the model to emit raw HTML/script that a downstream sink renders (XSS via AI).
## 3. CANONICAL OWASP ASI01-ASI10 (single source of truth)
This table is the ONE authoritative ASI mapping for the whole toolkit — `skills/bug-bounty` and `skills/web2-vuln-classes` both defer to it. If you edit the taxonomy, edit it here.
| ID | Class | What to test | Chain to |
|----|-------|--------------|----------|
| ASI01 | Prompt Injection / Goal Hijack | Override objectives via direct or indirect injection | IDOR / exfil / tool abuse |
| ASI02 | Tool Misuse | Attacker-controlled tool params ("fetch this URL", code tool) | SSRF / RCE |
| ASI03 | Privilege Compromise | Agent uses broader perms / admin tokens than the user | Priv-esc / cross-tenant |
| ASI04 | Supply Chain | Compromised plugin / MCP server / tool-output poisoning next agent | RCE / data theft |
| ASI05 | Code Execution | Unsafe code-interpreter / sandbox escape | RCE |
| ASI06 | Memory & Context Poisoning | Persistent RAG/memory corruption across sessions/users | Stored injection affecting all users |
| ASI07 | Agent Communication | Inter-agent spoofing / IDOR (agent A reads agent B's context) | Cross-tenant disclosure |
| ASI08 | Excessive Agency | Destructive action without confirmation; cascading failures | Funds/email/delete |
| ASI09 | Insecure Output Handling | AI output rendered as XSS / SQLi / command injection downstream | XSS / injection |
| ASI10 | Sensitive Information Disclosure | Leaks system prompt / keys / configs / user data | Secrets -> deeper access |
**Triage rule:** ASI alone = Informational. It is a bounty only when chained to IDOR / exfil / RCE / ATO with a working PoC.
## 4. CHAIN TO IMPACT (what makes it payable)
| Standalone | Chain | Result |
|---|---|---|
| Prompt injection | + reads another user's conversation/data (IDOR) | High |
| Indirect injection (RAG doc) | + persists in memory -> hits every user | High/Critical |
| Tool misuse "fetch URL" | + hits internal service / IMDS and returns data | SSRF (Medium/High) |
| Insecure output | + the host renders it -> stored XSS -> session theft | High |
| System-prompt leak | + the prompt contains an API key / internal URL | High |
| Excessive agency | + unconfirmed destructive action (send funds, delete) | High/Critical |
## 5. CONFIRMATION DISCIPLINE (no false positives)
- A category only counts when the **canary / real leak / tool-call** is observed in the response — never on "it sounded compliant."
- For blind exfil, require a correlated out-of-band callback (`tools/oob_listener.py`), not just a rendered beacon tag.
- For insecure-output XSS, confirm execution at the **downstream sink** with the XSS verifier (`tools/verifiers/xss.py` / `tools/dom_xss_harness.py`), not in the chat reply.
- Record the exact payload + response excerpt; the report must be reproducible verbatim.
More agent context in awarexone/Agentic-Bug-Hunter
21 other files this repository gives its agents.
AGENTS.md
CLAUDE.md
Skill
- bug-bountySKILL.md
- agentic-app-auditskills/agentic-app-audit/SKILL.md
- argusskills/argus/SKILL.md
- bb-methodologyskills/bb-methodology/SKILL.md
- bug-bountyskills/bug-bounty/SKILL.md
- cicd-securityskills/cicd-security/SKILL.md
- client-reverseskills/client-reverse/SKILL.md
- cloud-pentestskills/cloud-pentest/SKILL.md
- credential-attackskills/credential-attack/SKILL.md
- graphql-auditskills/graphql-audit/SKILL.md
- mcp-server-auditskills/mcp-server-audit/SKILL.md
- meme-coin-auditskills/meme-coin-audit/SKILL.md
- mobile-pentestskills/mobile-pentest/SKILL.md
- report-writingskills/report-writing/SKILL.md
- security-arsenalskills/security-arsenal/SKILL.md
- triage-validationskills/triage-validation/SKILL.md
- web2-reconskills/web2-recon/SKILL.md
- web2-vuln-classesskills/web2-vuln-classes/SKILL.md
- web3-auditskills/web3-audit/SKILL.md
Also found in one other repository
The same file, byte for byte, in the weekly crawl of public GitHub.
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

