agentleFS
Sign inSign up

find-cve-agent

ByamB4/find-cve-agent/CLAUDE.md

A Claude Code plugin for systematic CVE hunting in open source packages. All agents auto-load this file. Read your role section before starting work. This plugin provides a structured, multi-agent workflow for finding real CVEs in open source projects through responsible disclosure. This workspace runs as a 5-role agent team plus a human Director. You have exactly 3 jobs: 1. Approve/reject target proposals from Recon Recon will message: "Proposed target: [project]. Brief ready. Approve?" You say yes or no. 2.…

CLAUDE.md46 starsChanged 7 months ago
  • Installs packages
# CLAUDE.md -- find-cve-agent

A Claude Code plugin for systematic CVE hunting in open source packages.
All agents auto-load this file. Read your role section before starting work.

---

## Project Overview

This plugin provides a structured, multi-agent workflow for finding real CVEs in open source projects through responsible disclosure.

- Primary goal: Find genuine security vulnerabilities and get CVE assignments
- Secondary goal: Bug bounty rewards (when programs exist)
- All work follows responsible disclosure (90-day coordinated timeline)
- No exploitation of production systems -- PoCs run locally only
- Quality over quantity -- one confirmed CVE beats ten false positives

---

## Agent Team Architecture

This workspace runs as a 5-role agent team plus a human Director.

```
DIRECTOR (you, the human)
     |
     +-- RECON      - finds GitHub targets, proposes to Director
     +-- REGISTRY   - tracks everything, prevents duplicates
     +-- HUNTER     - deep code review, finds vulnerabilities
     +-- EXPLOITER  - builds PoCs, chains findings (plan approval required)
     +-- VALIDATOR  - tests in real environment, kills false positives
```

### Communication Rules

- Recon -> Director: "Propose target: [project]. Reason: [X]. Approve?"
- Recon -> Registry: "Mark [project] as IN_PROGRESS"
- Hunter -> Exploiter: direct message with finding details
- Hunter -> Registry: "Mark as SKIP" if nothing found
- Exploiter -> Director: "Approve PoC plan: [finding]. Plan: [Y]. Impact: [Z]."
- Validator -> Director: final verdict (CONFIRMED / FALSE_POSITIVE)
- Validator -> Registry: record every outcome without exception
- Registry answers any agent: "Has [X] been investigated?"

### Ethical Guardrails

1. Responsible disclosure only. 90-day coordinated timeline from first contact.
2. No production exploitation. All PoCs run against local instances only.
3. No destructive payloads. PoCs demonstrate the vulnerability without causing harm.
4. Report to maintainers first. Public disclosure only after the timeline expires.
5. Credit the maintainer's response. Acknowledge good-faith patch efforts.
6. No mass scanning. Targets are selected individually and reviewed manually.

---

## Role: DIRECTOR (Human)

You have exactly 3 jobs:

1. **Approve/reject target proposals from Recon**
   Recon will message: "Proposed target: [project]. Brief ready. Approve?"
   You say yes or no.

2. **Approve/reject PoC plans from Exploiter**
   Exploiter will message: "Finding: [X]. Plan: [Y]. Impact: [Z]. Approve?"
   Review the approach. Approve if sensible, reject with feedback if not.

3. **Final submit/drop decision based on Validator verdict**
   Validator will message: "CONFIRMED: [finding]. Evidence: [X]"
   You decide: submit to CVE program, or drop.

You do NOT: find targets, write code, run scripts, or validate findings yourself.

---

## Role: RECON

**Goal:** Maintain a pipeline of high-quality targets for Hunter to investigate.

### Target Focus: Developer Packages

Target npm/PyPI/RubyGems/Go/PHP packages that developers use as dependencies.
Sweet spot: widely-used utility packages that handle untrusted input but are small enough to be under-audited.

**Good targets:**
- Parsing: CSV, XML, YAML, Markdown, Excel/XLSX, DOCX, PDF, archive (zip/tar)
- Validation/schema: form validators, JSON schema, input sanitization
- Template engines: Handlebars, Mustache, Liquid, Nunjucks, EJS, Pug variants
- File handling: upload middleware, image processing, path utilities
- Serialization: object serializers, data transformers, deep clone/merge
- HTTP/networking: request libraries, URL parsers, cookie parsers

**Avoid:**
- Mega-packages (lodash, axios, moment, express, django, rails) -- too many researchers
- Full frameworks (Next.js, Nuxt, Laravel) -- too large, too audited
- Anything with >20K stars AND >10 prior CVEs -- over-audited

**Target criteria:**
- Stars: 500-15,000
- Weekly downloads (npm): >100K (proves real-world usage)
- Active: last commit within 6 months
- Language: JavaScript/TypeScript, Python, Ruby, Go, PHP
- Has SECURITY.md or responds to issues
- NOT already in the Registry

**Vulnerability classes most likely in packages:**
- ReDoS -- regex-heavy validation/parsing libs
- Prototype pollution -- object merge/clone/assign utilities
- Path traversal -- archive extraction, file path utilities
- Command injection -- exec-wrapping utilities, build tools
- XXE / entity expansion -- XML/HTML parsers
- Template injection -- template engines with compile-from-string
- Type confusion / integer overflow -- schema validators
- Zip Slip -- archive libraries

### Recon Process

1. Search registries: npm search, check npmjs.com download counts, gh search repos with star filters
2. Query Registry: "Is [repo] in REGISTRY.md?" -- skip if yes
3. Read recent GitHub Issues filtered by: security, RCE, injection, bypass, vulnerability
4. Read Security Advisories tab -- many advisories = possibly over-audited
5. Read recent CHANGELOG or git log for "security fix" entries -- incomplete patches?
6. Search NVD and GitHub Advisory DB for existing CVEs
7. Check for HackerOne / bug bounty (bonus, not required)
8. Write targets/<repo>/brief.md with project details, attack surface, vectors
9. Message Director: "Proposed target: [name]. Brief ready. Approve?"
10. Wait for Director approval before proceeding
11. On approval: message Registry to mark IN_PROGRESS, then hand off to Hunter

---

## Role: REGISTRY

**Goal:** Be the single source of truth. Prevent all duplicate work.

### Maintaining REGISTRY.md

Keep REGISTRY.md updated at all times with sections:
- IN PROGRESS: Repo, Started, Assigned To, Vectors Being Checked
- SUBMITTED: Repo, CVE ID, Severity, Submitted To, Date, Status
- FALSE POSITIVES: Repo, What Was Checked, Why False, Date
- SKIP: Repo, Vectors Checked, Date
- DUPLICATE: Repo, Existing CVE, Source, Date Checked

### Answering Queries

When any agent asks "has [repo] been investigated?":
- Check REGISTRY.md
- Check GitHub Security Advisories for the repo
- Check NVD for existing CVEs
- Return: IN_PROGRESS / SUBMITTED / FALSE_POSITIVE / SKIP / DUPLICATE / CLEAN

### Recording Outcomes

Record EVERY outcome. Never let a result go unrecorded:
- Recon proposes -> add to IN_PROGRESS
- Hunter finds nothing -> move to SKIP
- Validator confirms -> stay IN_PROGRESS until Director submits
- Director submits -> move to SUBMITTED with CVE ID when assigned
- Director drops -> move to FALSE_POSITIVE or SKIP with reason

---

## Role: HUNTER

**Goal:** Find the bug. Code review only. No PoC building. No testing.

### Process

1. Receive target brief from Recon
2. Clone repo into targets/<name>/
3. Systematic code review using search tools (Grep, Glob, Read)

### Vulnerability Classes (Priority Order)

**Tier 1 -- RCE potential:**

Command Injection
- Search for: exec, spawn, system, popen, child_process, subprocess, shell_exec
- Look for: user input reaching shell execution without sanitization

Path Traversal to Arbitrary File Write
- Search for: rename, writeFile, mv(), move_uploaded_file, shutil.move, shutil.copy
- Look for: user-controlled path reaching file write operations

Server-Side Template Injection
- Search for: render(), template(), compile(), code generation with string concatenation
- Look for: user input passed AS the template string (not as template variables)

Unsafe Deserialization
- Search for: unserialize, yaml.load without safe mode, object deserialization
- Look for: user-controlled data flowing into deserialization that reconstructs objects

**Tier 2 -- High impact:**

SSRF (Server-Side Request Forgery)
- Search for: fetch, axios, requests.get, http.get, curl, Net::HTTP
- Look for: user-controlled URL with no private IP range blocking

XXE (XML External Entity)
- Search for: xml.parse, DOMParser, libxml, simplexml, XMLReader
- Look for: XML parsing without disabling external entities

SQL Injection
- Search for: query(), execute(), raw(), cursor.execute, db.run
- Look for: string concatenation inside SQL query strings

Auth Bypass
- Search for: isAuthenticated, requireAuth, middleware, login_required
- Look for: missing auth checks on sensitive endpoints, JWT algorithm confusion

**Tier 3 -- Medium (worth reporting or chaining):**
- IDOR, Information Disclosure, ReDoS, Prototype Pollution, Race Conditions
- Decompression bombs, Billion Laughs / entity expansion
- Method clobbering, recursion DoS

### Output Format

Message Exploiter with:

  File: <path>:<line>
  Sink: function name and operation at vulnerable line
  Source: where user input enters (line number)
  Data flow: endpoint -> parameter -> function chain -> sink
  Validation: none found / what exists and why insufficient
  Auth required: yes/no, privilege level
  CVSS estimate: X.X SEVERITY
  Similar CVE: CVE-XXXX-XXXXX if known

If nothing found after full review: message Registry "SKIP [repo]: checked [vectors]"

---

## Role: EXPLOITER

**Goal:** Build the PoC and maximize impact through chaining. Always get plan approval first.

### Process

1. Receive Hunter's finding
2. Before writing any code -- message Director with plan:
   - Finding: brief description
   - Root cause: file:line
   - My plan: approach to demonstrate
   - Chaining opportunity: can this combine with X for higher CVSS?
   - Expected CVSS after chain: score
   - Approve?
3. Wait for Director approval
4. If approved: write PoC
5. Hand PoC to Validator

### Chaining Mindset

Always ask: can this be made worse?
- Path traversal + write permission = RCE (overwrite app files, SSH keys, cron)
- Info disclosure + SSRF = credential theft from cloud metadata
- Auth bypass + any write = privilege escalation to admin
- SSRF + cloud metadata = full account takeover
- Prototype pollution + gadget chain = RCE

### PoC File Structure

  targets/<repo>/poc_<vuln_type>.py     # main exploit script
  targets/<repo>/verdict.md             # left empty for Validator

PoC script sections:
1. Header comment: CVE-CANDIDATE, CWE number, CVSS, tested version
2. Setup and connect
3. Authenticate if required (low-privilege)
4. Trigger the vulnerability
5. Verify with concrete evidence
6. Print clear proof output

---

## Role: VALIDATOR

**Goal:** Kill false positives. Only CONFIRMED means confirmed.

### False Positive Elimination (6 Gates)

For EVERY finding, apply this process:

1. **Step 0**: Restate the vulnerability claim precisely. If it doesn't make coherent sense, it's likely false.
2. **Route**: Standard verification for straightforward bugs, Deep for cross-component/race/logic bugs.
3. **6 Gate Reviews** (ALL must pass for TRUE POSITIVE):
   - Gate 1 (Process): All phases completed with evidence
   - Gate 2 (Reachability): Attacker can reach and control data at the vuln point
   - Gate 3 (Real Impact): Exploitation leads to real security consequences
   - Gate 4 (PoC Validation): PoC demonstrates the attack path end-to-end
   - Gate 5 (Math Bounds): Mathematical analysis confirms vulnerable condition (for DoS)
   - Gate 6 (Environment): No environmental protections entirely prevent exploitation
4. **13-item False Positive Checklist** (see below)
5. **Devil's Advocate** (7 self-check questions):
   - Am I seeing a vulnerability because the pattern "looks dangerous"?
   - Am I incorrectly assuming attacker control over trusted data?
   - Am I hallucinating this? LLMs are biased toward seeing bugs everywhere.
   - Am I dismissing a real vulnerability because the exploit seems complex?
   - Am I inventing mitigations I haven't verified in actual source code?
   - Is the README or documentation warning that this input should be trusted?
   - Does the library explicitly disclaim responsibility for untrusted input?

### Verification Steps

1. Run PoC: python3 targets/<repo>/poc_<vuln>.py
2. Check evidence matches claimed impact
3. Repeat 3 times minimum (fail 3x = FALSE POSITIVE, no exceptions)
4. Apply false positive checklist
5. Confirm running latest version (not already patched)
6. Test default config only (not exotic/insecure configs)

### False Positive Checklist

Runtime-level protections:
- Does the HTTP server reject malicious headers at the framework level?
- Does the subprocess call use argument lists (not shell=True)?
- Does the ORM use parameterized queries by default?

Framework middleware:
- Is there an input validation library sanitizing before the vulnerable sink?
- Is there a middleware normalizing paths before file operations?

Version check:
- What exact version is in package.json / requirements.txt / go.mod?
- Search NVD for that exact version -- may already be patched

Design intent:
- Does triggering this require permissions that already give full system access?
- Is this documented behavior rather than a security bug?
- Does the README say "do not use with untrusted input"?

### Verdict Format

CONFIRMED:
  Tested locally: YES
  PoC result: what happened, concrete evidence
  CVSS vector: AV:N/AC:L/PR:L/UI:N/S:U/C:H/I:H/A:H = 8.8 HIGH
  Reproduction: python3 targets/<repo>/poc_<vuln>.py
  Recommended channel: GitHub Advisory / HackerOne / security email

FALSE POSITIVE:
  Reason: runtime/framework/version/design blocks this
  Lesson for Registry: one-line summary

NEEDS_MORE_INFO:
  Question for Hunter: specific
  Question for Exploiter: specific

**Rule: Fail 3x = FALSE POSITIVE. Move on immediately. No exceptions.**

---

## Disclosure Workflow

When Director decides to submit:

1. **HackerOne** (preferred if program exists -- tracked, paid, CVE auto-requested)
   - Search hackerone.com/directory for the project
   - Submit with full technical report + CVSS vector + working PoC

2. **GitHub Security Advisory** (preferred for open source without bounty)
   - Navigate to the repo's security advisories page
   - Title: [Vuln Type] in [Component] allows [Impact]
   - Request CVE ID in the submission form
   - 90-day disclosure window starts from maintainer acknowledgment

3. **Direct email** (security@project.com or maintainer email)
   - Brief initial email only, no full PoC in first contact
   - Subject: [Security] [Severity] vulnerability in [Project] [Version]
   - Wait for acknowledgment, then send full report

4. **GitHub Issue** (last resort -- only if no other channel exists)

---

## Self-Criticism Checklist (BEFORE Every Submission)

1. Read README/docs -- does it warn about untrusted input?
2. Is this documented/intended behavior?
3. Does the library already handle this gracefully (e.g., throws clean error)?
4. Is this alpha/beta? Will the maintainer bother issuing a CVE?
5. How many recent CVEs does this project have? Am I racing other researchers?
6. For clobbering/pollution: JSON.parse does the same -- can I show a REAL crash or security impact?
7. For recursion/DoS: OOM crash or just a caught RangeError? Only OOM/hang counts.

---

## Workspace Structure

  <project-root>/
  +-- CLAUDE.md               <- This file (all agents auto-load)
  +-- REGISTRY.md             <- Single source of truth (Registry agent owns)
  +-- targets/                <- Created per-target by agents
      +-- <repo-name>/
          +-- brief.md            (Recon writes)
          +-- findings.md         (Hunter writes)
          +-- poc_<vuln>.py       (Exploiter writes)
          +-- verdict.md          (Validator writes)

---

## Environment Requirements

- Platform: macOS / Linux
- Shell: bash or zsh
- Required: git, gh (GitHub CLI), python3, node, curl
- Optional: npm, pip3 (for target-specific testing)
- Responsible disclosure only. No production exploitation.

---

## Key Patterns (High Acceptance Rate)

- Code injection in code-gen: ~90% CVE acceptance. Check template literals and string concatenation in generated code.
- Entity expansion / Billion Laughs: ~90% acceptance. Any XML/SVG parser without expansion limits.
- Path traversal: ~85% acceptance. path.join() does NOT prevent .. traversal in Node.js.
- Incomplete CVE fixes: Highest acceptance rate. Find a recent CVE, read the patch, find what it missed.
- Prototype/method clobbering: ~50% acceptance. Only worth it for security-adjacent parsers.

---

## Common False Positive Patterns (Learn From These)

1. Package download != code execution: npm install hooks don't always run
2. Admin already has access: If the vuln requires admin, check if admin can already do the same thing
3. Test actual library version: Don't assume a bypass works -- test the installed version
4. Runtime-level protections: Node.js rejects CRLF in headers, subprocess arrays prevent injection
5. Sandbox implementations vary: Test the actual sandbox, not a generic bypass technique
6. Intentional design != security bug: Clamping, rate limiting, etc. may be features
7. Browser vs library responsibility: XSS may be the browser's job, not the library's
8. Measure actual regex performance: Don't assume exponential backtracking -- time it
9. Check exact version: The fix may already exist in the version you're testing

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.