agentleFS
Sign inSign up

sage

gendigitalinc/sage/CLAUDE.md

Sage (Safety for Agents) — a lightweight Agent Detection & Response (ADR) layer for AI agents that guards commands, files, and web requests. Supports Claude Code, Cursor, OpenClaw, and OpenCode. See docs/developer-guide.md for the full architecture and development reference. TypeScript monorepo with five packages: E2E is split into layers: Layer 1 (deterministic detection + host I/O contract, in pnpm test), Layer 2 (real agents + Sage — wiring + payload/tool-name drift, never re-asserting detection; containerized or, for claude/opencode/cursor/copilot, native/no-Docker), and…

CLAUDE.md310 starsChanged 15 days ago
  • Reads credentials
  • Installs packages
  • Commits and pushes
# CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

## Project

Sage (Safety for Agents) — a lightweight Agent Detection & Response (ADR) layer for AI agents that guards commands, files, and web requests. Supports Claude Code, Cursor, OpenClaw, and OpenCode. See `docs/developer-guide.md` for the full architecture and development reference.

## Commands

```bash
pnpm install                               # Install all dependencies
pnpm test                                    # Run all tests (builds automatically, E2E excluded)
pnpm test -- --reporter=verbose              # Verbose test output
pnpm test -- packages/core/src/__tests__/extractors.test.ts  # Single test file
pnpm test -- -t "test name"                  # Run single test by name
e2e/run.sh <agent>                           # Layer 2 containerized live E2E (claude|copilot|opencode|cursor|openclaw|vscode|all); needs Docker + e2e/.env
SAGE_E2E_RUNNER=native pnpm test:e2e:<agent> # Layer 2 native live E2E, no Docker (claude|opencode|cursor|copilot-cli); needs the CLI installed + authenticated
pnpm test:e2e:cursor                         # Layer 3 desktop Extension Host (installed Cursor binary); :vscode for VS Code
pnpm build                                  # Build all packages (tsc + esbuild bundle)
pnpm build:sea                              # Build standalone SEA binaries (requires official Node.js, not Homebrew)
pnpm lint                                   # Lint with Biome
pnpm lint:fix                               # Lint + auto-fix
pnpm check                                  # Type check all packages
pnpm changeset                              # Create a changeset for your changes
pnpm run version                            # Apply changesets: bump versions, generate changelogs, sync manifests
bash scripts/pre-release-audit.sh           # Pre-release content audit (requires .gitleaks.toml symlink)
pnpm eval:pi                                 # PI (prompt-injection) accuracy benchmark (manual; see Test Tiers)
```

## Repository Structure

TypeScript monorepo with five packages:

- `packages/core/` — `@gendigital/sage-core`: Platform-agnostic detection engine (extractors, heuristics, URL check client, decision engine, caching)
- `packages/claude-code/` — `@gendigital/sage-claude-code`: Claude Code MCP hook tools and session-start entry point
- `packages/openclaw/` — `@gendigital/sage-openclaw`: OpenClaw plugin connector (in-process hook, native approval). `resources/` is gitignored — generated by `pnpm build` via `sync-assets.mjs`.
- `packages/opencode/` — `@gendigital/sage-opencode`: OpenCode plugin connector
- `packages/extension/` — `sage-cursor`: Cursor and VS Code extensions
- `threats/` — YAML threat definitions (data, not code). `dummy.yaml` contains harmless canary rules for E2E testing
- `trusted-domains/` — Trusted domain allowlists
- `hooks/` — `hooks.json` registering hooks with Claude Code

## Test Tiers

E2E is split into layers: **Layer 1** (deterministic detection + host I/O contract, in `pnpm test`), **Layer 2** (real agents + Sage — wiring + payload/tool-name drift, never re-asserting detection; containerized or, for claude/opencode/cursor/copilot, native/no-Docker), and **Layer 3** (desktop GUI Extension Host). See `docs/developer-guide.md#test-tiers`.

| Tier | Scope | Files | Requires |
|------|-------|-------|----------|
| 1 — Unit | Core library functions | `packages/core/src/__tests__/*.test.ts` | pnpm dev deps |
| 2 — Layer 1 (contract/integration) | Detection + host contract + tool-name maps + connector behaviors (registration, prompt injection), via mock API | `packages/*/src/__tests__/integration.test.ts`, `*contract*.test.ts`, `packages/openclaw/.../e2e-integration.test.ts` | pnpm dev deps |
| 3 — Layer 2 (containerized live E2E) | Real agent + Sage in Docker; canary-deny wiring + drift | `packages/{claude-code,openclaw,opencode}/src/__tests__/e2e.test.ts`, `packages/extension/src/__tests__/e2e-copilot-cli.test.ts`, cursor-headless + vscode-container blocks of `packages/extension/src/__tests__/e2e.test.ts` | Docker + `e2e/.env`; `e2e/run.sh <agent>` |
| 3 — Layer 2 (native live E2E) | Same canary-deny wiring, no Docker; drift/tool-catalog checks stay container-only | Same files, minus `openclaw` (container-only) | Installed + authenticated CLI; `SAGE_E2E_RUNNER=native pnpm test:e2e:<agent>` |
| 3 — Layer 3 (desktop GUI) | Sage extension in installed Cursor / VS Code Extension Host | Block A of `packages/extension/src/__tests__/e2e.test.ts` | Installed Cursor / VS Code binary |
| 4 — ML accuracy benchmark | PI model recall + FP rate on benign/injection/IOC fixtures | `packages/core/scripts/eval-pi-accuracy.mjs` + `packages/core/src/__tests__/fixtures/pi-{benign,injection,ioc-snippets}*.json` | pnpm dev deps + model present at `~/.sage/models/<schema>/pi-model/` |

`pnpm test` runs tiers 1–2 automatically (builds via `globalSetup` before running). Tier 3 is excluded — run Layer 2 with `e2e/run.sh <agent>` (or `all`) or natively with `SAGE_E2E_RUNNER=native pnpm test:e2e:<agent>`, Layer 3 with `pnpm test:e2e:cursor` / `:vscode`.

**Layer 2 prerequisites (container):** Docker + a populated `e2e/.env` (copy `e2e/.env.example`). Suites are gated on `SAGE_E2E_RUNNER=container` (set by `run.sh`) and skip otherwise. Auth via `e2e/.env`: claude/opencode/openclaw use **Vertex ADC** (no API key — never `ANTHROPIC_API_KEY`/`GEMINI_API_KEY`), copilot uses a Copilot-entitled `GITHUB_TOKEN`, cursor uses `CURSOR_API_KEY`, vscode needs nothing. See `docs/developer-guide.md#layer-2--containerized-live-e2e`.

**Layer 2 prerequisites (native):** no Docker, no `e2e/.env` — just the CLI installed and already authenticated (claude, opencode, cursor-agent, or copilot; `openclaw` has no native path). Gated on `SAGE_E2E_RUNNER=native`, isolated to a temp HOME, reuses ambient auth (never the real `~/.sage`/`~/.claude`/`~/.cursor`/`~/.copilot`). First Windows-facing E2E path in this project — no Windows E2E CI exists for any layer, so verify locally before relying on it there. See `docs/developer-guide.md#layer-2--native-live-e2e-no-docker`.

**Layer 3 prerequisites:** Installed Cursor / VS Code executable (skips if absent). Optional overrides: `SAGE_CURSOR_PATH`, `SAGE_VSCODE_PATH`, `VSCODE_EXECUTABLE_PATH`.

**Tier 4 (ML accuracy):** `pnpm eval:pi` runs the cached PI classifier against three fixtures (50 benign, 50 synthetic injections, 10 sanitized real-world IOC snippets from Unit 42's IDPI research) and prints per-suite recall, FP rate, and per-category breakdown. Pure observability — no assertions, never gates CI. The model is no longer in the repo; the script reads it from `~/.sage/models/<schema>/pi-model/`. Run a Sage session with `pi_check.enabled = true` once to populate that directory, or place the files manually. Use it before/after model swaps or threshold tuning to see how detection shifts.

## Architecture

This is a **multi-platform plugin** with three connectors and a shared core:

1. **Claude Code Connector** (`packages/claude-code/src/`) — `mcp-server.ts` exposes long-lived MCP tools for PreToolUse/PostToolUse hook events, backed by `mcp-hook-tools.ts` and shared hook handlers. `session-start.ts` scans installed plugins for threats. Bundled by esbuild into single CJS files.

2. **OpenClaw Connector** (`packages/openclaw/src/`) — In-process plugin using `api.on('before_tool_call')`. Flagged actions use OpenClaw's native `requireApproval` mechanism with an `onResolution` callback that persists allowlist entries. Bundled into a single CJS file.

3. **Core Library** (`packages/core/src/`) — Platform-agnostic TypeScript library:
   - `extractors.ts` — Extracts URLs, commands, file paths from tool inputs
   - `threat-loader.ts` — Loads YAML threat definitions from `threats/`
   - `heuristics.ts` — Matches extracted artifacts against threat patterns
   - `tool-names.ts` — Canonical tool vocabulary and generic canonicalization helper
   - `clients/url-check.ts` — HTTP client for URL reputation checking and endpoint resolver
   - `engine.ts` — Decision engine combining signals into a Verdict
   - `cache.ts` — JSON file verdict cache with TTLs
   - `allowlist.ts` — User allowlist management
   - `trusted-domains.ts` — Trusted domain loading and matching
   - `audit-log.ts` — JSONL audit log for verdicts
   - `installation-id.ts` — Persistent installation UUID (`~/.sage/installation-id`)
   - `version-check.ts` — Version check via POST with environment context
   - `session-start.ts` — Session start orchestrator (scan + version check)
   - `plugin-scanner.ts` — Plugin file scanning for threats
   - `plugin-scan-cache.ts` — Plugin scan result cache
   - `content-snapshot.ts` — Structured `content` snapshot builder (per-field caps + home-path scrubbing) shared by audit log, detection telemetry, and FP reporting
   - `extended-info.ts` — `~/.sage/extended-info.json` loader/sanitizer + `mergeExtendedInfo` helper
   - `product-version.ts` — Platform-agnostic `product.json` version reader used by hook runner and MCP server child processes

4. **Plugin Config** — `hooks/hooks.json` registers hooks, `.claude-plugin/plugin.json` is the plugin manifest, `skills/security-awareness/SKILL.md` provides security knowledge to Claude.

**Data flow (Claude Code):** `mcp_tool` hook call → `sage_claude_pre_tool_use` / `sage_claude_post_tool_use` on the Sage MCP server → normalize hook input → canonicalize tool name → extract artifacts → check allowlist → check cache → heuristics + URL check → DecisionEngine → cache result → audit log → verdict JSON returned as MCP tool result

**Data flow (OpenClaw):** `before_tool_call` event → canonicalize tool name → extract artifacts → check allowlist → check cache → heuristics + URL check → DecisionEngine → cache result → audit log → block/pass

## Key Conventions

- **Never skip git hooks.** Do not use `--no-verify` on `git commit` or `git push`. The hooks run gitleaks, lint, build, typecheck, and tests. If a hook fails, fix the underlying issue.
- All detection patterns are data (YAML in `threats/`), not hardcoded
- Fail-open: every internal error path returns an allow verdict; extension hooks always exit with code `0`
- Verdicts use three decisions: `allow`, `deny`, `ask` with confidence scoring
- Sensitivity presets: paranoid (0.70), balanced (0.85), relaxed (0.95)
- Merge precedence when multiple signals fire: `deny > ask > allow`
- URL check works without API key for basic malware/phishing detection
- Standalone binary via Node.js SEA for distribution (no runtime deps)
- Naming: YAML/JSON data uses `snake_case` (`threat_id`, `source_file`); TypeScript interfaces use `camelCase` (`threatId`, `sourceFile`). Conversion functions bridge the two at serialization boundaries.
- **Connectors own tool name canonicalization.** Core defines the canonical vocabulary (`CanonicalToolType`) but has no knowledge of platform-specific names. Each connector maps its raw tool names to canonical form before calling the evaluator.

## Pre-PR Checklist

If you use Claude Code, run `/simplify` before creating a pull request to review changed files for clarity, consistency, and maintainability improvements. Apply any suggested changes before opening the PR.

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.