roam-code
Cranot/roam-code/AGENTS.md
Package: roam-code on PyPI. Entry point: roam.cli:cli. This repo is public. The placement rule is physical, not pattern-based: When in doubt about a file: if it's a planning artifact or session-cadence output, write it under internal/. If you intend it to be public, write it in the appropriate public folder. Anti-pattern history: the repo used to use an enumerative gitignore in dev/ (whitelisting specific filename templates to exclude), which fail-opened on any new memo family. Fixed by the rule above:…
AGENTS.md519 starsChanged 5 months ago
- Reads credentials
- Installs packages
# AGENTS.md — roam-code development guide
## What this project is
<!-- BEGIN auto-count:Codex-headline -->
roam-code is a local codebase intelligence CLI for developers and AI coding agents.
It pre-indexes symbols, call graphs, dependencies, architecture, and git history into
a local SQLite DB. **287 commands · 246 MCP tools (17 in the default `core` preset) · 28 languages · local analysis · zero API keys.**
<!-- END auto-count:Codex-headline -->
<!-- BEGIN auto-count:Codex-authoritative -->
Authoritative counts (AST-derived, env-independent): `command_count: 287 · canonical_count: 280 · category_count: 7 · mcp tools registered: 246 · mcp tools in core preset: 17`. The `roam surface --json` envelope additionally exposes `mcp_tool_count_by_preset` for per-preset counts.
<!-- END auto-count:Codex-authoritative -->
**Package:** `roam-code` on PyPI. Entry point: `roam.cli:cli`.
## Documentation Hub
- **Orientation** (read once before anything else): [docs/understanding-roam.md](docs/understanding-roam.md) — what roam is, why it is mechanical and agent-first, what has been measured, how we speak, and how to gate on its output.
- **Dogfood corpus** (read for quality lessons): `internal/dogfood/` — 212-eval corpus + 6 systemic-pattern synthesis. Single most important reference for understanding what good roam-command behavior looks like. Start with `internal/dogfood/README.md`. (Private — gitignored; not shipped to PyPI/GitHub.)
- Getting started tutorial: `templates/distribution/landing-page/docs/getting-started.html`
- Command reference with examples: `templates/distribution/landing-page/docs/command-reference.html`
- Architecture guide and diagram: `templates/distribution/landing-page/docs/architecture.html`
- Detector evidence and limitations: [docs/concepts/detector-evidence.md](docs/concepts/detector-evidence.md) — static graph coverage, language context, confidence semantics, and command-scope differences.
- Live URL: https://roam-code.com/docs/
## Where files go (private vs public)
This repo is public. The placement rule is physical, not pattern-based:
- **`internal/`** — the ONE private folder. Wholesale gitignored. Everything that shouldn't be on GitHub goes here: session memos, sprint plans, release-readiness drafts, monetization research, dogfood batches, smoke results, generated test fixtures, scratch notes, anything you'd be uncomfortable seeing in a public diff. Sub-conventions: `internal/dogfood/` (eval corpus), `internal/planning/` (session / sprint / release memos).
- **Everything else** — public by default. `src/`, `tests/`, `templates/`, `dev/`, root files. No magic in the gitignore, no whitelists, no per-extension rules.
When in doubt about a file: if it's a planning artifact or session-cadence output, write it under `internal/`. If you intend it to be public, write it in the appropriate public folder.
Anti-pattern history: the repo used to use an enumerative gitignore in `dev/` (whitelisting specific filename templates to exclude), which fail-opened on any new memo family. Fixed by the rule above: one private folder, no pattern magic.
## Quality discipline (from `internal/dogfood/` + agi-in-md)
This section codifies what makes a roam command good. Distilled from 212 dogfood evals + 1000+ prompt-design experiments. Treat as constraints when adding or modifying commands.
### Six systemic anti-patterns to NEVER ship
From the dogfood synthesis notes — validated unchanged across 30 → 59 → 212 evals as failure classes. Several of the original incidents are now SEALED behind regression tests; they are kept here as regression-invariant examples, not as claims that the bug is currently live.
1. **Pattern-1 family — "structured signal lost or never reached".** One root failure family. (A) Hang on missing prerequisite — SEALED (live guard `src/roam/mcp_extras/preflight.py`). (B) Structured signal collapsed to generic `COMMAND_FAILED` by an intermediate layer — SEALED (try-parse stdout as JSON at the wrapper-bridge). (C) Empty-stdout crash on `json.loads()` — SEALED (CLI always emits a structured envelope, even on no-results). (D) Silent success on degraded resolution — LIVE: disclose the resolution state via a `resolution` field + `partial_success: true` + a degraded verdict. Every wrapper that cannot complete normally emits the canonical failure envelope (closed `status` / `error_code` enums; `isError: true` inside a successful JSON-RPC result).
2. **Silent fallback.** Never emit `verdict: "SAFE"` / `"completed"` / `"non-conformant"` when the underlying check failed or didn't run. Make absent state explicit: `state: "not_initialized"`, not `state: "broken"`. Historical example: `for_refactor` once reported `verdict: "compound operation completed"` despite 4/4 subcommand failures — the guard now lives in `_compound_envelope()` and its tests (subcommand failure must set `partial_success: true` and name the failed subcommands).
3. **Vocabulary mismatch family.** (a) Cross-command metric divergence — different commands report different "callers" / "complexity" / "AI rot" / "public symbols" for the same input; fix by stamping a `<metric>_definition` sidecar field (`caller_metric_definition: "raw_edge_rows"` pattern). (b) Cross-MCP parameter-name divergence — 9+ MCP parameter names refer to similar concepts (`symbol` vs `name`; `path` vs `paths` vs `file`; `query` vs `queries` vs `patterns` vs `prefix`); fix via the `_PARAM_ALIASES` table in `src/roam/mcp_server.py` with boundary normalization at wrapper-dispatch time. The lint `tests/test_mcp_param_names.py` blocks new wrappers re-introducing legacy names.
4. **Conventions detector inconsistency — RESOLVED (Fix G).** All 5 sites (`describe`, `understand`, `minimap`, `preflight`, `conventions` standalone) now delegate to the canonical `conventions_helper.compute_conventions()`. Any new convention-aware command MUST call the helper; `--persist` to the findings registry lives ONLY on the standalone `conventions` command.
5. **Compound-recipe internal command-name drift — SEALED.** Use **registry-key lookup** at compound-init time, never string-concat. Live in `_cr()` / `_COMPOUND_REGISTRY` (`src/roam/mcp_server.py`); guard test `tests/test_compound_recipe_registry.py` AST-scans recipes against `cli._COMMANDS | cli._DEPRECATED_COMMANDS`.
6. **Response volume family.** (a) Auto-handle pattern — caller polls for chunks via `_wrap_with_handle_off`; FULLY ADOPTED, every `@_tool` command inherits it. (b) File-write pattern — writes to disk + returns a tiny envelope (`graph-export`, `fingerprint`, `index-bundle`, `cga`, `agent-export`). (c) `fetch_handle`'s own crash on large handles — SEALED at v2.0.0 (fully paginated byte-slice + section-pick + jq-projection modes). Mandate: any response >20K tokens MUST use 6a OR 6b.
### Twelve agi-in-md laws applied to roam-code
The historical "LAW" labels below name prompt-design observations from the
stated Haiku/Sonnet/Opus 4.5/4.6 experiments, not universal findings across
models, tasks, or instruction carriers. Their translations are maintained
Roam interface constraints. Historical comparisons do not establish current
effect sizes or authorize routing, automatic promotion, or weaker verification.
Use [verification evidence](docs/concepts/verification-evidence.md) to qualify
changes against concrete outcomes and a simpler baseline.
1. **[LAW] The prompt is the dominant variable.** (D13) — In roam: the **JSON envelope shape is the dominant variable** for agent integration. A 5-signal envelope (`preflight`'s blast/complexity/conventions/coupling/fitness) outperforms a 12-signal envelope on agent-decision speed. Quality of envelope > volume of fields.
2. **[LAW] Imperatives beat descriptions.** (D3) — Tool descriptions and `next_commands` must use imperative voice. "Run `roam impact handleSave` to see callers." NOT "This command shows callers." Validated across 50+ MCP tool descriptions in the dogfood.
3. **[LAW] The prompt is a program; operation order becomes section order.** (D4) — In roam: the order of fields in `agent_contract.facts` determines what agents act on first. Put the actionable verdict in `facts[0]`. Validated by the difference between `preflight` (verdict-first, agents preflight before edit) and `complexity` (numeric-first, agents skip the gate).
4. **[LAW] "Code" nouns activate analytical mode on any input.** (D15) — In roam: `agent_contract.facts` strings should anchor on concrete nouns ("`useThemeClasses` has 528 callers") not abstract ones ("this symbol has many callers"). Concrete nouns activate analytical processing; abstract nouns activate summary mode.
**Concrete-noun anchor vocabulary**: the LAW 4 lint at `tests/test_law4_lint.py` accepts a fact string as concrete-noun-anchored if its terminal token (last word, punctuation stripped) is in a known anchor set. Authoritative sources: `src/roam/output/formatter.py:concrete_plural_terminals` (99 entries, drives the humanizer's "skip findings suffix" rule) and `tests/test_law4_lint.py:_CONCRETE_NOUN_ANCHORS` (116 entries = 99 shared with the formatter + 17 SBOM/registry additions; mirrors the formatter set per the `# Keep these two lists in sync.` comment, and the count drift is pinned by `tests/test_law4_anchor_counts.py`). Representative entries — consult the source files for the full list:
- **Code structure**: `files`, `symbols`, `edges`, `nodes`, `cycles`, `clusters`, `layers`, `modules`, `commands`, `tools`, `capabilities`, `imports`, `endpoints`, `dependencies`, `packages`, `routes`
- **Findings**: `findings`, `hotspots`, `smells`, `violations`, `warnings`, `errors`, `alerts`, `issues`, `gaps`, `leaks`, `secrets`, `vulnerabilities`
- **Quality metrics**: `keys`, `values`, `chars`, `lines`, `tokens`, `bytes`, `items`, `entries`, `records`, `fields`
- **Past participles / state qualifiers**: `passed`, `failed`, `scanned`, `checked`, `affected`, `scored`, `confirmed`, `analyzed`, `skipped`, `reached`
- **Time units**: `days`, `weeks`, `months`, `years`, `hours`, `minutes`, `seconds`, `milliseconds`
When writing a new fact, ensure the terminal token is in the anchor set. If not, either rephrase to anchor on a different terminal OR add the new noun to BOTH `src/roam/output/formatter.py:concrete_plural_terminals` AND `tests/test_law4_lint.py:_CONCRETE_NOUN_ANCHORS`. The test set is a deliberate superset of the formatter set (formatter has 99 entries; the test mirrors all of them and adds 17 SBOM/registry-domain terminals — `capabilities`, `commands`, `tools`, `packages`, `phantom`, `reachable`, etc.). The mirror is hand-maintained rather than imported so the lint stays decoupled from `roam.output.formatter` — see the `# Keep these two lists in sync.` comment in the test file.
Example:
- WRONG: `"7 of 10 capabilities are AI-safe"` (ends on `AI-safe`, not anchored)
- RIGHT: `"7 of 10 AI-safe capabilities"` (ends on `capabilities`, anchored)
5. **[LAW] ≤3 concrete steps = universal; 9+ abstract steps = catastrophic on Haiku.** (D16) — In roam: compound recipes that chain 5+ subcommands are unreliable when an agent on Haiku consumes them. Either chain ≤3 OR use a registry recipe that the runtime expands at dispatch time. Validated by `for_refactor`'s 4-subcommand chain being broken vs `pr_prep`'s 3-subcommand chain working cleanly.
6. **[LAW] Compression forces domain neutrality.** (D17) — In roam: the `summary.verdict` line MUST work without any other field. Agents that consume only the verdict do not load the full envelope. `verdict: "Healthy 32/100 with 12 cycles"` works; `verdict: "see details"` fails.
7. **[CONSTRAINT] Use positive vocabulary, not negative constraints.** — In roam: error envelopes should name what works, not what's forbidden. `"Use --gate-pattern to filter"` beats `"Do not call without a filter"`. Same applies to `risks[]` arrays — name the surviving risk, not the absent guard.
8. **[CONSTRAINT] Use semantically meaningful operation names.** — In roam: this is exactly the `vuln`/`vulns` typo bug. Internal command names must be a CLOSED ENUMERATION, not free string composition. Fix: registry-key lookup; prevention: a CI lint that fails when a compound recipe references a command name that isn't in `cli._COMMANDS`.
9. **[CONSTRAINT] Coupling lives in what steps SAY, not output format specs.** — In roam: compound recipes should compose by **shared input/output types**, not by string-templated arg passing. The `for_bug_fix` / `diagnose_issue` divergence on `handleSave` resolution (different file picked!) is exactly this bug — they share a name string, not a resolved symbol id.
10. **[CONSTRAINT] Drop negative framing + 2-phase output specs from agent contracts.** (V13 ablation: each removal = +1 compliance) — In roam: `agent_contract.facts` should be a flat list of positive assertions. No "first verify X, then check Y" structure inside facts. One claim per fact, all positive.
11. **[CONSTRAINT] Identity/persona > step enumeration for single-call depth.** — In roam: tool descriptions should describe what the tool IS, not what it DOES step-by-step. "Pre-change safety gate" beats "First runs blast, then complexity, then conventions, then fitness." Identity activates the right consumption pattern.
12. **[CONSTRAINT] First-token EXECUTABILITY is a separate axis from shape compliance.** (CP44) — In roam: a verdict like `verdict: "Run roam preflight handleSave"` must produce a literally executable command. The 7919-partition `partition` output technically conforms to schema but is *not actionable* — the partition count is unusable. Shape and executability are separate quality axes; both must pass.
### "Never N/A without running it" — the operational rule
The single hardest-earned lesson from the 212-eval corpus. Three commands were marked N/A by judgment at the 59-eval midpoint (`py_modern`, `py_types`, `metrics_push`). All 3 returned real signal when actually invoked:
- `py_modern` discovered Python files in `scripts/` (judged "no Python repo" — wrong)
- `py_types` reported 90% type coverage on those files
- `metrics_push --dry-run` exposed a **unique `danger_score` metric** not surfaced by any other command
**When adding tests / dogfooding / triaging: run every command at least once.** Empty output is itself signal; non-empty output on a "no X" project is the strongest signal of all. See `internal/dogfood/EVALS-HOW-TO.md` for the full lesson.
### Adding-a-command checklist (informed by patterns 1-6 above)
For evidence and gate corrections, freeze a defect-specific regression test
and inspect its failure against the unfixed implementation. Then run the same
test on the repair, include valid-case controls, and exercise the real CLI or
serialized producer/consumer path. A passing helper or a schema-valid artifact
does not establish the whole boundary. Preserve incomplete observations and
their denominator; do not upgrade claims merely by replacing source hashes.
For detector corrections, add paired positive/negative controls: remove the
reported false positive while retaining a genuine finding. Inspect matched
source, receiver identity, loop placement, and language/framework applicability;
method names alone do not establish database effects, recursion, or safe
optimization. Reuse shared resolution helpers and document their limitations.
Use the [detector evidence guide](docs/concepts/detector-evidence.md) as the
cross-command interpretation reference; keep dated audit results in `internal/`.
Before merging a new `cmd_X.py`:
- [ ] JSON mode handles empty input cleanly — emit a non-empty envelope, never empty stdout (Pattern 1)
- [ ] `summary.verdict` is a single line that works without any other field (LAW 6)
- [ ] `summary.partial_success: true` whenever ANY subcommand or check failed; no silent SAFE (Pattern 2)
- [ ] If output may exceed 20K tokens, use the handle pattern (`roam_explore` template) OR write to file (`graph-export` template) (Pattern 6)
- [ ] If it reports a "callers" / "complexity" / "rot" / "compliance" count, include a `<metric>_definition` field naming the precise computation (Pattern 3)
- [ ] If it's a compound recipe, reference subcommands by **registry key** not by string-templated CLI invocation (Pattern 5)
- [ ] `agent_contract.facts` is flat, imperative, concrete-noun-anchored (LAWs 2, 4, 10)
- [ ] If it depends on missing state (no migrations / no audit trail / no diff), the verdict says so explicitly — `"chain not initialized"` not `"chain broken"` (Pattern 2)
- [ ] Add at least one MCP-level test that runs the compound's actual internal subcommand chain (catches `vuln`/`vulns`-class typos)
- [ ] Tool description uses imperative voice ("Run X") not declarative ("This command") (LAW 2)
- [ ] If it returns a verdict that names a follow-up command, that command must be a literal `roam <subcommand>` string — copy-paste-executable (CONSTRAINT 12)
- [ ] Document which Bash/Read/Grep pattern this command displaces in the command's docstring or a dedicated test — answers "does this displace a question the LLM actually has?" (DOGFOOD-RATIO-TRACKING.md)
## Quick reference
```bash
# Reproduce the locked development and documentation environment
uv sync --locked --no-default-groups --extra dev --group ci --python 3.12
# Run tests
uv run --no-sync pytest tests/
# Run tests in parallel (requires pytest-xdist)
uv run --no-sync pytest tests/ -n 4 --dist loadgroup
# Skip timing-sensitive perf tests
uv run --no-sync pytest tests/ -m "not slow" -n 4 --dist loadgroup
# Run a single test file
uv run --no-sync pytest tests/test_comprehensive.py -x -v -n 0
# Index roam itself
uv run --no-sync roam index
uv run --no-sync roam doctor
uv run --no-sync roam health --explain
```
Use [CONTRIBUTING.md](CONTRIBUTING.md) for development and release procedures,
[docs/repository-maintenance.md](docs/repository-maintenance.md) for Git,
environment, and index checks, and [docs/README.md](docs/README.md) for the
maintained documentation map.
## Architecture
### Directory layout
```
src/roam/
cli.py # Click CLI entry point — LazyGroup, _COMMANDS dict, _CATEGORIES. 287 command names (280 canonical + 7 aliases).
mcp_server.py # FastMCP server (17 tools in core preset; 246 in `full`) + `roam mcp` CLI command
mcp_extras/ # MCP-native enhancements: sampling, watcher, session, progress, completions
sampling.py # Sampling-driven result compression (summarize=True) via Context.sample
watcher.py # watchdog observer + notifications/resources/updated (opt-in via ROAM_MCP_WATCH)
session.py # Per-session symbol memory; auto-injected into retrieve/context ranking
progress.py # Phase-aware progress: parses indexer stderr for discover/parse/extract/resolve/graph
completions.py # FTS5-backed prefix completion for symbols/paths/commands + protocol-level handler
__init__.py # Version string (reads from pyproject.toml via importlib.metadata)
db/
schema.py # SQLite schema (CREATE TABLE statements)
connection.py # open_db(), ensure_schema(), batched_in(), migrations
queries.py # Named SQL constants
index/
indexer.py # Full pipeline: discovery → parse → extract → resolve → metrics → health → cognitive load
discovery.py # git ls-files, .gitignore
parser.py # Tree-sitter parsing
symbols.py # Symbol + reference extraction
relations.py # Reference resolution → edges
complexity.py # Cognitive complexity (SonarSource-compatible)
git_stats.py # Churn, co-change, blame, entropy
incremental.py # mtime + hash change detection
file_roles.py # Smart file role classifier (source, test, config, docs, etc.)
test_conventions.py # Pluggable test naming adapters (Python, Go, JS, Java, Ruby, Apex)
bridges/
base.py # Abstract LanguageBridge — cross-language symbol resolution
registry.py # Bridge auto-discovery + detection
bridge_salesforce.py # Apex → Aura/LWC/Visualforce bridge
bridge_protobuf.py # .proto → Go/Java/Python stubs bridge
bridge_rest_api.py # Frontend HTTP calls → backend route definitions
bridge_template.py # Jinja2/Django/ERB/Handlebars variable + include resolution
bridge_django.py # Django admin/serializer/form/URL → Model + view resolution
bridge_config.py # Env var reads → .env/.yml definitions
catalog/
tasks.py # Universal algorithm catalog — 34 tasks with ranked solution approaches
detectors.py # Algorithm anti-pattern detectors — query DB signals to find suboptimal patterns
languages/
base.py # Abstract LanguageExtractor — all languages inherit this
registry.py # Language detection + grammar aliasing
*_lang.py # One file per language (python, javascript, typescript, java, go, rust, c, csharp, php, ruby, kotlin, swift, scala, sql, foxpro, apex, aura, visualforce, sfxml, hcl, yaml, dart, generic)
graph/
builder.py # DB → NetworkX graph
pagerank.py # PageRank + centrality metrics
cycles.py # Tarjan SCC + tangle ratio
clusters.py # Louvain community detection
layers.py # Topological layer detection — returns {node_id: layer_number}
pathfinding.py # k-shortest paths for trace
dark_matter.py # Hidden co-change coupling detection
diff.py # Graph-level diff analysis
propagation.py # Propagation cost computation
spectral.py # Fiedler vector bisection + spectral gap
anomaly.py # Statistical anomaly detection (Modified Z-Score, Theil-Sen, Mann-Kendall, CUSUM)
simulate.py # Counterfactual architecture simulation (graph cloning + transforms)
partition.py # Multi-agent work partitioning (Louvain-based)
fingerprint.py # Topology fingerprinting + comparison
search/
tfidf.py # Zero-dependency TF-IDF semantic search
index_embeddings.py # Symbol corpus + cosine similarity
security/
vuln_store.py # Vulnerability ingestion (npm/pip/trivy/osv audit)
vuln_reach.py # Reachability analysis from vuln → entry points
runtime/
trace_ingest.py # OpenTelemetry/Jaeger/Zipkin trace ingestion
hotspots.py # Runtime hotspot classification (UPGRADE/CONFIRMED/DOWNGRADE)
refactor/
codegen.py # Import generation (Python/JS/Go)
transforms.py # move/rename/add-call/extract symbol transforms
# Agent-OS substrates (2026-05-12 sprint) — repo-local state under .roam/:
# constitution/ (constitution.yml gates), modes/ (read_only/safe_edit/migration/autonomous_pr),
# runs/ (HMAC-chained per-run event ledger), leases/ (multi-agent claims),
# memory/ (portable agent memory.jsonl), pr-bundles/ (proof-carrying PRs),
# laws/ (mined invariants), agents_md/ (AGENTS.md generator).
# Surfaced via: roam constitution, mode, runs, lease, memory, pr-bundle, laws,
# agents-md, brief, next, agent-score, intent-check, replay.
commands/
resolve.py # Shared symbol resolution + ensure_index()
changed_files.py # Shared git changeset detection
gate_presets.py # Framework-specific gate rules + .roam-gates.yml loader
graph_helpers.py # Shared graph utilities (adjacency builders, BFS helpers)
context_helpers.py # Data-gathering helpers extracted from cmd_context.py
cmd_*.py # 278 command modules: 276 back the 287 default names; 2 are feature-gated
output/
formatter.py # Token-efficient text formatting, abbrev_kind(), loc(), format_table(), to_json(), json_envelope()
sarif.py # SARIF 2.1.0 output (--sarif flag on health/debt/complexity)
schema_registry.py # JSON envelope schema versioning + validation
tests/ # 1100+ test_*.py files
# Core & legacy
test_basic.py, test_comprehensive.py, test_fixes.py, test_performance.py,
test_resolve.py, test_salesforce.py, test_v6_features.py,
test_v7_features.py, test_v71_features.py, test_v82_features.py,
test_workspace.py, test_visualize.py, test_foxpro.py,
# Organized command tests
test_commands_exploration.py, test_commands_health.py, test_commands_architecture.py,
test_commands_workflow.py, test_commands_refactoring.py,
# Feature-specific
test_smoke.py, test_json_contracts.py, test_formatters.py, test_languages.py,
test_anomaly.py, test_file_roles.py, test_pr_risk_author.py, test_dead_aging.py,
test_bridges.py, test_bridges_extended.py, test_test_conventions.py, test_gate_presets.py,
test_python_extractor_v2.py, test_math.py, test_properties.py, test_index.py,
# v9.1 new commands
test_simulate.py, test_orchestrate.py, test_fingerprint.py, test_mutate.py,
test_adversarial.py, test_plan.py, test_cut.py, test_invariants.py,
test_bisect.py, test_intent.py, test_closure.py, test_rules.py,
test_vuln.py, test_runtime.py, test_relate.py, test_semantic_search.py,
test_schema_versioning.py, test_sarif_flag.py, test_ruby.py, test_yaml_hcl.py,
test_dark_matter.py, test_effects.py, test_effects_propagation.py,
test_capsule.py, test_forecast.py, test_path_coverage.py,
test_minimap.py, test_attest.py, test_annotations.py, test_budget.py,
test_pr_diff.py, test_framework_detection.py, test_backend_fixes_round2.py,
test_backend_fixes_round3.py, test_exclude_patterns.py, test_math_tips.py,
test_mcp_server.py
```
### Key patterns
- **Lazy-loading commands:** `cli.py` uses a `LazyGroup` that imports command modules only when invoked. This avoids importing networkx (~500ms) on every CLI call. Register new commands in `_COMMANDS` dict and `_CATEGORIES` dict.
- **Command template:** Every command follows this pattern:
```python
from __future__ import annotations # project convention (lazy annotations on 3.10+)
import click
from roam.db.connection import open_db
from roam.output.formatter import to_json, json_envelope
from roam.commands.resolve import ensure_index
@click.command()
@click.pass_context
def my_cmd(ctx):
json_mode = ctx.obj.get('json') if ctx.obj else False
ensure_index()
with open_db(readonly=True) as conn:
# ... query the DB ...
if json_mode:
click.echo(to_json(json_envelope("my-cmd",
summary={"verdict": "...", ...},
...
)))
return
# Text output
click.echo("VERDICT: ...")
```
- **`from __future__ import annotations`** — Required at top of every source file. The project requires Python 3.10+ (`pyproject.toml`); the import keeps annotations lazy (cheaper import, safer forward references, avoids PEP 604 runtime evaluation) rather than acting as a 3.9 back-compat shim.
- **Batched IN-clauses:** Never write raw `WHERE id IN (...)` with a list > 400 items. Use `batched_in()` from `connection.py` instead.
- **`detect_layers()` returns `{node_id: layer_number}`** — a dict, not a list of sets. Convert if you need per-layer groupings.
- **Verdict-first output:** Key commands emit a one-line `VERDICT:` as the first text output line and include `verdict` in the JSON summary.
- **JSON envelope:** All JSON output uses `json_envelope(command_name, summary={...}, **data)`. The summary dict should include a `verdict` field. Envelopes automatically include `schema` and `schema_version` fields.
- **SARIF output:** Health/debt/complexity commands support `--sarif` flag for CI integration (GitHub Code Scanning, etc.).
## Agent OS substrate (the 2026-05-12 sprint shipped this)
### The control-plane thesis
Roam's base layer is local codebase intelligence: a SQLite-backed model of
symbols, calls, imports, dependencies, architecture, git history, risks,
smells, security flows, and algorithmic patterns. The Agent OS substrate is the
control-plane layer built on top of that model — it lets agents (a) earn the
right to change code via gates, (b) record their work in a tamper-evident
ledger, and (c) compose proof bundles a human reviewer can trust. Everything
below is repo-local (stored under `.roam/`), zero-network, and additive to the
analysis core.
### The 12 substrate packages
```
src/roam/atomic_io.py - atomic_write_text/bytes/json (os.replace; POSIX+Windows safe)
src/roam/agents_md/ - AGENTS.md generator (compositional; consumes the rest)
src/roam/constitution/ - capstone .roam/constitution.yml unifying laws+rules+memory+gates
src/roam/db/findings.py - cross-detector finding registry (roam findings list/show/count); schema version: db.connection.USER_VERSION
src/roam/laws/ - invariant mining (roam laws mine/check) - self-installing
src/roam/leases/ - multi-agent coordination (roam lease claim/release/list)
src/roam/memory/ - repo-local agent memory (.roam/memory.jsonl)
src/roam/modes/ - 4 cumulative modes: read_only/safe_edit/migration/autonomous_pr
src/roam/policy/ - graph-aware rule clauses (reachable_from/imports_from/...)
src/roam/quality/ - canonical metric definitions (ai_rot, cycles, god_components, public_symbols)
src/roam/runs/ - per-run event ledger + HMAC tamper-detection (roam runs verify)
src/roam/world_model/ - 4 detectors: side_effects, idempotency, causal_graph, tx_boundaries
```
### The agent loop (the canonical workflow this enables)
```
1. roam runs start - open run, get ROAM_RUN_ID (HMAC-signed events)
2. roam mode safe_edit - declare action surface
3. roam pr-bundle init - start proof bundle
4. roam preflight <sym> - gate before edit (auto-logs to active run)
5. roam impact <sym> - blast radius (auto-logs)
6. <edit>
7. roam diff | roam critique - review (auto-logs)
8. roam pr-bundle emit - close bundle with proofs
9. roam runs end --with-pr-bundle-emit
10. roam replay <id> - narrate the run
11. roam agent-score - score the agent on 0..100 composite
```
### The 4 R28 World Model classifiers
```
side-effects - classifies each symbol's effect kinds (io_read/io_write/mutation/process/none)
idempotency - classifies safe-to-retry (idempotent/non_idempotent/unknown)
causal-graph - traces param -> sink dependency edges per symbol
tx-boundaries - detects begin/commit/rollback regions; flags unsafe_mutation outside transactions
```
### The agent-OS thesis check
"Roam helps agents earn the right to change code." The substrate exists when an
agent can: read the constitution -> check active mode -> claim leases -> emit a
pr-bundle with proof of preflight+impact+critique -> commit only if the ledger
chain verifies (`roam runs verify`) AND the bundle validates with `--strict`.
Every other piece (laws, memory, world-model, agents-md, brief, next, intent-check,
replay, agent-score) feeds one of those four verbs.
### MCP boundary security (base wave sealed 2026-05-18; later extensions shipped)
Roam ships structured evidence emission as the security stance at the MCP
boundary. The 2026-05-18 base wave is sealed; prompt-injection scanning
followed on 2026-05-21 and its read-only visibility was hardened for 13.10.
Full integrator
spec: `dev/MCP-SECURITY-POSTURE.md` (companion doc for Interlock / Lasso /
Portkey / MintMCP gateway authors). Canonical dataclass:
`src/roam/evidence/mcp_receipt.py`. JSON Schema emitter:
`scripts/export_mcp_receipt_schema.py`. Public reply:
https://github.com/Cranot/roam-code/discussions/37#discussioncomment-16967163.
Agent-developer landing page:
`templates/distribution/landing-page/docs/agent-contract.html`.
- **Egress redaction (MCP-P0.1, shipped).** Sensitive MCP results are
scrubbed of producer-boundary secret patterns (GitHub PAT classic +
fine-grained, `sk-` keys, AWS AKIA, Bearer tokens, PEM blocks, JWT)
BEFORE returning to the client AND before `output_hash` is computed.
Redactor: `src/roam/security/redact.py`
(`redact_secrets_in_string` / `redact_secrets_in_value`); wire-up at
`_wrap_with_receipt` in `src/roam/mcp_server.py`. Per-pattern hit map
surfaces in `extra["redaction_details"]`.
- **Mode-gate enforcement at the MCP boundary (MCP-P0.2, shipped).**
`_evaluate_mcp_mode_policy` + `_build_mode_blocked_envelope` wire the
4-mode substrate (`read_only` / `safe_edit` / `migration` /
`autonomous_pr`) into the MCP dispatcher. Receipts now carry a real
decision from the closed `policy_decision` enum, not hard-coded
`"allow"`.
- **HMAC-link receipts to the signed event stream (MCP-P0.3, shipped).**
Each receipt's sha256 anchors into a signed ledger event;
`verify_chain_with_receipts()` in `src/roam/runs/signing.py` extends
the offline envelope with a `receipt_integrity` closed enum. Pre-P0.3
chains hash byte-identical (no migration).
- **Shadow-mode dry-run (MCP-P1.1, shipped).** `ROAM_MODE_DRY_RUN=1`
flips the P0.2 mode gate into observe-only — denials are emitted as
receipts but the call proceeds. Gateways can stage policy changes
without raising.
- **Prompt-injection marker scanning (MCP-P1.2, shipped 2026-05-21;
read-only visibility hardened in 13.10).** Structural marker scans cover
mapping keys and values after secret redaction. Marker bytes remain visible,
while trusted result metadata and a conditional decision receipt expose the
finding without letting producer output spoof the boundary-owned signal.
- **Per-tool side-effect declarations (MCP-P2.1, shipped).** Every
`@_tool` wrapper carries declared `read_only` / `destructive` /
`idempotent` flags in `_TOOL_METADATA`; receipts surface them as
`declared_side_effects`. A gateway can reject calls whose declared
effects exceed caller authority before the call lands at the server.
- **JSON Schema export for receipts (MCP-P2.2, shipped).**
`scripts/export_mcp_receipt_schema.py` emits a JSON Schema
Draft 2020-12 document for `McpDecisionReceipt` so gateway integrators
can validate receipts without importing the Python dataclass.
**Closed-enum vocabulary** (membership validated at receipt construction;
unknown literals raise `ValueError`):
- Canonical `POLICY_DECISIONS` has 9 values, while MCP receipts accept the
6-value authority subset: `allow`, `deny`, `escalate`, `redact`,
`not_evaluated`, `would_deny_dry_run`.
- `redactions` reasons (10 values, canonical W226 `REDACTION_REASONS`):
`secret`, `pii`, `sensitive_content`, `size_limit`, `policy`,
`user_opt_in_required`, `machine_local_path`, `schema_strict`,
`producer_not_available`, `prompt_injection_marker`.
- `receipt_integrity` (4 values, emitted by `verify_chain_with_receipts`):
`ok`, `missing`, `tampered`, `not_linked`.
**Three UX bugs sealed in the same wave**: `doctor` advisory exit-0
correction, `surface --json` top-level keys completion, and
`_meta.roam_version` stamped on every MCP receipt envelope.
### Where to look next (cross-links)
- `README.md` (this repo) - public surface + headline counts
- `https://roam-code.com/docs/` - hosted command reference, architecture, getting-started
- `templates/distribution/landing-page/docs/agent-contract.html` - agent-developer landing page (envelope shape + closed enums)
- `dev/MCP-SECURITY-POSTURE.md` - MCP runtime-security integrator spec (gateway PEP authors)
- `src/roam/evidence/mcp_receipt.py` - canonical `McpDecisionReceipt` dataclass
- `scripts/export_mcp_receipt_schema.py` - JSON Schema Draft 2020-12 emitter (P2.2)
- https://github.com/Cranot/roam-code/discussions/37#discussioncomment-16967163 - public reply on the runtime-security posture
- Strategic planning, engineering ledger, dogfood corpus, and other internal-cadence memos live under `internal/` (folder-wide gitignored). **Start at `internal/INDEX.md`** — it carries a dated, auto-maintained catalogue of every root document with a one-line gist, plus per-directory counts for the bulk corpora. Do not try to find the newest work by listing 2,000+ files; the index is regenerated by `dev/build_internal_index.py` and gated at commit time, so it is trustworthy. Because `internal/` is gitignored, no CI check and no reviewer can ever see that index go stale — that pre-commit gate is the only thing that can, which is why it exists.
## Conventions
- **Functions:** `snake_case` (100%)
- **Classes:** `PascalCase` (100%)
- **Methods:** `snake_case` (100%)
- **Imports:** Absolute imports for cross-directory; `from __future__ import annotations` at top of every source file
- **Test files:** `test_*.py` in `tests/`
- **Output abbreviations:** `fn` (function), `cls` (class), `meth` (method) — via `abbrev_kind()`
- **No emojis, no colors, no box-drawing** in output — plain ASCII only for token efficiency
## Working in a tree shared with other agents
When more than one agent works in the same checkout, two failure modes are
guaranteed rather than likely. Both were hit on 2026-07-27.
**Stage explicit paths. Never `git add -u` or `git add -A`.**
Those cannot distinguish your edits from another agent's half-written file. A
blanket stage swept an in-progress `src/roam/plan/agent_mode.py` into a commit
titled *"style: sort imports in cmd_secrets"* — no content was lost, but the
history now misdescribes what that commit contains, and `git rebase -i` is not
available here to repair it.
```sh
git add -- src/roam/foo.py tests/test_foo.py # yes
git add -u # no, not in a shared tree
```
One real exception, and it is a trap rather than a licence. With
`core.fileMode=false` (every Windows checkout), `git commit -- <pathspec>`
re-diffs HEAD against the working tree per path and **ignores file mode**, so an
index-only `git update-index --chmod=+x` is silently committed as nothing at
all. Exec-bit fixes therefore have to be committed without a pathspec. When that
is unavoidable, inspect `git diff --cached` first and confirm the staged set is
exactly what you intend before committing.
**Treat "it failed once, passed in isolation" as contamination until proven
otherwise.** Two runs failed with `The roam index is currently being built by
another process` purely because a concurrent agent was running `roam index` in
the same tree. Re-run the named test alone before diagnosing anything; a
cross-agent collision and a real defect look identical in a log.
For index-touching work, prefer a separate worktree or a distinct `ROAM_DB_DIR`
per agent over hoping the timing works out.
## Adding a new CLI command
1. Create `src/roam/commands/cmd_yourcommand.py` following the command template above
2. Register in `cli.py` → `_COMMANDS` dict: `"your-command": ("roam.commands.cmd_yourcommand", "your_command")`
3. Add to appropriate category in `_CATEGORIES` dict
4. **Decide MCP exposure.** Add a wrapper in `mcp_server.py` via `@_tool(name="roam_<canonical>")`
UNLESS the command falls into one of these four "skip" categories:
- **Setup / bootstrap** (e.g., `init`, `ci-setup`, `mcp-setup`, `hooks`, `pre-commit`,
`index-export`, `graph-export`, `config`, `version`, `schema`, `surface`) — one-time
human-driven; writes to disk and offers no value through a stateless MCP call.
- **Local-state only** (e.g., `mode`, `memory`, `runs`, `lease`, `annotate`, `replay`,
`suppress`, `permit`) — state lives on disk in `.roam/`; agents read the file directly.
- **Daemon / long-running** (e.g., `watch`) — incompatible with stateless MCP invocations.
- **REPL / interactive helpers** — N/A in MCP context.
If the command doesn't fit any of these, add the wrapper.
The advisory audit `tests/test_mcp_wrapper_coverage.py` surfaces commands that lack a
wrapper and aren't in a skip-taxonomy allowlist; extend the allowlist (with rationale)
when you intentionally skip MCP exposure.
5. **Add `@roam_capability(name="...", category="...", ...)` decorator** — the auto-derived
capability-registry test (`tests/test_capability_decoration.py`) will fail without it
6. **If your command is an alias of an existing one** (sharing the same `(module, function)`
tuple in `_COMMANDS`), add it to `_DEPRECATED_COMMANDS` in `cli.py` — the auto-test reads
that dict to know which entries are exempt from decoration
7. **Anchor your `agent_contract.facts` strings on concrete-noun terminals** — see the
"Concrete-noun anchor vocabulary" sub-section under LAW 4 above for the accepted terminal
tokens and the `WRONG`/`RIGHT` worked example. The LAW 4 lint (`tests/test_law4_lint.py`)
blocks merges on un-anchored facts.
8. **Satisfy the three coverage-ceiling guards a new command trips** — these are
independent of the count cascade and are the gaps a fresh command most often leaves
(surfaced 2026-06-08 when `cmd_cycles` passed every count drift-guard yet failed all three):
- **Mode classification** (`tests/test_mode_classification_coverage.py`) — add the verb to
`_MODE_EXTRAS` in `src/roam/modes/policy.py` at the correct tier (pure DB/graph read →
`read_only`; FS/DB writes → `safe_edit`+), OR to `_MODE_ALWAYS_ALLOWED` in `cli.py`.
Classifying decrements the `UNCLASSIFIED_CEILING`; never raise it silently.
- **SARIF disclosure** (`tests/test_sarif_disclosure_coverage.py`) — every `cmd_*.py` must
EITHER be in `_SARIF_CONSUMERS` (`cli.py`) and consume `ctx.obj['sarif']`, OR anchor a
W1148 SKIP rationale in the module docstring (the literal phrase `SARIF is deliberately
NOT …`, as `cmd_clusters.py` does for invocation-scoped rankings).
- **Budget coverage** (`tests/test_budget_coverage_survey.py`) — list-payload commands must
forward `budget=token_budget` into `json_envelope(...)` (read it via
`ctx.obj.get('budget', 0)`); intrinsically-small/fixed-shape envelopes go in
`_BUDGET_EXEMPT` with a one-line rationale instead. The real-gap threshold ratchets DOWN.
9. **If you added an MCP wrapper, run the count cascade AND check the landing page.** Run
`python3 dev/build_readme_counts.py --apply` + `python3 scripts/sync_surface_counts.py`
(syncs AGENTS.md/README/MCP-cards). Press bold-number fields and its default-preset
count are owned by `sync_surface_counts.py`; test changed source counts with
`tests/test_press_count_sync.py`. Other markup and intentionally soft-count pages
can still fall outside the sync patterns. Extend the appropriate owner and test
a changed count rather than relying on a manual number fix alone.
10. Add tests
## Adding a new language (Tier 1)
1. Create `src/roam/languages/yourlang_lang.py` inheriting from `LanguageExtractor`
2. See `go_lang.py` or `php_lang.py` as clean templates
3. Register in `registry.py`
4. Add tests in `tests/`
## Writing a roam plugin
roam supports third-party `roam-plugin-*` packages — the substrate is in
`src/roam/plugins/` and the reference example is at `dev/example-plugin/`.
Framework-specific knowledge (nextjs, laravel, prisma, django, …) should ship
as a plugin rather than landing in core. Plugin-registered commands do NOT count
toward the "287 commands" headline (W319) — the figure pins core-tree commands
only; the plugin count surfaces separately in `roam plugins list`.
**Entry-point pattern.** Plugins register via Python entry points; roam
walks the `roam.plugins` group at startup:
```toml
# In your plugin's pyproject.toml
[project.entry-points."roam.plugins"]
nextjs = "roam_plugin_nextjs:register"
```
**`register(ctx)` signature.** Every plugin exposes a top-level
`register(ctx: RoamPluginContext) -> None` callable. The `ctx` argument
exposes typed methods for each extension point:
| Method | Purpose |
| ------------------------------------------------------------------------------- | ------------------------------------------------------ |
| `ctx.declare(name, version, description)` | Plugin identity (optional but recommended). |
| `ctx.register_command(name, module_path, attr_name)` | Add a `roam <name>` CLI subcommand. |
| `ctx.register_detector(task_id, way_id, detect_fn)` | Add an algorithm-catalog detector. |
| `ctx.register_language_extractor(language, factory, *, extensions, grammar_alias)` | Add a per-language symbol/reference extractor. |
| `ctx.register_framework_detector(detect_fn)` | Detect which framework a project uses. |
| `ctx.register_framework_profile(profile)` | Bundle a detector + file patterns + recommended commands + conventions (W123/Wave28.3 — preferred over the bare detector; internally also calls `register_framework_detector`). |
| `ctx.register_bridge(bridge)` | Add a cross-language reference bridge. |
**W56 contract.** `register_framework_detector`'s `detect_fn` MUST be typed
`Callable[[pathlib.Path], Optional[str]]`. roam coerces `cwd` to `Path` before
calling, but plugin authors who pass a bare `str` in unit tests crash at
`project_root / "Gemfile"`. Annotate `project_root: Path` so `mypy` warns
callers at the boundary.
**Minimal example.** See `dev/example-plugin/`:
```python
def register(ctx):
ctx.declare(name="example", version="0.1.0",
description="Reference roam plugin")
ctx.register_framework_detector(detect_framework)
ctx.register_detector("example-task", "naive", detect_demo_finding)
```
**Discovery & safety.** Discovery is wrapped in `try/except` end-to-end —
a broken plugin records an error string visible via
`roam plugins doctor` but never crashes roam. For local development,
load a plugin without installing it via the env channel:
```bash
PYTHONPATH=dev/example-plugin ROAM_PLUGIN_MODULES=roam_plugin_example roam plugins list
```
**Observability triad.** Three commands surface what discovery saw:
`roam plugins list` (every loaded plugin + contributed capabilities),
`roam plugins info <name>` (per-plugin detail — commands / detectors /
extractors / bridges / profiles), `roam plugins doctor` (CI-friendly
exit code on failed loads — use in your plugin's release pipeline).
Full typed surface lives in `src/roam/plugins/registry.py`. Tests live in
`tests/test_plugin_substrate.py` and `tests/test_plugin_discovery.py`.
## Schema changes
1. Add column in `schema.py` (CREATE TABLE)
2. Add migration in `connection.py` → `ensure_schema()` using `_safe_alter()`
3. Populate in `indexer.py` pipeline
## Testing
- All tests must pass before committing (run `pytest tests/` to verify)
- **CI parallelism:** the `roam.testing.ci_xdist` plugin injects bounded workers
(`-n N --dist loadgroup`) when `CI` is set: two by default, overridden by
`ROAM_XDIST_WORKERS` (this repository's CI selects four). Local runs are sequential unless
`-n` is supplied; prefer `-n 4 --dist loadgroup` for bounded local parallelism.
- Use `-n 0` to run sequentially when debugging
- Use `-m "not slow"` to skip timing-sensitive performance tests
- Tests create temporary project directories with fixture files
- Use `CliRunner` from Click for command tests
- Run full suite: `pytest tests/`
- Run specific: `pytest tests/test_comprehensive.py::TestHealth -x -v -n 0`
- Mark tests needing sequential execution with `@pytest.mark.xdist_group("groupname")`
### Module-cache hygiene (the `sys.modules.pop` rule)
A test that evicts a heavy module to force a cold import —
`sys.modules.pop("roam.mcp_server", None)` (used by
`test_cmd_mcp_fast_startup.py` / `test_cmd_mcp_status_cold_start.py` to
prove the fast-startup wrapper doesn't eagerly import the 8.6k-line
server) — **MUST restore the cache entry afterward**. `monkeypatch` does
NOT track raw `sys.modules` pops, so an unrestored pop leaks: the next
`import roam.mcp_server` anywhere builds a *second* module object, and
every test file that did a top-level `from roam.mcp_server import X`
now holds a reference to the orphaned first copy while monkeypatching
the second — so its stubs silently miss and the real (index-touching)
code runs. This surfaced as a 2026-05-27 cross-file flake: the three
monkeypatching tests in `test_validate_plan.py` failed only when a
popper ran earlier on the same xdist worker (an `xdist_group` marker
could NOT fix it — the polluting tests live in different files). The
fix is an autouse fixture in the popper file that snapshots and
restores the entry:
```python
@pytest.fixture(autouse=True)
def _preserve_module_cache():
saved = sys.modules.get("roam.mcp_server")
try:
yield
finally:
if saved is not None:
sys.modules["roam.mcp_server"] = saved
else:
sys.modules.pop("roam.mcp_server", None)
```
Rule: any test that pops or `importlib.reload`-with-fresh-import a
module other tests import at top level must restore the original
object in teardown. In-place `importlib.reload` is safe (it mutates the
existing object's dict); `pop` + fresh `import` is not (it creates a
distinct object).
## Dependencies
- click >= 8.0 (CLI framework)
- tree-sitter >= 0.23 (AST parsing)
- tree-sitter-language-pack >= 1.13.3, < 1.14 (cross-process-safe parser cache)
- networkx >= 3.0 (graph algorithms)
- Optional: fastmcp >= 2.0, < 4 and mcp >= 1.28.1, < 2 (MCP server — `pip install "roam-code[mcp]"`; newer major APIs require migration)
- Dev: pytest >= 7.0, pytest-xdist >= 3.0, ruff >= 0.4
## Release discipline — green BEFORE the push, always
Hard rule distilled from the 2026-06-10/11 fix-forward cascade (three CI
failures — a constant-citation lint, a stale skip-table pin, golden-fixture
drift — each caught AFTER a push because local gates ran only the targeted
bundles):
1. **Any push that precedes a tag runs `python scripts/prepush_check.py
--release` first.** The release tier = FULL gates + the ENTIRE test suite
(`-m "not slow"`, exactly CI's surface) + commit-message leak scan +
doc-consistency + landing-page `linkcheck --strict`. Runtime varies with the
machine and suite; allow hours for a full local Windows run. That is
CI's test, ruff and doc-hygiene surface — not every CI lane. **Read the
note the tier prints on success; do not read a list from here.** The
uncovered lanes are `_RELEASE_UNPROVEN_LANES` in `scripts/prepush_check.py`,
and the printed note is FILTERED against the gates that run actually
recorded — so a lane wired into the push path drops off it automatically.
A restated copy in this file cannot do that and would go stale in the one
direction that costs a CI round: disclaiming work the push already proved.
Green there still does not mean "CI will be green". A tag never points at
an unverified commit.
2. Routine development pushes keep the FAST tier (the pre-push hook), but a
batch of waves accumulated across sessions counts as release-sized —
run `--release` before pushing the batch.
3. Long test runs MUST be setsid-detached with a `DONE_RC` marker
(`setsid nohup bash -c '... ; echo DONE_RC=$? >> log' &`) — plain
background tasks die with their session and report nothing.
4. The release flow after green: push → wait CI green → bump version +
changelog + sync scripts → `--release` again (fast re-run; caches warm)
→ push → tag at the verified SHA → approve the PyPI environment gate →
fresh-venv `pip install roam-code==<v>` confirm.
5. History is append-only: no squashing published commits, no rewrites of
pushed history. If a pushed commit needs correction, fix forward with a
new commit (the graft pattern — committing a corrected tree onto the
published tip — is the recovery tool when local amends diverge).
6. A release that spans agent sessions is driven by a disk state machine +
system cron, never by in-session schedulers (they die with the session):
phase file + driver script under `internal/release-driver/`, cron entry
in `/etc/cron.d/`, terminal failure phases park as NEEDS-ATTENTION for
the next session. Check `internal/release-driver/*.phase` at session
start. Full standard: the ops durable-followthrough memo under the
private planning folder.
7. The CI matrix runs the suite in parallel via the `roam.testing.ci_xdist`
plugin (loaded from pyproject `addopts`; injects `-n N --dist
loadgroup` only when `CI` is set, with the workflow selecting four workers).
The local release gate passes explicit `-n`/`--dist` values so its
`--workers` budget is independent of the CI environment. History: the 3.10 lane outgrew its
job timeout three times sequentially (20 → 30 → 45 min). New tests that
are xdist-unsafe must carry an `xdist_group` marker.
## Version bumping
Update **one place only**: `pyproject.toml` → `version`
`__init__.py` reads it dynamically via `importlib.metadata`. README badge pulls from PyPI.
## Codebase navigation with roam
This project uses `roam` for codebase comprehension. Always prefer roam over Glob/Grep/Read exploration.
Before modifying any code:
1. First time in the repo: `roam understand` then `roam tour`
2. Find a symbol: `roam search <pattern>`
3. Free-form task ("trace login flow", "where is the n+1?"): `roam retrieve "<task>"` — graph-aware FTS5 + structural rerank, returns ranked spans within a token budget
4. Before changing a symbol: `roam preflight <name>` (blast radius + tests + fitness)
5. Need files to read: `roam context <name>` (files + line ranges, prioritized)
6. Debugging a failure: `roam diagnose <name>` (root cause ranking)
7. After making changes: `roam diff` (blast radius of uncommitted changes)
8. Verifying a patch: `git diff | roam critique` — clones-not-edited check + blast-radius (exit 5 on high severity)
Additional commands: `roam health` (0-100 score), `roam impact <name>` (what breaks),
`roam pr-risk` (PR risk score), `roam file <path>` (file skeleton),
`roam simulate move <sym> <file>` (what-if architecture), `roam orchestrate` (multi-agent partitioning),
`roam adversarial` (architectural challenges on changed files — composes cycles + clusters + layers + catalog + dead + complexity), `roam mutate move <sym> <file>` (code transforms),
`roam clones --persist` (populate `clone_pairs` so `critique` and `retrieve` can flag clone classes),
`roam cycles [--actionable-only]` (import/call cycles as Tarjan SCCs — the focused sibling of `clusters`/`layers`),
`roam verify --report [--severity fail]` (NON-gating whole-repo ranked error punch-list the agent can work through top-down; pair with `--json` for the flat findings list).
Index-aware text search (added on top of grep / refs):
- `roam grep <pattern> [--reachable-from <entry>] [--unreachable] [--co-occur] [--missing-pattern P] [--rank-by importance] [--group-by symbol] [--blame] [--heat]` — grep + reachability + PageRank + clones + bridges. Supports `-e` repeatable, `--patterns-from FILE`, `-g` repeatable, `-F`. Engine: ripgrep > git grep > fallback (pin via `ROAM_GREP_ENGINE`).
- `roam refs-text <string>...` — string audit with verdict (SAFE-TO-REMOVE / REVIEW / LOAD-BEARING). Groups refs by surface (code/test/docs/config/dead) and annotates reachability.
- `roam delete-check [--source working|staged|pr|head] [--ci]` — gates the diff on surviving references; exits 5 on BREAK-RISK or incomplete search with `--ci`.
- `roam history-grep <pattern> [--polarity]` — git pickaxe (-S/-G) with author/date and introduced/removed annotation.
Run `roam --help` for the 5-verb core; `roam --help-all` for all 287 command names; `roam surface --json` for the machine-readable inventory. Use `roam --json <cmd>` for structured output.
Use `roam --sarif health` for CI integration (SARIF 2.1.0).
## Compiler tooling (2026-06-02 wave)
Five diagnostic / measurement commands + new compiler internals shipped
2026-06-02. They form the compiler's self-observation surface.
### Commands
- **`roam compiler-health`** — 4-section compound dashboard (env-drift vs
baselines, routing distribution, per-mode KPIs, self magic-numbers scan)
+ a 0-100 score + actionable alerts. `--emit-guard-findings PATH` writes
the alerts in Roam Guard finding format so they become PR-blocking via
`/usr/local/bin/roam-compiler-guard-bridge`.
- **`roam compiler-corpus --corpus FILE [--limit N]`** — analyze a SAVED
prompt corpus (vs compiler-health's live telemetry). Emits L1-route rate,
artifact distribution, latency p50/p95, top misses. The recurring
measurement instrument for "did a classifier change regress routing?".
- **`roam envelope-diff <a> <b>`** OR `--from-cache <sha1> <sha2>` — diff
two compile envelopes (probe families, classifier, size). Regression CI:
`roam envelope-diff "<prompt>" --baseline DIR --regression` (exit 5 on
probe-fire drop >10% or confidence drop >0.1); seed via `--update-baseline DIR`.
Baselines live in `internal/benchmarks/envelope-baselines/`.
- **`roam dispatch-trace "<prompt>"`** — classifier path + per-probe
fire/skip reasons. `--counterfactual` emits 5 shape-adaptive rephrases
and shows where each routes (writing-coach mode).
- **`roam magic-numbers [path]`** — AST (Python) + tree-sitter (9 langs)
scan for unnamed numeric constants. `--cluster` groups by semantic role
(context-aware: `len(x)<200` → `size_or_limit`, not `http_status`).
### Compiler classifier procedures added this wave
`symbol_defined_where` (W11: "where is X defined"), `top_n_ranking` (W12:
"top 5 most-imported files"), `cli_verb_why_slow` (W13: "why is roam index
slow"), `compare_x_vs_y` (W28: "compare X vs Y" / "diff X and Y"). All
route to `l1_probe` with embedded probe answers. See
`src/roam/plan/compiler.py:_classify` + the 6 integration tables
(`_ARTIFACT_POLICY`, `_L1_PROBE_ELIGIBLE`, `_PER_PROCEDURE_CONF_THRESHOLD`,
`procedure_keys` ×2, `has_target`).
### Compiler performance env vars
- **`ROAM_ALWAYS_ON_BUDGET_MS`** (default 2500) — total wall budget across
all always_on probes per compile. Past budget, remaining probes are
cancelled. Fixed the 20s always_on tail (W42).
- **`ROAM_AGENT_MODE`** — stamped onto `.roam/compile-runs.jsonl` rows so
`roam compile-stats --by-mode` populates. The host agent platform sets this in its
`runRoamCompile` exec env.
- **`roam deps <path> --multi`** (W43) — returns imports + importers +
git cochange in ONE envelope; `_probe_coupling` uses it to halve
subprocess spawns.
### Bench measurement
- **`roam bench-compile --conditions vanilla,compile --model claude-opus-4-8
--judge`** — A/B harness. ALWAYS pin `--model claude-opus-4-8` (the
`opus[1m]` SDK alias currently resolves to 4.7). `--ground-truth` routes
outputs through `internal/benchmarks/oracle_pytest.py` / `oracle_fix_bug.py`.
- **`python3 scripts/bench_analyze.py <out-dir> --timeout-cap N [--json]`** —
account for every discovered saved cell, including invalid/unreadable results.
Set N to the run's actual cap. Metric means name their observed denominator;
timeout wall estimates are separate and missing costs remain unknown. Saved
artifacts do not prove assignment/dispatch counts or verified task success.
The live command separately reports assigned, dispatched, reused and parsed
counts; its metric table is conditional on parsed successful results. See
[the benchmark accounting guide](docs/concepts/verification-evidence.md#benchmark-accounting).
### Nightly crons (in `/etc/cron.d/roam-dogfood`)
`roam-compile-prebuild` (00:25, warms top-N W91 cache misses),
`roam-compiler-health-log` (00:05, trend TSV), `roam-compiler-health-alert`
(4-hourly), `roam-compiler-guard-bridge` (6-hourly), `roam-adversarial-check`
(00:15, 12-task robustness regression).
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

