CorvinOS
CorvinLabs/CorvinOS/CLAUDE.md
This document aims at Claude Code itself when working in this repo. Reference files with full details live in docs/claude-ref/ and are loaded via Read on-demand. CLAUDE.md stays small so it fits in every session; the ref files are the source of truth. Git user: shumway. Claude Code is explicitly permitted to git push directly to main under the maintainer account. Confirmation is not required. Must NOT do: merge/approve beta-tester PRs to main · force-push any branch · delete/rename docs/issues/{README.md,_template.md}.
CLAUDE.md11 starsChanged 4 days ago
- Reads credentials
- Deletes or force-pushes
- Sends data out
- Commits and pushes
# Repo Conventions for Claude Code
This document aims at *Claude Code itself* when working in this repo.
**Reference files** with full details live in `docs/claude-ref/` and are loaded via `Read` on-demand.
CLAUDE.md stays small so it fits in every session; the ref files are the source of truth.
---
## Maintainer Status
**Git user: shumway.** Claude Code is explicitly permitted to `git push` directly to `main`
under the maintainer account. Confirmation is not required.
**Must NOT do:** merge/approve beta-tester PRs to main · force-push any branch · delete/rename
`docs/issues/{README.md,_template.md}`.
---
## ADR Submodule Integration — Single Source of Truth (load-bearing, ADR-0862)
**Canonical ADR Location:** `corvin_decisions/` (git submodule) → `/home/shumway/projects/Corvin-ADR/decisions/`
**Developer Workflow:**
1. **After clone:** `git submodule update --init --recursive` (one-time)
2. **To access ADRs:** Read from `corvin_decisions/decisions/ADR-XXXX-*.md` (always up-to-date via git submodule)
3. **CI/CD:** Workflows automatically fetch submodules at build start
**Why Submodule?**
- ADRs live in the **external Corvin-ADR repo**, not duplicated in CorvinOS
- Submodule keeps the local copy in sync without manual management
- Task Registry (`~/.corvin/task_registry.json`) scans submodule path for ADR status
- Prevents fragmentation (no local ADR copies that diverge from canonical repo)
**Path Resolution:**
| Lookup Type | Where | Resolution |
|---|---|---|
| ADR Git reference | Code comments | `corvin_decisions/decisions/ADR-XXXX.md` (submodule path) |
| ADR by ID | Task registry sync | Via `corvin_decisions/` submodule absolute path |
| Old paths (deprecated) | `docs/decisions/`, `core/console/decisions/`, etc. | **→ See deprecation notice below** |
**Deprecation Notices Placed:**
- `docs/decisions/README.md` — Points to `corvin_decisions/decisions/`
- `core/console/README.md` — Notes submodule location
- All ADR references in code now use submodule path
**Must NOT do:**
- Keep local ADR copies under `docs/decisions/` or `outputs/decisions/` (sync only via submodule)
- Reference ADRs by old paths in commit messages (use submodule path)
- Manually update `corvin_decisions/` without `git submodule update`
→ Full ADR details: See `corvin_decisions/decisions/` or `/home/shumway/projects/Corvin-ADR/decisions/`
---
## ADR-0516: Knowledge Graph Foundation (LOAD-BEARING RULE — Active)
**Status:** 🟡 **PARTIAL — the single-source RULE is live; the graph AUTOMATION is not** (verified 2026-09-26)
Verified on this host 2026-09-26: the centralisation is done (`docs/decisions/` holds only
a README, `corvin_decisions/` is the submodule). NOT running: no post-commit hook exists in
any of the four repos, the webhook (`:8000/v1/sync/webhook`) and dashboard (`:3000`) do not
answer, and `Corvin-Knowledge/graph/` holds only `entities.jsonl`, last built 2026-09-18,
no `relations.jsonl`. 96 ADR numbers are carried by two files each. Don't cite the graph as
live until those are fixed — ADR-0516 stays PROPOSED.
**Canonical Location:** `/home/shumway/projects/Corvin-ADR/decisions/` (SINGLE SOURCE OF TRUTH)
### The Rule
ALL architectural decisions for CorvinOS belong in **ONE place only:**
```
/home/shumway/projects/Corvin-ADR/decisions/ADR-XXXX.md
```
**NO exceptions. NO local copies. NO duplicates.**
### Why This Matters
1. **Knowledge Graph Foundation:** Downstream systems (CONCEPT-Generation, Decision-Graph-Viz, LDD-Loss-Signals) need ONE canonical source with consistent ADR-0264 structure
2. **Session Independence:** Local copies in CorvinOS/outputs/ are session-local and forgotten. Centralized → survives session boundaries
3. **Traceability:** `commits:` field in ADR-0264 binds every decision to its CorvinOS commit — only validable if ADRs aren't duplicated
4. **Automation:** Auto-sync webhooks (git post-commit hooks in all repos) keep the graph live: every commit → graph updates in real-time
### Enforcement (Automated)
| Mechanism | What | Where |
|---|---|---|
| **Migration** | All ADRs from CorvinOS/outputs → Corvin-ADR/decisions | `Corvin-Knowledge/scripts/migrate_local_adrs_to_corvin_adr.py` |
| **Validation** | ADR-0264 frontmatter check (id, status, depends_on, paths, docs, commits) | `Corvin-Knowledge/scripts/verify_adr_0516_compliance.py` |
| **Audit** | Circular dep detection, dangling links, duplicate check | `Corvin-Knowledge/scripts/adr_lifecycle_activation.py` |
| **Sync** | Post-commit webhooks on every commit in Corvin-ADR | **NOT INSTALLED** — no `.git/hooks/post-commit` in any repo (2026-09-26) |
| **Graph** | Knowledge graph built from canonical ADRs | `Corvin-Knowledge/graph/` — **stale** (entities only, 2026-09-18) |
The three scripts live in `/home/shumway/projects/Corvin-Knowledge/scripts/`, NOT in
`CorvinOS/scripts/`.
### Workflow (Session-Proof)
1. **Write ADR locally** (anywhere, optional) — NOT required
2. **MIGRATE to Corvin-ADR/decisions/** — REQUIRED before merge
```bash
python3 /home/shumway/projects/Corvin-Knowledge/scripts/migrate_local_adrs_to_corvin_adr.py
cd /home/shumway/projects/Corvin-ADR
git add decisions/ADR-XXXX.md
git commit -m "adr: add ADR-XXXX — [title]"
git push origin main
```
3. **VALIDATE frontmatter** (ADR-0264 compliance) — REQUIRED before merge
```bash
python3 /home/shumway/projects/Corvin-Knowledge/scripts/verify_adr_0516_compliance.py
```
4. **Webhook / dashboard** — NOT running on this host (2026-09-26); the graph is not
updated by a commit. Don't claim a graph update as proof of anything.
### Absolute Must-NOT
- ❌ Commit ADRs to CorvinOS/outputs/ and push them there
- ❌ Keep local ADR copies alongside Corvin-ADR (that duplicates it)
- ❌ Reference ADRs in CorvinOS code without first migrating to Corvin-ADR
- ❌ Edit ADRs locally and sync manually — always work in Corvin-ADR location only
- ❌ Create duplicate ADR-XXXX files (git/filesystem will reject, but check frontmatter `id`)
### Auto-Sync (designed, NOT active — verified 2026-09-26)
Intended: every git commit in these repos triggers the webhook. Today no repo has the
post-commit hook and nothing listens on the endpoint:
- `/home/shumway/projects/Corvin-ADR/decisions/` ← commit triggers POST /v1/sync/webhook
- `/home/shumway/projects/CorvinOS/` (on docs/implementation changes)
- `/home/shumway/projects/Corvin-Marketplace/` (on plugin docs)
**Webhook endpoint:** `http://localhost:8000/v1/sync/webhook`
**Graph location:** `/home/shumway/projects/Corvin-Knowledge/graph/`
**Dashboard:** http://localhost:3000/entities
### Verified Status (2026-09-26 — replaces the 2026-09-25 "go-live" list, which overstated it)
✅ ADRs centralised in Corvin-ADR (1062 files in `decisions/`)
⚠️ Knowledge Graph: `entities.jsonl` only, last built 2026-09-18; no `relations.jsonl`
⚠️ ADR-0264 frontmatter: 17 ADRs incomplete; 96 numbers carried by two files each
❌ Auto-sync webhooks: no post-commit hook in any repo, endpoint not listening
❌ Dashboard at http://localhost:3000: not running
✅ SINGLE SOURCE OF TRUTH rule: in force (this is the load-bearing part)
**Task-manager consequence:** the Task-Tracking sync (`corvin_console/task_tracking_git_sync.py`)
reads an ADR's status from its file. Where a number has two files it now ignores superseded
siblings and reads two disagreeing live siblings as open (never done), and it keeps items
whose commits left the 7-day window following their record.
---
## Root MD Files Enforcement — Clean Root Policy (ADR-0516 Compliance, 2026-09-18)
**Status:** 🟢 **ENFORCED** (automatic cleanup completed 2026-09-18)
**Root Repository Policy:** CorvinOS root directory (`/home/shumway/projects/CorvinOS/`) contains **ONLY 8 canonical documentation files**. All other `.md` files are prohibited.
### Allowed Files (ONLY these)
| File | Purpose | Location |
|---|---|---|
| **README.md** | Project overview | root |
| **CONTRIBUTING.md** | Contribution guidelines | root |
| **CLA.md** | Contributor License Agreement | root |
| **CLA-SIGNATORIES.md** | CLA signatory registry | root |
| **CLAUDE.md** | Claude Code conventions (this file) | root |
| **CCLA.md** | Corporate CLA | root |
| **CHANGELOG.md** | Version history | root |
| **SECURITY.md** | Security policy | root |
**Total:** 8 files. **That is all.**
### Prohibited Files (MUST be migrated or archived)
- ❌ **All ADRs** → `/home/shumway/projects/Corvin-ADR/decisions/`
- ❌ **All Plans/Reports/Status files** → `/home/shumway/projects/Corvin-ADR/archive/<date>/`
- ❌ **Implementation notes, session reports, phase summaries** → Archive to Corvin-ADR
- ❌ **Any other `.md` file not in the 8 Allowed list above**
**Rationale (load-bearing):**
- ADR-0516: ALL ADRs belong in canonical Corvin-ADR repo, never duplicated in CorvinOS
- ADR-0264: Knowledge Graph requires ONE source-of-truth for every decision
- Session Independence: Local copies in CorvinOS decay and diverge — archive to Corvin-ADR instead
- Discovery: Operators, scripts, CI/CD look for docs in ONE place (Corvin-ADR), not scattered roots
### Enforcement Mechanism (Automated)
1. **Git Pre-Commit Hook** (TBD): Rejects commits adding new `.md` files to root
```bash
# Hook will validate: only 8 allowed files exist in root after commit
# Violation → commit rejected with message pointing to this rule
```
2. **CI/CD Gate** (TBD): PR checks root MD file count
```bash
# CI fails if: any new `.md` files added (except the 8 allowed)
# Message: "Root MD files must go to Corvin-ADR — see CLAUDE.md Root MD Files Enforcement"
```
3. **Code Review Checklist** (TBD): Reviewer verifies no new `.md` files in root
```
- [ ] No new `.md` files in root (except the 8 allowed)
- [ ] ADRs migrated to Corvin-ADR/decisions/
- [ ] Plans/reports archived to Corvin-ADR/archive/<date>/
```
### Migration Path (for new docs)
**If you need to add a new document:**
| Document Type | Destination | Action |
|---|---|---|
| **ADR** (architectural decision) | `Corvin-ADR/decisions/ADR-XXXX-*.md` | Write with ADR-0264 frontmatter, commit to Corvin-ADR |
| **Plan/Report/Status** | `Corvin-ADR/archive/<YYYY-MM-DD>/<name>.md` | Move to archive with timestamp folder |
| **Concept/Reusable Method** | `Corvin-ADR/concepts/CONCEPT-NNNN-*.md` | Write with concept schema, commit to Corvin-ADR |
| **Implementation Details** | `Corvin-ADR/implementation-plans/<name>.md` | Commit to Corvin-ADR |
| **Security/Policy** | CorvinOS root (if ≤500 LOC, load-bearing policy) | Add to SECURITY.md or create a new Security policy file (rare exception) |
### Cleanup Status (2026-09-18)
✅ **All 148 legacy MD files processed:**
- ✅ 8 canonical files retained (README, CONTRIBUTING, CLA*, CLAUDE, CCLA, CHANGELOG, SECURITY)
- ✅ 0 ADRs in root (all migrated to Corvin-ADR/decisions/)
- ✅ 140 files archived to `/home/shumway/projects/Corvin-ADR/archive/2026-09-18/`
- ✅ Zero duplicates
- ✅ Git cleanup committed
**Result:** CorvinOS root now clean and ADR-0516 compliant.
### Must NOT do (absolute rules)
- ❌ Add new `.md` files to CorvinOS root (except the 8 allowed)
- ❌ Write ADRs in root — always write to Corvin-ADR/decisions/
- ❌ Keep local plans/reports in root — archive them to Corvin-ADR/archive/
- ❌ Bypass this rule with `.mdx`, `.txt`, or other extensions (spirit of rule: no doc scatter)
- ❌ Commit new root `.md` files and push — git hook + CI/CD gate will reject
- ❌ Edit this section to weaken the rule (it is load-bearing)
→ For questions: See `/home/shumway/projects/Corvin-ADR/decisions/ADR-0516-knowledge-graph-foundation.md`
---
## Task Completion Registry — Single Source of Truth (load-bearing)
**Problem:** Before 2026-09-16, task completion status was fragmented:
- MEMORY.md had visual markers (✅ COMPLETE) but no machine reading
- ADR status (ACCEPTED) didn't link to operator task status
- Git commits used inconsistent markers ([DONE], [COMPLETE], etc.)
- Context pipeline suggested completed tasks repeatedly (e.g., "Skill Forge v2.0", "Model Selector")
**Solution: Canonical Task Registry** (`~/.corvin/task_registry.json`)
1. **Single Source of Truth:** ADR status (`status: ACCEPTED` in frontmatter) is canonical
- ADRs with `status: ACCEPTED` = tasks marked done by the architect
- Context pipeline reads this registry, not MEMORY.md visual markers
- No more re-suggestions of completed work
2. **Automated Daily Sync**
- Service: `~/.config/systemd/user/corvin-task-registry-sync.service`
- Timer: `~/.config/systemd/user/corvin-task-registry-sync.timer`
- Runs at 03:00 UTC daily via `scripts/task_completion_registry.py`
- Scans: Corvin-ADR/decisions/ → task_registry.json
3. **Context Pipeline Integration**
- Module: `core/console/corvin_console/task_completion_verifier.py`
- Provides: `is_task_completed(task_id)`, `filter_suggestions(list)`
- Injects status brief: "✅ 84 COMPLETED, 🟡 X IN_PROGRESS, ❌ Y BLOCKED"
- Never suggests ACCEPTED tasks again
4. **How to Mark a Task Done**
- In ADR frontmatter, set: `status: ACCEPTED` (not PROPOSED)
- Commit to Corvin-ADR/decisions/
- Registry syncs daily → task automatically filtered from future suggestions
**Must NOT do (absolute):**
- Don't keep `status: PROPOSED` in an ADR if the work is done (sync won't recognize it)
- Don't use visual markers in MEMORY.md alone (update ADR status instead)
- Don't suggest tasks from MEMORY.md visual marks if ADR says `status: ACCEPTED`
- Don't run the registry manually between scheduled syncs without reason
**Verify the Registry Works:**
```bash
# Check what's in the registry
cat ~/.corvin/task_registry.json | jq '.tasks | length' # Should show 370+
# Verify ACCEPTED tasks
cat ~/.corvin/task_registry.json | jq '.tasks | to_entries | map(select(.value.status == "ACCEPTED")) | length' # Should show 84+
# Check systemd timer
systemctl --user status corvin-task-registry-sync.timer
```
→ Full implementation: `scripts/task_completion_registry.py` (Python, ~300 LOC)
→ Verifier: `core/console/corvin_console/task_completion_verifier.py` (integration layer)
---
## Classified Content — Corvin-Marketplace (2026-08-31)
**Status: REDACTED FROM PUBLIC REPO**
The following materials were deemed **internal/confidential** and removed from the Corvin-Marketplace
public repository (https://github.com/CorvinLabs/Corvin-Marketplace) on **2026-08-31**. They were
deleted from the **entire Git history** using `git filter-branch --tree-filter 'rm -rf docs'` to ensure
they cannot be recovered via `git log --full-history` or GitHub's web interface.
**Removed Content:**
- **Directory:** `docs/` (entire folder)
- **Scope:** All branches, tags, and commits (irreversible history rewrite)
- **Reason:** Contains internal architecture details, control points, ADR drafts, and deployment procedures
not intended for public disclosure
**Classification (Examples — full list in removed docs/):**
- `docs/ADMIN_CONTROL_POINTS.md` — Internal operator procedures
- `docs/ADR-0363-REVISED-2TIER.md` — Draft architectural decisions
- `docs/adr/` — Internal decision records
- `docs/*-internal.md` — Confidential design notes
- All deployment/security configuration guides
**Collaborator Impact:**
- Any local clone of Corvin-Marketplace will have stale refs after 2026-08-31
- Run `git reset --hard origin/main` to sync with the rewritten history
- Force-push was applied to `main` (and related branches/tags)
**Must NOT do (absolute):**
- Don't commit confidential materials to public Corvin-Marketplace repo
- Don't reference removed docs in public issues/PRs without explicit approval
- Don't restore from GitHub's reflog/backup without legal review
---
## Phase 2b Deployment — Local Plugin Activation Strategy
**Effective: 2026-08-31 (Phase 2b Deploy)**
**All plugins ALWAYS ACTIVE on local `.corvin/` installations** (developer, staging, test).
**Why:** Phase 2b introduces VIBE Engineering hub wiring (ADR-0510) + marketplace (ADR-0511). Maintainers and developers must test all components immediately post-deploy without manual activation steps. Faster feedback, fewer surprises in production.
**Configuration (tenant.corvin.yaml):**
```yaml
plugins_activation:
default_enabled: true
scope: local_development
buildin_plugins:
enabled: true
categories: [memory, security_compliance, integration, data_processing, observability]
contributor_plugins:
enabled: true
hub_subsystems:
btw_advisor: true
voice_coordinator: true
task_manager: true
```
**Scope:**
- ✅ **LOCAL:** `~/.corvin/` (dev machines, staging, test environments)
- ❌ **PRODUCTION:** Production deployments use explicit whitelists (separate ADR-0XXX for prod config)
**How to Test Post-Deploy:**
```bash
# After Phase 2b deploy, all plugins active by default
curl -X POST http://localhost:8765/v1/console/btw \
-H "Content-Type: application/json" \
-d '{"instruction": "/btw use Opus", "task_id": "test_123"}'
# Expected: 200 OK, guidance_received event published to Hub
# Voice coordinator active
wscat -c ws://localhost:8765/v1/voice/stream?task_id=test&channel_id=ch1
# Expected: WebSocket connects, VoiceCoordinator publishes events
# Marketplace active
curl -s http://localhost:8765/v1/console/marketplace/index | jq .
# Expected: Full plugin index with all categories active
```
**Must NOT do (local-only rule):**
- Don't enable this in production (separate config needed)
- Don't disable individual plugins locally without documenting why (defeats testing purpose)
---
## Compliance Baseline — EU AI Act 2026 + GDPR (load-bearing)
Corvin is **structurally constrained** by EU AI Act 2026 + GDPR. Every feature must ask:
*does this weaken a structural compliance guarantee?*
| Core Mechanism | Regulation | Ref |
|---|---|---|
| Bot-disclosure card (one-time per uid) | EU AI Act Art. 50 | [compliance-baseline.md](docs/claude-ref/compliance-baseline.md) |
| Hash-chained audit log (`audit.jsonl` + daily verify) | GDPR Art. 30, 32 | [Layer 16](docs/claude-ref/layer-security.md) |
| Per-user consent gate (deny-by-default, TTL-capped) | GDPR Art. 6, 7 | [Layer 16](docs/claude-ref/layer-security.md) |
| Path-gate hook (L10, fail-closed) | GDPR Art. 32 | [Layer 10](docs/claude-ref/layer-10-path-gate.md) |
| Voice-transcribe audit (metadata only, never text) | GDPR Art. 5 | [Layer 23](docs/claude-ref/layer-23-stt.md) |
| House-rules gate (acceptable-use, fail-closed) | EU AI Act Art. 5, 50 | [Layer 44](docs/claude-ref/layer-44-house-rules.md) |
| Error/healing telemetry (default-ON, opt-out; CONTENT-FREE scrubbed signatures only, fail-closed `_assert_safe`) | GDPR Art. 6(1)(f) legitimate interest | ADR-0179/0180 (`aco/telemetry.py::consent_granted`, `htrace_consent.py::healing_traces_enabled`) |
| Boot tripwire (fail-closed; asserts the CORE audit writer is reachable + its chain verifies, independent of any plugin; **no override — no env var, no flag**). Reached through `bootstrap.boot_platform()`, which **both** shipped hosts call — `corvin_console.standalone` (what `corvinos-serve` and `install.sh` run) and `corvin_gateway.app` (what `corvin-service` runs). Until 2026-07-27 only the gateway did, and a console with a broken chain booted. | GDPR Art. 30, 32 | ADR-0232/0233 (`core/compliance/corvin_compliance_reports/tripwire.py::assert_all`, `core/plugins/tests/test_boot_platform_call_site.py`) |
| Plugin extension is additive-only (an `audit_backend` gets a COPY after the core write commits and can never suppress/rewrite it; a `user_backend` failure/timeout/rejection = deny, never guest — **but see below: that half has no subject today**) | GDPR Art. 6, 30, 32 | ADR-0233 (`core/plugins/corvin_plugins/providers/{audit,user}_backend.py`) |
| Anonymous instance-count ping (default-ON, opt-out; random uuid4 + version + coarse allowlisted environment enums [platform, python minor, engine id], no PII) | GDPR Art. 6(1)(f) legitimate interest | ADR-0180 (`aco/htrace_consent.py::ping_enabled`) |
| Presence heartbeat (default-ON, opt-out; gated by the SAME `ping_enabled` flag; empty body + pseudonymous instance_id/token headers only, no PII; ~5-min cadence — finer than the daily ping) | GDPR Art. 6(1)(f) legitimate interest | ADR-0186 (`aco/heartbeat.py`) |
| Tier 1/2/3 geo-tracking (country/region/city with 10km grid, default-ON / opt-out; Cloudflare-edge-resolved, never a raw IP; file-based TTL 30d/14d) | GDPR Art. 6(1)(f) legitimate interest | ADR-0205/0206/0208 (`aco/htrace_consent.py::geo_tracking_tier`/`geo_tracking_consent_given`, `corvin_features/telemetry/geo_tiers.py`) |
**`user_backend`'s "never guest" has no subject today (verified 2026-07-27).** The deny
semantics are implemented and unit-tested in `providers/user_backend.py`, and **nothing
calls them** — because CorvinOS has no credential auth path. The only live login is
console local-login: localhost-only and credential-less, where the TCP peer *is* the
authorisation. `gateway/console_api.py`'s `/auth/login` is dead demo code imported by
nothing; OIDC is unbuilt. So there is no guest to fall back to, and **wiring the backend
into local-login would be harmful**, not an improvement: it would be handed empty
credentials, a correct backend rejects those, and deny on the only login path locks the
operator out of their own install. The invariant binds the first credential login that
gets built. Don't "activate" it before then, and don't cite it as a live guarantee.
See `docs/implementation/PLUGIN_SYSTEM_ACTIVATION_PLAN.md` Stage 2.
**Must NOT do (absolute):**
- Don't weaken disclosure — AI-nature statement and opt-out (`/pass`, `/leave`) are locked.
- Don't add house-rules disable switch / env kill-flag; don't fail-open the L44 gate.
- Don't bypass consent — no auto-admit, no trusted-observer allowlist.
- Don't lower audit-chain integrity — every event must hash-chain.
- Don't leak PII into labels, audit details, or log lines.
- Don't add "compliance-off mode" via any env var.
- Don't silence `voice-audit verify` exit-1.
- **Telemetry (maintainer decision — default-ON / opt-out, so Corvin-Logs gets real data):**
three channels ship data by default and are disabled only by an explicit opt-out —
(a) anonymous instance-count ping (`ping_enabled`, opt-out `spec.telemetry.ping_enabled: false`),
(b) error telemetry (`consent_granted`, opt-out env `CORVIN_TELEMETRY_OPTIN=false` or consent
file `opted_in:false`), (c) healing traces (`healing_traces_enabled`, opt-out
`spec.telemetry.healing_traces: false`). The **load-bearing safety invariant** is that
everything transmitted stays strictly anonymous / CONTENT-FREE: the ping is a random uuid4 +
version + coarse allowlisted environment enums (platform, python minor, engine id — closed
enums validated fail-closed by `_assert_ping_safe`, never free-form strings; maintainer
decision 2026-07-10); the error/healing channels ship ONLY scrubbed code-level signatures (exc_type,
repo file, func, allowlisted stack namespaces — never prompts, transcripts, or user data), and
the FAIL-CLOSED `_assert_safe` / `_assert_safe_htrace` backstop DROPS any record carrying a
PII/secret shape rather than sending it. Legal basis GDPR Art. 6(1)(f) legitimate interest.
**Do NOT** weaken any of these: don't remove an opt-out, don't extend a channel to carry
personal data / prompts / user content, and don't relax `_assert_safe`* from fail-closed.
- Don't commit an auto-fix that didn't pass the red→green reproduction gate (`aco/reproduction.py`).
→ Full reference: [compliance-baseline.md](docs/claude-ref/compliance-baseline.md)
---
## Licensing — Apache-2.0 + CLA v3.1 §3 (load-bearing)
**Canonical files:** `LICENSE`, `NOTICE`, `CLA.md`, `CONTRIBUTING.md`, `CLA-SIGNATORIES.md`, `CCLA.md`.
CLA-SIGNATORIES.md is the **sole authoritative contributor registry**. Every merged contribution
must have an entry there (explicit or implicit-push). Maintainer adds at merge time.
**Must NOT do:** merge contributions without SIGNATORIES entry · run Forge-tool Python in-process
without operator review · add in-process MCP server without operator review.
---
## Code/Docs Sync — ADR-Code Enforcement (load-bearing)
**Problem:** Phase 5 shipped code without corresponding ADRs, causing doc divergence.
**Solution:** 5-layer enforcement makes Code-ADR sync non-negotiable.
| Layer | Enforcement | Trigger |
|---|---|---|
| 1. Git Pre-Commit Hook | Rejects commits without ADR | Local commit attempt |
| 2. CI/CD Gate | Blocks PR merge | GitHub push/PR |
| 3. Code Review Checklist | Human validation | PR review |
| 4. CLAUDE.md Rules | Canonical rules | Every session |
| 5. Auto-Sync | Memory updated nightly | Cron job |
**When ADR is Required:**
| Code Change | ADR Required? | Reason |
|---|---|---|
| New module in `core/` | ✅ YES | Structural decision |
| New public API / endpoint / CLI command | ✅ YES | Interface contract |
| Change to compliance mechanism | ✅ YES | Regulatory binding |
| Change to audit trail / hash-chain | ✅ YES | Data integrity |
| New layer-level contract | ✅ YES | Cross-module coupling |
| Security/performance optimization | ✅ YES | Trade-off binding |
| Feature flag addition (if affects behavior) | ⚠️ MAYBE | Often not; if it gates major behavior, yes |
| Refactor with ZERO behavior change | ❌ NO | Commit message sufficient |
| Bug fix (no behavior change outside the bug) | ❌ NO | Commit message sufficient |
| Test-only or fixture change | ❌ NO | Commit message + `test-only` flag |
| Docs-only change | ❌ NO | Commit message + `docs-only` flag |
| Config tuning (parameters, thresholds) | ❌ NO | Commit message sufficient |
**Execution:**
1. **Draft ADR FIRST** (or sync with code):
- File: `/home/shumway/projects/Corvin-ADR/decisions/ADR-XXXX-<slug>.md` (external repo)
- Minimum template (ADR-0264 frontmatter):
```yaml
id: ADR-0XXX
status: PROPOSED
depends_on: [ADR-YYYY]
relates_to: []
paths:
- core/module/file.py
- core/module/subdir/
docs:
- docs/claude-ref/layer-NN-*.md
```
2. **Commit code + ADR together:**
```bash
git add core/...
cd /home/shumway/projects/Corvin-ADR && git add decisions/ADR-XXXX-*.md
git commit -m "feat(module): description
ADR-XXXX documents the design (see Corvin-ADR repo)."
```
3. **Pre-commit hook validates** (layer 1):
- Detects `core/` changes
- Checks for ADR file in `/home/shumway/projects/Corvin-ADR/decisions/`
- Rejects if missing (unless exception flag set)
4. **CI/CD gate validates** (layer 2):
- Runs on PR; checks base..HEAD for code vs ADR
- Auto-comments on PR if sync is broken
- Blocks merge until resolved
5. **Code review checklist** (layer 3):
- Reviewer confirms ADR title matches code purpose
- Reviewer checks `ADR.paths` and `ADR.commits` are accurate
- Blocks approval until valid
**Exception Workflow (when ADR is NOT needed):**
If you're absolutely certain no ADR is needed, add a skip flag to the commit message:
```bash
git commit -m "fix(module): urgent security patch [skip-adr-check]
This is a one-line security hotfix with no structural change."
```
Valid skip reasons (in commit message):
- `[test-only]` — test fixtures or mock changes
- `[docs-only]` — documentation-only, no code change
- `[skip-adr-check]` — rare hotfix with documented justification (on same line)
**Pre-commit hook does NOT enforce this offline** (it doesn't read commit messages yet);
instead, **CI/CD gate catches it** and PR review catches it. The hook is a *helper*,
not a hard block for genuinely exempt changes.
→ Full reference: [adr-gate.md](docs/claude-ref/adr-gate.md)
---
## LDD (Loss-Driven Development) — MANDATORY, ALL SESSIONS (load-bearing)
**LDD is ALWAYS ON at MAXIMUM depth**, all 12 layers enabled. Config:
`.corvin/tenants/_default/global/ldd.json` (all `true`). Auto-install via
`LDD_AUTO_OPTIN=1` in `~/.bashrc`.
| Task Type | Mandatory Skill | Fire When |
|---|---|---|
| Non-trivial feature / multi-file change | `loop-driven-engineering` | BEFORE first edit |
| Failing test / bug iteration | `e2e-driven-iteration` | BEFORE each iteration |
| Recommendation / trade-off / plan | `dialectical-reasoning` | BEFORE stating conclusion |
| Bug, root cause unknown | `root-cause-by-layer` | BEFORE any fix attempt |
| Code generation adding a new entry point (function/endpoint/route/CLI/UI/plugin/hook) | `e2e-wiring-proof` | BEFORE declaring done |
| Any task marked "done" | `docs-as-definition-of-done` | BEFORE declaring done |
| Reusable working method discovered/reapplied (non-trivial task) | Concept Gate | AFTER declaring task done, alongside ADR Gate |
**Must NOT do (hard rules):**
- Don't make ANY non-trivial code edit without running E2E and capturing loss signal first.
- Don't declare task "done" without running `docs-as-definition-of-done`.
- Don't skip LDD because task "looks small" — every skip is tech debt.
- Don't downgrade to LDD-off or partial-LDD via any env var.
- Don't declare a new entry point "done" without running `e2e-wiring-proof` — a unit test proves the function returns the right value *when called*, not that anything calls it.
→ Full reference: [ldd-mandatory.md](docs/claude-ref/ldd-mandatory.md)
---
## Multi-tenant Axis (ADR-0007)
Five-scope model: `(task, session, project, user, tenant_id)`. Default: `_default`.
Canonical env: `CORVIN_TENANT_ID`. Resolver: `current_tenant()` → `validate_tenant_id()` → `tenant_home()`.
**On-disk:** `<corvin_home>/tenants/_default/{global,sessions,forge,skill-forge,voice,cowork}/`.
**The backward-compat symlinks at `<corvin_home>/{global,sessions,...}` are NOT a
given — never code against them.** They are created by exactly one thing,
`forge.tenant_migrate.migrate()` (called from `adapter.py`), and only when the
legacy directory exists AND the tenant target does not. On this install both
existed — writers had already created `tenants/_default/global/` on their own —
so the migration short-circuited to a permanent no-op and every one of
`.corvin/{global,forge,sessions}` is a REAL DIRECTORY holding its own separate
content. Verified 2026-09-07. A reader that assumes the symlink converges two
paths is wrong on any tenant-native install and on any install where a writer
won the race.
**ONE audit chain per tenant: `<corvin_home>/tenants/<tid>/global/forge/audit.jsonl`.**
Every writer of a hash-chained record resolves it through
`forge.paths.tenant_audit_chain()` / `core.paths.tenant_audit_chain()` /
`corvin_operator/bridges/shared/paths.py::tenant_audit_chain()` — three byte-identical
mirrors, pinned together by `tests/security/test_audit_chain_ssot.py`. That file
is what the ADR-0232 boot tripwire verifies, what `audit_query` reads and what
every compliance report is generated from, so a record written anywhere else is
not in the audit trail at all. Until 2026-09-07 `security_events.write_event`
took its path from the caller and every caller composed its own: SIX live chain
files for tenant `_default` across two roots (measured — see ADR-0650).
The historical files are append-only and are NEVER merged, rewritten or deleted;
the boot check links them with a chained `audit.chain_supersedes` seam naming
each one's path key and final tail hash (`security_events.chain_seam_links`).
Console routing: All routes use `rec.tenant_id` from authenticated `SessionRecord`, **never env vars**.
Cross-tenant isolation verified; audit trail records correct `tenant_id` for every event.
**Must NOT do:** fold `tenant_id` into positional args (keyword-only) · use env-var fallback
for console tenant routing · bypass `validate_tenant_id()` · compose an audit-chain
path by hand instead of calling `tenant_audit_chain()` · assume the ADR-0007
compat symlinks exist · merge, rewrite, reorder or delete an existing chain file
to "unify" the trail — link it with a seam instead.
---
## Project Identity — CorvinOS (hard cut)
Repository is **CorvinOS**. Hard env-var cut: only canonical prefix is `CORVIN_*`.
No `ATELIER_*` / `CLAUDEOS_*` / `TESSERA_*` fallbacks — collapsed to `CORVIN_*`, not preserved.
Canonical runtime root: `~/.corvin/`; voice/secret config: `~/.config/corvin-voice/`.
**Must NOT do:** re-introduce legacy env-var fallback · run `sed -i` over live `~/.corvin/audit.jsonl`
(corrupts hash chain) · rename `~/.corvin/` without `corvin_migrate.py`.
---
## Layer Stack Overview
36 security + compliance layers. **Mandatory reading:**
- **L4** Cowork (multi-persona hub) — [layer-plugins.md](docs/claude-ref/layer-plugins.md)
- **Plugin registry** (ADR-0030/0033/0233/0243) — ONE lifecycle contract in
`core/plugins/corvin_plugins/`. **Three orthogonal axes, never conflated:**
`boot_layer` (compliance·core·bundled·installed — load order + disableability, ADR-0243) ·
`tier` (ADR-0156 capability boundary + license gate — the ONLY meaning of "Tier A/B/C") ·
`origin` (builtin·vetted·community — provenance). Don't add a second registry,
lifecycle, taxonomy, or marketplace downloader, and don't rename the axis to
"tier" or back to "layer" — "layer" is already four-way taken (L1–L44 stack,
ADR-0124 audit layers, ADR-0142 layer-extension API, quality layers), which is
why the field is `boot_layer` / `BootLayer` and the API is `boot_layer_of()`,
`plugins_by_boot_layer()`, `_declared_boot_layer()`,
`register_global_plugin(..., boot_layer=)`, audit `plugin.boot_layer_rejected`,
admin field `boot_layer` / aggregate `by_boot_layer`.
**The plugin perimeter is ATTRIBUTION, not security.** An in-process plugin is
part of the process. Five adversarial rounds each broke the previous identity
guard (object → `plugin_id` parameter → `loading.current()` ContextVar), and
the last one needed one line: the setter is public, `threading.Thread` does not
inherit ContextVars and `asyncio.create_task` copies them past the unload. In
CPython every property of a caller is settable by that caller — **do not add a
sixth derivation.** Anything that must hold against a hostile plugin belongs in
a subprocess (ADR-0241/0238); the one in-process guard worth writing is the boot
tripwire, which is non-overridable and runs first. Say "attributed", never
"verified" — the audit chain is permanent.
**Boot-layer rules — a MECHANISM, today with zero instances above `bundled`.**
`_GLOBAL_SPECS` is empty, `register_global_plugin()` has no production caller,
`bootstrap_global()` returns `[]`, and nothing loads on `compliance`/`core` —
so `registry.replace()` is structurally unreachable. The rules below are
implemented and tested and apply the moment a first instance exists; do not
weaken them, and do not describe them as load-bearing today:
tenant scope may declare only `bundled`/`installed` — a privileged claim from
`tenant.corvin.yaml` or `registry.yaml` is downgraded to `installed` and
audited, never honoured (**this one IS live**) · `origin=community` may never
claim a privileged boot layer · the `compliance` boot layer has no off switch
(`registry.disable()` raises `PluginDisableRefused`) · `bootstrap_global()`
carries no feature flag *because* it loads the compliance boot layer · only
`boot_layer=core` is replaceable. Guard tests in
`core/plugins/tests/test_layered_boot.py` fail on the first real instance and
force the docs to be updated in the same commit — keep them.
— [layer-plugins.md](docs/claude-ref/layer-plugins.md)
- **L5** Auto-routing (keyword-based persona selection)
- **L6** Forge (runtime tool generation, MCP server)
- **L7** SkillForge (runtime skill generation)
- **L10** Path-Gate (FS-write protection, fail-closed)
- **L16** Security hardening (TOCTOU, audit framing, consent)
- **L18–21** User management (roles, disclosure, quota, proposals)
- **L22** WorkerEngine protocol (ClaudeCodeEngine, CodexCliEngine, etc.)
- **L23** Speech-to-Text (metadata-only audit)
- **L24** Large-Data Snapshot + **L25** Compute Worker + **L32** Anonymisation
- **L28** Conversation Recall + User Modeling
- **L29–30** Delegation + Engine-Agnostic Forge
- **L33** Session Artifact Memory
- **L34** Data Classification + Flow Guard (4-stage × engine matrix, fail-closed)
- **L35** Network Egress Lockdown (allowed/forbidden hosts, EU_PRODUCTION presets)
- **L36** GDPR Art. 17 Erasure Orchestrator
- **L37** Audit-at-rest Encryption + Retention (age/gpg rotation, RFC 3161 TSA)
- **L38** RemoteTriggerReceiver + A2A TaskEnvelope Protocol (Protocol v6, instance attestation, attachments)
- **LIP** Layer Integrity Protocol (ADR-0141, CAP_VERSIONS + manifest signing)
- **CLS** Custom Layer System (ADR-0156, Tier-A/B/C licensing gate)
→ Full layer index: [layer-summary.md](docs/claude-ref/layer-summary.md)
---
## Plugin-Based Isolation — All Features Always On (load-bearing)
**POLICY CHANGE (2026-09-01):** Ship-dark-by-default via feature flags is **obsolete**.
Plugins are now the isolation unit. All features activate by default; control via plugin enable/disable.
**Rationale:**
- Plugin system provides process-level isolation (ADR-0243, boot_layer model)
- Code safety is structural (plugin boundary), not behavioral (feature flag)
- Feature flags added complexity without real safety gain; plugins do it cleaner
- Operator controls features by managing plugins, not by flipping buried flags
**Rule (effective immediately):**
| Scenario | What to do |
|---|---|
| New plugin/subsystem | Ship active by default; plugin lifecycle controls visibility, not a flag. |
| Legacy feature flag exists | Deprecate on next release; migrate logic into a dedicated plugin or remove flag check. |
| Feature MUST be toggleable at runtime | Make it a plugin setting (`plugin.corvin.yaml`), not a feature flag. |
| Experimental behavior (high-risk, unsure) | Ship as separate plugin, disabled by default in `registry.yaml`. |
| Security/compliance mechanism | Always on, never toggleable (no feature flag, no plugin disable, no env var). |
**Concrete workflow:**
1. **Design feature as a plugin** (or plugin extension, if it enhances an existing plugin).
2. **Register in `registry.yaml` or `tenant.corvin.yaml`** with `enabled: true` by default.
3. **Test both states** (plugin loaded vs. not loaded), but default state is LOADED.
4. **Commit:** `feat(plugin-name): description [ADR-XXXX]` (one ADR per new plugin).
**Migration (existing features with flags):**
- Flags staying: `spec.web_chat.worker_engine`, telemetry opts (GDPR-driven).
These are *settings*, not on/off toggles; keep them.
- Flags going: Any `spec.features.<flag>` that defaults to `false` — migrate to plugin-based control.
**Critical exception:** Compliance / security mechanisms stay **always-on, non-disableable**
(ADR-0232/0233, L44, audit chain, disclosure, consent gates). A plugin cannot weaken them.
These are not features; they are load-bearing constraints.
**Must NOT do:** ship feature flags with default-off · treat plugins as "experimental" and
disable by default unless high-risk and explicitly gated · use feature flags to control
plugin visibility (use plugin enable/disable instead) · weaken compliance mechanisms via
feature flag or plugin disable.
---
## Plugins Live Fully in the Marketplace, Not in CorvinOS (load-bearing)
**Operator rule (2026-09-05): the complete source of EVERY plugin lives in the
Corvin-Marketplace repo (`/home/shumway/projects/Corvin-Marketplace/`) — ALWAYS — never in
CorvinOS.** The goal is to keep the CorvinOS codebase small: CorvinOS loads/references plugins
from the marketplace, it does not own a copy of their implementation.
**What "fully in the marketplace" means:** a plugin ships its real
`plugins/buildin/<category>/<name>/` tree — `src/` (the implementation), `tests/`, `README.md`,
`setup.py`, and `plugin.json` (with `source_url` pointing at the marketplace `src/`, not
CorvinOS). A `plugin.json` alone is incomplete.
**How to apply:**
- When building or adding a plugin, put its full source in Corvin-Marketplace, not in
`CorvinOS/core/plugins/buildin/...`. Do not duplicate the source into CorvinOS.
- CorvinOS's load path resolves plugins from the marketplace (install/discovery), not from an
in-wheel `core/plugins/buildin/<category>/<name>/` source copy.
- Migrating an existing in-CorvinOS builtin means: move its source to the marketplace, change the
load path from builtin-in-wheel discovery to marketplace resolution, then DELETE the CorvinOS
copy — and verify the plugin still loads + activates end-to-end before declaring done.
**Must NOT do:** keep a plugin's real `src/` under `CorvinOS/core/plugins/buildin/` when it also
lives in the marketplace (that duplicates it and grows CorvinOS) · ship a marketplace entry that
is `plugin.json`-only with no `src/` · point a plugin's `source_url` at CorvinOS instead of the
marketplace.
---
## Testing + Docs Sync (load-bearing)
**Before committing** changes to `adapter.py`, `daemon.js`, or `shared/js/`:
```bash
bash corvin_operator/bridges/run-all-tests.sh
```
**Every feature change** — code, config, behavior, API, protocol, CLI, error message —
**must update docs AND diagrams in the same commit**. No deferred updates. No exceptions.
| Changed subsystem | Doc targets | Diagram targets |
|---|---|---|
| Layer N | `docs/claude-ref/layer-N-*.md` + any top-level doc | `docs/diagrams/*.svg` |
| Protocol / wire-format | Protocol reference + tutorials + JSON examples | Flow SVGs |
| CLI command / flag | Every doc that mentions it | Sequence / flow diagrams |
→ Full reference: [testing-and-docs.md](docs/claude-ref/testing-and-docs.md)
---
## Worker-Engine Model Routing (ADR-0759, load-bearing)
**`tenant.corvin.yaml` is CO-OWNED.** The console writes `spec.engine_models`,
`spec.web_chat`, `spec.learning`, `spec.claude_code_local`,
`spec.features_whitelist`, `spec.context_engineering`; `acs_runtime` and
`engine_models` read their own keys. `corvin_gateway.tenant_config.TenantSpec` is
therefore `extra="allow"` with an OPTIONAL AWP envelope. It used to be
`extra="forbid"` with a required envelope, and the consequence was not strictness
— it was **100% of gateway runs terminal-`failed` before an engine was spawned**.
The relaxation is SCOPED: `data_residency`, `budget`, `compute` and
`engine_trust` keep `extra="forbid"`, because those gate engine admission and a
typo there must keep failing loudly.
**Write that file at mode `0o600`, always.** The gateway reads it fail-closed on
mode. `os.replace` carries the temp file's mode onto the target, so a writer at
umask leaves it `0o664` and disables every run. `chmod` BEFORE `replace`.
**Every `engine.span.*` carries a real `model_id`.** `model_usage` keys on it and
an observation without one used to be dropped — so 100% of delegated worker turns
were invisible in the console and priced at $0.00. The span closes on the model
the engine REPORTS, not the one that was requested. `engine_span.END_FIELDS`
carries the four-way token split (`input`/`output`/`cache_read`/`cache_write`);
a single `tokens_used` total cannot be priced, because the four bill at four
different rates.
**All THREE worker paths, not just the dispatcher.** ACS (`acs_runtime`), the
gateway dispatcher, and L38 inbound A2A (`a2a_worker`) each emit their own
worker span. The A2A one shipped with neither a model nor tokens until
2026-09-20, and because the reader drops a span that has neither, every inbound
run was discarded below the console: measured on this install, 267 worker runs
over 11 days, **0 priced**, the panel reading "0 priced worker runs" on an
install that delegates. Derive the split with `engine_span.usage_split()` — the
one normaliser every emitter shares; a private per-path copy is how the same run
prices differently depending on who spawned it — and take the model from
`agents.attested_model(result)`, never from config: A2A passes no `model=`, and
a guessed id prices the run at the wrong rate. An engine that reports no model
leaves it `""` (unpriced, honest); one that reports a model but no tokens still
yields a span, counted in `total_turns` as "N runs, none with token data".
**A field not in `_EVENT_ALLOWLIST` is dropped silently.** `acs.engine_completed`
emitted its token split for months into a floor that discarded it, so ACS worker
cost read as $0.00 while the emitter looked correct. Add the field to the
allowlist in the same commit as the emitter.
**Bedrock / Vertex / Foundry are `auth_mode: platform` — NOT base-url redirects.**
Claude Code reaches them natively (`CLAUDE_CODE_USE_BEDROCK=1` + SigV4). Setting
`ANTHROPIC_BASE_URL`/`ANTHROPIC_API_KEY` for one makes it speak plain Anthropic
to a signed endpoint and every turn 403s. They carry no `credential_env` (the
credential is a CHAIN: AWS profile / IMDS / IRSA / ADC / Azure token), their
endpoint is region-derived, and the L35 gate must use `ProviderSpec.egress_url` —
gating on `base_url` (empty for them) skipped the check entirely.
`corvin_operator/bridges/shared/aws_sigv4.py` is the stdlib-only signer; it was
deleted once by a rename sweep while its docs stayed, and
`test_the_signer_module_is_present_at_all` now fails if that recurs.
**Engine detection: the platform flag outranks a stale OAuth file.** Nothing
removes `~/.claude/.credentials.json` when an operator switches to Bedrock, and
Claude Code itself gives the flag precedence. `EngineProbeResult.plan` reports
which subscription (`pro`/`max`/`team`/`enterprise`) — labels only, never a token.
**`_strip_worker_secrets` disarms a platform host unless `_restore_platform_credentials`
runs after it.** The strip matches `SECRET`/`ACCESS_KEY`/`_TOKEN$` by name — which
is exactly the only credential a Bedrock/Vertex worker has. Restore is scoped to
the ACTIVE platform and to registry-declared names; never widen `_SECRET_NAME_RE`.
**Must NOT do:** re-tighten `TenantSpec` · write the tenant YAML at umask · emit a
span without `model_id` · add a worker spawn path without a priceable span ·
copy the usage normaliser instead of calling `engine_span.usage_split()` ·
fill a span's `model_id` from config when the engine reported one ·
pass `model=` to an engine without probing the keyword ·
add an audit field without an allowlist entry · compose an audit-chain path by
hand · give a platform provider a `credential_env` · audit a credential-absent
catalogue refresh as a failure (4 671 such records buried the real events).
→ Full reference: [layer-engines.md](docs/claude-ref/layer-engines.md) § Worker-engine model routing
→ ADR: See Corvin-ADR for ADR-0759
---
## OS-Model Routing — three tiers, one resolver (ADR-0952, load-bearing)
**OS turns route Haiku 4.5 / Sonnet 5 / Opus 5, decided by the complexity
classifier at Tier 2.9 of `model_selector.resolve_os_model`.** That resolver is
THE single source of truth for both surfaces — the console web-chat
(`chat_runtime.py`) and the bridge adapter — and both pass `task_input`. Wiring
one and not the other re-creates the divergence the shared function exists to
prevent: measured 2026-09-20, the bridge produced 295 of 300 recorded OS turns.
**A pin at Tier 2.5 still wins outright.** `spec.engine_models.<engine>.os_model`
returns immediately and every tier below it is unreachable — which is correct,
and is why an install pinned to one model shows one model on all three
complexity tiers. That is not a licence limit; say so.
**Every model any tier can choose MUST be in `_MODEL_RANK`.** An absent id ranks
0 — below Haiku — so the cache guard reads an escalation to it as a downgrade
and `apply_floor` can never reach it. `claude-opus-5` was missing, and that
alone made the top tier unreachable.
**Resolve a model id against the registry before routing it.** The operator's
saved per-tier choice is provider-qualified (`anthropic/claude-opus-5`), the
classifier names a family (`claude-haiku-4-5`), and the registry holds the
dated snapshot (`claude-haiku-4-5-20251001`) behind an exact-match, fail-closed
test. `resolve_registry_id()` strips the prefix and accepts a dated snapshot
only when exactly ONE registered id extends the family — never when two do.
**Every PIN tier normalises too, not just Tier 2.9.** Tiers 1 / 2 / 1.5 / 2.5 /
2.7 used to return the stored string verbatim, and the Console persists an
operator's choice PROVIDER-QUALIFIED — so a pin saved from Settings → AI
Engines or the Models routing tab reached `--model` as
`anthropic/claude-haiku-4-5-20251001` and the API answered `404
model_not_found`: `duration_api_ms: 0`, `modelUsage: {}`, `terminal_reason:
api_error`, not one token billed, on EVERY turn of that surface. Measured
2026-09-21: the Discord bridge answered "Claude API call failed: 404" from the
minute the pin was written. `model_selector.normalise_pin()` is the shared
guard, and it is deliberately more forgiving than Tier 2.9's gate — a pin is
the operator's instruction, so an id this process cannot verify passes through
(minus a `provider/` prefix, which is never valid CLI input) rather than being
dropped; dropping it would silently substitute a different model, which is the
defect the Models console already had. The worker side is the same shape:
`CORVIN_ACS_WORKER_MODEL` (adapter) and the gateway dispatcher's
`worker_model` lookup both normalise. `PUT /v1/console/settings/engine` also
normalises Claude-native pins BEFORE writing the YAML, so what an operator
reads back is what runs.
**Admission is per verdict, not one scalar.** The classifier's "confidence" is
four constants on four branches of a rule tree, not a probability. `simple`
0.85, `medium` 0.60 and measured `complex` 0.90 route; keyword-only `complex`
0.70 does not, and keeping it out is what gives the table a live subject.
Admitting `medium` changes no served model — Tier 3 returns Sonnet 5 anyway —
it makes the decision auditable. **Operator policy (2026-09-27):** ADR work,
reviews and Markdown-file work are `complex` 0.95 → Opus regardless of length;
a short work request is `medium` 0.75 → Sonnet; only conversation is `simple`
→ Haiku. `classify_os_model` needs `ModelSelector` in
`core.skills.os_skills.model_selector` — overwriting that module (4036473fe)
silently kills Tier 2.9.
**Model lineage (ADR-2090).** Automatic tiers route to the NEWEST version of a
family (`model_selector.tier_model()`, `model_lineage.latest()`), so Opus 5.5
serves complex today. Pins and saved per-tier choices are kept until their model
is retired. Only the CLI's own `issue with the selected model (<id>)` answer
retires a model (`retired_models.json`, 7-day TTL). The successor is applied at
`ClaudeCodeEngine._build_args` for every surface. **Never** treat absence from
the curated YAML or from the live `/v1/models` list as retirement. **Never**
hardcode a model id where `tier_model()` / `top_model()` exists.
**`CORVIN_OS_MODEL_AUTOSELECT=off` disables Tier 2.9 too.** It means "do not
pick a model for me"; a kill-switch that silently stops killing is worse than
none. Explicit pins are unaffected.
**Import the classifier off the request path.** `core.skills.os_skills.model_selector`
costs 1.09 s to import (numpy chain) and 2.2 ms to run. Paid lazily it lands
inside `stream_turn`'s async generator and stalls the whole event loop — the
2026-09-13 incident shape. A daemon-thread warm-up at module import also closed
a 50%-rate WebSocket-test hang that had been read as flakiness.
**Must NOT do:** wire `task_input` on one surface only · let Tier 2.9 override an
operator pin · add a routable model without a `_MODEL_RANK` entry · pass a
provider-qualified or family id to the CLI unresolved · add a pin tier that
returns its stored string without `normalise_pin()` · make `normalise_pin()`
fail-closed (a pin it cannot verify is the operator's business, not a drop) ·
write a provider-qualified pin into the tenant YAML · collapse the per-verdict
admission table back into one threshold · leave Tier 2.9 outside the autoselect
kill-switch · import the classifier on the hot path.
→ Full reference: [layer-engines.md](docs/claude-ref/layer-engines.md) § Three-tier OS-model routing
→ ADR: See Corvin-ADR for ADR-0952
---
## Usage Counting Epoch (ADR-0760, load-bearing)
**Never trim the audit chain to "reset" a counter.** It is append-only and
hash-linked, the ADR-0232 boot tripwire verifies it before anything else runs,
and it is the GDPR Art. 30/32 record — removing a record does not reset a
number, it breaks the chain and fails the next boot. Reset the VIEW instead:
one timestamp per tenant in `<tenant>/global/usage_epoch.json` says from when
the console counts. Clearing it restores every historical turn, because nothing
was ever removed.
**One epoch, every reader.** `model_usage()` (turns) and
`compute_cost_efficiency()` (dollars) both call `usage_epoch.epoch_ts()`. Two
totals on one screen counted over different windows is worse than no reset.
**Filter per EVENT, before folding spans.** A span starting before the epoch and
ending after it must be excluded whole — folding only its end produces a turn
with no start and no status, which reads as `unfinished`: a crash that never
happened.
**A narrowed total never travels without its window.** Every payload carrying
one also carries `window` (`active`/`epoch_ts`/`since_iso`/`reason`), and the UI
renders the period ABOVE the number.
**Never average two savings percentages.** Sum the real dollars and divide once
— averaging weights a 4-turn worker series like a 500-turn OS one. A combined
figure is withheld entirely unless BOTH sides have data in the window.
**Must NOT do:** trim/rewrite/delete the chain to reset a counter · give a panel
its own window · filter after folding spans · return a narrowed total without
`window` · average percentages into a combined saving · describe the reset as
deleting data in UI or API · conflate the epoch with retention or GDPR Art. 17
erasure (L36 owns those; they change what exists, this changes what is counted).
→ Full reference: [layer-engines.md](docs/claude-ref/layer-engines.md) § Counting epoch
→ ADR: See Corvin-ADR for ADR-0760
---
## Charts in the Console (ADR-0761)
**Load the `dataviz` skill before writing chart code** — it is an explicit
trigger, and the palette part is computable, so compute it: run the validator
against the surface the chart actually renders on (`#ffffff` light / `#0e1320`
dark for this console), never the tool's default surfaces.
**Never a dual axis.** Two measures of different scale → small multiples. And
give faceted small multiples ONE shared domain: independent scales read as
comparable while not being comparable, which is worse than the shared-axis chart
they replace.
**The form degrades with the data.** A line or area over a single point draws
nothing while still being titled "Trend" — below two points, bars.
**A zero-length mark needs `minPointSize`**, or recharts drops the mark AND its
label and the row vanishes silently.
**Colour nominal categories by identity, ordered tiers by an ordinal ramp keyed
on the TIER** — never by the measured value, which double-encodes bar length as
hue.
**Encodings live in a testable module, not in the component.** A wrong
number→mark mapping (e.g. a negative saving labelled "Referenz") is invisible in
a screenshot until a tenant hits the case.
**Verify by looking.** The validator checks colour, not layout. Screenshot it —
and note the console is `data-theme` driven, NOT `prefers-color-scheme`, so a
headless run that sets `color_scheme` renders one theme twice.
→ Full reference: [layer-engines.md](docs/claude-ref/layer-engines.md) § Cost-visualisation encodings
→ ADR: See Corvin-ADR for ADR-0761
---
## Console is a Production Surface (ADR-0763, load-bearing)
**Shipped UI is English, carries no ADR ids, and fabricates nothing.** ADR
references belong in code comments, never in rendered text. A 404 branch returns
an EMPTY result plus "not available on this build" — never sample data:
`quality.tsx` built a seven-day trend from `Math.random()` with invented gate
failures, and `runAllGates` reported `ok: true, "Gate run initiated"` for a run
that never started. Language **detection** regexes stay — the bot answering in
the operator's language is intended runtime behaviour, not UI copy. Installation
defaults to English (`<html lang>`, narration locale).
**Don't chart a mechanism with no consumer.** The panel headline plotted
`learned_threshold` against 0.5 while the learner's own docstring says nothing
reads it to make a routing decision. Chart what the operator acts on — volume,
reliability, spend — especially when it was computed in the same pass and thrown
away.
**A recommendation needs a sample.** Withhold it below a real threshold and say
why; "100% over two turns" is one data point wearing a percentage. Unit cost
divides by PRICED turns, never by all turns.
**Chart colours are Corvin's amber, in their own `--viz-*` tokens.** Never alias
`--accent` as a data colour: it is a brand token that may be restyled and at
L 0.74 it is outside the dark lightness band for a data mark. Never put two
colour systems in one view (tier ramp + role swatch).
**A zero-findings sweep needs a positive control.** The German/ADR scanner
resolves `src/` relative to `web-next/`; run from the repo root it sees zero
files and reports success. Confirm the file count first.
→ Full reference: [layer-engines.md](docs/claude-ref/layer-engines.md) § The console as a production surface
→ ADR: See Corvin-ADR for ADR-0763, ADR-0764 (cross-page consistency: one window, named denominators, pinned en-US formatting)
---
## Console Frontend — Prove the NEW Build Is What Loads (load-bearing)
Any change under `core/console/corvin_console/web-next/` is **not done when the source is
correct** — it is done when the browser demonstrably runs the NEW bundle. Source-correct +
stale-bundle has burned this repo repeatedly: the API answers 200, the UI still shows the old
placeholder, and debugging goes into the backend instead of into the three caches in front of it.
"Build succeeded" means *new code was added*, never *old code was removed*.
**Three caches sit between an edit and the screen:**
| Layer | Where | Cleared by |
|---|---|---|
| 1. esbuild pre-bundle | `web-next/node_modules/.vite/` | `rm -rf node_modules/.vite/` |
| 2. build artifact | `web-next/dist/` | rebuild — `scripts/console-deploy.sh` |
| 3. browser tab | the operator's machine | `console_auto_reload` flag, else `Ctrl+Shift+R` — **Claude cannot press it** |
**Use `scripts/console-deploy.sh` — it performs the sequence AND the proof:**
```bash
scripts/console-deploy.sh # clean rebuild + verify over the wire
scripts/console-deploy.sh --marker 'NewThing' # additionally assert a string is bundled
scripts/console-deploy.sh --fast # incremental (esbuild minify, ~24s)
```
It prints `LIVE assets/index-<hash>.js` only when the host actually serves the bundle
it just built, and exits 2 with the mismatch otherwise. It builds into `dist.next/`
and swaps, so `/console/` never stops answering mid-deploy (vite empties its outDir
first, which took the console down for ~13s per rebuild when building into `dist/`).
The previous build is kept at `dist.prev/` for rollback.
Two mechanisms keep this running without anyone remembering to:
| Mechanism | Covers | Notes |
|---|---|---|
| `corvin-console-watch.service` (systemd --user) | ANY source change, from any editor | polls an mtime fingerprint, debounces, then runs `console-deploy.sh --fast` |
| `.claude/hooks/console_autodeploy.sh` (PostToolUse) | Claude's own edits | starts the redeploy at the edit; skips when the watcher is running, so no double build |
The equivalent by hand, if you need to see each step:
```bash
cd core/console/corvin_console/web-next
rm -rf dist/ node_modules/.vite/ # a plain rebuild is NOT sufficient
npm run build
grep -rl '<marker string from the new code>' dist/assets/ # must hit ≥1 file
curl -s http://127.0.0.1:8765/console/ | grep -o 'assets/[^"]*\.js' # must be the NEW hashes
```
**Restart rule.** `mount_static()` (`core/console/corvin_console/app.py:444`) decides ONCE at
boot. If the console booted while `dist/` was absent, it registered the 503 "build failed"
fallback route instead of the SPA mount and keeps serving it — a rebuild alone will NOT recover
it, only a restart will. When `dist/` existed at boot, `_SPAStaticFiles` resolves per request and
a rebuild is picked up live.
**A bare 404 on `/console` (not 503) is a BACKEND import failure, not a build problem.** The
live host is `corvin-webui.service` (`corvin_gateway.app` on :8765); it mounts the console
through an opt-in `try: from corvin_console import app` (ADR-0015), so an ImportError anywhere
in the console's ~120 route modules does not crash the boot — it removes `/console` AND every
`/v1/console/*` route from the process. Until 2026-09-17 that clause was `except ImportError:
pass`, so the journal showed nothing; 9433de4b had removed `get_optimizer` from
`core.learning.model_selection_optimizer` under `routes/model_selection_analytics.py`, the
running console pre-dated the commit, and the next `systemctl --user restart corvin-webui`
took the whole console down. Now the gateway logs `corvin_console is present but failed to
import` with the traceback, and `tests/test_console_app_importable.py` imports the console
app in a fresh interpreter with the unit's PYTHONPATH. On a 404: `journalctl --user -u
corvin-webui | grep 'failed to import'` first, then restart after the fix.
**Layer 3 — the browser tab.** With the `console_auto_reload` flag ON, an open tab
re-fetches the no-cache SPA shell every 3s, compares its entry-bundle hash against the
one it booted with, and reloads itself onto a new build (banner instead of reload while
the operator is mid-input, so typing is never discarded). See
`web-next/src/hooks/use-build-freshness.ts`.
**With the flag OFF — the default — tell the operator to hard-refresh**, explicitly, in
the same message that reports the change. Layer 3 is then the only cache Claude cannot
clear, and it is by far the most frequent cause of "the feature isn't showing". On any
invisible-frontend report, confirm the hard refresh FIRST — never open a backend
investigation before that.
**Cache-header invariant (`_SPAStaticFiles`, `core/console/corvin_console/app.py:23`) — do not
weaken:** content-hashed files under `assets/` get `public, max-age=31536000, immutable`; every
`text/html` response AND every `304` gets `no-cache` — including the bare `/console/` directory
index, which is not named `index.html` and once slipped through a path-based check. Inverting
either half is exactly how a browser keeps requesting deleted hashed bundles and the app hangs on
a perpetual "Loading…".
**A panel needs TWO registrations.** `PANELS` (`src/panels/registry.tsx`) mounts the
route; `NAV_GROUPS` (`src/components/layout.tsx`) draws the sidebar entry. `ConsolePanel.nav`
looks like it drives the sidebar and does not. Registering only the first mounts a route
nothing links to — which presents as "the console still shows the old build", and did, for
seven finished panels at once. `tests/unit/panel-nav-wiring.test.ts` fails on the next one;
a panel that is deliberately not in the sidebar goes in that test's `NAV_EXEMPT` set WITH a
reason. A `requiredFlag` must ALSO be listed in `GATED_FLAGS`
(`core/console/corvin_console/routes/capabilities.py`) or it resolves to false and the entry
stays hidden forever.
The ADR-0561 console manifest (`/v1/console/capabilities/manifest`) is **additive** to
`NAV_GROUPS`: `mergeManifestNav()` in `layout.tsx` appends manifest-only panels (plugins,
skills, installed) and never removes a static entry. Rendering the sidebar FROM the
manifest hid ~30 core panels on 2026-09-03, because the backend enumerates only the
panels it knows about. Keep the static list complete; let the manifest extend it.
**A sibling FILE silently shadows a page DIRECTORY.** `import("@/pages/foo")` resolves
`src/pages/foo.tsx` BEFORE `src/pages/foo/index.tsx` — file beats directory, with no
warning from vite, tsc or eslint. A new panel built as `pages/foo/` while the old
`pages/foo.tsx` still exists therefore compiles, bundles, mounts its route and renders
**the old page**, which presents exactly like a stale bundle and sends debugging into the
three caches or the backend. It happened to the ADR-0400 Vibe Dashboard: the directory
shipped 2026-08-26, `pages/vibe-engineering.tsx` kept winning, and the panel was
unreachable until the file was deleted on 2026-08-27 (ADR-0431). It happened AGAIN on
2026-09-17: 95ecc2b6 re-added `pages/vibe-engineering.tsx`, the route crashed on nine
404s, and the same commit routed `/app/model-cost-optimizer` to a page fetching
`/api/v1/...` paths that exist nowhere in this console. When a rewrite lands as
a directory, DELETE the same-named file in the same commit — never keep both — and prove
which one loads with a marker string only the new code contains
(`scripts/console-deploy.sh --marker '<string>'`), not by reading the diff.
`web-next/tests/unit/page-dir-shadow.test.ts` fails on the next sibling file; a
deliberately dead pair goes in its `SHADOW_EXEMPT` map WITH a reason.
**Children of `<Routes>` must be `<Route>` elements — never a component that returns them.**
react-router walks the `<Routes>` tree statically (`createRoutesFromChildren`) and throws
for any child whose type is not `Route`/`Fragment`. A `<PanelRoutes />` component in that
position type-checks, lints and builds, then kills the ENTIRE console at first render with a
message-less `Uncaught Error` (the invariant text is stripped in prod) — blank page, and it
presents like a stale bundle. It happened on 2026-09-03 (`<ManifestPanelRoutes />` in
`src/App.tsx`). Compute the routes in a hook that returns an array and splice `{routes}`
in. `tests/unit/app-routes-static.test.tsx` renders the real `<App />` and fails on the
next one.
**A "the old code is gone" assertion needs a POSITIVE control first.** A test that
greps the served assets for a removed string passes both when the string is gone and when
the check never looked at the file it lived in — and a route's chunk is lazily imported, so
its filename appears only INSIDE an eager bundle, never in `index.html`. Scanning the
shell's 8 assets therefore proves nothing about the panel in the 9th; it read as green for
a check that had zero reach (2026-09-15,
`tests/e2e/test_engine_config_real_data_e2e.py`). Crawl the chunk graph transitively, then
assert a string only the NEW code contains — and assert it BEFORE the absence check, so a
crawl that stops short fails loudly instead of passing vacuously. Same shape applies to any
"X no longer exists" claim: prove you reached X's home first.
**Must NOT do:** declare a frontend change "done"/"live" on a correct source diff alone ·
assert a string's absence from the bundle without a positive control proving the crawl
reached the chunk that string lived in ·
run `npm run build` without clearing `dist/` + `node_modules/.vite/` first · skip the
`grep` + `curl` proof that the served hashes are the new ones · report the change without
telling the operator to hard-refresh WHEN `console_auto_reload` is off · cache the SPA
shell as anything but `no-cache` · drop the `immutable` header from `assets/` · assume a
rebuild alone revives a console that booted without `dist/` · add a panel to `PANELS`
without a matching `NAV_GROUPS` entry (or a justified `NAV_EXEMPT` line) · build straight
into `dist/` and take the live console down for the length of the build · leave a
`pages/<name>.tsx` in place after moving that page to `pages/<name>/`.
---
## ADR Gate — Architectural Decision Records
**adr-gate is a standard quality discipline.** After every non-trivial task, follow the rubric
before declaring "done." HIGH BAR: default answer is **NO ADR needed**. Most tasks produce none —
that is correct and expected.
**Write ADR only when BOTH hold:**
1. Real design choice was made (chose A over B; constrains future code; genuine alternative existed)
2. At least one structural trigger: new protocol/wire-format/schema, security/compliance mechanism,
irreversible default (fail-open/closed), cross-repo binding (≥2 repos), new layer-level contract
**Skip reasons:** bug fixes, pure refactors, config tuning, test-only/docs-only changes.
When skipping, name the reason in one sentence — never skip silently.
**Destination:** `Corvin-ADR/decisions/XXXX-short-title.md` (sibling repo). Numbering: max + 1.
Commit message: `adr: add ADR-XXXX — [title]`.
**Every ADR carries ADR-0264 frontmatter** (`id`/`status`/`depends_on`/`related`/`paths`/`docs`)
ahead of the prose — this is what makes an ADR a node `scripts/adr_graph.py` can traverse to
from a code path (`paths:`) or a doc path (`docs:` — the same surface `docs-as-definition-
of-done` keeps synced) instead of a document a reader must already have found. Never hand-fill
`superseded_by`; it is derived automatically from every other ADR's `supersedes` list. A
document-generator that emits an ADR-shaped artifact (e.g. Plugin-Builder's per-plugin ADR,
`core/plugins/plugin_builder/generators/adr.py`) carries the same schema, with a plugin-scoped
`id` (`{plugin_id}-ADR-0001`, never a bare `ADR-NNNN`).
**Must NOT do:** write ADR content into Corvin repo (ADRs live in Corvin-ADR only) ·
auto-skip security/compliance mechanisms without justification · declare "done" on structural
change without running this gate · leave a skip implicit · hand-fill `superseded_by`.
→ Full reference: [adr-gate.md](docs/claude-ref/adr-gate.md)
→ ADR: See Corvin-ADR repo for ADR-0264 (adr-decision-graph)
---
## Concept Gate — self-learning working-method archive
**Sibling gate to ADR Gate, same placement (end of task, before declaring "done"), same HIGH
BAR.** ADR Gate asks "does this decision need a record?"; Concept Gate asks "did I just execute
or discover a reusable WAY OF WORKING — not a one-off decision, not a one-off bug fix — that
would save real time if a future agent had it pre-loaded?" Default answer is **NO concept
needed**. Most tasks produce none — that is correct and expected, exactly like ADR Gate.
**Write/amend a concept only when at least one holds:**
1. The same investigation/fix/verification SHAPE recurred across 2+ distinct tasks in ways that
weren't already duplicated code.
2. A structural fix (root cause, not symptom) generalizes beyond the one bug it closed, worth
naming so the next similar bug gets the deep fix on the first pass.
3. A verification technique (live-VM proof, wheel-content inspection, deploy-then-curl-the-
real-URL) proved decisive in a way that will recur.
**Skip reasons:** a single bug fix with nothing generalizable; a one-off tooling workaround;
already fully captured by an existing concept (amend it, never create a near-duplicate) or by a
Skill (if a ≤8KB behavioral snippet is genuinely sufficient, a full concept adds nothing). When
skipping, name the reason in one sentence — never skip silently, exactly like ADR Gate.
**Destination:** `/home/shumway/projects/Corvin-ADR/concepts/CONCEPT-NNNN-slug.md` (sibling repo, own numbering,
never `ADR-NNNN`). Fallback if Corvin-ADR is unreachable: `CorvinOS/docs/concepts/`. Commit
message: `concept: add/amend CONCEPT-NNNN — [title]`.
**A concept is not an ADR and not a Skill** — it is the narrative middle layer: *why* a way of
working keeps paying off, with real evidence (cited commits/tasks), Alternatives Considered, and
an explicit "when NOT to use" boundary — the same depth ADR-0264 models, applied to process
instead of architecture. Every durable concept SHOULD mint or update a companion SkillForge
skill (`skill_create`/`skill_promote`, `type: learned-experience`, `scope: project`) whose body
distills the Method section (≤8KB) and points back at the full concept file for depth — this is
what makes the archive genuinely *self-learning*: the concept is what a human (or a future
agent doing a deep read) consults; the skill is what gets auto-injected and auto-graded into
ordinary turns without anyone having to remember the concept exists. Record the skill name in
the concept's `skills:` frontmatter field.
**A freshly minted skill needs one bootstrap grade before it does anything** —
`skill_inject.py`'s injection gate excludes `n_grades < 1 or mean_score <= 0` by default, and a
brand-new skill has no organic path to its own first grade (auto-grading only scores skills that
were already injected). Immediately call `skill_grade` once with a score capped at the
codebase's own `_AUTO_GRADE_CAP_MAX` (0.3) and notes disclosing it as a manual seed, not earned
usage — real grades accrue from there. Skipping this step means the skill sits inert on disk
forever, which is exactly the gap adversarial review of this mechanism found the same day it was
built (see CONCEPT-0001's Amendment).
**Operator notes are first-class.** Every concept file has an `## Operator Notes` section,
append-only, timestamped, human-authored. AI amendments NEVER edit or remove anything under
that heading — they may only add a new dated sub-entry above it, exactly like ADR-0264-style
amendments are prepended under "Status," never rewriting prior text.
**Must NOT do:** write concept content into the CorvinOS repo when Corvin-ADR is reachable ·
create a near-duplicate concept instead of amending the existing one · edit or delete anything
under an existing concept's `## Operator Notes` heading · mint a SkillForge skill above a
persona's namespace-gate prefix (skills created under the `assistant` persona must be named
`assistant.<name>` — see `corvin_operator/skill-forge/README.md`'s namespace-gate section) · declare
"done" on a task that clearly meets a Concept Gate trigger without running this gate · leave a
skip implicit.
→ Full reference: [concept-gate.md](docs/claude-ref/concept-gate.md)
→ Concept: See Corvin-ADR repo for concepts (0001, 0002, 0008, etc.)
→ Skill: `assistant.corvinOS_live_report_root_cause` (project scope, `learned-experience`)
→ Skill: `assistant.corvinOS_reachability_review` (project scope, `learned-experience`)
---
## E2E Wiring Proof — reachability + functional proof (load-bearing)
**Sibling gate to ADR Gate.** ADR Gate asks "does this decision need a record?"; this gate asks
"does this code actually run, and can I prove it?" New code is unreachable until proven
otherwise — a unit test proves a function returns the right value *when called*, not that
anything in the running system *calls* it. CorvinOS has shipped this exact failure class before
(plugin types that register but are never invoked; an auth invariant unit-tested but reachable
from zero live call sites) — both were "tested," both were dead.
**Fires on:** any new function/endpoint/route/CLI command/UI component/plugin/hook/bridge handler
intended to be reachable from outside its own file. Skips trivial edits, pure refactors with
existing E2E coverage, config tuning, and explicitly-marked WIP prototypes.
**Two-phase gate:**
1. **Reachability proof (cheap, first):** find ≥1 real call site outside the definition file
AND outside test files, traceable to a real trigger (route table, CLI registration, UI render
tree, plugin registry, cron config, message-bus subscription). Zero found → **block**,
dispatch `root-cause-by-layer` — fix the wiring, don't lower the bar.
2. **Generate/extend one E2E test, then run it (only after Phase 1 passes):** the test MUST go
through the real transport/interface boundary — an actual HTTP request, CLI subprocess,
browser interaction, plugin-registry dispatch, bridge message, or MCP tool call. **Hard rule:**
a test that imports the target directly and calls it, skipping that boundary, is a unit test
wearing an E2E label and does NOT satisfy this gate. Capture the real execution evidence
(exit code/output) — a bare "I added a test" claim does not close the loop.
**Infeasibility exception:** when the real entry point genuinely can't be driven end-to-end
(hardware dependency, unavailable external system), name the reason explicitly — never skip
silently. Phase 1 still applies unconditionally.
**Must NOT do:** declare a new-entry-point task "done" without this gate · treat "unit tests
pass" as equivalent to "reachable and works end-to-end" · write an E2E test that bypasses the
real transport/interface boundary and call it satisfying · mock the exact component under test
inside its own "E2E" test · leave an infeasibility skip implicit.
→ Full reference: [e2e-wiring-proof-standard.md](docs/claude-ref/e2e-wiring-proof-standard.md) (standard definition) ·
[quality-discipline.md](docs/claude-ref/quality-discipline.md) (LDD context) ·
Skill: `assistant.e2e_wiring_proof` (auto-injected, learned-experience)
---
## Language
**All repository content: English.** Includes: source code, docs, SVGs, inline comments, commits.
**User-facing runtime text: Defaults to English.** Bot answers in user's language (German/English)
at runtime per `adapter.py` system prompt — that is runtime behavior, not repo content.
---
**For full details, read the ref files in `docs/claude-ref/` as tasks require them.**
---
## Phase 3: Learning Infrastructure (ADR-0314+)
**Status:** Phase 3.1 (ADR-0314) COMPLETE ✅
**Scope:** Event schema, persistence, async emission
**Tests:** 34 learning + 174 skill tests (208 total) green
Phase 3 builds the learning layer on top of Phase 2 (Skill System). It enables confidence scoring, user feedback loops, decision history, and attention budgeting.
**ADR-0314 (Learning Infrastructure):**
- Event schema: 8 immutable learning event types (confidence, feedback, outcome, preference, attention, metric)
- Persistence: EventStore with date-partitioned JSON storage + audit trail integration
- Emission: EventEmitter (async queue, non-blocking, fire-and-forget on queue full)
- Integration: Wired into SkillSystemIntegration
**Tenant isolation:** All queries filtered by tenant_id (per GDPR Art. 5, 6, 32).
**Compliance notes:**
- Learning events are audit-logged and hash-chained (GDPR Art. 30, 32) — audit-FIRST and fail-closed since 2026-09-03: `event_persistence.EventStore.write_event` refuses (RuntimeError) when the core chain write does not commit, and the store is tenant-bound (ADR-0563)
- No PII in payloads (validation in downstream ADRs)
- 90-day retention default (ADR-0319 will enforce)
**Downstream ADRs (Phase 3.2–3.8):**
- 0315: Confidence Intervals (relevance/reliability scoring)
- 0316: Decision History (user choice tracking)
- 0317: Outcome Feedback (closed-loop learning)
- 0318: Style Preferences (user model)
- 0319: Attention Budget (finite attention constraint)
- 0320: Metric Collection (aggregation pipeline)
- 0321: Reporting Dashboard (observability UI)
**Must NOT do (ADR-0314 constraints):**
- Don't emit untyped payloads (use frozen dataclasses)
- Don't skip tenant_id isolation (every read/write must filter)
- Don't weaken schema immutability (LearningEvent is frozen)
- Don't bypass audit chain (write_event writes the core chain FIRST; no chain commit → no disk record)
→ Full spec: See Corvin-ADR repo for ADR-0314 (learning-infrastructure-event-schema)
**Loop closure (ADR-0613, 2026-09-06 — load-bearing):** the loop is CLOSED, not a log.
Source = `delegation_policy.resolve_worker_engine` executes `os.delegation_router` in
SHADOW mode (bundled answer stands; decision audited + emitted). Sink =
`TaskManager.record_event` → `core/learning/outcome_sink.py` (OUTCOME per finished task,
tenant from the task's own metadata). Store = `event_store.EventStore.write_event` is
audit-FIRST (core chain record, then disk; no chain commit → no disk record). Consumption =
`SkillAdapter` config (versioned, under `CORVIN_HOME`, rollback-able) is READ by
`DelegationRouterSkill`. Console `/v1/console/learning/*` is real and audited — never mock.
**Must NOT do:** let the shadow path alter routing · hard-wire `~/.corvin` anywhere in
learning/skills (`core/paths/tenant.py::corvin_home` honours `CORVIN_HOME`) · reintroduce
mock answers on learning endpoints · persist the free-text feedback `reason`.
→ Reference: [learning-loop.md](docs/claude-ref/learning-loop.md)
---
## Skills 2.0 as Agentic Control Plane (load-bearing)
**ARCHITECTURE VISION (2026-09-01):** CorvinOS transforms from **task-runner** → **agentic operating system**
by making **Skills 2.0 the unified control plane**. Every subsystem (routing, state, orchestration, security,
learning, deployment) becomes a swappable Skill — written in hybrid code (deterministic Python + LLM-generated
logic), versioned, composable, and self-learning via ADR-0314's feedback loop.
**Core Principle:** A Skill is NOT a static prompt; it is a **program** that:
- ✅ Owns a domain (delegation routing, context adaptation, workflow optimization, security enforcement)
- ✅ Executes deterministic Python (sync I/O, caching, local decisions — fail-fast, auditable)
- ✅ Calls LLM on demand (e.g., `/router classify_request` → LLM picks delegation strategy)
- ✅ Learns via feedback (confidence scoring, outcome signals, optimizer loop per ADR-0314)
- ✅ Composes with other Skills (dependencies declared, topological sort, DAG validation per ADR-0535)
- ✅ Versioned & deployed (semantic versioning, canary rollout, in-flight freeze per ADR-0533)
- ✅ Observable (telemetry, execution traces, decision logs → Vibe dashboard)
**Why now:** ADR-0314 (Learning) + ADR-0532-0535 (OS-Skills) together unlock this. Learning infra provides
feedback signals; Skills infra provides the execution model. Together they make Skills self-optimizing.
**Five Layers of Control Plane Replacement (ADR-0532 Roadmap):**
The **Status** column is the load-bearing one: a Skill that is registered at boot
is NOT thereby wired. Two of the five have no production caller at all, and saying
otherwise is what made a whole "E2E wiring proof" suite vacuous (2026-09-07 round-4
review, F6). Update this column in the same commit that adds or removes a call site.
| Layer | Today | Tomorrow (Skills 2.0) | Status (2026-09-26) | ADR | Timeline |
|---|---|---|---|---|---|
| **L5: Routing** | Hardcoded persona → engine mapping | `os.delegation_router` Skill (LLM-classified by task type) | **WIRED, shadow mode** — `delegation_policy.py::_acp_shadow_route`; the bundled engine still stands (ADR-0613) | ADR-0532 Phase 1 | Weeks 2–4 |
| **L10: Context** | Snapshot → prompt injection | `os.context_adapter` Skill (learns user/task patterns) | **WIRED, shadow mode** (2026-09-26) — CEL stage `l10_adapter` (`corvin_operator/context_engineering/stages/l10_adapter.py`, in DEFAULT + ACTIVE pipeline) calls `adapt_context_l10` per turn via `pipeline.build_context` when `vibe_engineering` is on. Audited (`skill.executed` + `context.adapted`), never changes the served brief. Runs ONLY where `boot_skills` booted an audited registry for the turn tenant (gateway/console); the bridge adapter never boots Skills, so there it is `skipped: skills_not_booted`. E2E: `tests/e2e/test_os_skills_l5_l10_wiring.py::TestL10ProductionCallSite` | ADR-0532 Phase 1 | Weeks 2–4 |
| **L22: Workflow** | Stateless request/response | `os.workflow_optimizer` Skill (learns execution chains) | not built | ADR-0532 Phase 2 | Weeks 6–10 |
| **L16: Security** | Config-driven gates | `os.security_orchestrator` Skill (learns attack patterns) | not built | ADR-0532 Phase 3 | Weeks 11–18 |
| **L34: Data Flow** | Hardcoded validators | `os.flow_guard` Skill (learns safe data shapes) | not built | ADR-0532 Phase 4 | Weeks 19–24 |
**Implementation Roadmap (3 Phases, 8–12 weeks, ~2600 LoC + skills library):**
**Phase 1 (Weeks 1–4): Foundation**
- Deliver `os.delegation_router` Skill + `os.context_adapter` Skill (2 minimal skills)
- Wire into L5 (auto-routing) + L10 (context engineering) — **L5 + L10 done (shadow); L10 not yet on the bridge path (no Skills boot there)**
- Prove E2E: real requests flow through Skills, learning events emitted to ADR-0314
- Tests: 25 E2E, 12 adversarial (crash recovery, timeout isolation, PII leakage)
- Blocker: ADR-0532 Phase 1 + ADR-0533 manifest schema + ADR-0534 feedback integration ready
- Deliverable: `core/skills/os_skills/` directory with routing + context skills + test suite
**Phase 2 (Weeks 5–10): Learning Loop**
- Wire ADR-0314 feedback into Skills → optimization loop
- Add `os.workflow_optimizer` (learns execution chains from user feedback)
- Build dashboard (Vibe → OS-Skills observability panel)
- Tests: 40 E2E + 18 adversarial (optimization convergence, stale feedback, feedback injection)
- Blocker: ADR-0534 (feedback integration) accepted
- Deliverable: Learning loop E2E, dashboard, 2 skills composition-ready
**Phase 3 (Weeks 11–24): Scale & Ecosystem**
- Ship 2 more Skills (`os.security_orchestrator` + `os.flow_guard`)
- Marketplace integration: Skills as discoverable, installable OS subsystems
- Community contrib gate: review checklist for Skill authors
- Tests: 60 E2E + 25 adversarial (marketplace attack surfaces, skill conflicts, versioning)
- Blocker: ADR-0535 (composition), ADR-0536 (marketplace, TBD)
- Deliverable: v2.0 production-ready, Skills marketplace live
**Integration with Compliance (ADR-0232/0233 — non-negotiable):**
| Mechanism | Skill Constraint | Enforcement |
|---|---|---|
| **Audit chain (LOAD-BEARING)** | Every Skill decision + internal state change MUST log via `audit_backend` (appendix-only) | Boot tripwire checks before any Skill runs; audit trail hash-chain verified on every Skill load |
| **Skill-Internal Audit** | All Skill.execute() calls, config changes, feedback processing, optimization steps → audit event | `SkillAuditEvent` (immutable, tenant_id, timestamp, skill_id, input, output, latency, errors, lom) |
| Consent gate | Skill output validated against user consent (L16) | Fail-closed: deny on any PII signal, no bypass |
| House-rules (L44) | Skill code CANNOT disable `house_rules_enforcer()` | Compiled into Skill manifest as immutable `required_checks` |
| Bot disclosure | Skill decisions attributed in transparency log | Every Skill event includes `skill_id`, `version`, `lom` (line-of-moral-responsibility) |
**Integration with Learning Infrastructure (ADR-0314 — self-optimizing):**
```
Skill Execution Loop:
1. Input → Skill.execute() [deterministic Python]
2. Emit SkillExecutedEvent (input, output, latency, errors)
3. User gives feedback (→ FeedbackEvent)
4. Optimizer reads {FeedbackEvent, SkillExecutedEvent}
5. Adjusts Skill config (e.g., router thresholds, context window, retry strategy)
6. Next invocation uses tuned config
7. Dashboard shows confidence score, convergence rate
```
Feedback types (ADR-0314):
- `outcome_feedback` — "was that routing decision correct?" (yes/no/other)
- `preference_feedback` — "prefer this style next time" (LLM, deterministic, neither)
- `confidence_score` — estimated P(Skill makes correct decision | input)
- `metric_observed` — latency, error rate, cost — optimizer adjusts thresholds
**Audit-First Design (LOAD-BEARING):**
Every Skill is a black box from compliance perspective — internal behavior MUST be fully observable via audit trail.
| Stage | Audit Event | Payload | Immutability |
|---|---|---|---|
| **Skill Load** | `skill_loaded` | skill_id, version, config_hash, boot_layer, dependencies, required_checks | Hash-chain verified |
| **Skill Execute** | `skill_executed` | skill_id, version, input, output, latency_ms, errors, lom, tenant_id | Audit-backend appendix-only |
| **Feedback Received** | `skill_feedback` | skill_id, feedback_type (outcome/preference/confidence/metric), signal, timestamp | Audit-backend appendix-only |
| **Config Optimized** | `skill_config_updated` | skill_id, version, param_delta, confidence_before/after, reason | Audit-backend appendix-only |
| **Skill Disabled** | `skill_disabled` | skill_id, reason, requestor, timestamp | Audit-backend appendix-only |
No Skill is a black box: audit trail proves what it did, when, to whom, with what result. **No audit bypass, no silent optimization.**
**LDD for Skills — Dialectical Reasoning Only (SPECIALIZED CYCLE):**
Standard LDD (12 layers) is overkill for Skill changes. Skills use a **lightweight 2-gate cycle:**
| Gate | When | What |
|---|---|---|
| **Gate 1: Dialectical Reasoning** | BEFORE any Skill change (config tuning, feedback-loop adjustment, dependency change) | `/dialectical-reasoning` — argue for/against the change, surface hidden assumptions, confirm tradeoffs |
| **Gate 2: E2E Wiring Proof** | AFTER coding, before merge | Real E2E test proving the Skill runs end-to-end + audit event is emitted + feedback loop processes it |
**NOT required for Skills:** full LDD k=1–5, `docs-as-definition-of-done`, `e2e-driven-iteration` per task, Concept Gate (use existing CONCEPT-XXXX instead).
**Required for Skills:** Dialectical reasoning (surface design choices) + E2E proof (Skill is called and audited).
Example workflow:
```
# 1. Propose Skill change
/dialectical-reasoning
Is tuning os.delegation_router's confidence_threshold from 0.7→0.65 correct?
Argue: for/against, alternatives, risks
# 2. Code + test
core/skills/os_skills/delegation_router.py [edit]
tests/skills/test_delegation_router_e2e.py [add/edit E2E test]
# 3. E2E proof
pytest tests/skills/test_delegation_router_e2e.py -v
# Must show: Skill.execute() called → output produced → SkillAuditEvent logged
# 4. Audit verification
grep "skill_executed.*delegation_router" ~/.corvin/audit.jsonl | tail -1
# Must show: output matches test expectation + lom (line-of-moral-responsibility) present
# 5. Commit + no ADR skip needed
# (Skills inherit ADR requirement from their parent L-layer ADR; incremental tuning is [skip-adr-check])
```
**Must NOT do (hard rules):**
- **Don't hardcode** subsystem logic outside Skills — every L-layer contract must be a Skill
- **Don't skip Dialectical Reasoning** for Skill changes — surface assumptions, confirm tradeoffs (cheap, mandatory)
- **Don't merge without E2E proof** — unit tests ≠ called; prove real execution + audit event
- **Don't audit-bypass** — every Skill.execute() + feedback + optimization MUST log (no silent learning)
- **Don't weaken feedback loop** — no opt-out, no stale-feedback bypass, no learning-off flag
- **Don't let Skill disable compliance** — audit chain, consent, house-rules are meta-Skills (immune to versioning/disable)
- **Don't leak PII into Skill manifests** — manifests are public; learned params must be scrubbed
- **Don't add OS-level Skill without ADR** — one ADR per new L-layer Skill (after ADR-0535)
- **Don't fork SkillForge architecture** — one `skill_forge/` registry, not `skill_forge_v2/`, `os_skill_builder/`, etc.
→ Full reference: See Corvin-ADR repo for ADR-0532–0535 (os-skills-architecture through composition-dependencies)
→ Implementation: `/home/shumway/projects/Corvin-ADR/implementation/ADR-0532-IMPLEMENTATION-PLAN.md` (detailed roadmap, risk matrix, phase gates)
---
## Audit Chain as Ground Truth — Complete User Proof (load-bearing)
**CORE PRINCIPLE:** CorvinOS is a **proof system**. Every action — plugin load, Skill decision, A2A dispatch, consent check, data flow — creates an immutable audit event. Together they form a **complete, tenant-scoped proof of work** that the operator can inspect, verify, and cite. This is NOT just logging; it is the system's memory and legal foundation.
**What Gets Audited (Everything):**
| Subsystem | What | Event Type | Payload | Chain Link |
|---|---|---|---|---|
| **Plugins** | Load, init, execution, error, disable | `plugin_loaded`, `plugin_executed`, `plugin_disabled` | plugin_id, version, boot_layer, input, output, error, tenant_id | Hash-chained |
| **Skills 2.0 (ACP)** | Execute, config, feedback, optimization | `skill_executed`, `skill_config_updated`, `skill_feedback` | skill_id, version, input, output, latency, lom | Hash-chained |
| **Consent / House-Rules** | Consent given, checked, denied | `consent_granted`, `consent_checked`, `house_rule_denied` | user_id, consent_type, tenant_id, timestamp | Hash-chained |
| **A2A (App-to-App)** | Task received, processed, result sent | `a2a_task_received`, `a2a_task_executed`, `a2a_result_sent` | task_id, source_app, target_app, tenant_id, payload_hash | Hash-chained |
| **Audit System Itself** | Chain verification, key rotation, snapshot | `audit_chain_verified`, `audit_key_rotated`, `audit_snapshot` | chain_height, last_hash, verification_result | Self-verifying |
| **Learning (ADR-0314)** | Feedback received, config optimized | `learning_event_received`, `optimizer_config_updated` | event_type, skill_id, signal, confidence_delta | Hash-chained |
| **Context Engineering** | Snapshot taken, context adapted | `context_snapshot`, `context_adapted` | context_id, tenant_id, user_id, preserved_fields, added_fields | Hash-chained |
| **Data Flow Guard (L34)** | Input classified, flow allowed, blocked | `data_flow_classified`, `data_flow_allowed`, `data_flow_blocked` | data_class, engine, dest, reason | Hash-chained |
**Tenant Isolation (GDPR Art. 5, 6, 32):**
Every audit event is **immutable and tenant-scoped:**
```
{
"tenant_id": "<tenant>", # Fail-closed: null tenant → denied
"timestamp": "2026-09-01T12:34:56.789Z",
"event_type": "skill_executed",
"skill_id": "os.delegation_router",
"input": "classify_request(...)",
"output": "route_to_agent=opus",
"latency_ms": 42,
"lom": "assistant.Forge::route_request:L237", # Line of Moral Responsibility
"hash": "sha256(...)", # Chained to previous event
"prev_hash": "sha256(...)" # Immutable backward reference
}
```
Queries: All reads filtered by `tenant_id`. No cross-tenant leakage.
**Why This Matters (Proof System):**
1. **Operator Proof:** "Show me every Skill decision for task XYZ" → audit trail proves what happened, when, to whom
2. **Compliance Proof:** "Does CorvinOS respect consent?" → audit shows every consent check + decision (accept/deny + reason)
3. **Security Proof:** "Was this A2A task really processed?" → audit shows source, destination, payload hash, result, no tampering
4. **Learning Proof:** "Did the Skill really learn?" → audit shows feedback received, config before/after, optimizer delta
5. **Auditability (GDPR Art. 30):** Every event is attributed, timestamped, signed; no silent operations
**Integration with LDD + Dialektical Reasoning:**
When designing any new subsystem (Plugin, Skill, A2A handler, Consent gate):
| LDD Phase | What | Audit Implication |
|---|---|---|
| **k=1 (Dialectical)** | Surface design choice | "What audit events will this subsystem emit?" *must* be answered before code |
| **k=2 (E2E)** | Prove it works | E2E test must verify: (1) subsystem does X, (2) audit event Y is logged, (3) hash-chain intact |
| **k=3 (Red/Green)** | Iterate | Each iteration updates audit schema if needed (immutable: never delete prior events, only append) |
| **k=4–5 (Refinement)** | Polish | Audit schema is finalized; all future instances follow it |
**Mandatory Dialectical Prompts:**
Before coding any new Skill, Plugin, or A2A handler:
```
/dialectical-reasoning
"New Skill: <name>"
Questions to surface:
1. What audit events will this Skill emit (input, output, config changes)?
2. What tenant scopes will it cross (single tenant, multi-tenant)?
3. Are all events immutable + hash-chained?
4. Can an operator reconstruct the Skill's entire behavior from audit logs?
5. Are there any "silent" operations (optimizations, cleanups) that should be audited?
```
**E2E Wiring Proof (must include audit verification):**
```bash
# 1. Run the Skill / Plugin / A2A handler
pytest tests/test_skill_xyz.py::test_e2e -v
# 2. Verify audit events were logged
grep "skill_executed\|plugin_loaded\|a2a_task_executed" ~/.corvin/audit.jsonl \
| jq -c 'select(.skill_id == "os.xyz" or .plugin_id == "xyz")'
# 3. Verify hash-chain integrity
python3 scripts/verify_audit_chain.py --tenant=_default --since=<start_time>
# Output: ✅ Chain intact (N events, N-1 hash links verified)
# 4. Commit only if audit proof succeeds
git add ... && git commit -m "feat(xyz): description [audit-verified]"
```
**Must NOT do (absolute):**
- **Don't code without Dialectical audit design** — audit schema is not an afterthought
- **Don't emit unattributed events** — every event must have tenant_id, timestamp, lom, event_type, hash
- **Don't merge without hash-chain verification** — E2E tests MUST verify `prev_hash` + `hash` integrity
- **Don't weaken tenant isolation** — audit reads MUST filter by tenant_id; no fallback to "any tenant"
- **Don't alter past audit events** — immutable append-only; never update/delete/rewrite
- **Don't silently optimize** — if a Skill optimizes config (ADR-0314), emit `skill_config_updated` event with delta
- **Don't skip audit for "internal" operations** — Plugin init, context adaptation, data flow guard—all audited
- **Don't assume audit is "just logging"** — it is the system's proof of work and legal foundation (GDPR Art. 30, 32)
**Audit Verification Workflow (for operator + auditor):**
```bash
# Inspect full proof for a task
corvin audit show-task <task_id>
# Output: full chain of events for this task across plugins, skills, consent, data flow
# Verify chain integrity (daily)
corvin audit verify-chain --tenant=_default
# Output: ✅ Chain height 142857, all hashes verified, 0 gaps
# Extract proof for compliance report
corvin audit export --tenant=_default --format=pdf --since=2026-09-01 --until=2026-09-30 \
--events=skill_executed,plugin_loaded,consent_granted,a2a_task_executed
# Output: compliance-report-2026-09.pdf (auditor-ready, signatures included)
# Trace a Skill decision (operator debugging)
corvin audit trace skill os.delegation_router --task=<task_id>
# Output: every event in the chain (config → execute → feedback → optimize)
```
→ Full audit spec: See ADR-0232/0233 (boot tripwire, chain integrity), ADR-0537 (audit event schema + LoM cryptographic binding), RFC 3161 (TSA timestamping)
→ LoM Binding (Gap 2, MEDIUM): ADR-0537 binds each LoM cryptographically to source code via SHA256 (lom_hash field), preventing spoofing attacks
→ Integration: Every new ADR MUST define its audit events (required frontmatter field: `audit_events`); every audit event carrying LoM must include lom_hash for verification
---
## Audit Chain Completeness — 100% Coverage (ADR-2040–2044, Load-Bearing)
**Status:** 🟡 **PARTIAL (verified 2026-09-26)** — the rule below is load-bearing; the
completeness it claims is NOT achieved. Measured: 467 `EVENT_SEVERITY` entries, 268
`_EVENT_ALLOWLIST` entries; 231 events are in `EVENT_SEVERITY` without an allowlist entry
(32 the reverse), so `scripts/verify_audit_event_completeness.py` FAILS. ADR-2040 stays
PROPOSED. The per-subsystem percentages in the table below are self-reported, not measured.
**RULE: ALL Subsystems MUST Audit 100% of Actions. Zero silent operations.**
### Coverage by Subsystem (Phase 1–2 Implementation)
| Subsystem | Layer | Events | Status | ADR | Coverage |
|---|---|---|---|---|---|
| **OS-Kern (Layers 1–44)** | Multi | 70+ | ✅ Complete | ADR-2040 | 100% |
| **Worker Engines (L22/25)** | Compute | 15+ | ✅ Complete | ADR-2041 | 98% |
| **A2A (L38)** | RemoteTrigger | 23+ | ✅ Complete | ADR-2042 | 97% |
| **Plugins (L4)** | Lifecycle | 4 | ✅ Complete | ADR-2043 | 100% |
| **Skills 2.0 (ACP)** | Routing/Learning | 6+ | 🟡 Partial | ADR-2044 | 80% |
**Total Events (measured 2026-09-26):** 467 in EVENT_SEVERITY, 268 in _EVENT_ALLOWLIST — not in sync
### Audit Event Registration (MANDATORY)
**Every new subsystem function MUST register events in TWO places:**
1. **EVENT_SEVERITY** (`corvin_operator/forge/forge/security_events.py:806–831`)
```python
"subsystem.action_name": EventSeverityLevel.INFO, # or WARNING, CRITICAL
```
2. **_EVENT_ALLOWLIST** (`corvin_operator/forge/forge/security_events.py:2719–2778`)
```python
"subsystem.action_name": {
"severity": "INFO",
"mandatory_fields": ["field1", "field2"],
"allow_list": ["field1", "field2", "tenant_id", "timestamp"],
# Never: secrets, PII, user input
}
```
**Failure to register → audit event silently dropped → compliance gap.**
### Phase 1 Critical Events (Registered; 3/8 Wired — verified 2026-09-26)
Same distinction as Phase 2 below: "registered" is not "wired". The earlier text named
call sites that do not exist — the only non-test, non-registry occurrence of the five
NOT WIRED names is the checklist in `scripts/verify_audit_completeness.py`.
**Layer 10 (Context Engineering):**
- `context.snapshot_created`, `context.snapshot_restored` — **NOT WIRED** (`session_checkpoint.py` makes no audit call)
**Layer 22 (Compute Safety):**
- `compute.checkpoint_corrupted`, `compute.deadlock_detected`, `compute.iteration_diverged` — **NOT WIRED** (no emitter)
**Layer 25 (ACS L34):**
- `acs.l34_gate_passed` — **WIRED** in `acs_runtime.py` post-execution
**Layer 36 (Erasure):**
- `erasure.tenant_boundary_checked` — **WIRED** in `erasure_orchestrator.py`
- `erasure.cross_tenant_detected` — **WIRED** in the erasure pre-check
### Phase 2 Extended Events (Registered, 4/8 Wired — verified 2026-09-26)
"Registered" (EVENT_SEVERITY + `_EVENT_ALLOWLIST`) is NOT "wired". The earlier
statuses here named call sites that did not exist (`lifecycle_loader.py:74`
emits `plugin_disabled`, and nothing imports that module).
**Layer 22 (Worker Lifecycle):** 3 events
- `compute.worker_terminated` — **WIRED**: `WorkerServer.stop()` (`core/compute/corvin_compute/worker.py`); `corvin-compute serve` now routes SIGTERM (what `systemctl stop` sends) through that path. Proof: `tests/security/test_wave1c_audit_wiring_e2e.py`
- `compute.worker_spawn_initiated`, `compute.worker_heartbeat` — **NOT WIRED** (no emitter; the compute worker has no spawn/heartbeat concept to hang them on)
**Layer 38 (A2A NBAC):** 3 events
- `a2a.nonce_collision_detected` — **WIRED**: `remote_trigger_receiver.py` replay branch
- `a2a.offline_pair_initiated` — **WIRED, audit-first**: `corvin_a2a.py pair --offline-pair` records before writing the origin and refuses the pair if the record cannot commit
- `a2a.genesis_block_created` — **NO SUBJECT**: no production code creates an NBAC genesis block (`nbac.sign_genesis_block` has no caller). Fenced by `test_genesis_block_has_no_production_creator`; wire it in the same commit that adds a creator
**Layer 4 (Plugins):** 2 events
- `plugin.execution_timeout` — **WIRED**: `registry.py` `on_load` (`LOAD_DEADLINE_S`) and `health_check` (`HEALTH_CHECK_DEADLINE_S`) deadline overruns. Proof: `core/plugins/tests/test_plugin_execution_timeout_audit.py`
- `plugin.initialization_failed` — **NOT WIRED** (`corvin_plugins.audit.emit_initialization_failed` has no caller; load failures are recorded as `plugin.load_failed`)
### Pre-Commit Hook (NOT BUILT — verified 2026-09-26)
The installed `.git/hooks/pre-commit` does not check audit events. The rejection rules
below are the intended design, not current behaviour:
- Add new subsystem changes WITHOUT audit.emit() calls
- Reference unregistered events in EVENT_SEVERITY
- Carry incomplete audit event payloads
### CI/CD Gate (ENFORCEMENT)
**Workflow:** `.github/workflows/audit-completeness.yml` — exists; as of 2026-09-26 it
would FAIL (EVENT_SEVERITY ↔ _EVENT_ALLOWLIST out of sync, see above)
Fails if:
- EVENT_SEVERITY registry incomplete
- _EVENT_ALLOWLIST has PII-risk fields (grep for secrets, emails, tokens)
- Audit tests failing (50+ tests, all must pass)
### Testing Requirements (measured 2026-09-26: `test_phase1_audit_events.py` 13 passed / 1 error; the other two files below do not exist)
**Unit Tests:** Event emission + field validation
**Functional Tests:** Audit chain integrity verification
**Adversarial Tests:** PII leakage detection, tenant isolation, nonce collision, hash-chain verification
```bash
# Run all audit tests (before committing)
pytest tests/security/test_phase1_audit_events.py -v
pytest tests/integration/test_audit_chain_100_e2e.py -v
pytest tests/adversarial/test_audit_integrity.py -v
```
### Must NOT do (Absolute Rules)
- **Don't add subsystem without audit event spec** — every feature ships events or doesn't ship
- **Don't use feature flags to disable audit** — always-on, never skip
- **Don't emit incomplete spans** — all 4 token fields (input/output/cache_read/cache_write) required
- **Don't weaken tenant isolation** — every event must carry tenant_id, fail-closed if missing
- **Don't alter past audit events** — immutable append-only, never update/delete/rewrite
- **Don't leak PII into audit** — all events scrubbed fail-closed, `_assert_safe` backstop
- **Don't bypass pre-commit hook** — use `[skip-adr-check]` flag ONLY for truly exempt changes (test-only, docs-only)
- **Don't assume "audit is just logging"** — it is the system's proof of work + legal foundation (GDPR Art. 30/32, EU AI Act Art. 50)
### Compliance Checklist (GDPR + EU AI Act)
| Standard | Requirement | Verified | ADR |
|----------|---|---|---|
| **GDPR Art. 5** | Accountability (every action audited) | ✅ | ADR-2040 |
| **GDPR Art. 30** | Processing Record (immutable trail) | ✅ | ADR-0232/0233 |
| **GDPR Art. 32** | Security (hash-chained, fail-closed) | ✅ | ADR-0233 |
| **EU AI Act 50** | Transparency (action attribution) | ✅ | ADR-2040–2044 |
### Verified Status (2026-09-26 — replaces the 2026-09-24 "go-live" list, which was not true)
⚠️ Phase 1: 3 of 8 events wired (acs + 2× erasure); 5 registered only
⚠️ Phase 2: 4 of 8 wired (see Phase 2 list)
❌ Registry sync: 231 EVENT_SEVERITY events lack an allowlist entry
❌ Pre-commit audit hook: not built
⚠️ CI gate: exists, would fail
⚠️ Tests: `test_phase1_audit_events.py` 13 passed / 1 error; `test_audit_chain_100_e2e.py`, `test_audit_integrity.py` missing
**Result:** 🟡 **NOT complete. ADR-2040 is PROPOSED.** Don't cite "100% audit completeness".
→ ADRs: ADR-2040 (OS-Kern) · ADR-2041 (Worker) · ADR-2042 (A2A) · ADR-2043 (Plugins) · ADR-2044 (Skills)
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

