agentleFS
Sign inSign up

headroom-desktop

gglucass/headroom-desktop/CLAUDE.md

This file holds only what the model cannot derive: project invariants, environment quirks, and concrete commands. Generic coding-style coaching (write simply, read before editing, state the bug and stop) was removed on 2026-08-03 — modern models do it unprompted and the harness enforces the rest. Do not re-add it.

CLAUDE.md578 starsChanged 24 days ago
  • Installs packages
# CLAUDE.md - headroom-desktop

This file holds only what the model cannot derive: project invariants, environment
quirks, and concrete commands. Generic coding-style coaching (write simply, read
before editing, state the bug and stop) was removed on 2026-08-03 — modern models do
it unprompted and the harness enforces the rest. Do not re-add it.

## Testing Rules
- After any code change, run the relevant tests/checks before declaring the task done. Do not ask the user to verify.
- Rust changes: `cargo test --manifest-path src-tauri/Cargo.toml --lib <filter>` for the affected module, plus `cargo check --manifest-path src-tauri/Cargo.toml` if the change crosses module boundaries.
- Frontend changes: `npx tsc --noEmit` and any relevant Vitest suite. For visual changes, see Styling Rules.
- If a test cannot be run in this environment, say so explicitly rather than skipping silently.

## Wheel Bump Rules
- Before changing `HEADROOM_PINNED_VERSION`: diff upstream `savings_tracker.py`, `prometheus_metrics.py`, and `server.py` between the old and new pins, and check every consumed field against `stats_contract_pins_every_consumed_path` in state.rs. Upstream has silently redefined persisted savings fields before (0.36.0 widened `compression_savings_usd`); the savings-rate canary in state.rs is the runtime tripwire, this diff is the compile-time one.
- Re-pick every platform's wheel URL/sha256 when bumping (see the pin comment in tool_manager.rs).
- Diff `rollout.py`'s FEATURES registry between the old and new pins. The app declares `HEADROOM_ROLLOUT_CHANNEL=beta` (tool_manager.rs), so any new feature with `default_enabled_in` at or below beta auto-enables for every user on the bump; any feature the app requests whose `available_in` moved above beta silently turns off (this is how 0.37.0 disabled the output shaper on stable).
- Diff `ProxyConfig` (`proxy/models.py`) and `cli/proxy.py` defaults between the pins, and grep the handler diffs for new early 4xx/429 returns. A "fix" that starts ENFORCING a config value nobody felt before is a default-behaviour change the FEATURES diff, the vendor checks and the `/stats` diff all miss: 0.39.0's #3350 began consuming a 100k TPM bucket that caps at its own rate, so every request over 100k input tokens got a permanent 429 ("Token rate limited. Retry after Ns") and every long 1M-context session bricked on rc7/rc8. The desktop now passes `--no-rate-limit` (rc9); upstream fix is PR #3806.
- Diff `headroom/transforms/` between the pins. This is the compression ENGINE and nothing else in these rules covers it; the 0.35.0 -> 0.37.0 bump changed it by +421 lines and shipped unexamined. Diff it, but do NOT assume a diff means a behaviour change. Bisected 2026-09-02 with a fixed workload replayed on both wheels installed from PyPI (`uv venv` + `uv pip install headroom-ai==<v>`, about 3s per version): `ContentRouter.compress()` AND the full `ContentRouter.apply()` path returned BYTE-IDENTICAL output on 0.35.0 and 0.37.0 (tok_before 62,384 -> tok_after 57,068, identical transforms list, on both). The engine is EXONERATED. The real production drop (fable-5 500k+: 4.64% / 25,631 tokens per request down to 0.45% / 2,463 at matched model and size) is therefore caused by the INPUT reaching the compressor, not the compressor. Six hypotheses are refuted and must not be restated as fact: the prefix-replay guard (identical on full-replay and floor builds), workload/model mix (matched), the TTL-aware net-cost gate (flag-gated behind `HEADROOM_NET_COST_POLICY=1`, not set, zero `netcost:skip` markers across 925 requests), protection rate (0.35 vs 0.32 per request), router acceptance thresholds (`min_ratio_*` identical), and content granularity (today's messages are BIGGER, 3,053 vs 1,982 tok/msg, which should compress better). What the router logs do show: more messages compressed per request (4.0 -> 5.1) each yielding far less, at an unchanged frozen fraction (86.8% -> 85.2%). Investigate session shape next, not the wheel.
- The pin is `0.39.0` (2026-09-26). That wheel shipped seven vendors the 0.38.0 sitecustomize carried, each deleted, not re-pinned: request-log body window (`RequestLogger.MESSAGE_WINDOW`), Kompress request deadline #3693 (`shares_request_deadline`), feed `?include_messages=0` #3672 (the native feed also carries the per-request uncached/cache_write split `models.rs` reads), rollup cache-read cost #3734 (`_empty_cache_delta`), streaming metering headers #3769, Codex exec JS-literal reads #3737, and the client-side tool-search repair (`strip_unsupported_tool_search_blocks` now has a client-side `tool_result` pass that keeps deferred-but-present refs, so no rc.5 regression). `sitecustomize_drops_vendors_the_wheel_ships` fails if any of these, or the nine 0.38.0 dropped (#3380 prefix floor, #3414, #3480, #3482, #3483, #3484, #3685, #3613, #3106), comes back. Re-pinning one would double-apply the fix: the feed vendor, for one, has no self-neutralization check and would re-register the route.
- Six vendors are still owed upstream and are exact-pin gated to `0.39.0`: transient-system lineage (HEADROOM_TRANSIENT_SYSTEM_LINEAGE; `prefix_tracker.py` is byte-identical to 0.38.0), cache-integrity observer (HEADROOM_CACHE_INTEGRITY, probe `scripts/verify-cache-integrity.py --sitecustomize`), responses shared budget (HEADROOM_RESPONSES_SHARED_BUDGET, probe `scripts/verify-responses-budget.py --sitecustomize`; the native #3693 deadline is per `ContentRouter.apply`, and with the vendor off the Responses unit fan-out still times out with a continuing worker), tool-ref 400 hint (HEADROOM_TOOL_REF_HINT; the wheel still appends no user hint), quarantine spare capacity (HEADROOM_QUARANTINE_SPARE_CAPACITY, probe `scripts/verify-quarantine-spare-capacity.py`; `_run_compression_in_executor` is byte-identical to 0.38.0), and proxied guarded upstreams (HEADROOM_ALLOW_PROXIED_GUARDED_UPSTREAMS, which the desktop sets at launch, probe `scripts/verify-proxied-guarded-upstream.py`; self-neutralizes on `upstream_pinning.proxied_guarded_upstreams_allowed`, see the pinning bullet below, upstream PR #3804). Attribute-gated guards (#2942 context limit, response-cache poisoning, #3170 dollar unfold, #2668 chained reads, #3379 replay guard, #3166 cc-switch reset) bind on 0.39.0 unchanged; the replay guard stays inert. 0.39.0 defaults periodic malloc trim on for linux as well as darwin, so the desktop's `HEADROOM_MALLOC_TRIM=1` is now redundant but harmless.
- 0.39.0 added connect-time upstream pinning (`proxy/upstream_pinning.py`): a caller-supplied `x-headroom-base-url` upstream (Grok and OpenCode routes, `proxy_intercept.rs`) is REFUSED with `UnpinnableUpstreamError` when the backend has `HTTP(S)_PROXY`/`ALL_PROXY` set, because the proxy resolves the target itself. Operator-configured upstreams (Anthropic, OpenAI) are unaffected. `HEADROOM_ALLOWED_BASE_URLS` cannot fix it: it flips the guard to allow-only mode, which would break OpenCode's arbitrary gateways. httpx takes proxies from the macOS/Windows system settings too, so no env var is needed to hit this. The desktop sets `HEADROOM_ALLOW_PROXIED_GUARDED_UPSTREAMS=1`, vendored on 0.39.0: a proxy route forwards on the guard's name verdict (the 0.38.0 behaviour), direct routes stay pinned. Upstream PR #3804.
- On every bump: (0) CI's test-macos step "Vendored patches against the pinned wheel" builds a runtime from the lock plus the NEW pin and fails on any `skipping:`, so a vendor that stops binding fails the bump PR; it proves binding, not the fleet-level effect, and does not replace (1)-(3). (1) re-run the eight `*_behaves_against_the_installed_wheel` tests against the NEW wheel before the runtime is upgraded, by installing it into a scratch venv (`uv venv v && uv pip install -p v/bin/python "headroom-ai[proxy,code,spreadsheet]==<v>" truststore`), symlinking it to `<dir>/headroom/runtime/venv`, and running `HEADROOM_DATA_DIR=<dir> cargo test --lib against_the_installed_wheel`; every one of those tests self-skips when its vendor does not bind, so a green run on the OLD runtime proves nothing. (2) Check each owed vendor against the new wheel's SOURCE, not only its test: a vendor without a self-neutralization check still binds and passes when the wheel already ships the fix. Diff `prefix_tracker.py` `resolve_tracker` (lineage, cache integrity), `server.py` `_run_compression_in_executor` (quarantine), `handlers/openai.py` Responses executors (shared budget, and run its probe with `--disabled`), and grep the Anthropic handlers for a tool-reference hint; drop each vendor the wheel carries, re-pin the rest. (3) Boot the new wheel in the scratch venv with the shipped sitecustomize on PYTHONPATH and diff `/stats` key sets against the old wheel (a 0.38.0 boot removed nothing and added `savings.by_layer.output_shaping` at zero traffic with `method: "inactive"`, which `parse_output_reduction` now treats as no claim).
- #3460 (conversation `clusters` in the output-savings ledger) shipped in 0.38.0 and needs NO ledger purge: the merged accumulator keeps conversation-qualified observations in separate `qn`/`qsum`/`qsumsq` fields, so legacy rows never enter the measured mean and both arms restart together on the bump. `output_savings.rs` mirrors that gate (`MEASURED_MIN_CLUSTERS`), so the holdout boost (`output_holdout_for`) stays at 10% until both arms hold five conversations; without the mirror the Rust estimator would have promoted to 3% on legacy data the wheel itself refuses to score.
- #3699 (cache-aware counterfactual, savings schema v6) shipped in 0.38.0. `/stats` `cost.compression_savings_usd` (what the desktop's daily buckets delta) is still list-priced via `merge_cost_stats`; only `savings.breakdown.compression_savings_usd`, `cost.cache_aware_savings_usd` and the persisted `lifetime.compression_savings_usd` (read only as the pre-rollup fallback in `parse_savings_breakdown`) moved to the cache-aware basis. `savings_basis` and `compression_savings_list_usd` sit beside them. The savings-basis canary is token-based and untouched.
- 0.38.0 moved `proxy_output_shaper` to `available_in=STABLE` (still not default-enabled, so `HEADROOM_OUTPUT_SHAPER=1` is what turns it on); nothing else moved in the FEATURES registry, and `HEADROOM_ROLLOUT_CHANNEL=beta` stays for read_maturation. `HEADROOM_BEACON=off` still silences the v2 beacon (#3253 widened the payload, not the gate). Locks: anyio topped up to 4.14.2 (#3662 CVEs); the litellm floor moved to 1.96.2 for Python 3.14 wheels only and the pins stay at 1.90.1/1.88.1 (see the lock lineage note).
- 0.38.0 (#3204) writes the wheel runtime log as `~/.headroom/logs/proxy-<port>.log` and stops writing `proxy.log`; an upgraded machine keeps the stale 0.37.0 `proxy.log` forever. Anything scanning that log (the Kompress status markers in `newest_wheel_proxy_log`) must pick the newest `proxy*.log`, never the fixed name. Read protection for shell reads is `HEADROOM_PROTECT_READS`, switched on by the `coding` profile the desktop runs, not by an env the desktop sets; #3621 extends it to Codex `exec` custom_tool_call reads, so expect Codex tool-output compression to drop a little on this bump while re-reads drop with it.

## Compression / Cache Change Rules
The 0.9.4 prefix-replay regression cost every upgraded user ~17pp of their input savings rate for ~18 hours across 89 installs. All three rules below are things that, done, would have caught it before release.
- **Verify both sides of a trade.** Any change that buys one metric with another (the prefix-replay guard buys provider cache hits by spending compression) is only verified when BOTH are measured. The 0.9.4 sign-off measured cache reads and busts and never looked at `tok_saved`, so it reported "better on every metric" while compression had collapsed. Minimum evidence: `cache_read/forwarded` AND `tok_saved/tok_before`, on requests in the size band the change affects.
- **A ratio is not a measurement.** `cache_read/forwarded` falling can mean the cache broke or the conversations grew. Check the ABSOLUTE per-request figure before blaming either. A flat `cache_read` per request against a growing denominator is workload; a `cache_read` that stops scaling with conversation size is a bug (that plateau was the 0.9.4 fingerprint).
- **Soak before promoting.** Anything touching the compression or cache path waits one FULL day on staging, then `bin/rails savings:did` in headroom-web is the promotion gate (`DAY=<first full day> VERSION_PREFIX=<version>`). Release-day runs prove nothing: the day bucket is ~24h, so a same-day release is at most ~22% of it and dilutes a real effect ~4x. 0.9.4 went rc.1 to stable in six hours, which made detection structurally impossible.
- Do NOT add a client-side savings-rate canary that compares a machine against its own history. It has no control group, so a user switching models or growing conversations trips it; that was tried and produced 12 false events across 9 hosts (Sentry RUST-89/8C). The fleet DiD has a control arm, which is the entire reason it works.
- Do NOT reimplement upstream #3380 against the pinned wheel. The 0.9.4-rc.4 failure was a splice reimplementation (replay to the confirmed floor, then stitch on this turn's fresh output) that shipped every turn's beyond-floor pipeline drift and lost 22% of fleet cache coverage (1.20 -> 0.94 reads/sent, n=17, p=0.007). The sanctioned form is the sitecustomize prefix-floor VENDOR (0.9.6): the PR's overlay_cached_prefix + finalize_turn exec'd verbatim, exact-pin gated to wheel 0.37.0, pre-clamp floor bridged via prepare_turn's tracker_frozen, full-replay fallback for floorless callers, kill switch HEADROOM_PR3380_VENDOR=0, functionally tested against the installed wheel. Anything in between - partial hunks, rewritten logic, a floor derived at the overlay call site - is banned. The vendor was dropped in 0.9.20-rc.1 because the 0.38.0 wheel ships #3380 natively; if a future wheel loses it, vendor it the same way again.

## Release Cadence / Update Loudness
- Stable releases batch weekly; out-of-band stable releases are for regressions and security only (user feedback 2026-09-03: daily update prompts read as churn). Daily rc's on the beta channel are fine.
- A stable can be built without being offered to anyone: put `[publish: download-only]` in the release commit (or dispatch release-macos.yml with publish=download-only). The release stays a prerelease, `releases/latest` keeps the previous stable, and only new downloads reach it via `HEADROOM_DOWNLOAD_TAG=vX.Y.Z` on the headroom-web Railway `web` service. Promote with `gh release edit vX.Y.Z --prerelease=false --latest` and clear the variable. New installs have no pre-period, so this is NOT a substitute for the staging soak + savings:did gate on compression changes.
- Updates are quiet by default: no dialog/notification, and on macOS they silently download+install with only a passive "Restart to update" affordance. To make a release loud (old interrupting flow, for must-take updates), put `<!-- headroom:loud -->` anywhere in `.github/release-notes/<VERSION>.md` - the marker flows through latest.json `notes` and is stripped from the displayed notes.

## Persistence Rules
Most stability bugs in this codebase's history were violations of one of these five. Follow them for any new code; treat violations found in existing code as bugs.
- Anything persisted uses `client_adapters::atomic_write` (tmp+rename), never plain `fs::write`. Crash mid-write must not truncate state, and a rewrite must keep the file's mode and write THROUGH a symlink (a hand-rolled rename replaced dotfiles-managed `~/.zprofile` and `~/.claude.json` links with regular files, silently). CI's `scripts/check-direct-writes.py` fails any other `fs::write`/`fs::rename`/`File::create` in non-test code unless it carries `// direct-write: <reason>`; only Headroom's own files qualify.
- Anything versioned/deserialized carries `#[serde(default)]` (container-level where possible). One added required field must not wipe a user's history. On parse/schema failure: back the file up and log, never silently overwrite; salvage format-agnostic fields where possible.
- Anything appended (logs, JSONL) has a size cap or rotation from day one.
- Never kill a pid resolved from a port without verifying its identity (argv/process name) first.
- Day/hour bucket keys must state their timezone. User-facing "days" are local (`local_day_key`); if a source is UTC-bucketed (backend rollups), key it by its UTC date and say so - never relabel one as the other.

## Formatting
- No em dashes, smart quotes, or decorative Unicode. Plain hyphens and straight quotes, so output stays copy-paste safe. Accented letters and CJK are fine when the content needs them.

## Styling Rules
- Never hardcode colors in component CSS. Use the semantic tokens defined in `:root` in `src/styles.css` (`--surface-*`, `--text-*`, `--border-*`, `--fill-*`, `--accent*`, `--warning*`, `--danger*`, `--chip-*`).
- If a needed color does not exist as a token, add it to both the `:root` block and the `@media (prefers-color-scheme: dark)` override — do not inline a hex/rgba in a component rule.
- Exceptions: pure `#fff`/`#000`, brand gradients, and launcher/splash-only one-offs that are intentionally theme-invariant. Comment the exception inline.
- When adding or modifying a component, visually verify both light and dark mode before declaring done. If dark mode cannot be tested, say so explicitly.
- Run `npm run check:colors` (or `./scripts/check-colors.sh`) on any CSS you touch. It flags raw hex/rgba in component rules. Migrate any new offenders to tokens before committing. Existing offenders are the Stage 4 migration backlog — don't add to them.

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.