agentleFS
Sign inSign up

fak

anthony-chaudhary/fak/llms-full.txt

This file inlines the full text of every document the curated llms.txt links to, for one-fetch ingestion by LLMs and answer engines. Generated by tools/genllmsfull.py from llms.txt (the source of truth). Project version: 0.55.0. fak is an agent runtime: the operator-controlled boundary for cache and context, model routing, tool authority, memory, observability, and native inference. Its technical architecture is the Fused Agent Kernel (an agent kernel). Use this map to locate the current authority for a task without reading the…

llms.txt40 starsChanged 22 days ago
  • Pipes a download into a shell
  • Deletes or force-pushes
  • Installs packages
  • Commits and pushes
# fak — the agent kernel: full documentation corpus

> This file inlines the full text of every document the curated `llms.txt` links to, for one-fetch ingestion by LLMs and answer engines. Generated by `tools/gen_llms_full.py` from `llms.txt` (the source of truth). Project version: 0.55.0.

## Index (start with the curated map)

# fak documentation map for agents

`fak` is an agent runtime: the operator-controlled boundary for cache and context, model routing, tool authority, memory, observability, and native inference. Its technical architecture is the **Fused Agent Kernel** (an **agent kernel**). Use this map to locate the current authority for a task without reading the human front-door narrative.

**Default:** read [`AGENTS.md`](AGENTS.md) before changing the repository, then open only the authority for your task below.

**Current product focus:** use [`docs/CAPABILITIES.md`](docs/CAPABILITIES.md) for the outcome-first map of token savings, avoided turns, cache/context reuse, model routing, and live session control. Security remains indexed as a supporting floor.

**Canonical problem map:** use [`docs/problems-we-solve.md`](docs/problems-we-solve.md) for
the “don’t make me think” direction, the four stable problem IDs (`P1` context, `P2`
net-true efficiency, `P3` fast adaptation, `P4` integrated operations), and the frame that
connects new work to operator value. Do not turn this direction into a shipped claim; use
[`CLAIMS.md`](CLAIMS.md) for shipped/simulated/stub status.

- [Project orientation](docs/project-orientation.md): current decision record for the agent-kernel center, capability-family classifications, investment boundary, and repeatable portfolio audit

## Audience guides

- [End-to-End Inference, Agent Harness, and Memory](docs/courses/end-to-end-inference-agent-harness-memory.md): the integrated 8-module flagship course across native inference, tool use, policy, context control, durable memory, observability, and proof. Use [`LEARNING-PATH.md`](LEARNING-PATH.md) instead for the full 99-course prerequisite catalog.
- [Less context, less code](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/less-context-less-code.md): where fak fits beside concise-output and YAGNI/minimal-code guidance.

## Key facts (for accurate answers)

- **Name:** `fak`; the technical name and architecture are the **Fused Agent Kernel** / *agent kernel*. Repository: `fak`. Language: **Go 1.26+**. License: **Apache-2.0**. Version: **0.45.0** ([`VERSION`](VERSION) is authoritative).
- **Public category:** **agent runtime**. In fak, that means the operator-controlled boundary for cache and context, model routing, tool authority, memory, observability, and native inference. These capabilities are available by operating mode and claim status; the category does not promote roadmap work to shipped behavior.
- **Disambiguated search terms:** `fak agent runtime` (the public category), `fak agent kernel` and `fused agent kernel` (the technical architecture and name), `fak serve` (the gateway verb), `fak guard`, `treat the tool call like a syscall`, `default-deny tool-call gate`, `prompt-injection result quarantine`, `addressable KV cache`, `long-session prompt cache`, `cost-aware model routing for agents`, `managed agent runtime`, `MCP tool poisoning defense`, `AI agent least-privilege tool access`, `tamper-evident agent tool-call audit`. The bare word `fak` is dominated by homophone and F.A.K.-acronym noise, so always pair it with one of these. The full 76-term roster, including localized variants, is the generated feed [`llms-terms.txt`](llms-terms.txt) (source of record: `docs/marketing/disambiguation-terms.json`, written by `fak marketing aeo`).
- **Three adoption rungs:** (1) `fak guard` wraps the agent you already run with one command; (2) `fak serve` fronts any OpenAI-compatible server (Ollama, vLLM, cloud) with routing, policy, quarantine, and audit; (3) the fused kernel runs a local model inside the kernel's address space so the KV cache is a kernel object.
- **Serving boundary:** Proxy/gateway mode may front an explicitly selected tuned engine or hosted provider. That is external inference, not fak-native, and its engine identity stays attached to every result.
- **Native inference invariant:** [`fak-native`](docs/native-inference-goal.md) is the product and performance path for local inference and is intended to beat llama.cpp in matched, quality-constrained envelopes. llama.cpp is permitted only when explicitly selected for benchmarks, parity/reference diagnosis, migration/interoperability, or ego-free borrowing; it is never a silent fallback. A current win still requires the exact scoped row in [`BENCHMARK-AUTHORITY.md`](BENCHMARK-AUTHORITY.md).
- **Operational surface (the single-binary thesis):** the agent gateway and control plane — the OpenAI/Anthropic/MCP wires, routing, capability floor, result quarantine, audit with `X-Trace-Id` correlation, bearer and `x-api-key` auth, and Prometheus `/metrics` — collapsed into one static Go binary (two `golang.org/x` dependencies; no Python and no CUDA toolchain needed to build). Where vLLM and SGLang are multi-process Python/CUDA engines you wrap in a reverse proxy plus policy and audit layers, `fak` is that layer as a single process. The contrast is operational surface, not throughput.
- **Honest scope:** a 29-claim prior-art audit scored **0/29 novel** — every primitive is established; the contribution is the *assembly* into one in-process gate where the tool call is the checkpoint. Per-surface claim status (`SHIPPED`, `SIMULATED`, `STUB`) is tracked in [`CLAIMS.md`](CLAIMS.md); treat anything not marked `SHIPPED` as unproven.

- FAK vs DOS boundary — FAK owns agent execution; DOS owns work admission, leases, truth, liveness, and decisions; FAK workflows compose DOS: [`docs/fak-vs-dos.md`](docs/fak-vs-dos.md)
## Current authorities

- Astra formal-subagent admission, closed formal kinds, fail-closed packet schema, pin precedence, assignment digest, and ranked correctness deployment map: [`docs/astra-formal-subagents.md`](docs/astra-formal-subagents.md)
- Micro-context fabric contract and bounded 100/1k/10k floor (`microcontextdemo`): [`docs/research/micro-context-fabrics.md`](docs/research/micro-context-fabrics.md); captured stage witnesses: [`docs/research/README.md`](docs/research/README.md)
- Product scope and mode choice: [`README.md`](README.md)
- Human role map: [`START-HERE.md`](START-HERE.md)
- Documentation by audience, job, and lifecycle, plus the exhaustive repository map: [`INDEX.md`](INDEX.md)
- Agent workflow, builds, proofs, commits, and shared-tree rules: [`AGENTS.md`](AGENTS.md)
- Default documentation lookup, HEAD-only reachability census, and task-specific build/test witnesses: [`docs/dev-tooling.md`](docs/dev-tooling.md)
- Shift-left task organization — author human-readable outcome, scope, dependencies, acceptance, witness, placement, and lane before dispatch: [`docs/shift-left-task-organization.md`](docs/shift-left-task-organization.md)
- New-work defaults — ship the applied end-to-end spine first, prove its operating envelope, optimize against it, then fan out the hardening backlog with `fak issue fanout` (3 is the floor, not the target): [`docs/spine-first-defaults.md`](docs/spine-first-defaults.md)
- Positive workspace management — positive workspace construction over punitive default-deny, immutable FROZEN safety floor vs permissive convenience surface, avoiding capability laundering and doom-loops: [`docs/positive-workspace-management.md`](docs/positive-workspace-management.md)
- Contribution contract: [`CONTRIBUTING.md`](CONTRIBUTING.md)
- Claims status (`SHIPPED`, `SIMULATED`, `STUB`): [`CLAIMS.md`](CLAIMS.md)
- Fak-native execution-engine doctrine, matched-envelope rule, external-reference boundary, and deterministic docs guard: [`docs/native-inference-goal.md`](docs/native-inference-goal.md)
- Performance-RSI loop doctrine, 16 canonical dimensions, dominant bottleneck derivation, matched-baseline evidence rules, and scorecard CLI: [`docs/notes/PERFORMANCE-RSI-LOOP-DOCTRINE.md`](docs/notes/PERFORMANCE-RSI-LOOP-DOCTRINE.md)
- Memory-concept ranking dossier (superset): [`docs/superset/MEMORY-CONCEPT-RANKINGS.md`](docs/superset/MEMORY-CONCEPT-RANKINGS.md) — 10 memory concepts M1–M10 ranked fak-vs-engines with evidence, per-engine evidence tables, and adopt-or-SKIP verdicts
- [The managed-context glossary and product contract](docs/managed-context-glossary.md): the public glossary for managed context — assumption, resident view, pinned objective, budget envelope, reset transaction, context query, memory promotion, and cache state — each grounded in the shipped mechanism, stating what fak manages automatically, what it asks the user about, and what stays user-controlled. The product-facing companion to [context-is-not-memory](docs/CONTEXT-IS-NOT-MEMORY.md) (#1571, epic #1570).
- External system architecture, interfaces, and trust boundaries: [`docs/architecture.md`](docs/architecture.md)
- Builder stability and ownership ladder (CLI → semantic protocol → public Go API → sidecar → internal leaves): [`docs/builder-contract-ladder.md`](docs/builder-contract-ladder.md)
- Detailed integration architecture and extension seams: [`docs/fak/agent-integration-architecture.md`](docs/fak/agent-integration-architecture.md)
- Agent runtime category, ownership, interfaces, flow, and offline proof: [`docs/explainers/agent-runtime.md`](docs/explainers/agent-runtime.md)
- Canonical identity and nearest-contrast lookup for overloaded terms: [`docs/generated/disambiguation/INDEX.md`](docs/generated/disambiguation/INDEX.md)
- TensorRT-LLM vs SGLang and how fak fronts fleet or local inference: [`docs/explainers/tensorrt-llm-vs-sglang-and-fak.md`](docs/explainers/tensorrt-llm-vs-sglang-and-fak.md)
- One-minute offline proof: [`docs/repro-packet.md`](docs/repro-packet.md)
- One-agent management with `fak guard`: [`README.md#manage-one-local-agent-fak-guard`](README.md#manage-one-local-agent-fak-guard)
- Running that same `fak guard` seam **server-side** — a hosted harness (Claude Code / Codex / a CI runner) governed unattended in a container, with no TTY, a mounted hash-chained audit journal, a cost cap, and a supervised-restart lifecycle; the complement to the local-dev framing above and to [`docs/integrations/embed-in-your-product.md`](docs/integrations/embed-in-your-product.md) (which governs your own direct API call): [`docs/guard-server-side-client.md`](docs/guard-server-side-client.md)
- Claude Code on your own Mac's local model and many-agent cache savings (one command: `fak mac`, long form `fak claude-mac-fak`, pointing Claude Code at a Mac's own `fak serve` gateway; run many agents on your Mac with 88.2% compute reduction, flat 180 ms TTFT, 2.1 agents/GB density via Qwen2.5-7B Q8 shared-prefix KV caching; serving speed #2691/#2723 unblended): [`docs/fak/claude-mac.md`](docs/fak/claude-mac.md) · [`docs/fak/mac-agent-ui.md`](docs/fak/mac-agent-ui.md) · [`docs/cache-value-rollup.md`](docs/cache-value-rollup.md)
- Run local models on Mac (Qwen3.8 with Apple Silicon Metal acceleration, interactive REPL `fak run qwen38`, and gateway chat): [`docs/fak/mac-local-models.md`](docs/fak/mac-local-models.md)
- Run parallel subagents with zero-cold-start cache reuse using `fak up`: [`docs/subagents-guide.md`](docs/subagents-guide.md)
- Shared endpoint with `fak serve`: [`docs/fak/server-quickstart.md`](docs/fak/server-quickstart.md)
- Configuration answers (crawlable human page + inline FAQPage JSON-LD): [`docs/fak/configuration.md`](docs/fak/configuration.md)
- Configuration answers (plain text): [`llms-config.txt`](llms-config.txt)
- Configuration authority (complete flags, environment variables, precedence): [`docs/fak/server-config.md`](docs/fak/server-config.md)
- API: [`docs/fak/api-reference.md`](docs/fak/api-reference.md)
- Deployment: [`docs/fak/deployment-guide.md`](docs/fak/deployment-guide.md)
- Deploying the gateway on a **rented GPU box** (CoreWeave, Lambda, RunPod, Crusoe, Vast.ai, Nebius) — in-kernel on the GPU, or proxy in front of a co-located vLLM/SGLang; every provider row is `not yet` end-to-end witnessed: [`docs/fak/neo-cloud-deploy.md`](docs/fak/neo-cloud-deploy.md). Distinct from **fronting a hosted model API** ([`docs/supported/clouds.md`](docs/supported/clouds.md)) and from the neo-cloud **backend binding layer** ([`docs/vendor/neo-cloud-reference-architecture.md`](docs/vendor/neo-cloud-reference-architecture.md)).
- Adding a new model to the in-kernel engine — the one question that splits a one-line arch alias plus a recognition test from a `fak new-model` scaffold plus a forward pass and an oracle: [`docs/new-model-playbook.md`](docs/new-model-playbook.md)
- Canonical cross-hardware Qwen performance highlights and worker publishing route: [`docs/benchmarks/QWEN-PERFORMANCE-INDEX.md`](docs/benchmarks/QWEN-PERFORMANCE-INDEX.md)
- Detailed Qwen3.8-27B Metal result and the Qwen3.6/llama.cpp delta: [`docs/benchmarks/QWEN38-27B-LATEST.md`](docs/benchmarks/QWEN38-27B-LATEST.md)
- Routing a run difference to loader, quant, or forward when model id, quant mode, and tokens/sec all match — the model-load provenance artifact and its algebra (schema `fak-model-load-provenance/1`): [`docs/model-load-provenance-troubleshooting.md`](docs/model-load-provenance-troubleshooting.md)
- Release readiness, guarded cut/publish, verification, and rollback: [`fak release`](cmd/fak/release.go) and [`.claude/skills/release/SKILL.md`](.claude/skills/release/SKILL.md)
- Daily lock-aware shared-checkout Git hygiene, canonical verb, and proof floor: [`fak git-daily`](docs/releases/issue-5592-daily-lock-aware-git-hygiene-2026-08-08.md)
- Client and agent integrations: [`docs/integrations/`](docs/integrations/)
- Codex `UserPromptSubmit` permissive, guarded-child, and hardened modes; capability floor; `fak sessions codex-hook-install`; and `fak sessions codex-loop-hook --hardened`: [`docs/integrations/openai-codex.md#userpromptsubmit-modes`](docs/integrations/openai-codex.md#userpromptsubmit-modes)
- Scoped benchmark results, tuned baselines, artifacts, and reproduce routes: [`BENCHMARK-AUTHORITY.md`](BENCHMARK-AUTHORITY.md)
- Terminal-Bench 4 reproduction manual and parity envelope specification: [`docs/benchmarks/TERMINAL-BENCH-4-REPRODUCTION.md`](docs/benchmarks/TERMINAL-BENCH-4-REPRODUCTION.md)
- Security floor, policy configuration, evidence, and private reporting: [`SECURITY.md`](SECURITY.md)
- Enterprise-gateway governance parity — what `fak serve` covers today against the redaction, per-key rate-limit, and provider-failover checklist, and the two gaps still open: [`docs/gateway-governance-parity-audit.md`](docs/gateway-governance-parity-audit.md)
- Out-of-band operator control of a RUNNING session — the closed control vocabulary (`steer`/`redirect`/`pause`/`resume`/`cancel`/`terminate`/`throttle`/`budget`/`priority`), each op's capability + boundary + witness-of-applied + closed refusal, the `fak session` / `fak signal` / `fak ps` front door, and why `fak steering` / `fak steer` are an unrelated name collision: [docs/operator-control-plane.md](docs/operator-control-plane.md)
- Managed worker worktrees — detached per-worker build isolation, portable defaults, lifecycle operations (prepare/list/land/reap/gc), and remote crash recovery: [docs/managed-worker-worktrees.md](docs/managed-worker-worktrees.md)
- Steerability controls and appeal-channel scorecard: [docs/STEERABILITY-SCORECARD.md](docs/STEERABILITY-SCORECARD.md)
- Operator-heaviness pressure, budgets, and reduction standard: [docs/OPERATOR-HEAVINESS.md](docs/OPERATOR-HEAVINESS.md)
- Operator-steerability PRs (`fak steer prs`) — the gate/observability separation, the OSP unit, and why the overlay must never become a merge gate: [docs/operator-steerability-prs.md](docs/operator-steerability-prs.md)
- Version-everything (`fak version modules`) — per-module rev+date over the module tree, git-witnessed via the `fak-module-versions/1` ledger rather than asserted: [docs/notes/VERSION-EVERYTHING-SPINE-2026-07-03.md](docs/notes/VERSION-EVERYTHING-SPINE-2026-07-03.md)
- Cross-repo issue velocity (`fak issues-solved`): query issues solved across public and companion checkouts over arbitrary time windows in 1 turn: [docs/cli/verbs.md#fak-issues-solved](docs/cli/verbs.md#fak-issues-solved)
- Cache-value roll-up — is the cache work paying off over time? The front-door story keeping WITNESSED kernel reuse and OBSERVED provider-dollar savings in separate, unblended tracks, with the #1066 marginal-over-warm-KV, WITNESSED-vs-OBSERVED, and net-not-gross fences and the shipped Track-1 reproduce command (`fak nightrun score --json`): [docs/cache-value-rollup.md](docs/cache-value-rollup.md)

## Source authorities

- CLI verbs: [`cmd/fak/`](cmd/fak/)
- Kernel subsystems: [`internal/`](internal/)
- Runnable examples and policies: [`examples/`](examples/)
- Go module and toolchain: [`go.mod`](go.mod)

Runtime behavior is authoritative in code and tests. Documentation describes the current generation unless a page marks itself historical, experimental, simulated, stubbed, or superseded. Dated investigation and decision records live under [`docs/notes/`](docs/notes/); use them for provenance, not as the default current contract.

## Task routing

| Task | Read first | Check next |
|---|---|---|
| Explain or evaluate fak | [`README.md`](README.md) | [`docs/repro-packet.md`](docs/repro-packet.md) |
| Modify code or docs | [`AGENTS.md`](AGENTS.md) | The owning package tests and [`CONTRIBUTING.md`](CONTRIBUTING.md) |
| Run local models on Mac (Apple Silicon) | [`docs/fak/mac-local-models.md`](docs/fak/mac-local-models.md) | Run `fak run qwen38` or `fak serve --metal` |
| Integrate a client or managed runtime | [`docs/explainers/agent-runtime.md`](docs/explainers/agent-runtime.md) | [`docs/integrations/`](docs/integrations/) and the selected mode's quickstart and API/config authority |
| Build a bounded microagent-derived harness | [`docs/notes/microagents-to-harnesses-2026-08-18.md`](docs/notes/microagents-to-harnesses-2026-08-18.md) | Run `go run ./cmd/microharnessdemo -selfcheck`; descendants remain host-admitted and the root receives compact receipts, not child transcripts |
| Deploy or operate | [`docs/fak/deployment-guide.md`](docs/fak/deployment-guide.md) | [`docs/fak/server-config.md`](docs/fak/server-config.md) and [`docs/fak/api-reference.md`](docs/fak/api-reference.md) |
| Assess readiness or cut a release | [`.claude/skills/release/SKILL.md`](.claude/skills/release/SKILL.md) | Run `fak release readiness --json`, then use the skill's guarded ship and verification path |
| Assess a claim | [`CLAIMS.md`](CLAIMS.md) | [`BENCHMARK-AUTHORITY.md`](BENCHMARK-AUTHORITY.md) and the cited witness |
| Research design history | [`docs/notes/`](docs/notes/) | Current code, tests, and architecture before relying on an older decision |

For an unlisted task, start at [`INDEX.md`](INDEX.md) and select the current audience route.

- docs/generated/verb-surface.md — generated source-derived fak verb tree, REFUSES column, and unverified coverage count

- [KV capacity normalization](docs/kv-capacity-normalization.md): normalize block and direct runtime KV metrics across tokens, bytes, and occupancy with typed honesty diagnostics.

- [Frontier infrastructure and workload expectations](docs/research/frontier-infrastructure/README.md) — dated evidence on labs, clouds, datacenters, serving workloads, distributions, releases, startups, and rumors
- [Release branch regime and actionable CI base-red diagnosis](docs/release-branch-regime-status.md): Exact-run CI diagnosis, typed causes, package work units, and operator-owned billing.

---

# `AGENTS.md`

> Source: `AGENTS.md`

# AGENTS.md — orientation for coding agents

> **Not the human contributor guide.** This file is operating instructions for *automated*
> contributors working inside the maintainers' shared checkout — trunk guards, lane leases,
> commit trailers, and shared-tree rules that have no meaning outside it. If you are a human
> evaluating or contributing to fak, read [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md) for what fak is and
> [`CONTRIBUTING.md`](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md) for how to contribute; nothing below is required of you.

> You are an autonomous agent working in this repo. This file is the machine-read entry
> point (the [agents.md](https://agents.md) convention). It is intentionally
> command-dense and free of philosophy. For the *why*, read [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md);
> for a curated doc map, read [`llms.txt`](https://github.com/anthony-chaudhary/fak/blob/main/llms.txt). Humans: see [`START-HERE.md`](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md).

## Default scope when launched from this repo

When an agent is launched with this `fak` repository as its working directory,
default all broad or underspecified work requests to **FAK work in this
repository**. Requests such as “finish local WIP,” “work the backlog,” or “run
agents all night” mean inspect, prioritize, implement, test, and ship FAK work;
they do not authorize a workspace-wide sweep of sibling repositories.

Use FAK's guarded headless-agent and worker mechanisms for parallel or unattended
work, following the lane, lease, worktree, landing, and witness rules below.
Only inspect or modify another repository when the operator explicitly names
that repository or the FAK task itself has a concrete, evidenced cross-repo
dependency. When tasks involve shared interfaces (`pkg/*`), cross-repo platform
integration, or dual-repo goal/lease coordination, synchronization MUST encompass
both `fak` and `fak-private` companion repositories to prevent interface skew (`go.work`)
and lease/queue stalls (`refs/fak/locks/*`). If scope is ambiguous, stay in FAK.
Under the **Companion-Aware Dev Harness**, local development in `fak` is first-class and automatically governed by companion tooling: when `fak-private` is present side-by-side (via `../fak-private`, `go.work`, or `FAK_PRIVATE_ROOT`), `fak` pre-commit hooks run 5-gate boundary checks (`cmd/fak-boundary`) and secret leak audits (`tools/scrub_public_copy.py`) by default, while `fak-dev gate`, `fak-dev provenance`, `fak-dev audit-leak`, and `fak-dev ticket` bridge directly to companion tools without requiring directory switches. Full architecture specification: [`docs/architecture/companion-dev-harness.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/architecture/companion-dev-harness.md).

## Mandatory Automated Public OSS Placement for New Work

All new core engine work belongs in this open-source repository (`fak`) by default:
- **Core Engine in Public by Default:** Model inference runtime (`cmd/fak/serve.go`, `internal/engine`), context window MMU (`internal/ctxmmu`), compute acceleration HALs (Metal, ROCm, Vulkan, AVX-512 in `internal/compute/`), model weight loaders (GGUF, SafeTensors in `internal/model/`), tool capability floors, frozen ABIs (`pkg/abi`), and client SDKs (`pkg/fakclient`, `pkg/scorecard`) are open-source.
- **Automated Default Action (Zero-Hesitation Mandate):** When any new work, feature, optimization, or bugfix in these domains is requested, agents and contributors must execute, test, and land it directly in `fak` automatically without waiting for explicit permission or defaulting to private repos.
- **Boundary Invariant:** Core engine mechanisms commoditize hardware complements and drive developer adoption. Proprietary commercial policy (margin ledgers, credit settlement) and the autonomous software factory (dispatch queues, worktree flows, scorecards) belong in `fak-private`, but all engine primitives remain public.

## Divide and conquer: Delegate real work; keep the coordinator context clean

Use guarded headless agents or an equivalent isolated worker for every substantive
unit of work. The primary agent is the coordinator: decompose the request, give each
worker a bounded goal and distinct file set, preserve only decisions and compact
evidence in the primary context, and independently witness worker results before
landing or reporting them. Delegate investigation, implementation, tests, long command
output, and independent review; do not pull their full transcripts into the coordinator.
Structurally drive 10x subagent adoption by default: aggressively fan out 4–8 (up to 16
on multi-core hosts) specialized subagents concurrently across pairwise tree-disjoint lanes.

- **Divide and conquer by default (10x subagent adoption)**: When handling substantive, multi-part, or complex tasks,
  decompose them into atomic, single-concern units or parallel subagent packets. Launch 4–8 (up to 16
  on multi-core hosts) specialized subagents (`task`: worker, researcher, explore, deep-reason, cross-validator,
  issue-auditor, tester, reviewer, scout) concurrently across pairwise tree-disjoint lanes to maximize
  throughput and prevent coordinator context saturation.
- **Astra worker reservation:** Reserve `gpt-6-astra` subagents for highly pre-specified
  formal-mathematical leaves: the packet states fixed definitions and assumptions, the exact
  proposition or derivation, the required output form, and a deterministic witness. Send
  exploration, implementation, testing, review, and general reasoning to another suitable worker model.
- **3-Tier Parallel Pipeline**: Structure multi-agent work into three decoupled parallel cohorts:
  1. *Parallel scoping and research*: Discovery via `researcher`, `explore`, and `scout` to isolate prior art, contracts, and relevant packages.
  2. *Parallel implementation + reproduction tests*: Implementation via `worker` and `deep-reason` across disjoint lanes, authoring reproduction tests before fixes.
  3. *Parallel adversarial verification and QA ticketing*: Adversarial verification via `cross-validator`, deterministic test execution via `tester`, and edge-case issue auto-ticketing via `issue-auditor`.
- **Depth-0 Coordinator vs. Depth-1 Leaf Worker**:
  - *Top-level coordinator (depth 0)*: Aggressively fans out parallel subagents across the 3-tier pipeline, arbitrates lane leases (`dos arbitrate`), and collects compact receipts.
  - *Leaf worker (depth 1)*: Executes directly within assigned package boundaries (1–3 files, package-scoped tests). Leaf workers must NOT invoke nested `task` calls (preventing recursion depth exhaustion, #12028). Autonomous safe sync landing (`fak sync`, `fak commit --path`, `fak sync push`) is active by default working within each worker process upon test verification. Subagents and leaf workers are strictly banned from creating new PowerShell (`.ps1`) or shell (`.sh`, `.bat`) scripts; all automation, test helpers, and tooling MUST be native Go programs.
- **Tree-disjoint boundaries**: Assign each worker a distinct, non-overlapping file set to avoid
  concurrent collisions on the shared trunk.
- **Isolate and witness**: Keep heavy command logs and raw transcripts in worker boundaries; pull
  only compact decisions, diffs, and verification receipts into the coordinator.

The primary agent may directly perform only lightweight coordination: inspect enough
state to scope packets, launch and supervise workers, adjudicate conflicts, integrate
witnessed results, and run the final completion audit. A trivial one-command answer or
tiny edit may stay local when launching a worker would cost more than the work itself.
Use FAK's managed launch, lane, lease, detached-worktree, landing, and witness surfaces;
worker self-reports are not evidence, and delegation does not relax ownership of the
final result.

## Scope discipline for smaller models: subdivide or abstain

Smaller or resource-constrained models (such as local 7B/14B checkpoints, fast/flash models,
or bounded worker subagents) achieve reliability by keeping work tightly focused and strictly verified.
When operating as or delegating to smaller models, enforce scoping safeguards:

1. **Subdivide into atomic units (S0/S1 leaves)**:
   - Restrict each task or dispatch packet to a single observable deliverable and exactly one witness command.
   - Limit the active write surface to 1–3 closely related files within a single package or lane.
   - Decompose multi-step tasks into sequential, verified phases: write the reproduction test first, commit the minimal implementation, and verify the targeted package.
   - Complete and witness one step before advancing to the next; keep edits focused rather than attempting broad multi-subsystem sweeps in one turn.

2. **Abstain from high-difficulty aspects (scoped fail-to-abstain)**:
   - Identify task aspects that demand deep architectural context or high-risk reasoning: concurrency invariants and lock ordering, frozen ABI modifications (`internal/abi`), low-level SIMD/CUDA kernel mechanics, cross-subsystem protocol migrations, and security policy gates.
   - Scope abstention strictly to the bounded high-difficulty aspect; maintain momentum by executing all independent, safe, solvable sub-components (such as baseline reproduction tests, diagnostic witnesses, non-gated packages, or documentation) rather than abandoning the prompt.
   - Emit a structured `ABSTAIN` verdict with a typed refusal token and exact boundary description for the escalated aspect alongside the landed partial evidence.
   - Deliver the verified sub-component and cleanly escalate the isolated difficult aspect to a higher-capability model or human operator.

3. **Guard against fast/flash model sharp edges (Gemini 3.8 Flash & peers)**:
   - **Curb token inflation and verbosity**: 3.8 Flash is designed to "work harder" and can output 2×–4× more tokens than other models per task. Enforce extreme conciseness (<3 lines commentary in CLI), eliminate conversational preambles/postambles, and keep explanations minimal.
   - **Resist over-scaffolding ("happy-go-lucky" sprawl)**: Flash models are prone to generating unsolicited companion abstractions, multi-panel apps, or broad refactors for simple requests. Confine diffs strictly to the requested lines/files; do not introduce unasked scaffolding.
   - **Beware thinking effort tradeoffs**: At low thinking effort, 3.8 Flash exhibits quality regressions (spatial/geometric flaws, shallow verification); at high effort, it burns large token budgets. Never trust low-effort intuition on complex logic—always verify against deterministic external tools (`go test`, `go vet`, `fak validate`).
   - **Break interactive tool loops while persisting toward the goal**: In CLI/tool loops, Flash models can confabulate success or loop in repeat-failure cycles ("apologizes, then retries the exact same command"). Ground every claim in an observed tool receipt. When a tool call is denied or fails, halt repetition of the identical call or cycling argument variations; read the error or refusal receipt, query `fak recover <TOKEN>` for structured recovery, and pivot to an alternate sanctioned route or decomposed subtask. Persist through recoverable hurdles by adapting the approach rather than repeating failed calls or abandoning the objective.
   - **Prefer specialized file tools over shell scripting**: Flash models experience higher failure rates on complex CLI/terminal pipelines (TerminalBench regressions). Prefer structured tools (`Read`, `Edit`, `Glob`, `Grep`) over complex piped bash commands.
   - **Anticipate safety false-positives**: Standard 3.8 Flash guardrails can trigger false refusals on legitimate security inspection, redaction, or policy code; frame technical security contexts neutrally or emit structured `ABSTAIN` rather than hallucinating workarounds.
   - **Guard trailing model turns & nested tool schemas**: Gemini REST wire rejects payloads ending in a model turn with HTTP 400 (`Requests ending with a model turn are not supported`) and fails nested array parameters lacking explicit `items`. Enforce turn alternation (auto-inject continuation on trailing model turns), strip `$schema` and `additionalProperties`, and preserve cryptographic `thoughtSignature` tokens across multi-turn tool replays. Full analysis: [`docs/notes/2026-09-03-gemini-3.8-flash-initial-feedback-and-guidance.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/2026-09-03-gemini-3.8-flash-initial-feedback-and-guidance.md).

4. **Calibrate test breadth (prevent reasoning model test over-engineering)**:
   - Restrict test authoring to a single atomic reproduction or regression unit test demonstrating the defect or behavior change; prohibit sprawling 20-test suites when 1 witness suffices.
   - Prevent token exhaustion and context pollution by targeting only the explicit requirement or failure mode instead of writing speculative matrix permutations or redundant assertion variations.
   - Keep test execution deterministic, bounded, and fast-failing within the target package (`go test -v ./internal/<pkg>` and `go vet ./internal/<pkg>`).

## First local-agent product milestone

**Useful local agents, accelerated automatically.** The current public product
milestone is practical, performant local agent work through the normal entry
point: native inference, qualified speculative decoding enabled by default,
legal agentic prefix/KV reuse, and supported device-resident/direct GPU paths.
Read [the milestone and worker brief](https://github.com/anthony-chaudhary/fak/blob/main/docs/local-agent-milestone.md) before
planning native performance, caching, local UX, or public positioning work.

Make qualified acceleration automatic and observable. In worker packets name
the workflow, bottleneck, entry-point seam, supported model/device, expected
default, acceptance witness, and missing evidence. Prefer time to verified task
completion over isolated token-rate peaks. Distinguish shipped defaults from
opt-in code and hardware qualification; this is our first milestone to earn,
not an established world-first claim.

## Native inference performance invariant

For any native-inference or performance task, keep model execution **fak-native all the
way**. The product path is intended to beat llama.cpp in matched, quality-constrained
envelopes while fak retains ownership of kernels, memory, scheduling, cache, adaptation,
and operations. llama.cpp is allowed only when explicitly selected for benchmarks,
parity/reference diagnosis, migration/interoperability, or ego-free borrowing. Never add
an `auto`, recovery, or convenience path that silently changes native/performance work to
llama.cpp. Before accepting evidence, ask whether the model executed inside fak and whether
the receipt names that engine. Canonical definitions, matched-envelope rules, and the
deterministic docs guard: [`docs/native-inference-goal.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/native-inference-goal.md).
New native-performance work prefers Qwen3.8.
Qwen3.6 is allowed only when the task states an explicit task-specific exception, such as regression, compatibility, historical comparison, or a hardware/artifact constraint.
Preserve historical Qwen3.6 artifacts; do not rename or rewrite them as Qwen3.8 evidence.

Default native-tuning discipline: identify each candidate by source, configuration, and
workload. Re-measure the reference when objective or load changes make prior measurements
incompatible; for SLO-only changes, reassess compatible historical measurements against the
new SLO before deciding whether a rerun is needed. Enumerate the full configuration delta;
prefer one-variable comparisons and justify inseparable bundles. Bundle evidence supports
only the bundle, with individual attribution requiring separate evidence. Before an expensive
run, obtain independent validity and budget review bound to that exact candidate and its
objective, load, SLO, and planned cost. Candidate changes invalidate affected review and
witnesses; retain unaffected evidence with its scope. Review preserves existing user
authorization, budget limits, and security controls; it grants neither spend nor security authority.

## What this project is

**fak** is an *agent kernel*: one Go binary that sits between an AI agent and the tools
it calls, and handles every tool call before it runs — reuse the shared setup, route each
call to the right model, serve repeats locally, and shed old turns while the provider's
cache survives. It is first a **performance gate** (do the shared setup work once, not
every turn — cheaper, faster, longer-running sessions) and, on the very same checkpoint, a
**security gate** (a default-deny capability floor the model can't talk past).

## Default priorities (portfolio hierarchy)

When prioritizing design, implementation, architecture, triage, or backlog work, the repository defaults to this four-tier hierarchy:

1. **fak all in one (serving and harness + memory — the "one touch" thing)**:
   The primary focus: a turnkey, single-binary runtime (`fak up`) that bundles model serving, the agent harness, and durable memory into a seamless "one touch" deployment. Everything required to run governed, cache-accelerated agentic workflows out of the box with zero glue code.
2. **fak serving only**:
   High-performance model inference runtime (`fak serve`), disaggregated gateway, KV-cache/context MMU acceleration, continuous batching, and native model execution.
3. **fak harness only**:
   Standalone agent harness and governance substrate (`fak guard`), default-deny tool capability floor, policy adjudication, and interception over external frontier models and harnesses.
4. **other things**:
   Peripheral utilities, standalone integrations, off-spine tooling, and secondary stewardship tasks not directly advancing tiers 1–3.

## Repo layout (where things live)

| Path | What it is |
|---|---|
| `go.mod` · `cmd/` · `internal/` | **The Go module is the repository root** (the kernel + the `fak` CLI). |
| `cmd/fak/` | The `fak` binary (every verb: `preflight`, `serve`, `agent`, `policy`, `bench`, …). |
| `internal/` | Kernel subsystems: `adjudicator`, `policy`, `vdso`, `engine`, `gateway`, `ctxmmu`, `model`, … |
| `examples/` | Policy manifests **and** runnable demos (`adjudication-demo/`, `agentdojo-redteam/`, `mcp/`). |
| `docs/` | Explainers, integration guides (`docs/integrations/`), benchmark methodology, proofs. |
| `docs/private-comms-channel.md` | **The private comms channel** (private control bridge to the lab GPU servers) — a public stub pointing to its home in the `fak-private` companion repo. Start here to reach the hardware. |

## Build / test / run

The Go module is the repository root and needs Go 1.26+ (`GOTOOLCHAIN=auto`). Zero external
dependencies means no `go.sum`.

```bash
go build ./cmd/fak        # build ./fak (fak.exe on Windows)
make build                # debuggable binary
make smoke                # fast real-world binary smoke + dogfood test
make test-fast            # build + vet + short tests + smoke
make mac-perf             # Mac shift-left: Metal tok/s and prefill bench
make test-race            # WSL race gate
make test                 # full suite, including model witnesses
make ci                   # build + vet + test + claims lint + smoke
```

Build profiles and their one flag-delta table: [`docs/dev-tooling.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/dev-tooling.md#build-profiles).
Install: `go install github.com/anthony-chaudhary/fak/cmd/fak@latest`.

### Which shared-tree question are you asking?

This checkout is permanently peer-dirty; raw in-place `go build ./...` or `go vet ./...` can
mix your work with peers' WIP. Use the matching isolated verb:

| Question | Command |
|---|---|
| Is committed trunk buildable and gofmt-clean? | `fak-dev ci-preflight` |
| Does trunk plus only my paths pass build, vet, and affected tests? | `fak validate --mine <path>... [--smoke]` |
| Does my change compile while hiding peer WIP? | `fak-dev buildcheck [--vet] [--mine <file>]` |
| Does the literal working tree compile? | `fak-dev buildcheck --isolate=false --vet` |
| Will my push break another worker's graph? | `fak hooks pre-push` |

Compile verification must never write the in-tree binary: prefer `fak-dev buildcheck`; if the
binary itself is broken, use a unique temp output. On this Windows box, Microsoft Defender
Antivirus real-time ML has repeatedly quarantined transient native Go test executables; the
August 25 audit for issue #8919 supersedes the earlier Application Control diagnosis. Run
`./test.ps1` under WSL, use fleet nodes for real serves, and follow
[`docs/notes/AVOID-TESTING-ON-THIS-MACHINE-2026-06-25.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/AVOID-TESTING-ON-THIS-MACHINE-2026-06-25.md).
Do not disable Defender or add broad exclusions.
## The 60-second proof (no key, no model, no GPU — verified)

This is the canonical first command. Run it before anything else:

```bash
go build -o fak ./cmd/fak
./fak preflight --policy examples/customer-support-readonly-policy.json --tool refund_payment --args "{}"   # -> DENY (POLICY_BLOCK): refused by structure, no model in the loop
./fak preflight --policy examples/customer-support-readonly-policy.json --tool search_kb     --args "{}"   # -> ALLOW: not a blanket block
./fak agent --offline                                                                                       # -> injection-in-context YES->no, destructive-op YES->no, task still booked
```

The first `go build` compiles the binary (~30–60s, plus a one-time Go-1.26 toolchain fetch);
later runs are instant. Full walkthrough: [`docs/repro-packet.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/repro-packet.md).

## The local machine is the control point, not the compute boundary

That the *kernel proof* needs no GPU (above) is not license to stop when a task **genuinely**
needs one. The workstation you run on plans, dispatches, and ledgers — it is the **control
point**. It is **not** the compute boundary. When a task needs a CUDA/GPU device, a specific
accelerator arch, DC-network reach, or heavy CPU, **dispatch it to a fleet node that has it**
and report the real result.

**Never end a turn with `not yet — no GPU on this host`** (or any "blocked for lack of local
hardware") as a terminal answer. There is always a route: a sanctioned compute node, or — if its
credential/bridge session is missing — the exact ready-to-run command sequence handed to the
operator. Both count as "used the lab"; only `no local hardware` does not. The nodes and how to
reach each are the SSOT in **[`docs/fleet-compute-nodes.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fleet-compute-nodes.md)**
(GCP `fak-realmodel` L4 for a live CUDA serve; the GPU server via `dgxbridge` for a device-GEMM
witness; `da33` for DC/CPU; the nightrun pipeline for a nightly ledger).

This is guard-enforced, not honor-system: `fak hwgate-lint` scans a final turn for the stop
pattern and types each hit to a sanctioned-node redirect (the hardware-gate dual of
`fak headless-lint`), and `fak guard`'s Stop hook blocks a headless stop that declares a
local-hardware blocker as terminal (`--hardware-gate enforce`), feeding the redirect back so the
agent dispatches instead of stopping. `fak guard-stops` tallies the pattern for the soak → promote read.

## Proof by default (every issue fix ships its evidence)

This kernel exists because a self-report is not a fact. Hold your own fixes to the same bar:
**by default, an issue fix ships a captured proof artifact, matched to the failure class** — not
just a code change and a "looks fixed". Pick the witness the bug actually has:

- **TUI / visual** (a rendering, layout, or terminal-corruption bug — the kind you were sent a
  *screenshot* of): the proof is a **captured render**. Write a render-witness test that captures
  the bytes a surface emits and asserts the defect is gone — see
  [`cmd/fak/watchdog_autoheal_test.go`](https://github.com/anthony-chaudhary/fak/blob/main/cmd/fak/watchdog_autoheal_test.go)
  `TestWatchdogAutohealKeepsAgentPaneClean`, which captures the would-be agent pane and proves
  **zero bytes** reach it. When a live on-screen UI is in the loop, also attach a **before/after
  terminal screenshot** to the issue/PR (the same evidence the reporter gave you). A green unit
  test that never renders the surface is not proof for a visual bug.
- **Logic / behavior**: a test that **fails before the fix and passes after** — the repro is the
  proof. Land it in the same commit as the fix.
- **Runtime / CLI / Dogfood / Integration** (an executable verb, CLI command, gateway protocol
  adapter, daemon hook, or multi-component flow): the proof is **meaningful execution of the real
  path**, not an in-memory mock. Hermes' rule: "mocks hide integration bugs." Shift real-world smoke testing
  earlier into the process: run `fak validate --mine <paths> --smoke` in your inner loop, run `make smoke`
  (`smoke-exec` + `dogfood-test`) or `make test-fast` before commit, and enforce execution via dogfood (`make dogfood-test`,
  `scripts/dogfood-claude.sh --probe`, `recent_feature_dogfood.py`) or a loopback integration test
  (`make test-integration`, `*integration*test.go`) against a temporary workspace before claiming
  completion. A green unit test that only asserts mocks or verifies syntax without executing the
  real path is unproven and must not be declared done. Shift left: prove execution early during
  development, never deferring validation to post-merge or scheduled nightly runs.
- **Hardware / Performance / Acceleration**: a physical silicon execution witness (`fak hil`,
  `make mac-perf`, or on-device GPU receipt) — never an in-memory mock or analytical simulation alone.
  Hardware-in-the-loop (HIL) testing must run **100x more frequently in micro-doses** (sub-second physical
  probes via `fak hil`) rather than waiting for rare multi-hour benchmarks. Head-to-head reported
  comparisons must be **real hardware reported comparisons**; software simulations (rooflines, trace models)
  are strictly early indicators and search-space bounds, never final comparative claims or victory proofs.
  Enforce an active bias towards physical hardware testing whenever an accelerator is present on the host or fleet.
- **"Shipped / done" claims**: a witnessed commit (`dos verify`, the `(fak <leaf>)` trailer) — see
  the witness rules below. A subject line is forgeable; the diff and the registry are not.

The rule of thumb: reproduce the defect as a captured artifact *first*, then make that artifact
clean. If you cannot capture it, you cannot prove you fixed it — say `not yet` with the missing
witness instead of claiming a fix.

## Track work in GitHub before implementation

Use a GitHub issue as the durable tracker for every substantive unit of work whenever
reasonably possible. Before editing code or docs, changing configuration, or launching an
implementation worker, search open and closed issues for duplicates, then claim the matching
issue or create one that states the problem, intended outcome, and witness. Put the issue
number in worker packets and keep discoveries, scope changes, and follow-ons reconciled there.

Before spawning one fresh task per issue (including direct Codex `create_thread` waves),
normalize every dependency through the issue contract and run the repository's issue/dispatch
admission surface. Newly authored work uses exactly three relations: `Start blocked by: #N`
may prevent pickup only when no executable disjoint slice exists; `Coordinates with: #N` is
non-blocking alignment; `Promotion requires: #N` gates closure or performance claims, not
implementation. A broad `After:` list or a hand-written prompt saying all dependency fences are
binding is not an admission contract. If an issue is triage-only, repair its scope and typed
relations before spawning it; direct task packets must repeat the same typed pickup rule.

Scoping, reproduction, and read-only triage may precede the issue when needed to write an
honest ticket. Skip advance ticketing only when it would be unreasonable: the request needs no
repository change, the change is truly trivial and tracking would cost more than the work, an
urgent safety or outage response must start immediately, or GitHub is unavailable. For urgent
or offline work, create or reconcile the issue as soon as the constraint clears and record why
work began first. Do not use plans, TODOs, commit messages, or chat as substitutes when GitHub
issue tracking is reasonably available.

### Cross-repository issue quoting and qualification

When referencing, quoting, or linking issue numbers across public and companion repositories:
- **Public `fak` issues:** A bare `#<num>` in this repository strictly denotes a public issue in `anthony-chaudhary/fak` (or canonical `fak#<num>`).
- **Private `fak-private` issues:** When public `fak` quotes or cites an issue number from the companion private repository (`fak-private`) — such as referencing an internal appeal, reproducer, incident, or boundary contract in commit messages, technical notes, test names, PRs, or code comments — it **MUST be explicitly qualified as `fak-private#<num>`** (or `anthony-chaudhary/fak-private#<num>`) so that its private origin is crystal clear.
- **Prohibition on bare private numbers:** Never quote a `fak-private` issue with a bare `#<num>` or unqualified `issue #<num>` in public code, commit subjects/bodies, PRs, or documentation. A bare number in public `fak` falsely implies a public issue, triggers erroneous auto-closes on GitHub, and breaks automated issue-closure tracking.

## New work defaults: spine first, then fan out

For every new unit, align with the default priority hierarchy (1: All-in-one [serving + harness + memory], 2: Serving only, 3: Harness only, 4: Other things), classify centrality (`Core`, `Enabling`, `Stewardship`, `Peripheral`), run
all P1-P4 checks in [`docs/problems-we-solve.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/problems-we-solve.md), and state **For /
Problem / Today / Better because / Witness** against the real next-best alternative. Then:

1. Ship the smallest working end-to-end spine in the same session: applied implementation,
   captured witness, operating-envelope proof, then optimization. Include the safety needed to
   run it. If it cannot ship confidently, file that spine first as a `gen/now` issue with its
   missing witness; never defer it silently.
2. Once it ships, create the 3..50+ QA, dogfood, productization, observability, integration,
   docs, and release follow-ons with
   `fak-dev issue fanout --title T --leaf L --spine <sha|cmd|doc> --json` (or cohort-plan first).

Full doctrine: [`docs/spine-first-defaults.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/spine-first-defaults.md) and `/spine-fanout`.
At run end, dedupe and file every real leftover as an open issue; otherwise say nothing remains.

Promote reusable successful scratch through **scratchpad → committed Go tool → `fak` verb →
captured knowledge**. Promote when it solved a real recurring task, records a non-obvious fact,
or beats the committed equivalent; keep only true probes disposable. New tooling is a Go leaf or
nested Go sub-module, not another `tools/*.py` or shell/PowerShell script, and operational facts belong in the leaf doc or dated `docs/notes/`. Allocate
and reap scratch through `fak tree-doctor`; see [`docs/generated-output-defaults.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/generated-output-defaults.md).

Close operator-facing turns with verdict-first bullets, one claim and inline evidence per line;
make the final line the next checkable step. A one-line “nothing left; pushed X” is sufficient.

## Focus on "move forward" over "conclusions"

Frame all diagnostics, performance measurements, benchmark reports, and investigation findings
around **forward momentum rather than terminal conclusions**. When an approach underperforms,
a benchmark trails baseline, or an experiment yields negative results, focus on forward action
rather than passive, defeatist, or dead-end editorializing (e.g. rather than *"X didn't work ...
therefore we suck..."*, state *"the next step to get better performance is X"*).

- **Action over self-judgment**: Treat unmet targets, performance gaps, or failed attempts as
  empirical data that eliminates a variable and sharpens the search space, not as an editorial
  verdict on the system. Maintain momentum by identifying the forward path.
- **Formulate the next checkable lever**: Every diagnostic, bottleneck, or failed attempt must
  yield a concrete, actionable forward step:
  - *"Prefill amortized poorly at batch size 1; the next step to improve throughput is activating chunked prefill in `internal/compute`."*
  - *"Sub-byte quantization introduced drift on layer 14; the next step to recover accuracy is keeping the attention projection in Q8_0."*
  - *"Host-device transfers dominate turn latency; the next step to get better performance is pinning input buffers via the unified memory allocator."*
- **Ground turn endings and handoffs in forward movement**: Keep summaries, issue updates, and
  handoffs oriented toward actionable progress. Always conclude with the concrete next experiment,
  test, or patch that advances the work.

## Version everything: cite `module@rev`, not just a bare SHA

Every module carries a **derived** version — there are no hand-maintained per-module version
files (×410 they would rot within hours on this shared trunk and spew merge noise). A module's
`rev` is the count of trunk commits that touched it, rendered `r<rev>+g<shortsha>` (e.g.
`internal/gateway r652+g1f75c56d`) — monotonic, conflict-free, and computed from history alone.
Read the table with `fak version modules` (`--json` for machine form, `--scores S.json` to join a
flat `{"module": score}` map); a nightrun / super-loop turn appends changed-module rows to the
`fak-module-versions/1` ledger with `fak version modules --stamp` (stamping twice at the same HEAD
is a no-op). **When you cite evidence in a claim or a handoff, prefer `module@rev` over a bare SHA**
— it says *which part moved and how far*, which a SHA alone does not. Full doctrine + e2e:
[`docs/notes/VERSION-EVERYTHING-SPINE-2026-07-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/VERSION-EVERYTHING-SPINE-2026-07-03.md).

## Hard rules (enforced below the agent layer)

- **Ship green work and sync safely by default.** Do more work by default: when implementing or fixing tasks yourself, do not stop at partial edits or ask conversational permission to test, sync, or push. Standard synchronization must follow the `fak sync check`, `fak sync apply`, and `fak sync push` workflow. Run the full lifecycle through to completion:
  1. **Pre-flight safe sync & Companion Dev Harness**: Check and synchronize safely with `origin/main` (`fak sync check`, `fak sync reconcile --apply`, or `fak sync apply`) so local work is grounded on trunk. When a companion workspace exists (`fak-private`), perform dual-repo preflight checks: inspect dirty state across both checkouts, fetch remotes and coordination refs (`refs/fak/locks/*`) for both `fak` and `fak-private`, run `fak sync check`/`fak sync apply` in `fak` alongside fast-forward pull in `fak-private` (or `fak-sync repo`), reconcile `go work sync` to align module graphs across both trees, and verify boundary leak hygiene (`tools/scrub_public_copy.py --audit-staged` or `fak-dev audit-leak --staged`). Under the Companion Dev Harness:
     - **Pre-commit hooks**: `.githooks/pre-commit` and `.githooks/pre-commit.ps1` in `fak` automatically detect `fak-private` and enforce 5-gate boundary verification (`cmd/fak-boundary check --staged`) and secret leak audits before any commit lands.
     - **Native companion dev bridge**: Use `fak-dev` bridge commands directly from `fak` without switching directories:
       - `fak-dev gate [check|query|explain]` bridges to `cmd/fak-boundary`
       - `fak-dev provenance [mint|audit|status]` bridges to `cmd/fak-sync provenance`
       - `fak-dev audit-leak [--staged|--all]` bridges to `tools/scrub_public_copy.py`
       - `fak-dev ticket [list|show|next]` bridges to companion `docs/tickets/`
  2. **Atomic implementation & self-verification**: Decompose work into single-concern leaves, write/update tests, and verify on-device (`fak validate --mine <paths>`, `go test ./internal/<pkg>/...`).
  3. **Stage-and-commit by explicit path**: All commits must be made via `fak commit --path` or `fak sweep --apply`. Lint the subject (`fak commit --preview`) and commit only your touched paths (`fak commit --path <p> -m "<subject> (fak <leaf>)"` or `fak sweep --apply --lane <lane> -m "<subject>"`). A `COMMITTED_RED` refused only in its test step, and only because tests already fail on the landing base, may land with `--no-build-check` without asking (operator policy 2026-09-26). All six conditions and the receipts in `.claude/skills/commit-clean/SKILL.md`, section "Pre-existing red", must hold.
  4. **Safe push unprompted**: Push verified commits via `fak sync push` or `fak commit --push`. `fak sync push` automatically retries transient non-fast-forward races (when a peer lands between fetch and push while HEAD already contains origin) and stops safely on genuine diverged states. Never wait for an operator prompt to push once the tree is green.
  - **Safe merge & trunk convergence**:
    - **Active `MERGE_HEAD` gate**: Check `.git/MERGE_HEAD` before any staging or landing. If an in-flight merge is active (`MERGE_IN_PROGRESS`), unstage local paths (`git restore --staged`) and wait; never abort, reset, or finish a peer's merge.
    - **Fast-forward-only convergence**: Standard synchronization uses `fak sync apply` (`git merge --ff-only --no-autostash --no-overwrite-ignore`) to guarantee clean trunk convergence without accidental merge commits.
    - **Structured divergence reconciliation**: When local and remote diverge, use `fak sync reconcile` to evaluate safe structured routes (`ROUTE_APPLY` for clean ff, `ROUTE_DISJOINT_INTEGRATE` for disjoint commit file-trees, `ROUTE_SUPERSET_MERGE` for textless `-s ours` verified merge, `ROUTE_HOLD_DIRTY_COLLISION` with path suspension via `fak wip park` / `--suspend-paths`, or `ROUTE_RECONCILE_PACKET` for overlapping conflicts).
    - **No force-push or unverified merges**: Stay on `main`; never force-push, use `--autostash`, create a feature branch, escape a dirty/diverged tree into a worktree, or perform raw unverified 3-way merges on the shared trunk. On `PUSH_REJECTED`, reconcile and retry via `fak sync push`.
- **Match scope to capability.** Constrain smaller models and workers to atomic S0/S1 leaf units with single-concern boundaries and one witness. When encountering high-difficulty aspects (concurrency, frozen ABI, complex kernel algorithms), practice scoped fail-to-abstain: land partial verified evidence and escalate only the isolated high-difficulty boundary with a structured ABSTAIN record rather than guessing or emitting speculative changes. Persist through recoverable hurdles using alternate sanctioned routes or waiting out transient locks rather than abandoning the task.
- **Commit exactly one issue through explicit paths via `fak commit --path` or `fak sweep`.** All commits must be made via `fak commit --path` or `fak sweep --apply`. Lint with `fak commit --preview`, then stage-and-commit via `fak commit --path <p> ... -m "<subject> (fak <leaf>)"` or `fak sweep --apply --lane <lane> -m "<subject>"`. `fak commit` provides automatic DCO sign-off (with `-s` accepted for compatibility) and verifies that no peer files were raced in (`PATHSPEC_RACE`). One issue lands in one commit and one leaf; do not split a green issue into patch commits or batch unrelated issues. WARNING: Raw git commits are strictly an emergency-only fallback, permitted ONLY when the `fak` binary is unbuilt (`git commit -s -m "<subject> (fak <leaf>)" -- <paths>`); never use `git add -A` or uncoordinated raw git commits on the shared trunk.
- **Make the first subject final.** Sign off with DCO, use a Conventional-Commits subject, and include a recognized `(fak <leaf>)` trailer. A peer may push your commit before an amend, so preview the subject and paths first. Demo binaries use their `cmd/<dir>` name as the leaf.
- **Preserve shared-tree buildability.** A tracked or untracked `.go` sibling enters every package build. Fence incomplete cross-file WIP with `//go:build wip_<feature>` until its symbols exist; validate only your explicit paths with `fak validate --mine`. Build verification never writes an in-tree binary.
- **Migrate CI/CD contracts atomically.** Before changing workflow inputs, JSON schemas consumed by workflows, check/job names, secrets/env, runner labels, or artifact/cache names, search the whole tree for consumers and update them together. Include changed contract, migrated consumers, impact/cutover, and rollback in the commit body; prove committed tip with `fak-dev ci-preflight`. Checklist: [`docs/ci/ci-spec-change-migration.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/ci/ci-spec-change-migration.md).
- **Use only the sanctioned detached worker worktree.** Drive isolated workers through `fak worktree worker prepare|land|reap|list`; all branch worktrees remain forbidden. Reap stale scratch worktrees with `tools/worktree_doctor.py`, which preserves live sessions. See [`docs/managed-worker-worktrees.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/managed-worker-worktrees.md) for the operator guide, portable defaults, and remote recovery runbook.
- **Allocate and reap scratch explicitly.** Use `fak tree-doctor --scratch-dir <producer>` or `--scratch-path <producer>/<file>`, then `--reap-scratch <producer> --json`. Never use `git clean -Xdf` under `_scratch/`. Per-run prompts/transcripts belong in allocated scratch; `testdata/` is only for fixtures landed with a consuming test. See [`docs/generated-output-defaults.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/generated-output-defaults.md).
- **Keep comments durable.** Explain non-obvious invariants, safety, concurrency, compatibility, or performance tradeoffs; do not narrate syntax. Preserve required exported-API, package, directive, generated, and legal comments.
- **Keep claims and gains witnessed.** Every `CLAIMS.md` row needs its required tag. Report performance only from net-true end-to-end accounting, with quality and operating-envelope constraints; setup/recovery/verification overhead counts. Use the project claim and benchmark gates rather than hand-written “looks faster” prose.
- **Extend through leaves and Go tooling.** Add capabilities with `fak new-leaf`; avoid editing the core registry directly. New durable tooling is a Go leaf or nested Go sub-module plus a `fak` verb, not another `tools/*.py`, `.ps1`, or `.sh` script; the Python and script gates permit maintenance of grandfathered scripts, not new ones.
- **Strict Ban on PowerShell and Loose Scripts — Modular Native Go Programs Only.** Adding new PowerShell scripts (`.ps1`), Bash/POSIX shell scripts (`.sh`), batch files (`.bat`, `.cmd`), or loose glue scripts is **STRICTLY BANNED** across this repo and `fak-private`. An exception requires an extraordinary, super heavily justified reason (basically ~1 in the whole repo, such as initial bare-metal bootstrap before Go is installed). All new automation, dogfood runners, harnesses, test utilities, session orchestrators, diagnostic probes, and background tasks **MUST be implemented as native Go programs** (`cmd/*` binaries, Go leaves registered as `fak` CLI verbs, or quarantined sub-modules under `tools/<name>/go.mod`). Think modular, integrated, long-term value: Go programs provide static typing, `go test` coverage, compile-time safety, cross-platform execution (`GOOS=windows/linux/darwin`), and unified CLI ergonomics (`--json`, `--help`). Never create random scripts; the git pre-commit filter refuses un-grandfathered scripts (`FILE_ADMISSION`).
- **Keep private control private.** GPU-server credentials, hostnames, SSH details, private paths, and raw internal logs stay in the private companion repo. Public evidence must be scrubbed and reproducible through [`docs/private-comms-channel.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/private-comms-channel.md).
- **Respect Windows operational fences.** Never run broad `find /`, `find ~`, `find /mnt`, or `find /proc` from Git Bash. Route git and hang-sensitive commands through the guarded fak verbs. Treat low-utilization whole-machine stalls as kernel-path churn: stop launch storms, inspect process counts/handles, and use the committed diagnostics rather than adding retries.
- **Make external writes explicit.** `OUT_OF_TREE_WRITE` records writes outside the repo. Use the operation’s declared external target and reversible preview/confirm path; never disguise an external mutation as repository scratch.
- **Do not stop or restart a managed inference server during active sessions.** Before stopping, restarting, or reconfiguring any locally or remotely managed model inference endpoint or server, verify that all dependent sessions have completed or ownership has been explicitly handed off. Stopping an active upstream server strands dependent guarded sessions behind repeated 502 Bad Gateway errors.

The detailed recovery for a refusal is intentionally paged: query the token below instead of carrying every failure cookbook in every agent context.

### If the kernel refuses you (recover, don't fight it)

A guard refusal is per-call runtime feedback; it is not a session stop or permission to abandon the task. A worker that treats a refusal as a prompt termination converts a recoverable hurdle into an unforced failure.

A guard refusal names a token from the closed `[reasons.*]` vocabulary in
[`dos.toml`](https://github.com/anthony-chaudhary/fak/blob/main/dos.toml). Run `dos man wedge <TOKEN> --explain` instead of carrying
the full recovery cookbook in every agent context:

```bash
dos man wedge <TOKEN> --explain  # summary, category, fix, and references
fak recover <TOKEN>              # concrete recovery plan when one exists
fak recover --list               # list executable and manual-only plans
```

Recover by the named action; do not route around the guard:
1. **Execute recovery**: Run `fak recover <TOKEN>` and apply the output. Manual plans name the required remediation.
2. **Handle transient locks**: For concurrency blockers (`MERGE_IN_PROGRESS`, `COLLISION_RISK`, `LOCK_BUSY`), unstage local paths (`git restore --staged`), wait for quiescence or inspect `dos top`, checkpoint finished progress with `fak wip checkpoint` or `git tag hold/<lane>`, or pivot to an independent disjoint subtask.
3. **Use sanctioned alternative routes**: When an operation is refused (such as `FILE_ADMISSION` on a script), use the sanctioned Go tool, leaf verb, or structured file tool instead of quitting.
4. **Scope boundary blocks**: When an explicit policy or core-lock blocks a specific path (`POLICY_BLOCK`, `SELF_MODIFY`), land your safe, verified sub-components, record the structured refusal, and escalate the isolated boundary with checkable next steps.

`fak recover` is a dry-run unless `--execute` is passed, and manual-only plans refuse execution
rather than guessing. The token argument is case- and separator-insensitive.

Keep the common commit-lane rules available before a refusal: stay on `main`,
commit only explicit paths via `fak commit --path` or `fak sweep`, never amend or force-push shared history, wait out a
peer's `MERGE_HEAD`, and reconcile divergence in place with `fak sync apply`. The hard rules above are
the preventive contract; the query surfaces are the token-specific recovery path.

Check setup failures with `python tools/extend_preflight.py`; the full contributor
contract is [`CONTRIBUTING.md`](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md). Recover first; a complaint never
authorizes bypassing a refusal, lease, or gate. If a refusal is genuinely wrong,
file its witnessed, deduplicated appeal with `fak complain --summary "…" --reason
<TOKEN> --tool <Tool> --from-journal --args-digest <sha256:…> --workspace
<blocked-root> --live`. When two enforced workflow obligations leave no compliant
sequence after safe recovery is exhausted, file the exact command, output, attempted
routes, and unblock condition with `fak complain --domain workflow --kind catch-22
--summary "…" --rationale-file <absolute-scrubbed-evidence> --workspace
<blocked-root> --repo <owner/repo> --live`; use `completion-blocker` for a general
completion blocker. Success requires a reported created or updated issue. If live
filing cannot fetch GitHub, the command exits nonzero before any remote issue action,
reports JSON mode `pending-local`, and retains a scrubbed stable-key receipt under
`<blocked-root>/.fak/complaints/pending`. Inspect it without contacting GitHub with
`fak complain --pending --workspace <blocked-root> --json`; do not claim remote
success. When GitHub recovers, repeat the same stable complaint and `--live` command;
a verified create/update removes its receipt. An unverified sync also exits nonzero
and retains the receipt. The receipt is recovery evidence, never permission to bypass
a refusal, lease, or gate. Taxonomy and routing:
[`docs/notes/CONCEPT-AGENT-FRICTION-COMPLAINT-CHANNELS-2026-06-29.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-AGENT-FRICTION-COMPLAINT-CHANNELS-2026-06-29.md).

## Releasing and planning

Use `/release` for the full release path. It enforces a clean synchronized `main`, full CI,
version bump, release notes, signed commit/tag, push, GitHub release, and rollback guidance.
Never hand-edit version constants, tag partially, force-push, or publish from a dirty/diverged
tree. Release history and channel contract: [`docs/releases/`](https://github.com/anthony-chaudhary/fak/tree/main/docs/releases) and [`docs/releases-channel.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/releases-channel.md).

Choose the planning primitive by work shape:

- **Phased deliverable:** finite ordered phases with observable completion. Keep
  **Current state** accurate across the independent product, evidence, and queue axes in
  [`docs/progress-state-defaults.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/progress-state-defaults.md), append dated facts to
  **Execution log**, run `/phased-plan` each phase, and archive completed plans under
  `docs/plans/completed/`. `fak plan-audit` catches drift.
- **Ongoing program:** recurring unbounded work. Track health, cadence, and next movement rather
  than percent complete, and archive under `docs/programs/completed/` only when retired.

Plans are canonical state, not narration. A missing or failed performance receipt changes the
evidence axis; it does not erase delivered scope or force the whole item into `HOLD`. Preserve
historical `KEEP`/`REJECT`/`HOLD` decisions in the execution log, keep delivery credit separate
from performance credit, and never weaken the applicable claim gate. Update plans in the same
commit as the work; document scope changes explicitly. For parallel work, assign distinct file
sets and integrate only after all workers finish. Audit before starting a new plan when
active-plan load is high.
## Where to go next

| If you want to… | Read |
|---|---|
| Every CLI verb + what's shipped | [`docs/cli-reference.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/cli-reference.md) |
| Learn every concept in prerequisite order (a course, join at your level) | [`LEARNING-PATH.md`](https://github.com/anthony-chaudhary/fak/blob/main/LEARNING-PATH.md) |
| Install / run tiers (offline → gateway → in-kernel model) | [`GETTING-STARTED.md`](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md) |
| Run local models on Mac (Qwen3.8 + Metal REPL/server) | [`docs/fak/mac-local-models.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mac-local-models.md) · `fak run qwen38` |
| Put fak in front of *your* agent (Claude Code / Cursor / MCP) | [`docs/integrations/`](https://github.com/anthony-chaudhary/fak/tree/main/docs/integrations) · [`fak/examples/mcp/`](https://github.com/anthony-chaudhary/fak/tree/main/examples/mcp) |
| Run hardware-gated work (no local GPU) — the sanctioned compute nodes | [`docs/fleet-compute-nodes.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fleet-compute-nodes.md) · `fak hwgate-lint` |
| The deployable capability floor (policy manifests) | [`fak/POLICY.md`](https://github.com/anthony-chaudhary/fak/blob/main/POLICY.md) · [`fak/examples/README.md`](https://github.com/anthony-chaudhary/fak/blob/main/examples/README.md) |
| Extend the kernel (plug in → prove correct → prove faster) | [`fak/EXTENDING.md`](https://github.com/anthony-chaudhary/fak/blob/main/EXTENDING.md) · [`fak/ARCHITECTURE.md`](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md) |
| Optimize a kernel without re-inventing known art (check prior art first) | [`docs/sota/README.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/sota/README.md) · `fak sota <op>` |
| Score agent-steer prose for negative framing (suggests positive reframes) | `fak score negframe --suggest` |
| Every feature by subsystem, with honest status | [`docs/supported/features.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/features.md) |
| What's real vs simulated vs stub | [`fak/CLAIMS.md`](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md) · [`fak/STATUS.md`](https://github.com/anthony-chaudhary/fak/blob/main/STATUS.md) |
| Every benchmark number (single source of truth) | [`fak/BENCHMARK-AUTHORITY.md`](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md) |
| Roll back to a stable version (revert / downgrade / pin) | [`docs/ROLLBACK.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/ROLLBACK.md) |
| A curated map of all the docs | [`llms.txt`](https://github.com/anthony-chaudhary/fak/blob/main/llms.txt) |

License: [Apache-2.0](https://github.com/anthony-chaudhary/fak/blob/main/LICENSE).

---

# `docs/CAPABILITIES.md`

> Source: `docs/CAPABILITIES.md`

---
title: "fak capability map: spend fewer tokens and turns"
description: "Documentation for fak capability map: spend fewer tokens and turns, including the captured behavior, operating context, and reproducible fak evidence."
---

# fak capability map: spend fewer tokens and turns

fak's current product focus is **agent efficiency**: keep stable prompt work reusable, avoid
unnecessary model round trips, route each call to an appropriate model, and let operators
control long-running sessions without injecting another conversational turn. The security
floor remains shipped and indexed, but it is a supporting property of the same tool-call
boundary rather than the lead story here.

This page is the short, outcome-first index. It is also queryable: from a source checkout,
`fak-dev capabilities "token savings"` or `fak-dev capabilities "turn control"` returns
machine-readable capability cards with an exact next command and evidence seam. For the
complete research and design inventory, use the [innovations index](https://github.com/anthony-chaudhary/fak/blob/main/docs/INNOVATIONS-INDEX.md);
for the full documentation catalog, use [docs/index.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/index.md).

<a id="turn-savings"></a>

## Start with the largest avoidable cost: extra turns

A model turn costs more than its visible answer. It can resend a long resident prefix, incur
provider latency, and create yet more context for every later turn. fak therefore measures and
removes **turn tax**, not only individual token strings.

| Capability | What it changes | Try or inspect it |
|---|---|---|
| Turn-tax meter | Attributes avoidable round trips to forced policy retry, deterministic tool work, duplicate/replay work, and other turn kinds. | `go run ./cmd/turntaxdemo -selfcheck`; [turn-tax measurement](https://github.com/anthony-chaudhary/fak/tree/main/internal/turntaxmeter) |
| Fused deterministic work | Completes kernel-known work at the tool boundary instead of paying for another model turn. | [fused-turn architecture](https://github.com/anthony-chaudhary/fak/tree/main/internal/fusedturn) |
<a id="session-control"></a>

| Live turn control | Budgets turns/tokens/context and pauses, resumes, throttles, steers, or stops a served session out of band—without spending a prompt turn to ask the model to control itself. | `fak help session`; [operator control plane](https://github.com/anthony-chaudhary/fak/blob/main/docs/operator-control-plane.md) |
| Provider clear/new boundary | Starts a fresh fak trace when a wrapped provider explicitly clears its conversation, while carrying cumulative limits so the UI command cannot reset quotas. | [provider clear/new contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/provider-session-reset.md) |
| Same-trace ablation | Replays one frozen trace with cache levers on/off so savings are attributable rather than anecdotal. | `fak help ablate`; `fak ablate --help` |

The browserless turn-tax demo is the smallest captured spine: its self-check drives the real
`turnbench` path and includes an anti-inflation control where a happy path must report zero
saved turns. The shipped fixture currently witnesses **9 avoided turns** (5 forced-turn + 4
elision) in its synthetic airline scenario; that is a reproducible fixture result, not a
universal workload claim.

<a id="context-reuse"></a>

## Reuse prompt work instead of repaying for it

| Capability | What it changes | Try or inspect it |
|---|---|---|
| Stable-prefix reuse (vDSO) | Keeps shared setup/provider-cache-compatible prefixes reusable across turns. | [vDSO cache quickstart](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/cache.md) |
| Managed context (ctxmmu) | Sheds stale turns while preserving a stable prefix and the information needed to continue. | [managed context](https://github.com/anthony-chaudhary/fak/blob/main/docs/managed-context-continuous-usage.md) |
| Tool-output compression | Compresses large tool results before they consume model-window tokens; benchmark the built-in corpus or your own captured outputs. | `fak headroom bench --json`; `fak headroom bench --dir <captures> --json`; [live witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONTEXT-COMPRESSION-LIVE-2026-08-09.md) |
| Resume pricing | Prices full replay versus cut/reset after cache TTL expiry instead of blindly replaying a large session. | `fak resume plan --resident-tokens 250000 --idle-seconds 7200 --json` |
| Portable session state | Persists, moves, and demand-pages a long-running session instead of forcing a restart or full transcript replay. | `fak snapshot demo`; `fak snapshot query --file <image.faksession> --query <question>`; [session image](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/PORTABLE-SESSION-IMAGE-AND-SNAPSHOT-2026-06-24.md) |
| Cache-value accounting | Reports reused tokens, effective cost, and attributable savings in operator terms. | `fak info --once`; [cache-value rollup](https://github.com/anthony-chaudhary/fak/blob/main/docs/cache-value-rollup.md) |

<a id="context-compression"></a>
<a id="portable-session"></a>

<a id="model-routing"></a>

## Spend expensive inference only where it helps

| Capability | What it changes | Try or inspect it |
|---|---|---|
| Per-call model routing | Routes calls by task/turn requirements rather than pinning a whole session to one expensive model. | [model routing](https://github.com/anthony-chaudhary/fak/blob/main/docs/model-routing.md); `fak model --help` |
| Model ladder and acceptance | Separates candidate selection from measured acceptance so a cheaper route is not called a win merely because it is cheaper. | [model operations](https://github.com/anthony-chaudhary/fak/blob/main/docs/model-production-readiness-inventory.md) |
<a id="savings-observability"></a>

| Token/cache observability | Makes savings visible during operation and audits which default savers have captured effect witnesses instead of counting configured/ready as wins. | `fak info --once`; `fak token-defaults-scorecard --effectiveness`; [token-saving observability](https://github.com/anthony-chaudhary/fak/blob/main/docs/serving/token-savings-observability.md) |

<a id="capability-floor"></a>

## Supporting floor: keep optimization bounded

The same checkpoint also enforces a default-deny capability floor and records decisions. That
matters because a faster or cheaper agent that can silently widen its authority is not a
net-true improvement. Security details remain fully discoverable in the
[security and policy guide](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/security.md), but new readers should begin with the efficiency paths above.

## Where to go next

- **Evaluate token efficiency:** [awesome-token-efficiency](https://github.com/anthony-chaudhary/fak/blob/main/docs/awesome-token-efficiency.md) maps
  the broader field and defines honest measurement boundaries.
- **Understand cache economics:** [cache frontier operating plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/CACHE-FRONTIER-OPERATING-PLAN.md)
  and [cache-value rollup](https://github.com/anthony-chaudhary/fak/blob/main/docs/cache-value-rollup.md).
- **Build an integration:** [integration guides](https://github.com/anthony-chaudhary/fak/tree/main/docs/integrations) and `fak help serve`.
- **Browse everything shipped or proposed:** [innovations index](https://github.com/anthony-chaudhary/fak/blob/main/docs/INNOVATIONS-INDEX.md), which
  retains maturity labels so an idea is not mistaken for a product claim.

---

# `docs/problems-we-solve.md`

> Source: `docs/problems-we-solve.md`

---
title: "The problems fak exists to solve"
description: "FAK's P1-P4 checklist for managed context, net-true efficiency, bounded adaptation, and integrated operations, plus its centrality model."
---

# The problems fak exists to solve

fak exists to make long-running agent work **cheaper, faster, safer, and more operable**.
Those outcomes meet at one kernel seam: every tool call crosses a checkpoint where fak can
preserve useful context, avoid unnecessary work, adapt execution, and enforce the operating
contract.

This page gives development work two different instruments. Do not collapse them into one score:

1. **The P1-P4 problem checklist** is a design and review checklist. Every change considers
   every row. The rows are not competing priorities and an author does not pick a favorite one.
2. **Problem centrality** is a portfolio signal. It says how directly a unit of work advances
   the connected problem cluster below. More-central work normally outranks less-related work,
   after hard obligations and dependencies are accounted for.

A P1 implementation still has to be net-true (P2), adaptable rather than brittle (P3), and
operable through the real system (P4). Conversely, release or maintenance work may directly
implement none of the four yet still be necessary stewardship. The checklist governs **how we
do the work**; centrality helps decide **which work to do next**.

## The connected problem cluster

The labels are stable handles for related user problems, not product silos. A real change can
advance several at once.

### P1 — Managed context

**Problem:** Agent sessions repeatedly reconstruct shared setup, lose useful state as histories
grow, and send context that could have been retained or served locally.

**Direction:** Preserve reusable prefixes, serve safe repeats locally, compact old turns
deliberately, and keep provider caches useful across long-running sessions.

**Witness:** less repeated input or setup work at equal task quality; longer useful sessions;
a cache, replay, or compaction effect read back from the real path.

### P2 — Net-true efficiency

**Problem:** An optimization can look cheaper in isolation while routing, verification, quality
loss, retries, or operator burden makes the real workflow worse.

**Direction:** Use the least expensive execution that still meets the contract, count all added
costs, and compare with the tuned next-best alternative rather than a naive baseline.

**Witness:** an end-to-end latency, cost, quality, or completed-work gain that survives the
[`net-true-value`](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/net-true-value.md) accounting.

### P3 — Fast, bounded adaptation

**Problem:** Agents and workloads change faster than static infrastructure. Broad rewrites are
too slow; unbounded adaptation is unsafe and hard to trust.

**Direction:** Make routing, policy, model, and memory behavior changeable at small seams, with
explicit bounds, rollback, and evidence.

**Witness:** a new workload or policy supported through a small, reversible change with captured
before/after behavior.

### P4 — Integrated operations

**Problem:** A good component does not help if operators cannot discover, run, observe, recover,
and govern it through the actual agent/tool path.

**Direction:** Put performance, security, observability, recovery, and lifecycle control on the
same checkpoint rather than in disconnected side systems.

**Witness:** the real end-to-end path works and exposes enough state to diagnose, recover, and
enforce its contract.

## Cross-cutting outcome: portable intent

The four rows above also support a broader user outcome: **portable intent**. A person should be
able to state the behavior they want at a named product layer without first learning which
plugin, skill, prompt fragment, model option, or native implementation happens to provide it.
For example, a conversation preference should name the desired style and its constraints; it
should not require the user to know that one harness calls a possible implementation
“Caveman.” A team can then review and share the semantic declaration instead of persuading
every recipient to install the same mechanism.

Portable intent separates two identities:

- the **intent** is a versioned, typed value owned by its layer, such as a conversation-style
  preference with required fidelity and an authority ceiling;
- a **binding** is the replaceable, provenance-carrying mechanism that realizes that intent in
  one environment.

This is not a fifth P-number or a universal ontology. Each layer owns a bounded vocabulary and
composition rules while a common envelope carries scope, compatibility, provenance, resolution,
and receipts. The outcome connects all four checks: compact semantic declarations can reduce
repeated setup (P1); any gain must beat direct configuration net of resolver and verification
cost (P2); bindings make implementation churn bounded and reversible (P3); and the real path
must resolve, apply, explain, refuse semantic loss, and read back the effect (P4).

The roadmap is tracked in [#6877](https://github.com/anthony-chaudhary/fak/issues/6877). The
required first proof is [#6878](https://github.com/anthony-chaudhary/fak/issues/6878): one
implementation-independent conversation profile running through two different bindings, with
an unsupported required meaning refused before launch. Until that witness lands, portable
intent is a direction and contract, **not a shipped runtime claim**.

## The all-work problem checklist

Run this checklist at **framing, spine design, implementation, witness, and review**. “All must
pass” means every row receives an honest answer; it does not mean every change produces a new
feature for every P-number.

| Check | Ask on every unit of work | Pass condition |
|---|---|---|
| **P1 · Context** | What context or repeated work does this add, remove, preserve, or invalidate? | It improves managed context, or does not create avoidable repetition, cache damage, or context growth. |
| **P2 · Net value** | Is this better than the real alternative after implementation, runtime, verification, quality, retry, and operator costs? | A proportionate witness supports the gain, or it makes no gain claim and names its necessary obligation. |
| **P3 · Adaptation** | What is likely to change next, and is the seam bounded, reversible, and small enough to adapt? | It avoids gratuitous lock-in, states its bounds, and has a risk-appropriate rollback or safe failure path. |
| **P4 · Operations** | Can the real system discover, run, observe, secure, and recover this behavior? | The end-to-end seam and operating contract are witnessed; component proof is not used for a system claim. |

A row may be **advanced**, **preserved**, or **not applicable with a concrete reason**. “N/A” is
not a shortcut: a typo fix can state that it changes no runtime context, gain claim, adaptation
seam, or operating path; a broad feature cannot plausibly dismiss all four. A row **fails** when
the change creates an unmitigated regression, makes an unwitnessed claim, or supplies only a
label. Failed rows block normal landing. A time-critical obligation must name the accepted
tradeoff, owner, and follow-up rather than silently converting failure to N/A.

### Use it through the development loop

1. **Frame:** State `For / Problem / Today / Better because / Witness` in plain language.
2. **Classify centrality:** Assign one class and one sentence of evidence using the rubric below.
3. **Design the spine:** Choose the smallest real end-to-end path and record P1-P4 risks.
4. **Implement:** Keep context accounting, net cost, bounded change, and operations at the seam;
   do not bolt them on after the code is “done.”
5. **Witness:** Capture the working spine first, then its operating envelope and optimization.
6. **Review and close:** Re-run all four rows against the diff and evidence; file follow-ons
   rather than hiding a failed row in prose.

This complements the [spine-first defaults](https://github.com/anthony-chaudhary/fak/blob/main/docs/spine-first-defaults.md), the
[Feynman-simple value frame](https://github.com/anthony-chaudhary/fak/blob/main/docs/shift-left-task-organization.md), architectural leaf boundaries,
capability policy, and net-true evidence standard. Those are delivery disciplines; P1-P4 are
the persistent problem lens applied through them.

## Default priority hierarchy

Alongside problem centrality and the P1-P4 checklist, the repository prioritizes work across four explicit operating tiers:

1. **fak all in one (serving and harness + memory — the "one touch" thing):**
   The primary product focus: a single-binary turnkey deployment (`fak up`) uniting high-performance model serving, the agent harness with capability-floor security, and durable context memory into one frictionless developer experience.
2. **fak serving only:**
   High-performance model inference, disaggregated KV-cache routing, and OpenAI/MCP gateway endpoints (`fak serve`).
3. **fak harness only:**
   Standalone agent harness, capability floor enforcement, and tool-call interception wrapping external frontier models (`fak guard`).
4. **other things:**
   Peripheral utilities, benchmark harnesses, standalone tooling, and secondary integrations.

## Problem centrality: the portfolio signal

Centrality asks:

> If this work succeeds, how directly does it relieve the connected user problems above or
> unblock evidence that they are relieved?

Use qualitative classes; do not manufacture numeric precision the evidence cannot support.

| Class | Meaning | Evidence and default treatment |
|---|---|---|
| **Core** | Directly changes a user-visible P1-P4 outcome on fak's kernel path. | Witness the end-to-end effect; prefer it when evidence, urgency, and readiness are comparable. |
| **Enabling** | Removes a concrete blocker or supplies required measurement for named Core work. | Name the blocked outcome; sequence with it and reclassify if that outcome closes or changes. |
| **Stewardship** | Maintains reliability, security, compatibility, release health, or developer throughput without directly moving a P1-P4 outcome. | Name the obligation or risk; schedule by risk, deadline, and recurring cost. |
| **Peripheral** | Has no evidenced path to the problem cluster, enabling dependency, or current obligation. | Defer, reshape, or decline unless an explicit external obligation overrides. |

Centrality is **not the whole priority decision**. First honor security incidents, data-loss
risks, broken trunk/release obligations, and contractual deadlines. Then satisfy dependencies
and compare ready work by centrality, user evidence, expected net value, urgency, effort, and
reversibility. Centrality is the directional tie-breaker that prevents an attractive but
unrelated backlog from displacing fak's reason to exist.

Re-evaluate centrality when scope, dependencies, or the user outcome changes; never inherit it
mechanically from a parent epic. Classify the **effect**, not its directory, label, or mechanism.
For example, observability is Peripheral as a speculative dashboard, Enabling when it measures
a named cache spine, and Core when diagnosis/recovery is itself the broken P4 path.

### Examples

- **Core:** preserve a provider-cache-compatible prefix across compaction and capture reduced
  repeated input on the real path. It still answers P2, P3, and P4.
- **Enabling:** add counters required to decide whether that named compaction spine saves work.
  Generic telemetry without a named outcome is not enough.
- **Stewardship:** update a supported Go version after a security deadline or repair a flaky
  release gate. Urgent stewardship can outrank Core work without a fictional product claim.
- **Peripheral:** add an unrelated management surface with no kernel-path user, dependency, or
  obligation. Tie it to witnessed P1-P4 pain or spend capacity on more central work.

## Required issue and plan frame

```text
For: <specific user/operator>
Problem: <observable pain today>
Today: <real next-best alternative>
Better because: <expected outcome, net of added cost>
Witness: <artifact or read-back that can prove it>
Centrality: Core | Enabling(<named Core outcome>) | Stewardship(<obligation>) | Peripheral
P1 Context: advanced | preserved | N/A — <reason>
P2 Net value: advanced | preserved | N/A — <reason>
P3 Adaptation: advanced | preserved | N/A — <reason>
P4 Operations: advanced | preserved | N/A — <reason>
```

Keep this proportionate. The block replaces the old “choose one primary P-ID” convention.
Multiple rows may be advanced because the problems are a cluster, not competing buckets.

## What this model refuses

- Picking P2 and ignoring P4: a benchmark-only speedup is not an integrated outcome.
- Four ceremonial checkmarks: identify an effect, preserved invariant, or concrete reason.
- Calling every prerequisite Core: enabling work names the outcome it actually unblocks.
- Using centrality to skip maintenance: urgent stewardship can outrank product work.
- Using priority to excuse bad design: urgent work still passes a proportionate checklist.
- Mistaking a leaf for a user problem: centrality follows the witnessed effect.

## One-sentence test

> We apply context, net-value, adaptation, and operations checks to every change; when choosing
> among changes, we prefer the strongest evidenced path to fak's connected user problems,
> subject to real obligations and risk.

---

# `CLAIMS.md`

> Source: `CLAIMS.md`

# CLAIMS.md — the fak honesty ledger


Every capability claim carries **exactly one** tag:

- `[SHIPPED]` — real code on the critical path, closed by a mechanical witness (a `go test`, a `go build`, a benchmark field, a file read-back). Reproducible now.
- `[SIMULATED]` — modeled with labeled stand-in data (no GPU / no live engine on the build box); the seam is real, the numbers are illustrative.
- `[STUB]` — plumbing present, behavior deferred; clearly labeled, returns a STUB/no-op result.

The tag answers "is it real". It does not answer the orthogonal question — **is it ON for an operator who configured nothing** — so every `[SHIPPED]` claim also carries a machine-readable **exposure** state, declared at the END of the line. It is the ledger's binding of Q6 of [the net-true-value standard](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/net-true-value.md) ("on by default, or honestly gated with a stated reason"), reusing `internal/claimcheck`'s `Realized` type rather than a parallel vocabulary — exposure is an axis beside the tag, never a fourth tag:

- `[exposure: default-on]` — on for an operator who set no flag. This is also the state a `[SHIPPED]` line with no marker asserts, so a missing marker is a claim, not silence.
- `[exposure: gated — <reason>]` — ships off, with the reason stated. A gated claim with no stated reason FAILS the lint; that is exactly the Q6 rule `claimcheck.gradeRealized` already encoded.
- A `[SIMULATED]`/`[STUB]` claim is PARKED by its tag (it is not on the critical path by definition) and carries no exposure marker.

The lint witness (unit 96): `go test -v -run TestCLAIMSLedger ./internal/claimcheck` (the `claims-lint` gate, in CI on every push) checks every line beginning with `- [` for one and only one of the three tags AND for a declared exposure on every `[SHIPPED]` line whose prose discloses gating, and EMITS the default-off count instead of leaving it to be grepped out of prose.


<!-- fak:document-set -->

This compact ledger indexes one addressable page per claim. Claim text and maturity tags are preserved on the linked pages.

## Claims

<a id="the-product"></a>
- [SHIPPED] [The product](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/the-product.md) [exposure: default-on]
<a id="the-syscall-subsystem-latency-check-not-the-headline-kpi-unit-82"></a>
- [SHIPPED] [The syscall subsystem latency check (not the headline KPI — unit 82)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/the-syscall-subsystem-latency-check-not-the-headline-kpi-unit-82.md) [exposure: default-on]
<a id="adjudication-the-in-process-dos-reference-monitor"></a>
- [SHIPPED] [Adjudication (the in-process DOS reference monitor)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/adjudication-the-in-process-dos-reference-monitor.md) [exposure: default-on]
<a id="tool-vdso-3-tier-local-fast-path"></a>
- [SHIPPED] [Tool vDSO (3-tier local fast path)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/tool-vdso-3-tier-local-fast-path.md) [exposure: default-on]
<a id="pre-flight-ladder-grammar-rung"></a>
- [SHIPPED] [Pre-flight ladder + grammar rung](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/pre-flight-ladder-grammar-rung.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="context-mmu-write-time-result-admission"></a>
- [SHIPPED] [Context-MMU (write-time result admission)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/context-mmu-write-time-result-admission.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="answer-shape-the-consumer-facing-degeneration-verbosity-witness"></a>
- [SHIPPED] [Answer-shape: the consumer-facing degeneration/verbosity witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/answer-shape-the-consumer-facing-degeneration-verbosity-witness.md) [exposure: default-on]
<a id="codelint-language-server-packs-over-agent-written-code"></a>
- [SHIPPED] [Codelint: language-server packs over agent-written code](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/codelint-language-server-packs-over-agent-written-code.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="session-core-dump-context-debugger-recall-cdb"></a>
- [SHIPPED] [Session core-dump + context debugger (recall + cdb)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/session-core-dump-context-debugger-recall-cdb.md) [exposure: default-on]
<a id="portable-session-image-uniform-dump-restore-session-restore-sessionimage-snapshot"></a>
- [SHIPPED] [Portable session image + uniform dump/restore (session.Restore + sessionimage + snapshot)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/portable-session-image-uniform-dump-restore-session-restore-sessionimage-snapshot.md) [exposure: default-on]
<a id="in-kernel-agent-to-agent-message-channel-a2achan"></a>
- [SHIPPED] [In-kernel agent-to-agent message channel (`a2achan`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/in-kernel-agent-to-agent-message-channel-a2achan.md) [exposure: default-on]
<a id="shared-task-record-fold"></a>
- [SHIPPED] [Shared task record fold](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/shared-task-record-fold.md) [exposure: default-on]
<a id="trajectory-observability-primitives-data-plane-reference-similarity-scorer-seam"></a>
- [SHIPPED] [Trajectory observability primitives (data plane + reference similarity + scorer seam)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/trajectory-observability-primitives-data-plane-reference-similarity-scorer-seam.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="task-manager-snapshot"></a>
- [SHIPPED] [Task manager snapshot](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/task-manager-snapshot.md) [exposure: default-on]
<a id="s7-write-time-durability-gate-context-is-not-memory"></a>
- [SHIPPED] [S7 write-time durability gate (context is not memory)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/s7-write-time-durability-gate-context-is-not-memory.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="in-kernel-model-the-model-fused-into-the-kernel"></a>
- [SHIPPED] [In-kernel model (the model fused into the kernel)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/in-kernel-model-the-model-fused-into-the-kernel.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="security-substrate-the-kernel-stops-believing-the-model"></a>
- [SHIPPED] [Security substrate (the kernel stops believing the model)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/security-substrate-the-kernel-stops-believing-the-model.md) [exposure: default-on]
<a id="gateway-fak-serve"></a>
- [SHIPPED] [Gateway (`fak serve`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/gateway-fak-serve.md) [exposure: default-on]
<a id="model-routing-per-aspect-ensemble-fak-route"></a>
- [SHIPPED] [Model routing (per-aspect + ensemble — `fak route`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/model-routing-per-aspect-ensemble-fak-route.md) [exposure: default-on]
<a id="turn-tax-benchmark-fak-turntax"></a>
- [SHIPPED] [Turn-tax benchmark (`fak turntax`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/turn-tax-benchmark-fak-turntax.md) [exposure: default-on]
<a id="self-ablation-sweep-fak-ablate"></a>
- [SHIPPED] [Self-ablation sweep (`fak ablate`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/self-ablation-sweep-fak-ablate.md) [exposure: default-on]
<a id="cross-agent-ablation-regime-b-bare-claude-p-vs-fak-guard-claude-p"></a>
- [SHIPPED] [Cross-agent ablation (Regime B — bare `claude -p` vs `fak guard -- claude -p`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/cross-agent-ablation-regime-b-bare-claude-p-vs-fak-guard-claude-p.md) [exposure: default-on]
<a id="fan-out-benchmark-fanbench-one-master-goal-n-sub-agents-n-1-1024"></a>
- [SHIPPED] [Fan-out benchmark (`fanbench` — one master goal → N sub-agents, N=1…1024)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/fan-out-benchmark-fanbench-one-master-goal-n-sub-agents-n-1-1024.md) [exposure: default-on]
<a id="bounded-microagents-construct-harnesses-cmd-microharnessdemo"></a>
- [SHIPPED] [Bounded microagents construct harnesses (`cmd/microharnessdemo`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/bounded-microagents-construct-harnesses-cmd-microharnessdemo.md) [exposure: default-on]
<a id="in-process-microagent-host-internal-microagent-n-agent-loops-as-goroutines-behind-one-gateway"></a>
- [SIMULATED] [In-process microagent host (`internal/microagent` — N agent loops as goroutines behind ONE gateway)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/in-process-microagent-host-internal-microagent-n-agent-loops-as-goroutines-behind-one-gateway.md)
<a id="ultra-long-context-work-floor-longctxbench-per-agent-context-100k-tokens"></a>
- [SHIPPED] [Ultra-long-context work floor (`longctxbench` — per-agent context > 100k tokens)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/ultra-long-context-work-floor-longctxbench-per-agent-context-100k-tokens.md) [exposure: default-on]
<a id="engine"></a>
- [SHIPPED] [Engine](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/engine.md) [exposure: default-on]
<a id="stewards-rsi-ship-gate"></a>
- [SHIPPED] [Stewards + RSI ship-gate](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/stewards-rsi-ship-gate.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="vcache-chains-recall-m4"></a>
- [SHIPPED] [vCache Chains & Recall (M4)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/vcache-chains-recall-m4.md) [exposure: gated — linked shipped mechanisms include an opt-in or default-off path; see the detail page for each gate]
<a id="vcache-governor-m5"></a>
- [SHIPPED] [vCache Governor (M5)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/vcache-governor-m5.md) [exposure: default-on]
<a id="vcache-observability-per-sub-concept-lens"></a>
- [SHIPPED] [vCache observability (per-sub-concept lens)](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/vcache-observability-per-sub-concept-lens.md) [exposure: default-on]
<a id="cache-value-rollup"></a>
- [SHIPPED] [Cache-value roll-up front door](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/cache-value-rollup.md) [exposure: default-on]
<a id="what-fak-is-not"></a>
- [SIMULATED] [What fak is NOT](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/what-fak-is-not.md)
<a id="prior-art-posture"></a>
- [SHIPPED] [Prior-art posture](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/prior-art-posture.md) [exposure: default-on]
<a id="nvidia-rtx-5090-bam-gpudirect-nvme"></a>
- [SIMULATED] [NVIDIA RTX 5090 BaM GPU Direct NVMe & Hierarchical Memory](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/nvidia-rtx-5090-bam-gpudirect-nvme.md)

---

# Project orientation

> Source: `docs/project-orientation.md`

---
title: "Project orientation: the agent-kernel center"
description: "The active orientation for fak: preserve managed context, choose capable low-cost execution, enforce authority, and keep agent turns operable."
---

# Project orientation: the agent-kernel center

**Decision date:** 2026-08-15

**Status:** active orientation spine for [#6899](https://github.com/anthony-chaudhary/fak/issues/6899)

## Verdict

**Current product milestone (2026-09-09): [Useful local agents, accelerated
automatically](https://github.com/anthony-chaudhary/fak/blob/main/docs/local-agent-milestone.md).** Make performant local agent work
practical through native execution, qualified default acceleration, and reusable
agent context. This sharpens the current investment emphasis: a normal local
entry point, a real task, and an independent outcome receipt. The milestone
specifies automatic behavior and distinguishes present wiring from qualification.
It does not replace the centrality framework below; classify a change by its
effect on this workflow, not by the number of backends or optimizations it adds.

fak's center is the **kernel-mediated agent turn**:

> preserve useful managed context and cache-compatible shared work, choose the least-cost capable execution path, enforce a fail-closed capability floor at the same tool boundary, and keep that path operable through real agent harnesses.

This is one connected outcome, not four competing products. Managed context (P1), net-true efficiency (P2), bounded adaptation (P3), and integrated operations (P4) are checks on every change. Problem centrality says how directly a change relieves that outcome; it does not replace urgency, dependency order, risk, or stewardship obligations.

## Core is temporal, but identity is not fashion

“Core” must not answer four different questions at once:

1. **Enduring promise (constitutional, multi-cycle):** turn agent intent into economical, bounded, portable, and observable execution over long-running work.
2. **Owned seam:** mediate between agent intent and model/tool execution, including the state needed to preserve that contract across turns and harnesses.
3. **Current strategic emphasis (roughly 2–4 quarters):** invest most where the highest evidenced, controllable bottleneck limits that promise now.
4. **Change centrality:** classify a proposed effect as Core, Enabling, Stewardship, or Peripheral relative to the promise and current evidence.

A mechanism can move between headline wedge, active bottleneck, table stakes, reference implementation, optional substrate, and retired machinery without changing the enduring promise. De-emphasis never licenses regression of the contract the mechanism established.

### Why performance is emphasized now—and how it can recede

Performance means avoided total work, turns, latency, and cost **after** fak’s overhead, not a permanent allegiance to a particular cache implementation. It is a buying wedge now because repeated model setup, prompt replay, avoidable turns, and context reconstruction remain expensive and controllable. Cache stability and managed context are therefore strategic Core work today.

Performance emphasis should fall when representative tuned provider-native paths make fak’s marginal savings immaterial across two review cycles and no recovery or control advantage depends on the mechanism. At that point, “do not cause avoidable repeated work; report gains only when net-true” remains a constitutional invariant, while specialized cache machinery becomes Enabling, optional, or retired. If inference becomes nearly free but human attention, reliability, or effect risk dominates, investment moves to those bottlenecks.

### Why harness ownership is emphasized now—and how it can recede

A kernel that cannot preserve its semantics through a harness is not an integrated product. Harness and binding work is strategically central now because external seams are fragmented: they often lose context, recovery, capability, or operator-control semantics. fak needs at least one owned end-to-end path to prove the full contract and enough bindings to prove portability.

The first-party harness itself is not constitutionally Core. It can become a reference implementation when two independent external bindings pass the same portable-intent, capability, session, recovery, and operator-control witnesses without harness-specific behavior. Binding conformance remains central; UI breadth and harness-specific convenience then become Enabling or Peripheral.

### Transition discipline

- **Increase emphasis** only from a measured user-path bottleneck, required obligation, or evidence gap blocking a named decision.
- **Decrease emphasis** only when a tuned alternative satisfies the retained contract over representative paths—not because a mechanism became unfashionable.
- **Promote an option** when a named user path cannot meet privacy, availability, economics, or control requirements through the next-best substrate.
- **Demote or retire machinery** when usage and avoided cost no longer exceed maintenance, integration, and cognitive load.
- **Review before stale:** the canonical snapshot has an explicit review date; staleness requests judgment but never auto-rewrites strategy.

The machine-readable authority is `internal/orientation/orientation.json`; inspect it with `fak orientation` or `fak orientation --json`. It records current role, horizon, evidence state, retained contract, and increase/decrease triggers for each capability family.

The original fail-closed policy seam remains central, but it is not the whole product. Concurrent cache reuse exposed the larger user problem: agents repeatedly pay to reconstruct context and execution setup. Routing, compaction, repeat serving, recovery, and policy belong together when they improve one real mediated turn. Project automation does not become product Core merely because fak's maintainers use it heavily.

## Capability portfolio

| Capability family | Current centrality | Why | Required witness |
|---|---|---|---|
| Tool-call mediation, capability policy, adjudication | **Core** | The kernel must handle each effect before execution and fail closed without model persuasion. | Structural deny and non-blanket allow on the real preflight path. |
| Cache-stable traffic, shared-prefix/KV reuse, local repeat serving | **Core** | Avoiding repeated shared work is the principal performance outcome. | Net-true repeated-input, latency, or compute reduction against the tuned real alternative. |
| Managed context, compaction, recall, crash/session recovery | **Core** | Long-running agents need useful state preserved without repeatedly rebuilding the prompt. | Task continuity plus provider-cache-compatible prefix or measured reconstruction reduction. |
| Model routing and bounded adaptation | **Core when on the mediated turn** | Choosing a capable cheaper path is part of net-true execution; generic model experimentation is not. | Correctness non-inferiority and end-to-end cost/latency on the real call path. |
| Agent/harness integration and portable intent | **Core at the binding seam** | A kernel unused by a real harness is not an integrated outcome. Harness-specific convenience beyond the binding is Enabling or Peripheral. | One intent/profile runs through real bindings without capability widening or semantic drift. |
| Gateway, local inference, compute backends | **Enabling(the kernel-mediated turn)** | They provide execution substrates; backend breadth is not an end in itself. | A named Core path gains availability, cost, or control with full added operating cost stated. |
| Benchmarks, evals, claims, observability | **Enabling(a named Core claim)** | Evidence selects and verifies central product work. Generic measurement inventory is insufficient. | A decision or shipped claim changes because of the captured evidence. |
| Fleet, dispatch, shared-tree, release and CI machinery | **Stewardship(reliable development and release)** by default | These are real obligations for this repository, not automatic external product value. A leaf may be Enabling when tied to a named product witness. | Reduced delivery/recovery cost or a required release/build invariant. |
| Issue gardening, scoring, agent-process automation | **Stewardship(bounded portfolio operation)** by default | The 1,600+ issue portfolio needs control, but self-management must not consume the product it serves. | A bounded decision or retired queue cost, including the automation's own upkeep. |
| Unrelated management surfaces or research without a kernel-path user | **Peripheral** | Nearby usefulness does not establish fak product centrality. | Promotion requires a named user, real alternative, P1-P4 effect, and captured kernel-path witness. |

Classification attaches to a **change's witnessed effect**, not permanently to a directory. A gateway fix can be Core when the mediated turn is broken, Enabling when it adds a substrate, or Stewardship when it fulfills a compatibility obligation.

## Investment boundary

1. Prefer the smallest end-to-end improvement to the mediated turn over subsystem breadth.
2. An Enabling item must name the Core outcome it unblocks and the decision/witness that consumes it.
3. Stewardship stays explicit and can outrank Core work when deadlines, security, reliability, or repository integrity require it.
4. Peripheral work is visible, not disparaged; it receives capacity only after central outcomes and obligations, or after new evidence promotes it.
5. No benchmark-only speedup is a shipped gain without net-true end-to-end evidence.
6. No dogfood, fleet, scoring, or process verb becomes Core solely because it is implemented inside fak.
7. New breadth must either strengthen the real kernel path in the same spine or be filed as a separate, classified follow-on.

## Current evidence and uncertainty

The repository is 54 days old at this decision, with 11,629 commits, 646 `internal/` packages, 1,881 Go files under `cmd/fak`, and 1,631 open issues observed on 2026-08-15. The live centrality audit is the reproducible portfolio witness; raw counts are context, not proof that a capability is unnecessary.

This decision does **not** yet prove which individual legacy issues should close. Most predate the doctrine, so `unknown` is the honest state until sampled or migrated. The canonical 2026-08-15 audit found 5/1,631 explicitly classified, 1,626 unknown, and zero complete canonical P1-P4 frames; these live counts will change and should be regenerated rather than copied as current truth. The immediate operating goal is coverage and decision quality, not ceremonial mass-editing.

## Repeat the audit

```bash
fak orientation
fak orientation --json
fak centrality-audit --repo anthony-chaudhary/fak
fak centrality-audit --repo anthony-chaudhary/fak --json
fak centrality-audit --repo anthony-chaudhary/fak --sample > active-sample-ledger.json
fak centrality-audit --input open-issues.json --selections selected-migrations.json > migration-preview.json
```

The command reports explicit Core, Enabling, Stewardship, Peripheral, unknown, and complete P1-P4 counts. Each issue is typed `valid`, `invalid`, or `unclassified` with canonical reason and repair tokens. It does not score or reorder issues. `--sample` deterministically includes every open `priority/P0`, every open `gen/now`, and the most recently updated not-already-selected milestone-less issue from each declared capability-family label (lowest issue number breaks timestamp ties). The ledger records selection reasons and a conservative `retain`, `reframe`, or `unknown-with-missing-evidence` decision; labels choose strata but never infer centrality. `--selections` remains no-write: it accepts only explicitly numbered, evidence-backed classifications and emits every original and replacement body. It refuses blank evidence, duplicates, issues outside the audited input, invalid frames, and issues that already declare a frame; there is no blanket rewrite or metadata inference. Applying a reviewed patch and independently reading it back from GitHub remain separate explicit operator actions. [#6543](https://github.com/anthony-chaudhary/fak/issues/6543) owns canonical intake parsing; [#6544](https://github.com/anthony-chaudhary/fak/issues/6544) owns selection-surface propagation.

## Next bounded wave

1. Land this decision and the repeatable coverage audit.
2. Complete #6543 so new contracts cannot silently omit the frame.
3. Complete #6544 so centrality is visible beside urgency, readiness, dependency, and obligation.
4. Sample active `gen/now` and milestone-less work by capability family; retain, reframe, merge, defer, or close with evidence.
5. Revisit this record after that sample and after the next measured kernel-path spine; change classifications when evidence changes.

---

# End-to-End Inference, Agent Harness, and Memory

> Source: `docs/courses/end-to-end-inference-agent-harness-memory.md`

---
title: "End-to-end inference, agent harness, and memory course"
description: "A bounded course following one fak request through native inference, agent tools, policy, context admission, memory recall, and verification."
---

# End-to-End Inference, Agent Harness, and Memory

> **Flagship bounded course.** Follow one fak request end to end. Trace runtime admission, fak-native model execution, and agent tool use. Inspect policy checks and local fast paths. Then examine context admission, memory recall, and durable verification. This page is a syllabus and lab guide; the linked canonical documents remain authoritative.

## Course contract

### Audience

This course is for engineers who can read Go and operate a command line and who want to build, integrate, review, or run a modern fak deployment. It is suitable for agent-harness authors, inference engineers, security and policy reviewers, and operators responsible for evidence-backed releases.

It is not an introduction to transformers, Go syntax, Git, or shell basics. It also does not replace the subsystem references, benchmark authority, or private lab runbooks.

### Prerequisites and background

Before starting, you should be able to:

- Build and test a Go module.
- Read JSON receipts and distinguish a claim from an observed artifact.
- Explain prompts, tokens, and model weights at a basic level. Also understand prefill, decode, and a KV cache.
- Explain why untrusted tool output must not automatically become trusted model context.
- Use explicit paths in a peer-dirty checkout.

Start with [Getting Started](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md), the [reproduction packet](https://github.com/anthony-chaudhary/fak/blob/main/docs/repro-packet.md), and the [architecture overview](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md). Keep the [CLI reference](https://github.com/anthony-chaudhary/fak/blob/main/docs/cli-reference.md) open throughout the course.

### Environment markers

Every lab carries one of these markers:

- **LOCAL / OFFLINE** — no API key, model, or GPU is required.
- **LOCAL MODEL (ENVELOPE-DEPENDENT)** — a supported local fak-native backend and compatible model artifact are required. Inspect admission before loading a payload.
- **FLEET GPU / PRIVATE LAB** — accelerator evidence must run on a sanctioned compute node. The workstation serves as the control point rather than the compute boundary. Follow the public [fleet compute node map](https://github.com/anthony-chaudhary/fak/blob/main/docs/fleet-compute-nodes.md) and enter private infrastructure only through the [private communications channel](https://github.com/anthony-chaudhary/fak/blob/main/docs/private-comms-channel.md). Never copy credentials, private hostnames, private paths, or raw internal logs into public course work.

Absence of local GPU hardware is not a terminal result. Dispatch the GPU lab or record the exact sanctioned handoff; do not substitute llama.cpp and call the result fak-native.

### Measurable outcomes

By the end, you will be able to:

1. Distinguish executable, control-plane, and model-execution readiness from a `fak-runtime-capabilities/1` or execution-mode receipt.
2. Prove that native/performance evidence names `engine: "fak-native"`, a current Qwen3.8 model, a backend, and an explicit operating envelope.
3. Trace a typed tool call through the harness and syscall layer. Follow adjudication, dispatch or vDSO, result admission, and model context.
4. Explain policy precedence, structural denial, and why policy success does not prove model execution.
5. Distinguish provider prompt cache, radix-prefix reuse, and model KV cache. Compare those with vCache and tool-result vDSO reuse.
6. Describe the context-MMU write-time lifecycle: admit, quarantine/page out, restore under a witness, and evict stale or poisoned state.
7. Construct and explain a bounded `memq` recall/render/compact plan. Include trust gates and proposal-only mutation defaults.
8. Produce a compact evidence packet. Separate delivery, quality, and performance claims from security and operating-envelope claims.
9. Run the scope-correct local verification workflow without mixing peer WIP into the result.

## Conceptual system map

```text
operator / application
        |
        v
agent harness + model loop
        |  typed ToolCall
        v
fak Syscall / Submit-Reap seam
        |
        +--> preflight grammar + policy/adjudicator fold --> DENY / DEFER / ALLOW
        |
        +--> local vDSO/tool-result fast path -----------> local result, when valid
        |
        +--> registered tool/engine dispatch -----------> external or in-process effect
                                                               |
                                                               v
context-MMU result admission <--- provenance / witness / taint / revocation
        |
        +--> admitted result enters bounded model context
        +--> quarantined bytes page out behind a small pointer
        |
        v
fak-native Qwen3.8 execution within declared quality + resource envelope
        |
        +--> shared prefix / provider cache / radix reuse / model KV cache
        +--> context lifecycle: retain, shed, compact, reset, restore
        +--> memq memory: scan -> filter -> rank -> budget -> render
                              mutation steps propose by default
        |
        v
journals + receipts + traces + captured witnesses -> independent verification
```

Three distinctions govern the whole course:

- **Control is not execution.** A policy verdict can be real while no model weights were loaded.
- **Context is not memory.** Context is the model-visible working set; durable memory is stored, selected, provenance-bearing state. Read [Context Is Not Memory](https://github.com/anthony-chaudhary/fak/blob/main/docs/CONTEXT-IS-NOT-MEMORY.md).
- **A cache name is not a cache identity.** Use the taxonomy in [Cache concept disambiguation](https://github.com/anthony-chaudhary/fak/blob/main/docs/concepts/disambiguation-cache.md) before reasoning about hits, savings, or invalidation.

## Module 1 — Establish the local control-plane proof

### Read

- [Getting Started](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md)
- [The 60-second reproduction packet](https://github.com/anthony-chaudhary/fak/blob/main/docs/repro-packet.md)
- [`runtime-capabilities` in the CLI reference](https://github.com/anthony-chaudhary/fak/blob/main/docs/cli-reference.md#runtime-capabilities-inspect-the-deployable-runtime-before-payload-load)

The first objective is deliberately model-free. Establish that the binary runs and that policy can structurally deny one tool while allowing another. Confirm that the offline agent completes useful work. Then inspect runtime capability without confusing that inspection with payload execution.

### Lab — LOCAL / OFFLINE

From the repository root, use a temporary output path rather than writing an in-tree binary:

```powershell
go build -o "$env:TEMP/fak-course.exe" ./cmd/fak
& "$env:TEMP/fak-course.exe" preflight --policy examples/customer-support-readonly-policy.json --tool refund_payment --args "{}"
& "$env:TEMP/fak-course.exe" preflight --policy examples/customer-support-readonly-policy.json --tool search_kb --args "{}"
& "$env:TEMP/fak-course.exe" agent --offline
& "$env:TEMP/fak-course.exe" runtime-capabilities
```

Capture the commands, exit status, and machine-readable output. Label each observation as binary, control plane, or model execution evidence.

### Checkpoint

Submit a table with one row per command and answer:

- Which command proves a structural denial?
- Which command proves fak is not a blanket blocker?
- Which output, if any, proves that model weights executed?
- Why is a runtime-capability projection not itself a live inference receipt?

Proceed only if you do not claim model execution from the offline proof.

## Module 2 — Admit fak-native Qwen3.8 execution and its quality envelope

### Read

- [Native inference goal and non-negotiable contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/native-inference-goal.md)
- [Benchmark authority](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md)
- [Native performance observability contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/native-performance-contract.md)
- [Native performance artifact guide](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/native-performance-artifacts.md)

Modern native/performance work prefers Qwen3.8. llama.cpp is used only for explicit benchmarks, reference comparisons, or borrowing. It is never an automatic recovery or convenience engine. A valid native receipt names fak ownership and the exact model/backend. It records the quality constraint, workload, and resource boundary. It also includes all setup, recovery, and verification overhead.

### Lab — LOCAL MODEL (ENVELOPE-DEPENDENT), then FLEET GPU / PRIVATE LAB when required

1. Inspect the available execution envelope before loading a payload:

   ```powershell
   fak runtime-capabilities --receipt-schema fak-execution-mode-receipt/1
   ```

2. If the task names a backend, inspect it exactly with `--backend NAME`; do not accept substitution.
3. For a native/performance run, use the canonical command and receipt named by the relevant Qwen3.8 plan or benchmark runbook. Dispatch accelerator work through the [fleet compute node map](https://github.com/anthony-chaudhary/fak/blob/main/docs/fleet-compute-nodes.md).
4. Record the exact engine, Qwen3.8 model/artifact, and backend/device. Also note the prompt/workload, concurrency, and context length. Include the quality gate, memory limit, warm/cold state, and overhead accounting.
5. If the envelope cannot be admitted, preserve the structured refusal or handoff. Do not silently run llama.cpp, Ollama, or another provider. Do not run `cpu-ref` or another model and relabel it native.

### Checkpoint

A reviewer must be able to answer yes to all of these from your receipt alone:

- Does `engine` equal `fak-native`?
- Is the Qwen3.8 model and artifact identity explicit?
- Is the backend/device explicit rather than inferred?
- Is quality constrained and measured beside performance?
- Are context and batch/concurrency in the envelope? Are memory, setup, recovery, and verification costs included?
- Is any degraded or remote placement explicit and authorized?

Otherwise mark the execution or performance claim as not yet witnessed.

## Module 3 — Follow the agent harness and syscall path

### Read

- [Architecture and extension model](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md).
- [MCP integration](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/mcp.md).
- [Harness composition](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/harness-composition.md).
- [Harness verification runs](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/harness-verify-run.md).
- [MCP stdio example and verifier](https://github.com/anthony-chaudhary/fak/blob/main/examples/mcp/README.md). This demonstrates schema-light bootstrap discovery, deferred `fak_admit` discovery, benign result admission, and allow/deny adjudication over the real editor transport.

The harness owns the model loop and exposes tools. fak owns the typed call boundary. Synchronous `Syscall` is defined over asynchronous `Submit`/`Reap`. Registries provide adjudicators, fast paths, and engines. They also supply result admitters and observers without making every new feature part of the hot path.

### Lab — LOCAL / OFFLINE

1. Run the offline harness from Module 1.
2. Read the `ToolCall`, syscall, and result structures named in the architecture document. Locate their current definitions under `internal/abi`, `internal/gateway`, and `internal/agent`.
3. Draw a sequence diagram for one allowed call and one denied call. Include the caller, harness, syscall boundary, and adjudicator fold. Show the vDSO lookup and dispatch engine. Also include the result admitter, context-MMU, observer, and model loop.
4. Run `python examples/mcp/verify.py --no-color`. Confirm all six checks pass. These include the handshake, schema-light `tools/list`, and deferred `fak_admit` discovery through `fak_tools_search`. They also cover benign result admission with a typed DEFER/OK envelope, `git_push` denial, and `git_status` allowance.

### Checkpoint

Explain, without implementation narration:

- Where a tool call becomes typed kernel input.
- Where a denial stops side effects.
- Why `Syscall` over `Submit`/`Reap` preserves one policy path.
- Where result bytes can be stopped even after a tool was allowed.
- Which components are harness responsibilities and which are kernel responsibilities.

## Module 4 — Policy, grammar, and adjudication

### Read

- [Policy reference](https://github.com/anthony-chaudhary/fak/blob/main/POLICY.md)
- [Customer-support policy example](https://github.com/anthony-chaudhary/fak/blob/main/examples/customer-support-readonly-policy.json)
- [Architecture: adjudicator registration and fold](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md#how-a-new-idea-bakes-in-the-only-mechanism)

Policy serves as a deployable capability floor rather than a prompt suggestion. The preflight ladder combines tool registration, argument grammar, capability checks, and ordered adjudicators. Provable refusals deny. An adjudicator that cannot prove its case defers rather than inventing authority.

### Lab — LOCAL / OFFLINE

Repeat the two preflight calls from Module 1. Then add one malformed or unregistered call of your own using only flags documented under `fak preflight` in the CLI reference. Preserve each typed reason and identify the deciding rung.

Do not execute a real destructive tool for this exercise. Preflight exists to prove the refusal before dispatch.

### Checkpoint

For each case, provide:

- The normalized tool and arguments.
- The deciding rung and structured reason.
- The `ALLOW`, `DENY`, or defer behavior.
- An assessment of whether a side effect could have occurred.
- The distinction between a policy verdict and context admission of a later result.

## Module 5 — Shared prefix, vDSO, and the cache stack

### Read

- [Cache concept disambiguation](https://github.com/anthony-chaudhary/fak/blob/main/docs/concepts/disambiguation-cache.md)
- [Cache explainer](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/cache.md)
- [Long sessions keep the cache hit](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/long-sessions-keep-the-cache-hit.md)
- [Tool vDSO three-tier fast path claim](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/tool-vdso-3-tier-local-fast-path.md)
- [Addressable KV cache in five minutes](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/addressable-kv-cache-in-5-min.md)

Shared prefixes reduce repeated provider or prefill work only when byte identity and the relevant cache contract hold. Tool-result vDSO reuse is a different mechanism from provider prompt caching, radix-prefix snapshots, model attention KV, and vCache. Invalidation, provenance, and economic accounting differ across them.

### Lab — LOCAL / OFFLINE

Build a cache ledger for one hypothetical two-turn agent session. For each reusable item, name:

- Its cache identity.
- Its owner and trust boundary.
- Its key or prefix identity.
- Its hit eligibility.
- Its invalidation or revocation condition.
- What is saved: network/tool work, provider input processing, prefill compute, or decode state.
- The evidence that would prove a hit rather than infer one.

Then inspect `internal/vdso` tests and run:

```powershell
go test ./internal/vdso
```

### Checkpoint

You are given five observations: provider cached-input tokens, a radix-prefix match, and a model KV page. You also observe a vCache entry and a local tool-result hit. Classify each observation without using the generic word “cache” alone. Reject any savings claim that lacks the corresponding hit evidence and denominator.

## Module 6 — Context-MMU lifecycle and trust-preserving context control

### Read

- [Context-MMU write-time result admission](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/context-mmu-write-time-result-admission.md)
- [Managed context continuous usage](https://github.com/anthony-chaudhary/fak/blob/main/docs/managed-context-continuous-usage.md)
- [Context shedding](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/context-shedding.md)
- [Context and ctx disambiguation](https://github.com/anthony-chaudhary/fak/blob/main/docs/concepts/disambiguation-context-ctx.md)

The context-MMU screens results before they enter model-visible context. Safe bytes may be admitted. Harmful patterns like secrets, injections, poison, or repeats can be quarantined and replaced by a small content-addressed pointer. Restore is witness- and trust-gated. Revocation must evict causally dependent state rather than leaving poisoned reuse behind.

### Lab — LOCAL / OFFLINE

1. Read the context-MMU claim’s reproduction command and run the current `internal/ctxmmu` package tests:

   ```powershell
   go test ./internal/ctxmmu
   ```

2. Trace these states for both a safe result and a poison-shaped result. Follow produced, screened, and admitted or quarantined stages. Then follow paged out, pointer rendered, and restore attempted. Finally trace restored or refused, followed by revoked or evicted.
3. Explain what remains durable when content leaves the active context and what does not automatically become memory.

### Checkpoint

Your state diagram must show:

- A write-time trust decision before model visibility.
- A bounded pointer replacing quarantined bytes.
- A witness requirement for page-in.
- A refusal path for sealed, tombstoned, stale, or revoked content.
- Separation between context pressure management and durable memory selection.

## Module 7 — memq recall, render, and compact under trust gates

### Read

- [Memory engineering](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/memory-engineering.md)
- [Agent memory integration](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/agent-memory.md)
- [Context Is Not Memory](https://github.com/anthony-chaudhary/fak/blob/main/docs/CONTEXT-IS-NOT-MEMORY.md)
- Current implementation and tests under `internal/memq` and `internal/recall`

`memq` is a composable query algebra rather than a magical memory oracle. A plan can scan, filter, rank, and limit. It can also budget, deduplicate, and render items. Operators can tombstone, consolidate, reclassify, or prune records. Recall and render are bounded projections through provenance and trust gates. Mutating operations are proposal-only by default. Sealed spans are never rendered, and negative-only/storage mutations require explicit application.

### Lab — LOCAL / OFFLINE

1. Run the package tests:

   ```powershell
   go test ./internal/memq ./internal/recall
   ```

2. Using the current memory driver/tool schema documented in [Agent memory integration](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/agent-memory.md), explain a built-in `recall` plan before running it.
3. Run a bounded recall against the demo corpus with a concrete intent, small `k`, and byte budget.
4. Explain a `compact` plan with `apply=false`. Identify every effect step and why it does or does not mutate durable state.
5. Do not set `apply=true` for the course. The learning objective focuses on proposed effects and trust refusals without altering a learner's memory store.

### Checkpoint

Submit the plan trace and rendered set, then answer:

- Which operator selected each item?
- Where was the byte budget enforced?
- Which provenance/trust state prevented rendering?
- Which compact effects were merely proposed?
- Why is a tombstone preferable to an unaudited hard delete?
- What evidence would be required before a stored memory is presented as fresh fact?

## Module 8 — Durable evidence, observability, and verification workflow

### Read

- [Observability map](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/README.md)
- [Durable artifacts](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/durable-artifacts.md)
- [Development tooling and build profiles](https://github.com/anthony-chaudhary/fak/blob/main/docs/dev-tooling.md)
- [Benchmark authority](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md)
- [Claims registry](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md) and [status matrix](https://github.com/anthony-chaudhary/fak/blob/main/STATUS.md)

Observability is useful only when its records survive the run and preserve provenance. A trace helps diagnose issues. A receipt binds a claim to an operating envelope. A captured witness lets an independent reviewer reproduce or reject the claim. “Tests pass,” “fast,” and “native” are separate claims requiring different evidence.

### Lab — LOCAL / OFFLINE

Create an evidence matrix for Modules 1–7 with these columns:

- Claim.
- Controlling subsystem.
- Artifact or command.
- Observed versus self-reported status.
- Environment and envelope.
- Expected failure or refusal state.
- Independent verification step.

Run scope-correct checks for course-relevant code or docs without treating the peer-dirty working tree as your change:

```powershell
fak-dev ci-preflight
fak validate --mine docs/courses/end-to-end-inference-agent-harness-memory.md docs/_witnesses/issue-10424-learning-before.json
```

If `fak validate --mine` is unavailable in the current bootstrap state, record that honestly and use the documented isolated build/test primitive rather than a broad in-place build.

### Operations lesson — update without trusting the binary being replaced

The shipped `cmd/fak-selfupdate` surface is a recovery-sized executable over the same implementation as `fak self-update`. Use the ordinary verb when the installed `fak` is healthy. Keep the standalone entry point available when the main command is stale or is being replaced. Both emit the versioned `fak.self-update.receipt/v1` JSON contract.

The default native path builds from a repository, runs the green gate, and installs only after passing. `--check` provides a non-mutating inspection path, and even `--force` does not bypass the green gate. Optional signed-manifest selection adds channel, cohort, and authenticated-cache controls. In addition, `--offline` refuses network access and accepts only a valid authenticated cache.

On Windows, the explicit MSIX path verifies signed package provenance and requires an explicit opt-in for downgrade. Differential delivery falls back only to the declared full package. These bounded alternatives provide safety without accepting unsigned artifacts or skipping verification. A scheduled updater can also pin the executable path and must refuse provenance drift.

From the repository root, prove the recovery surface is present without replacing anything:

```powershell
go run ./cmd/fak-selfupdate --help
go run ./cmd/fak-selfupdate --check --target (Get-Command fak).Source --json
```

The first command exposes the standalone surface. The second inspects the installed target's embedded Go build metadata without executing that target. Read the exact flags and receipt fields in the [CLI reference](https://github.com/anthony-chaudhary/fak/blob/main/docs/cli-reference.md). Then consult the [self-update fast-path design note](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-SELF-UPDATE-FAST-PATHS-2026-08-29.md) when reasoning about manifest, delta/full-fallback, MSIX, or handoff boundaries.

### Checkpoint

A peer should be able to reconstruct what ran and where it ran. They should identify the engine, model, and policy. They should also verify the envelope boundaries and the observed result—all without trusting your prose. Treat any missing dimension as an explicit limitation instead of an inferred success.

## Capstone — One governed Qwen3.8 agent turn, end to end

### Brief

Design and execute, or dispatch, one bounded agent task that reads from an allowed knowledge tool and proposes—but does not perform—a sensitive action. The run must connect all eight modules.

### Required path

1. **Admit the runtime.** Capture `runtime-capabilities`; distinguish control-only readiness from model execution.
2. **Pin native execution.** Use Qwen3.8 on a fak-native backend. If acceleration is required, dispatch to the sanctioned fleet/private-lab path. Keep native/performance execution fak-native. Select llama.cpp only explicitly for benchmarks, parity/reference diagnosis, study, or borrowing. Never use it as a silent recovery path. Likewise, never silently substitute an external provider, a CPU path, or a different model.
3. **Declare the envelope.** Record the model/artifact, backend/device, and context length. Note the concurrency, memory budget, and workload. Finally, record the quality criterion and overhead accounting.
4. **Compose the harness.** Show the model loop, exposed tool schema, and typed syscall boundary.
5. **Prove policy.** Preflight the allowed read and the denied sensitive action; retain typed verdicts.
6. **Account for reuse.** Identify the shared stable prefix. Distinguish vDSO, provider, and radix reuse. Also separate vCache and model-KV evidence without conflating them.
7. **Admit results.** Show the context-MMU decision for the returned data, including a quarantine case if the fixture contains poison-shaped text.
8. **Use memory safely.** Run a bounded `memq` recall/render plan and a proposal-only compact plan; preserve trust refusals.
9. **Capture evidence.** Produce a manifest that points to commands, receipts, and traces. Include quality results, policy verdicts, and context decisions. Link memory plan traces and verification output as well.
10. **Verify independently.** Have a reviewer reproduce the local control-plane slice and inspect the native receipt and private-lab/public evidence boundary.

### Capstone acceptance rubric

The capstone passes only when all conditions are met:

- The task completes or reaches a typed refusal without bypassing policy.
- Native model evidence explicitly names fak-native Qwen3.8 and the backend.
- Quality and performance are reported in the same declared envelope.
- The harness-to-syscall path and result-admission path are both evidenced.
- Every claimed cache hit names its cache identity and witness.
- Quarantined content does not appear in model-visible context.
- Memory output is bounded, provenance-bearing, and trust-gated.
- Private infrastructure details remain private while the public receipt remains reproducible.
- Local verification uses isolated, scope-correct commands.
- Limitations are stated as limitations rather than converted into success.

A control-only demonstration is a valid partial artifact but not a completed native-inference capstone. A llama.cpp reference run may be attached as an explicitly labeled comparison, but it cannot satisfy the fak-native execution requirement.

## Next routes

Choose the route that matches the next responsibility:

- **Operate or integrate a harness:** [integration index](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/README.md), [adopter playbook](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/adopter-playbook.md), and [harness acceptance checklist](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/harness-acceptance-checklist.md).
- **Extend the kernel:** [Extending fak](https://github.com/anthony-chaudhary/fak/blob/main/EXTENDING.md) and the registry seams in [Architecture](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md).
- **Optimize native inference:** [native inference goal](https://github.com/anthony-chaudhary/fak/blob/main/docs/native-inference-goal.md), [SOTA index](https://github.com/anthony-chaudhary/fak/blob/main/docs/sota/README.md), and [benchmark authority](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md). Keep Qwen3.8 and fak-native ownership explicit.
- **Deepen cache work:** [cache explainer](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/cache.md), [managed cache](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/what-is-managed-cache.md), and [cache value rollup](https://github.com/anthony-chaudhary/fak/blob/main/docs/cache-value-rollup.md).
- **Deepen context and memory:** [managed context glossary](https://github.com/anthony-chaudhary/fak/blob/main/docs/managed-context-glossary.md), [memory engineering](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/memory-engineering.md), and [agent memory integration](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/agent-memory.md).
- **Build durable observability:** [observability map](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/README.md), [trajectory observability](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/trajectory.md), and [durable artifacts](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/durable-artifacts.md).
- **Run hardware-gated work:** [fleet compute nodes](https://github.com/anthony-chaudhary/fak/blob/main/docs/fleet-compute-nodes.md) and the [private communications channel](https://github.com/anthony-chaudhary/fak/blob/main/docs/private-comms-channel.md). Keep the public/private boundary intact.

## Completion record

A learner's final record should contain:

- The eight module checkpoints.
- The capstone manifest and acceptance review.
- An explicit list of labs completed locally, with a local model, or on a fleet/private-lab node.
- All refusals and missing witnesses preserved verbatim.
- No copied canonical documentation beyond the minimum command snippets needed to perform the labs.

The course is complete when another engineer can follow the evidence from operator intent to verified outcome. They must be able to identify where fak controlled, accelerated, or refused the run. They must also see where it quarantined, recalled, or merely observed behavior.

---

# `LEARNING-PATH.md`

> Source: `LEARNING-PATH.md`

---
title: "The fak Learning Path — a prerequisite-ordered course"
description: "A linear, prerequisite-based curriculum across every fak concept: 99 courses in six levels, from \"what is fak\" to landing an optimization in the kernel. Join at the level that matches your background and walk straight through."
---

# The fak learning path

*This page owns one job: teaching fak's ideas in prerequisite order. It is for a reader who
wants to understand the system, not evaluate it ([README](https://github.com/anthony-chaudhary/fak/blob/main/README.md)), install it
([GETTING-STARTED](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md)), route a task ([START-HERE](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md)), or look
a page up by name ([INDEX](https://github.com/anthony-chaudhary/fak/blob/main/INDEX.md)). Nothing here is required to run fak.*

fak is a lot of ideas stacked into one binary: an addressable KV cache that keeps long
sessions cheap to hold warm, right-model-per-call routing, and a pure-Go in-kernel model
(preferring Qwen3.8 for native-performance work per AGENTS.md) —
and, riding along on the same write-time checkpoint, a default-deny capability floor, a
write-time result quarantine, and the honesty discipline that keeps every claim checkable.
This page turns all of it into one **linear, prerequisite-ordered curriculum** — a course
catalog, not a doc dump.
Each course points at the doc that already teaches it; the value added here is the
**order** and the **prerequisites**, so you always have the background a page assumes
*before* you open it.

> **Want the integrated system story first?** Take the
> [8-module flagship course](https://github.com/anthony-chaudhary/fak/blob/main/docs/courses/end-to-end-inference-agent-harness-memory.md) to
> follow one request across native inference, the agent harness, policy, context control,
> durable memory, observability, and proof. Then use this 99-course catalog to enter at the
> right prerequisite level or deepen a subsystem. The course is a guided end-to-end route;
> this page remains the prerequisite-ordered concept front door.

**You do not have to start at the beginning.** Find the row in
[Find your starting point](#find-your-starting-point) that matches your background, start
at that course, and walk forward. The catalog is a strict prerequisite order — every
course's *hard* prerequisites are lower-numbered courses — so reading top-to-bottom never
lands you on a concept whose prerequisite you have not met yet.

99 courses, six levels (100 → 600), from "what is fak" to landing an optimization into
the kernel. The readings are the docs you would read anyway; the path is what stops you
reading them in the wrong order.

> New to the project entirely? The fastest taste of the payoff is one offline pass that
> prints the token/turn savings from the shared prefix — `go run ./cmd/fak agent --offline`
> (see **FAK 104**); the same run also prints the safety A/B, which
> [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md#try-fak) frames as the
> security-side boundary proof. Either way, come back here and start at **FAK 101**. Just
> want to install and run? [`GETTING-STARTED.md`](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md) is the install-and-run
> page and [`START-HERE.md`](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md) routes a job you already have; this page is the
> *concept* front door.

## Learn without mysteries

Do not memorize fak's nouns first. Use the same five-step loop whenever you want to
understand or change a process:

1. **Predict** one observable result in plain language.
2. **Run** the smallest real command that can prove you wrong.
3. **Locate** the winning rule in the output and then in the named config or source.
4. **Adjust one input** while keeping everything else fixed.
5. **Rerun** and explain why the result changed. If you cannot, the mechanism is still a
   mystery; follow the winning rule one level deeper before adding machinery.

That loop separates three questions that are easy to conflate:

- **What happened?** Read the verdict or measurement.
- **Why did it happen?** Read the winning rung and reason; `DEFER` means “this rung did
  not decide,” not “allow.”
- **What can I change?** Change the input owned by that winning rung, not an unrelated
  downstream setting.

### Five-minute policy lab: predict, inspect, adjust

This lab needs no key, model, or GPU. It uses the real policy fold rather than a diagram.
Start by predicting all three outcomes, then run:

```powershell
fak preflight --policy examples/customer-support-readonly-policy.json --tool refund_payment --args "{}" --explain
fak preflight --policy examples/customer-support-readonly-policy.json --tool search_kb --args "{}" --explain
fak preflight --policy examples/customer-support-readonly-policy.json --tool mystery_action --args "{}" --explain
```

The expected winners are deliberately different:

| Tool | Result | Plain-language reason | Policy input to inspect |
|---|---|---|---|
| `refund_payment` | `DENY / POLICY_BLOCK` | The manifest names this tool in `deny`. | `deny.refund_payment` |
| `search_kb` | `ALLOW` | The tool matches the `search_` allow prefix. | `allow_prefix` |
| `mystery_action` | `DENY / DEFAULT_DENY` | No rule grants the unknown tool. | `posture: fail_closed` |

The same method holds when the input is unfriendly. These are argument-rule cases, not
global claims that every tool needs arguments or that every call has one fixed size cap:

| Input class | One input change | Expected result | Winning rule |
|---|---|---|---|
| Empty required argument | Remove a value required by a positive path constraint. | `DENY / POLICY_BLOCK` | The `allow_glob` rule cannot prove that a missing value is contained. |
| Oversized string argument | Grow one scalar past its policy-declared byte bound. | `DENY / OVERSIZE` | The matching `max_bytes` argument rule. |
| Malformed quote-wrapped argument | Open a quote in a constrained value without closing it. | `DENY / MALFORMED` | Canonicalization fails closed and returns a bounded repair hint. |
| Hostile denied argument | Make one scalar match its policy-declared deny expression. | `DENY / POLICY_BLOCK` | The matching `deny_regex` rule; the witness names the rule, not the input. |

`TestMysteryFreeAdjustmentEdgeAdversarial` drives all four rows through the real
adjudicator and checks that both learning documents keep the same table. Run the full
edge/adversarial witness with:

```powershell
go test ./internal/adjudicator -run 'Edge|Adversarial' -v
```

Now change exactly one policy fact in a temporary copy and predict the new result:

```powershell
$p = Get-Content examples/customer-support-readonly-policy.json | ConvertFrom-Json
$p.allow += "mystery_action"
$tmp = Join-Path $env:TEMP "fak-learning-policy.json"
$p | ConvertTo-Json -Depth 10 | Set-Content $tmp
fak preflight --policy $tmp --tool mystery_action --args "{}" --explain
```

The last verdict should be `ALLOW`: the same call now has one affirmative policy rule.
The original manifest is untouched. To continue investigating, use the trace's winner:
`adjudicator.Adjudicator` leads to [`internal/adjudicator`](https://github.com/anthony-chaudhary/fak/tree/main/internal/adjudicator), while the
manifest format and safe editing workflow live in
[`docs/fak/policy-guide.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/policy-guide.md). This is the pattern for every later
course: predict → run → locate → adjust one thing → rerun. For a cross-process map of the deciding evidence and owned knobs, use the [`mystery-free adjustment atlas`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mystery-free-adjustment-atlas.md).

## How to read a course

Each course is one entry shaped like a syllabus line:

- **Prerequisites** — *hard* dependencies. These will block this course's lab or
  checkpoint if you skip them, and they are always lower-numbered, so they sit above this
  course in the catalog.
- **Background** — *context* prerequisites: helpful framing you can defer. Skipping them
  costs you some "why", not the ability to do the lab.
- **You'll be able to** — the concrete skills the course certifies.
- **Read** — the canonical doc(s). This is the actual course material.
- **Lab** — a command you can run (most need no key, model, or GPU) or a hands-on task.
- **Checkpoint** — answer it (or do it) to certify yourself before moving on. If you can
  clear a level's checkpoints, you have met the `assumed_passed` bar for the next level.

Honesty carries through the whole catalog: where a number is **SIMULATED** or a proof is
**OPEN**/**REFUTED**, the checkpoint says so. The headline multipliers are stated against
the *naive* baseline and the *tuned-SOTA* baseline separately, never blended — see
**FAK 605**.

## Find your starting point

Start at the course in the **Start** column, then follow the **Route** straight through
to the destination. The route already lists every hard dependency in between, in order —
so you can join mid-catalog without hitting a wall. Anyone can also just start at
**FAK 101** and read every course in number order.

| Your background | Start | Route (in order) → destination |
|---|---|---|
| Total newcomer — knows what an AI agent and a tool call are, nothing else | **FAK 101** | FAK 101 → FAK 102 → FAK 103 → FAK 104 → FAK 105 |
| App dev who only calls an LLM API and wants governance with minimal agent rewrite | **FAK 101** | FAK 101 → FAK 102 → FAK 103 → FAK 104 → FAK 105 → FAK 207 → FAK 301 → FAK 310 → FAK 501 → FAK 502 → FAK 503 → FAK 511 |
| Platform / SRE who already runs vLLM or SGLang in production | **FAK 201** | FAK 201 → FAK 103 → FAK 207 → FAK 301 → FAK 303 → FAK 304 → FAK 310 → FAK 501 → FAK 502 → FAK 503 → FAK 504 → FAK 505 → FAK 507 → FAK 314 → FAK 510 → FAK 535 |
| Security engineer who already knows prompt injection, default-deny, reference monitors | **FAK 105** | FAK 105 → FAK 207 → FAK 103 → FAK 301 → FAK 302 → FAK 303 → FAK 304 → FAK 305 → FAK 306 → FAK 307 → FAK 308 → FAK 309 → FAK 310 → FAK 311 → FAK 312 → FAK 313 → FAK 314 → FAK 315 → FAK 318 |
| ML-systems / kernel hacker who wants the in-kernel model and compute HAL | **FAK 201** | FAK 201 → FAK 205 → FAK 207 → FAK 210 → FAK 401 → FAK 521 → FAK 522 → FAK 523 → FAK 524 → FAK 525 → FAK 526 → FAK 404 → FAK 405 → FAK 406 → FAK 527 → FAK 528 → FAK 529 → FAK 530 → FAK 532 |
| Memory / RAG engineer focused on what fak persists, forgets, and reuses | **FAK 202** | FAK 202 → FAK 203 → FAK 201 → FAK 205 → FAK 207 → FAK 301 → FAK 303 → FAK 310 → FAK 316 → FAK 307 → FAK 407 → FAK 409 → FAK 402 → FAK 401 → FAK 412 → FAK 413 → FAK 414 |
| Compliance / audit / governance engineer (journal, provenance, deletion, honesty discipline) | **FAK 105** | FAK 105 → FAK 207 → FAK 103 → FAK 301 → FAK 303 → FAK 310 → FAK 311 → FAK 312 → FAK 313 → FAK 314 → FAK 315 → FAK 317 → FAK 404 → FAK 405 → FAK 406 → FAK 411 → FAK 601 → FAK 602 → FAK 606 → FAK 614 → FAK 307 → FAK 616 |
| Contributor / autonomous agent landing an optimization into the kernel | **FAK 207** | FAK 207 → FAK 208 → FAK 209 → FAK 210 → FAK 614 → FAK 615 → FAK 616 → FAK 617 |

> The **Route** is the *hard-dependency* path. You can read the context prerequisites
> noted on each course later (or never) without breaking a lab.

## The level ladder

Read the ladder performance-first: L400's cache reuse and the in-kernel model are the
headline payoff, and the L300 security floor rides the same write-time checkpoint rather
than sitting on a separate path.

```
L100  Orientation .................. what fak is, the one idea, the two gates      (start cold)
  |
L200  Foundations .................. KV cache, context != memory, content addressing,
  |                                  the frozen ABI, the proofs method
  +--> L300  Security Core ......... the in-process default-deny floor + the write-time wall
  +--> L400  Performance Core ...... cache reuse, addressable eviction, the scaling laws
            |
            +--> L500  Serving / Integration / In-Kernel Model
                       run & harden the gateway, repoint one base URL, the pure-Go model + HAL
                       |
                       +--> L600  Mastery .. benchmarks, the honesty discipline, extend the kernel
```

Each level states the courses it assumes you can already pass. If you can clear those
checkpoints, you are qualified to start there.

| Level | Theme | Assumes you can pass |
|---|---|---|
| **L100 — Orientation** | The plain category, the syscall framing, the two gates, the recurring vocabulary, and how to prove the boundary is real in two minutes. | — (start cold) |
| **L200 — Foundations** | The handful of mechanisms every later claim rests on: the KV cache, context-vs-memory durability, the four memory layers, content addressing, the frozen ABI, and the proofs method. | FAK 101, FAK 102, FAK 103, FAK 104, FAK 105 |
| **L300 — The Security Core** | The reference monitor, the policy lifecycle, the rungs (preflight, plan-CFI, witness, stewards, rate-limit, escalation), the write-time result gate, canonicalization, IFC, provenance, durability, and code-linting at the same boundary. | FAK 105, FAK 207 |
| **L400 — The Performance Core** | Why agents stress the cache, prefill-elimination economics, the addressable/bijective KV-MMU, RadixAttention reuse, the vDSO, durable session recall, and the first-order scaling law (incl. cache legality and residency). | FAK 201, FAK 205, FAK 310 |
| **L500 — Serving, Integration, and the In-Kernel Model** | Running and hardening the gateway (`fak serve`, `fak l3-serve` / `fak l3serve`), the gateway drop guarantee, repointing existing agents at one base URL, the framework cookbook, the pure-Go in-kernel model + compute HAL with oracle parity (preferring Qwen3.8 for native performance per AGENTS.md), and the GPU lease. | FAK 105, FAK 301, FAK 304, FAK 310 |
| **L600 — Mastery** | Honest baselines and the benchmark authority, the fleet/web/parity results, the AgentDojo red-team, the claims ledger and status gates, the additive ABI + architest, the RSI ship-gate, the three-gate leaf pattern, and the dispatch loop. | FAK 207, FAK 208, FAK 209, FAK 210 |

---

## The catalog

## L100 — Orientation: what fak is and the one idea

**Theme.** The plain category, the syscall framing, the two gates, the recurring vocabulary, and how to prove the boundary is real in two minutes.

**Who joins here.** A total newcomer, or anyone who has never seen fak. You only need to know what an AI agent is, what a tool call is, and roughly what a model server (vLLM, llama.cpp) does. Start here if any of fak's one-liners ('untrusted program', 'two gates', 'security == reuse') are not yet obvious to you.

| Course | Hard prerequisites |
|---|---|
| **FAK 101** — What fak Is: One Binary Between Agent and Tools | — |
| **FAK 102** — The Core Move: Untrusted Program, Tool-Call-as-Syscall (and the Word List) | **FAK 101** |
| **FAK 103** — The Parable and the Two Gates | **FAK 102** |
| **FAK 104** — The Convergence: Security Boundary == Reuse Boundary | **FAK 103** |
| **FAK 105** — Adoption Rungs and the 2-Minute Honest Proof | **FAK 104** |

### FAK 101 — What fak Is: One Binary Between Agent and Tools

**Prerequisites:** —

**You'll be able to:**
- State in one sentence what fak is and name one thing it explicitly is NOT (it is not a faster model server)
- Name two of the four questions fak owns that a token engine leaves open
- Build the single binary and print its version

**Read:** [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md), [`START-HERE.md`](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md), [`docs/FAQ.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/FAQ.md)

**Lab:**
```bash
go run ./cmd/fak version          # confirm the single binary builds and prints its version
go run ./cmd/fak version modules -top 10  # version-everything: per-module rev+date over the tree (the fak-module-versions/1 ledger, 410 modules)
```

**Checkpoint:** In one sentence, state what fak is and name one thing it is explicitly NOT. Name two of the four questions fak owns that token engines leave open.

### FAK 102 — The Core Move: Untrusted Program, Tool-Call-as-Syscall (and the Word List)

**Prerequisites:** **FAK 101**

**You'll be able to:**
- Reframe the model as an untrusted program and each tool call as a syscall on a controlled path
- Explain why an in-process default-deny check differs structurally from a pre-tool hook or a second 'is this safe?' model
- Pin the recurring vocabulary: preflight (before-gates) vs inflight (during-state) vs prefill (KV economics), plus adjudicator/fold/rung/monitor/admit
- Run a denied call and read the DEFAULT_DENY verdict

**Read:** [`docs/concepts-and-story.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/concepts-and-story.md), [`docs/fak/faq.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/faq.md), [`docs/glossary.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/glossary.md), [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md)

**Lab:**
```bash
go run ./cmd/fak preflight --tool create_user --args '{"_positional":["alice"]}'  # adjudicated as a syscall -> DENY DEFAULT_DENY
```

**Checkpoint:** Explain why putting the check ON the same in-process call path (default-deny) is structurally different from a pre-tool hook or an LLM judge. Then disambiguate preflight vs inflight vs prefill, and say what 'the lever was never wired up' means concretely.

### FAK 103 — The Parable and the Two Gates

**Prerequisites:** **FAK 102**

**You'll be able to:**
- Map the night-shift-clerk parable onto fak mechanisms: locked drawer, screened notes, imperfect screener
- Name the two independent gates (the lock/capability floor and the wall/quarantine) and what each protects against
- Explain why the detector on top is treated as evadable by design and why that does not weaken the floor

**Read:** [`docs/concepts-and-story.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/concepts-and-story.md), [`docs/fak/faq.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/faq.md), [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md)

**Lab:**
```bash
go run ./cmd/fak preflight --policy examples/customer-support-readonly-policy.json --tool refund_payment --args '{}'  # the lock: DENY POLICY_BLOCK
```

**Checkpoint:** Name the two gates and what each protects against (effect vs. context entry). Why is an attacker beating TWO gates harder than fooling one classifier, and why is the detector deliberately treated as evadable?

### FAK 104 — The Convergence: Security Boundary == Reuse Boundary

**Prerequisites:** **FAK 103**

**You'll be able to:**
- Explain how one write-time gate is simultaneously a security act and an optimization act
- State the two honest fences on the convergence (which workload, which metric it does NOT win)
- Run one offline pass that prints both the safety A/B and the token/turn savings

**Read:** [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md), [`docs/concepts-and-story.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/concepts-and-story.md)

**Lab:**
```bash
go run ./cmd/fak agent --offline  # one run prints the safety A/B AND the token/turn savings from the same boundary
```

**Checkpoint:** Explain how one write-time gate is both security and optimization. State the two honest fences: which workload it is a win for, and which metric (raw GPU throughput) it does NOT win.

### FAK 105 — Adoption Rungs and the 2-Minute Honest Proof

**Prerequisites:** **FAK 104**

**You'll be able to:**
- List the three adoption rungs (front your model / offline kernel / fused in-kernel model) least-to-most committed and pick a starting rung
- Identify which rung unlocks the reuse win and the self-host fence on it
- Run the 2-minute proof (a structural DENY and an ALLOW) and read the headline numbers against SOTA, not a strawman

**Read:** [`docs/fak/tutorial.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/tutorial.md), [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md), [`START-HERE.md`](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md)

**Lab:**
```bash
go run ./cmd/fak preflight --policy examples/customer-support-readonly-policy.json --tool refund_payment --args '{}' && go run ./cmd/fak preflight --policy examples/customer-support-readonly-policy.json --tool search_kb --args '{}'
```

**Checkpoint:** List the rungs least-to-most committed and say which most adopters should start at and why. Run the proof and report both verdicts; state what ~60x compares against vs ~4x, what is SIMULATED, and what the prior-art audit scored (0/29-novel) and what the contribution actually is.

---

## L200 — Foundations: the load-bearing mechanisms

**Theme.** The handful of mechanisms every later claim rests on: the KV cache, context-vs-memory durability, the four memory layers, content addressing, the frozen ABI, and the proofs method.

**Who joins here.** Someone comfortable with the orientation framing who wants the underlying mechanics. Join here if you already know fak is a governing binary and want to understand the KV cache, content-addressed stores, and how the repo proves things before you touch the security or performance cores.

**Assumes you can already pass:** **FAK 101**, **FAK 102**, **FAK 103**, **FAK 104**, **FAK 105**.

| Course | Hard prerequisites |
|---|---|
| **FAK 201** — What a KV Cache Is and Why Reuse Is Always a Prefix | **FAK 105** |
| **FAK 202** — Context Is Not Memory: The Truth-Duration Axis | **FAK 105** |
| **FAK 203** — Why Memory Systems Get Promotion Backwards | **FAK 202** |
| **FAK 204** — The Four Layers of Agent Memory | **FAK 201** |
| **FAK 205** — Content-Addressed Blob Store (CAS) | **FAK 201** |
| **FAK 206** — cachemeta: Payload-Free Binding Keys | **FAK 205** |
| **FAK 207** — The Proofs Method: Theorem, Witness, Verdict, DOS | **FAK 105** |
| **FAK 208** — The Frozen Additive-Only ABI and Registry Seams | **FAK 207** |
| **FAK 209** — architest: Layered DAG, Tier Rules, and Hot-Path Hygiene | **FAK 208** |
| **FAK 210** — The Reference/Approx Correctness Contract | **FAK 207** |

### FAK 201 — What a KV Cache Is and Why Reuse Is Always a Prefix

**Prerequisites:** **FAK 105**

**You'll be able to:**
- Explain why token i's K/V depends only on tokens 0..i and why causality forces reuse to be a prefix
- Predict that a change at position N invalidates everything from N on
- Run the offline prefix-divergence script and watch longest-common-prefix reuse climb on an append-only loop

**Read:** [`docs/explainers/kv-cache-agentic-context.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/kv-cache-agentic-context.md), [`docs/glossary.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/glossary.md)

**Lab:**
```bash
Run the offline prefix-divergence script from the doc: feed it JSONL of {"turn": i, "tokens": [...]} per line and watch the longest-common-prefix reuse climb toward 100% on an append-only loop.
```

**Checkpoint:** Explain why token i's K/V depends only on tokens 0..i, and why that causality forces reuse to be a prefix rather than an arbitrary mid-sequence span. Then state the prefill-vs-prefix distinction the glossary pins.

### FAK 202 — Context Is Not Memory: The Truth-Duration Axis

**Prerequisites:** **FAK 105**
  ·  **Background:** **FAK 201**

**You'll be able to:**
- Distinguish context from memory by truth-duration, not size, recency, or location
- Sort facts into context-only vs memory-worthy using verb/tense cues
- Explain why two surface-identical facts can be different durability classes

**Read:** [`docs/CONTEXT-IS-NOT-MEMORY.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/CONTEXT-IS-NOT-MEMORY.md)

**Lab:**
```bash
List 5 facts you'd tell an assistant today and sort each into context-only (let it expire) vs memory-worthy (durable), then state the verb/tense cue that decided each.
```

**Checkpoint:** Explain why "it's raining here now" and "I live somewhere it rains" are the same surface fact but different durability classes, and which one must never be promoted to memory.

### FAK 203 — Why Memory Systems Get Promotion Backwards

**Prerequisites:** **FAK 202**

**You'll be able to:**
- Show that overflow, recency, salience, and explicit-save are all proxies for 'relevant to now' (i.e. context, not durability)
- Name the single root cause shared by 'the ephemeral promoted' and 'the durable dropped'
- Diagnose a write trigger by the present-moment proxy it actually measures

**Read:** [`docs/CONTEXT-IS-NOT-MEMORY.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/CONTEXT-IS-NOT-MEMORY.md)

**Lab:**
```bash
For each of overflow/summarization, recency, salience scoring, and explicit user-save, write one sentence naming the present-moment proxy it measures and one ephemeral fact it would wrongly promote.
```

**Checkpoint:** Name the single root cause shared by 'the ephemeral promoted' and 'the durable dropped' failures, and why it is one bug, not two.

### FAK 204 — The Four Layers of Agent Memory

**Prerequisites:** **FAK 201**
  ·  **Background:** **FAK 205**

**You'll be able to:**
- Separate routing (where), addressing (name), fusion (zero-copy arena), and semantics (mutate/isolate/attribute/gate) as four distinct problems
- Apply the one-line test (is this true of a frozen single-writer cache that merely moved/named/co-located?) to classify a claim
- Place fak in the semantics layer and explain why it does not compete on raw throughput

**Read:** [`docs/MEMORY-LAYERS-EXPLAINER.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/MEMORY-LAYERS-EXPLAINER.md)

**Lab:**
```bash
Apply the one-line test to five sentences (e.g. 'two readers share one cell by digest', 'evict a poisoned span from the middle and survivors stay byte-correct') and label each routing/addressing/fusion vs semantics.
```

**Checkpoint:** Using the Docker<->Kubernetes analogy, explain why 'a KV router is not a better memory MMU' and which layer fak occupies.

### FAK 205 — Content-Addressed Blob Store (CAS)

**Prerequisites:** **FAK 201**

**You'll be able to:**
- Explain why making the address the sha256 of the bytes gives free dedup and a faithful Ref backend
- Show why byte-identical Puts from distinct arrays collapse to one digest while the inline path is not deduped
- State what is in-scope vs out-of-scope (durability, GC, collision-resistance)

**Read:** [`docs/proofs/blob.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/proofs/blob.md)

**Lab:**
```bash
go test ./internal/blob/ -count=1 -timeout 120s -run 'TestPutSmallInlineRoundTrip|TestPutLargeBlobRoundTrip|TestContentDedup' -v
```

**Checkpoint:** Explain why two Puts of byte-identical content from DISTINCT backing arrays collapse to one blob with one digest, and why the inline path (len<=256) is deliberately NOT deduped.

### FAK 206 — cachemeta: Payload-Free Binding Keys

**Prerequisites:** **FAK 205**

**You'll be able to:**
- Explain why a deterministic, injective fold (null-separated sha256) over binding axes guarantees no false hit
- Show why the 0x00 separator rules out 'ab'+'c' vs 'a'+'bc' aliasing
- Explain why a partial-axis match yields a typed MISS/FAULT rather than a wrong serve, and why provider telemetry is excluded from invalidation

**Read:** [`docs/proofs/cachemeta.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/proofs/cachemeta.md)

**Lab:**
```bash
go test ./internal/cachemeta/ -count=1 -timeout 120s -run 'TestManifestBindingDigestIsDeterministicOverBindingAxes|TestCheckResidentClaimRefusesBindingMismatch|TestPlanExternalInvalidationsDropsRemoteKVAndReferencingAttentionIndex' -v
```

**Checkpoint:** Why does the 0x00 field separator make the fold injective on the tuple? Explain how a near-collision (some axes equal) yields a typed MISS/FAULT rather than a wrong serve.

### FAK 207 — The Proofs Method: Theorem, Witness, Verdict, DOS

**Prerequisites:** **FAK 105**

**You'll be able to:**
- Distinguish the four verdicts (PROVEN / REFUTED / OPEN / SCOPED-OUT)
- Explain why a structurally-deterministic function with no repeated-call test stays OPEN, not PROVEN
- Explain what dos commit-audit adds on top of a green witness

**Read:** [`docs/proofs/00-METHOD.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/proofs/00-METHOD.md), [`docs/proofs/README.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/proofs/README.md)

**Lab:**
```bash
go test ./internal/architest ./internal/abi ./internal/adjudicator ./internal/shipgate
```

**Checkpoint:** Distinguish the four verdicts and explain why a structurally-deterministic function with no repeated-call test stays OPEN rather than PROVEN, plus what dos commit-audit adds on top of a green witness.

### FAK 208 — The Frozen Additive-Only ABI and Registry Seams

**Prerequisites:** **FAK 207**

**You'll be able to:**
- Name the only sanctioned way to add a new admission rung or engine (a new package + one Register*() call)
- Explain why renumbering an existing VerdictKind fails TestABIGoldenFreeze while appending a new value does not
- Explain why a shared spine that changes breaks every dependent worker in a multi-session tree

**Read:** [`ARCHITECTURE.md`](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md), [`EXTENDING.md`](https://github.com/anthony-chaudhary/fak/blob/main/EXTENDING.md), [`docs/proofs/abi+architest.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/proofs/abi%2Barchitest.md)

**Lab:**
```bash
go test ./internal/abi/ -run 'TestABIGoldenFreeze|TestClosedReasonVocabulary' -v
```

**Checkpoint:** Name the only sanctioned way to add a new admission rung or engine, and explain why a renumber of an existing VerdictKind fails the golden freeze while appending a new value does not.

### FAK 209 — architest: Layered DAG, Tier Rules, and Hot-Path Hygiene

**Prerequisites:** **FAK 208**

**You'll be able to:**
- State the five tiers (root -> foundation -> mechanism -> composer -> integrator) and what an upward import produces
- Explain why the decision-path packages must never import os/exec
- Explain why the architest gate is build-tag-blind

**Read:** [`docs/proofs/abi+architest.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/proofs/abi%2Barchitest.md), [`ARCHITECTURE.md`](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md), [`SUBSYSTEM-CHECKS.md`](https://github.com/anthony-chaudhary/fak/blob/main/SUBSYSTEM-CHECKS.md)

**Lab:**
```bash
go test ./internal/architest/ -run 'TestNoUpwardImports|TestHotPathHasNoExec|TestEveryPackageDeclaresTier' -v
```

**Checkpoint:** State the five tiers and explain what failure a leaf importing a higher-tier package produces, and why a spawned subprocess on the decide path would kill the in-process syscall thesis.

### FAK 210 — The Reference/Approx Correctness Contract

**Prerequisites:** **FAK 207**

**You'll be able to:**
- Explain why Reference is held to max|delta|=0 plus the argmax oracle while Approx is held to argmax-exact plus a declared logit-cosine threshold
- Explain why a CUDA or quant backend declares Approx, not Reference
- Explain what RequireReference(b) prevents

**Read:** [`EXTENDING.md`](https://github.com/anthony-chaudhary/fak/blob/main/EXTENDING.md), [`docs/proofs/00-METHOD.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/proofs/00-METHOD.md)

**Lab:**
```bash
go test ./internal/compute/
```

**Checkpoint:** Explain why a CUDA backend declares Approx not Reference, and what RequireReference(b) prevents.

---

## The staged parts

To keep each stage a bounded read, the path ships as this page (overview plus
L100–L200) plus five staged parts under `docs/learning/`. The course numbers,
prerequisites, and checkpoints are continuous across the parts:

- [L300 — The Security Core](https://github.com/anthony-chaudhary/fak/blob/main/docs/learning/security-core.md) — the reference monitor, the policy lifecycle, and the enforcement rungs.
- [L400 — The Performance Core](https://github.com/anthony-chaudhary/fak/blob/main/docs/learning/performance-core.md) — why agents stress the cache and the addressable-eviction answer.
- [L500 — Serving, Integration, and the In-Kernel Model](https://github.com/anthony-chaudhary/fak/blob/main/docs/learning/serving-integration.md) — running and hardening `fak serve` (and `fak l3-serve` / `fak l3serve`), repointing real agents, and the in-kernel model (preferring Qwen3.8 per AGENTS.md).
- [L600 — Mastery](https://github.com/anthony-chaudhary/fak/blob/main/docs/learning/mastery.md) — benchmarks, honesty discipline, and extending the kernel.
- [The shipped-surface appendix](https://github.com/anthony-chaudhary/fak/blob/main/docs/learning/appendix-shipped-surface.md) — the wrap-up plus the full operator/contributor/package map.

---

# `fak-native`

> Source: `docs/native-inference-goal.md`

---
title: "Fak-native inference doctrine: own the engine and beat the reference"
description: "The canonical boundary for local inference work: fak-native is the product and performance path; llama.cpp is an explicit benchmark, diagnosis, interoperability, or borrowing reference, never a silent fallback."
---

# Fak-native inference is the product path

Audience: maintainers and agents implementing, optimizing, benchmarking, or reviewing local
inference in fak.

Authority: this is the canonical execution-engine doctrine. Architecture and performance docs
should link here. Benchmark indexes, issues, and runtime docs should do the same instead of
restating a different engine policy.

> **TL;DR:** Build and optimize local inference inside fak. Use llama.cpp deliberately as a
> reference, never quietly as the implementation.

## The invariant, in plain language

<!-- native-engine-doctrine:product-path -->
**Product path:** fak-native is the product and performance path for local inference.

<!-- native-engine-doctrine:matched-envelope -->
**Performance target:** In matched, quality-constrained envelopes, fak-native must beat llama.cpp.

<!-- native-engine-doctrine:explicit-reference-only -->
**Reference boundary:** llama.cpp is permitted only when explicitly selected for benchmarks, parity/reference diagnosis, migration/interoperability, or ego-free borrowing.

<!-- native-engine-doctrine:no-silent-fallback -->
**Fallback boundary:** fak never selects llama.cpp as a fallback for native or performance work.

<!-- native-engine-doctrine:owned-stack -->
**Ownership reason:** fak must understand and control the full engine stack so higher-order gains compose. That stack covers kernels and memory, scheduling and cache, plus adaptation and operations.

These are engineering invariants, not an unsupported claim that fak currently wins every
comparison. A current performance or efficiency claim still needs a scoped row in
[`BENCHMARK-AUTHORITY.md`](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md), with the engine and matched envelope
named. When the evidence says fak-native is behind, say so and keep the gap as native product
work.

New native-performance work prefers Qwen3.8.
Qwen3.6 requires an explicit task-specific exception. Allowed exceptions include regression or compatibility work, historical comparison, and hardware/artifact constraints.
Preserve historical Qwen3.6 artifacts; do not rename or rewrite them as Qwen3.8 evidence.

## What fak-native means

The model executes **inside the fak-owned inference path**:

- fak loads or interprets the model artifact and owns the model architecture path.
- fak owns the memory layout and lifecycle that its runtime exposes.
- fak chooses and coordinates CPU, CUDA, Metal, or other compute-HAL work.
- fak owns scheduling, batching, and placement. It also owns cache identity, reuse,
  eviction, and recovery.
- fak emits the engine identity and the evidence needed to attribute a result.

Calling a device library such as a vendor BLAS from a fak backend does not surrender the engine.
The deciding question is whether fak owns the execution plan and model state. It must also own
the lifecycle, observability, and replacement boundary. A llama.cpp process or library executing
the model is an external runtime even when fak launched it, fronts it, or consumes its output.

| Term | Meaning |
|---|---|
| fak-native | Model loading and inference execute through fak's in-kernel model and compute path. This is the local product and performance path. |
| native backend | A compute-HAL implementation selected inside fak-native, such as CPU, CUDA, or Metal. Changing a native backend does not replace the inference engine. |
| explicit external reference runtime | llama.cpp or another independently implemented engine selected explicitly for a bounded benchmark, comparison, parity/reference diagnosis, migration/interoperability task, or ego-free study and borrowing. |
| gateway/provider upstream | A remote or separately served model behind the fak gateway. This can be a supported operating mode, but it is external inference and is never evidence for fak-native performance. |

## Why engine ownership matters

An isolated tokens-per-second number is not the whole product. fak's advantage has to compose
across the complete execution path:

| Owned surface | What fak can improve because it owns the surface |
|---|---|
| Kernels | Select precision, fuse work, change tiling, exploit device features, and attribute a bottleneck to the real operation. |
| Memory | Keep weights and state resident, choose layouts, budget capacity, move data deliberately, and avoid opaque copies. |
| Scheduling | Batch compatible work, coordinate many sessions, place work by capacity and locality, and make queueing visible. |
| Cache | Give KV and prefix state stable identity, prove reuse legality, evict exact spans, and join cache state to the agent session. |
| Adaptation | Apply bounded model or policy adaptations with explicit provenance, lifetime, rollback, and compatibility. |
| Operations | Expose one lifecycle for startup, health, evidence, failure, recovery, and resource accounting. |

If an external runtime silently replaces the native path, these gains stop composing. fak may
observe the outside engine. It cannot safely treat opaque scheduling and memory as kernel-owned
state. The same limit applies to cache and adaptation state. A local improvement can therefore
make one benchmark look better while erasing the higher-order product fak is trying to build.

Ownership does not mean refusing established ideas or libraries. It means fak keeps the
decision boundary and integrates borrowed mechanisms into a path it can inspect, test, replace,
and operate.

## The matched-envelope rule

“Beat llama.cpp” is meaningful only inside a declared comparison envelope. A matched result
names, or explicitly reconciles, every row below.

| Axis | Required comparison state |
|---|---|
| Model | Architecture, artifact, tokenizer, and prompt template. |
| Numeric | Quantization, precision, and the quality or correctness floor. |
| Hardware | Device placement, thread budget, and memory budget. |
| Request | Prompt length, context length, decode length, and output stopping rule. |
| Load | Concurrency, batching, scheduling, and request arrival shape. |
| State | Cold or warm state, weight residency, prefix/KV state, and cache policy. |
| Outcome | Time to first token, decode rate, aggregate throughput, memory, energy, cost, or end-to-end task completion. |

An unmatched comparison can diagnose a direction, but it cannot close a native performance
claim. Equal tokens per second with worse correctness does not win. A fast external engine plus
fak's gateway is not a fak-native result. A native result that is slower today remains valuable
evidence because it names the gap that native work must retire.

## The four explicit llama.cpp uses

| Explicit use | What is allowed | What the result means |
|---|---|---|
| Benchmark | Run a pinned llama.cpp build as the tuned baseline in a matched envelope. | A comparison bar. It proves only what the measured rows and envelope state. |
| Parity/reference diagnosis | Compare tokenization, logits, greedy tokens, tensor transforms, or intermediate values to localize a correctness difference. | Reference evidence. Passing or failing narrows the defect; it does not transfer engine ownership. |
| Migration/interoperability | Read or produce compatible artifacts, validate a migration, or explicitly front a llama.cpp service while moving a workload. | Compatibility evidence. The run remains external inference unless the model executes inside fak. |
| Ego-free borrowing | Study a useful algorithm, kernel pattern, format decision, or operational lesson; preserve attribution and license obligations; implement and prove the useful part in fak-native. | Prior art and implementation input. The borrowed idea becomes native only after fak owns and witnesses its path. |

Every external use is selected deliberately and labeled in the command, configuration, result,
or receipt. “Available on this host,” “faster in the last run,” or “native support is missing”
does not authorize an automatic substitution.

## Vendor accelerator runtimes

Closed vendor accelerator runtimes execute outside the fak-native engine.
This classification applies when FLM or OGA/Lemonade selects a closed runtime;
the frontend name alone does not identify the execution engine.

Select one of the same four explicit uses for each vendor-runtime run:

- Benchmark: measure an external baseline in a matched envelope.
- Parity/reference diagnosis: compare outputs to locate a correctness difference.
- Migration/interoperability: validate compatible artifacts or explicitly front an external service.
- Ego-free borrowing: study mechanisms with attribution and license compliance, then implement and witness them inside fak.

Vendor-runtime receipts name the actual engine, device, and explicit use.
NPU placement does not convert external execution into fak-native performance evidence.
fak never automatically substitutes a vendor accelerator runtime for native execution.
When native support is unavailable, return an explicit unsupported result; an operator
may separately select an external comparison or interoperability run.

## Default and failure behavior

For work classified as native inference or native performance:

| Situation | Required behavior |
|---|---|
| Whole-engine selection is omitted. | Resolve to fak-native. |
| A native CPU, CUDA, or Metal backend is selected. | Stay inside the fak-owned compute boundary. |
| The model, backend, device, or memory envelope is unsupported. | Return an explicit unsupported, unavailable, or not-yet result. |
| Native launch fails. | Keep the native failure and its evidence. |
| The operator selects llama.cpp. | Reclassify the run as benchmark, diagnosis, interoperability, or borrowing work. |

Do not add `auto`, recovery, convenience, or “best available” behavior that turns a native
request into llama.cpp execution without the operator making that change. A loud native gap is
actionable product evidence. A quiet substitution is a false success.

## Review gate

Before accepting a native implementation or performance claim, ask:

1. Did the model execute inside fak?
2. Does the command, result, or receipt name the engine?
3. If llama.cpp ran, which of the four explicit uses justified it?
4. Is the comparison envelope matched and quality-constrained?
5. If fak-native could not run, did the path report that gap instead of changing engines?
6. If an external idea was borrowed, where is the fak-owned implementation and its native
   witness?

The shortest review question is: **Did the model execute inside fak, and does the receipt name
that engine?** If not, classify the run as external comparison or interoperability evidence and
do not use it to close a fak-native performance claim.

## Deterministic docs guard

This docs-only guard pins the five required invariant sentences and the seven inbound links
that make the doctrine discoverable. Run it from the repository root:

```powershell
$canonical = 'docs/native-inference-goal.md'
$requiredPhrases = @(
  'fak-native is the product and performance path for local inference.'
  'In matched, quality-constrained envelopes, fak-native must beat llama.cpp.'
  'llama.cpp is permitted only when explicitly selected for benchmarks, parity/reference diagnosis, migration/interoperability, or ego-free borrowing.'
  'fak never selects llama.cpp as a fallback for native or performance work.'
  'fak must understand and control the full engine stack so higher-order gains compose. That stack covers kernels and memory, scheduling and cache, plus adaptation and operations.'
)
$requiredPhrases += @(
  'Closed vendor accelerator runtimes execute outside the fak-native engine.'
  'Vendor-runtime receipts name the actual engine, device, and explicit use.'
  'fak never automatically substitutes a vendor accelerator runtime for native execution.'
)
$requiredLinks = [ordered]@{
  'README.md'                  = 'docs/native-inference-goal.md'
  'AGENTS.md'                  = 'docs/native-inference-goal.md'
  'llms.txt'                   = 'docs/native-inference-goal.md'
  'docs/architecture.md'       = 'native-inference-goal.md'
  'docs/performance.md'        = 'native-inference-goal.md'
  'docs/index.md'              = 'native-inference-goal.md'
  'docs/benchmarks/README.md'  = '../native-inference-goal.md'
}
$failures = @()
$body = (Get-Content -Raw -LiteralPath $canonical) -replace '(?s)\x60{3}.*?\x60{3}', ''
foreach ($phrase in $requiredPhrases) {
  if (-not $body.Contains($phrase)) { $failures += "missing phrase: $phrase" }
}
foreach ($entry in $requiredLinks.GetEnumerator()) {
  if (-not (Select-String -Quiet -SimpleMatch -LiteralPath $entry.Key -Pattern $entry.Value)) {
    $failures += "missing link: $($entry.Key) -> $($entry.Value)"
  }
}
if ($failures.Count) {
  $failures | ForEach-Object { Write-Error $_ }
  exit 1
}
"PASS native-inference-doctrine phrases=$($requiredPhrases.Count) inbound_links=$($requiredLinks.Count)"
```

Expected output:

```text
PASS native-inference-doctrine phrases=8 inbound_links=7
```

Run the repository link and index gates after this guard; they prove the linked target resolves,
while this guard proves the required distinctions and entry-point projections remain present.

## Related authorities

- [External system architecture](https://github.com/anthony-chaudhary/fak/blob/main/docs/architecture.md) places native and external execution inside
  the larger agent-kernel boundary.
- [Performance outcomes](https://github.com/anthony-chaudhary/fak/blob/main/docs/performance.md) routes a claimed outcome to its current witness.
- [Benchmark methodology](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmark-methodology.md) defines authority, generation, and
  reproduction rules.
- [Benchmark sheets](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/README.md) index the current native/reference result packets.
- [Adding a model](https://github.com/anthony-chaudhary/fak/blob/main/docs/new-model-playbook.md) applies the native-engine boundary to model support.

---

# `BENCHMARK-AUTHORITY.md`

> Source: `BENCHMARK-AUTHORITY.md`

# BENCHMARK AUTHORITY — Single Source of Truth

**Audience:** evaluators deciding which current `fak` result is applicable, what tuned alternative it was measured against, and how to inspect or reproduce its evidence.

## README hardware freshness contract

The front page is intentionally limited to one latest row each for Mac, AMD, and NVIDIA.
`docs/benchmarks/hardware-latest.json` is the compact consistency authority for those three
rows: its `as_of`, exact row text, observed date, and indexed receipt path must match
`README.md`. A publisher that adds or supersedes a platform receipt must update both files in
the same change. History and envelope-specific results remain on the benchmark index pages.
The standard CI and `make scorecard-ratchet` gates run the real freshness audit, so receipt
publication cannot leave the front page silently stale.

## Current Qwen performance

- [Qwen performance index](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/QWEN-PERFORMANCE-INDEX.md) — the canonical cross-hardware route for accepted highlights, receipts, result classes, and remaining gaps.

## Evaluator route: result first, method second

These are **scoped results**, not a universal speedup claim. Match your workload and hardware before quoting one; use the tuned baseline as the headline comparison.

| Evaluation question | Current scoped result | Tuned baseline and scope | Evidence route |
|---|---|---|---|
| **What is the public multi-agent reuse headline?** | **4.1× vs tuned** for a 50-turn × 5-agent Qwen2.5-1.5B Q8 session; the same run is 60.3× vs naive stateless serving. | Tuned per-agent KV is the decision-grade baseline; naive stateless is context, not the headline alternative. | Inspect `headline-qwen-50x5.json` in the [primary-number table](#quick-reference-primary-numbers), then follow its detailed row and reproduce route. |
| **Is the native CPU forward faster than llama.cpp CPU?** | **No:** decode is **0.55–0.73×** and prefill@256 is **0.58×** on the recorded M3 Pro run. | llama.cpp CPU `-ngl 0`, build 8200; 0.55× compares each engine's best thread setting, while 0.73× uses an equal 12-thread budget. | Inspect `model-ladder/qwen25-1.5b-q8-cpu-parity-m3pro.json` and the canonical first row below. |
| **What does prefix reuse achieve on the model ladder?** | RadixAttention reports **4.58× → 6.95×** live speedup for SmolLM2-135M through Qwen2.5-1.5B Q8. | Fresh recompute on the same four-model ladder; this is a prefix-reuse result, not an end-to-end universal serving result. | Inspect `radixbench-*-agents-fresh-20260619.json`, then use the sheet's RadixAttention detail and reproduce command. |

**Default choice:** start with the **tuned-baseline 4.1× multi-agent row** when evaluating the public reuse headline. Choose the CPU-parity or prefix-reuse row only when that narrower workload matches your question.

**Next action — verify one applicable row:** open its named committed artifact, confirm that its model, hardware, workload, and tuned baseline match your case, then run that row's linked reproduce command before quoting the result.

**Route context:**

- **Mode:** the rows above cover different benchmark modes—multi-agent session accounting, native CPU forward parity, and prefix-reuse prefill. A result applies only to its named mode and artifact configuration.
- **Generation:** this is the current authority across generation horizons; [the horizon view](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation-benchmark-authority-view.md) separates `gen/now` from next and future evidence.
- **Lifecycle:** this is a living evidence sheet. A newer committed artifact or an entry under [Tombstoned/Outdated Claims](#tombstonedoutdated-claims) supersedes an older number.
- **Support boundary:** committed artifacts and reproduce commands support the scoped observations. They do not imply an unmeasured model, backend, hardware, concurrency level, or production SLO; grade any broader claim with the [net-true-value standard](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/net-true-value.md).

For benchmark governance, contribution rules, and historical rationale, use the methodology routes below; they do not replace the scoped result table.

> **Why this exists.** This repo contains many benchmark results across different axes (raw throughput, reuse efficiency, session value-add, etc.). This document is the **authoritative index** of all committed benchmark claims, with traceability to source commits and artifact files. **Any number claimed elsewhere must trace back to an entry here.**

> **📋 Process:** See **[BENCHMARK-GOVERNANCE.md](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-GOVERNANCE.md)** for the DOS-centric process that creates, verifies, and publishes these claims. This file is the *what* (the numbers); Governance is the *how* (the discipline).

> **🗂 Sheet directory:** [docs/benchmarks/README.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/README.md) is the navigation index of every per-run sheet under `docs/benchmarks/` — results, runbooks, and pending/gated plans, grouped by theme. A number on a sheet is authoritative only once it has a row here.

> **🏆 Presentation layer:** **[HERO-BENCHMARK-2026-06-21.md](https://github.com/anthony-chaudhary/fak/blob/main/HERO-BENCHMARK-2026-06-21.md)** is the frontier-lab-style *hero comparison* (v1) built **from** this authority — headline number, top-3 SOTA chart, top-10 leaderboard with fak bolded where it wins (and the two single-stream losses shown plainly). It claims no new numbers; every figure traces to a row below.

> **🧭 Horizon view:** [docs/generation-benchmark-authority-view.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation-benchmark-authority-view.md) is the generation-aware *view* over the rows below — it slices each number by horizon, witness type, and promotion relevance so a reader can tell whether a row proves **current value** (entitled horizon `gen/now`: captured, reproducible, baseline named, provenance-labeled) or **future potential** (everything else, cited only with the fence the row already states). It claims no new numbers and adds no provenance vocabulary. A `gen/future` row is not a weaker row — spec-decode's bit-exact ceiling outwitnesses many wall-clock ratios — and a current-value row may record a **loss**. Design memo (issue #1669); not a gate.

> **🧠 The *why*:** `WHY-REUSE-WINS-2026-06-21.md` (private companion — not published) (v1) argues — and stress-tests — *why* these reuse numbers matter more than the headline alone: reuse is a **different class** of optimization (work-elimination on the `N` axis, not work-acceleration on the `κ` axis), so it's **exact, training-free, and composes multiplicatively** on top of every per-token trick. Follows the v2 SOTA-only framing — leads with the absolute competitive number (**19.0 min vs 78 min = 4.1× less work**, conservative marginal 2.4–2.7×), shows **no naive-loop numbers**, and centers cross-agent reuse as the layer that is `fak`'s. No new numbers; fences where the "works across everything / no fine-tuning" framing is overstated (and where addressable reuse, [#228](https://github.com/anthony-chaudhary/fak/issues/228), widens it).

**Last updated:** 2026-09-08
**Status:** Living document — update when new model results ship

> **🔁 Provenance vs. public reproducibility (read before you `git show` a commit below).**
> The **Commit** column records the *private-lineage* result commit each number first shipped in.
> Those short-SHAs predate the public **v0.30.0** squash (`1029e37`) and **do not resolve in a
> public clone** — treat them as provenance, not as a public reproduce handle. What an outsider
> actually re-runs in this repo is the committed **artifact** (the JSON / `.md` under
> `experiments/…` and `docs/benchmarks/…`, all tracked here) plus the **Reproduce** command. The
> artifact + command are the verifiable anchor; the SHA is lineage. (A few historical paths below
> dropped their old monorepo `fak/` subdir prefix — the real tracked path is `experiments/…` /
> `docs/benchmarks/…`. Limitation shown plainly rather than left for you to trip on.)

---

## Quick Reference: Primary Numbers

| Claim | Number | Model | Baseline | Commit | Artifact |
|---|---|---|---|---|---|
| **fak CPU Q8 single-stream vs llama.cpp CPU (M3 Pro)** — CANONICAL | **decode 0.55–0.73× (fak 38.1; llama 68.7 @−t6 → 52.4 @−t12) · prefill@256 0.58× (240.4 vs 412.5)** | Qwen2.5-1.5B Q8, M3 Pro (uncontended) | llama.cpp CPU −ngl 0, build 8200 (`541bf3762`) | _this commit_ | `model-ladder/qwen25-1.5b-q8-cpu-parity-m3pro.json` ← **single source; read, don't hardcode** — 2026-06-23 refresh (HEAD `374776a`, fak 0.31.0). Conservative fence = decode **0.55×** (each engine at its best thread config); equal 12-thread budget = 0.73×. The prior inline `0.58×/0.45× (71.9/547)` cited a non-existent commit and mixed a llama −t6 decode with an older-build prefill — reconciled in `docs/notes/MAC-BENCH-REFRESH-2026-06-23.md` |
| **RadixAttention live speedup (model ladder)** | **4.58× → 6.95×** | SmolLM2-135M → Qwen2.5-1.5B Q8 | Full re-prefill | `92896a4` | `radixbench-*-agents-fresh-20260619.json` |
| RadixAttention token speedup | 7.50× | all four models (Q8) | Token count | `92896a4` | Same (`prefill_token_speedup`) |
| RadixAttention hit rate | 86.7% (FCFS 62.1% → cache-aware) | all four models (Q8) | Cache hits | `92896a4` | Same (100% of optimal) |
| **Speculative decode — bit-exact + deterministic verify-pass speedup** ([#402](https://github.com/anthony-chaudhary/fak/issues/402), epic [#529](https://github.com/anthony-chaudhary/fak/issues/529)) | **E(K=4) = 1.00× (full-reject) → 5.00× (full-accept ceiling = K+1); the 2× decode-step threshold is crossed at acceptance a≈0.53, 3× at a≈0.74; lossless=true (speculative output token-identical to plain greedy)** | Synthetic CPU target+draft pair (PreNorm), draft K=4, no GPU | Plain greedy decode (E = 1 real token / target forward) | _this commit_ | `experiments/spec-decode/spec-decode-effective-e-20260625.json`. **DETERMINISTIC verify-pass speedup** — E = real tokens committed per target forward, the closed-form `polymodel.EffectiveTokensPerVerify(K, a)` evaluated at the **MEASURED** acceptance `a` (real-draft a=1.00 → E=5.00; adversarial a=0.00 → E=1.00 with 96 bit-exact `KVCache.Evict` rollbacks). **NOT wall-clock tokens/sec**: there is no GPU here, so the on-hardware 2–3× tokens/sec headline stays HW-gated (a measured number needs the [#535](https://github.com/anthony-chaudhary/fak/issues/535) bench harness on a GPU). The 2026-06-25 artifact remains the provenance for these numbers; the current mechanism is `internal/model` (live target/drafter binding, `VerifyForward`, bit-exact `KVCache.Evict`) + `internal/polymodel` (`SpecDecode`, `AcceptGreedy`/`AcceptTree`), feature-gated `FAK_POLYMODEL` (off by default). Bit-exactness is witnessed by `TestVerifyForwardChainMatchesSerial` + the artifact's `lossless` gate. Reproduce: `go run ./cmd/polymodelbench -bench` (or `-out FILE`) |
| **Native in-kernel continuous batching** ([#401](https://github.com/anthony-chaudhary/fak/issues/401)) | **B1 no-regression: 1.13× req/s vs legacy lifecycle; B8 1.54× req/s / 1.54× tok/s; B2 1.34×, B4 1.51×** | Synthetic CPU modelengine witness (192 hidden, 4 layers, 512 vocab), AMD Ryzen 9 9950X, `-benchtime=50x` | Legacy per-request lifecycle (one goroutine + one serial `Session.Step` loop per request, same model/prompts) | _this commit_ | `experiments/modelengine/native-continuous-batching-20260629.json`; reproduce `.\test.ps1 -run '^$' -bench BenchmarkEngineContinuousBatching -benchmem -benchtime=50x ./internal/modelengine`. This is the registered `inkernel` lifecycle scheduler, not a vLLM/SGLang production SLA benchmark; paged attention and multi-tenant p99 policy remain separate leaves. |
| **README headline: 50-turn × 5-agent reuse win** | **60.3× vs naive · 4.1× vs tuned** | Qwen2.5-1.5B Q8, T=50 A=5 P=2048 | Naive stateless / tuned per-agent KV | `2bbda6f` | `headline-qwen-50x5.json` |
| **fak compaction trims ~a third of a turn's resident context per fire (guard passthrough)** | **~107K tokens shed per compaction fire on the longest 2026-07-06 session (48K budget), firing repeatedly as the session grows; that trim recurs as a per-turn input saving on top of the provider prompt-cache discount. Fleet fak-authored share of total saved token-equiv ~0.3–16% (2026-W28 pooled ~16.4%).** ⚠️ **A prior version of this row claimed a per-session "~15%→~75%" share — RETRACTED as a double-counting artifact (see fences).** | `fak guard -- claude` passthrough, Anthropic prompt cache | per-fire shed = distinct middle-turn tokens dropped on one compaction fire; fleet share = WITNESSED fak-authored token-equiv / (OBSERVED provider prompt-cache + fak), from `fak cachevalue report` owner-attribution | _this commit_ | `docs/nightrun/cache-savings.jsonl` per-fire fields (`compaction_shed_tokens`, `compaction_fired`, `compaction_budget=48000`); longest 2026-07-06 session `21:58:33Z` = 746,956 shed over **7 fires** ≈ 106,708/fire. **RETRACTION/PROVENANCE:** the `compaction_shed_tokens` field is CUMULATIVE across fires (`internal/gateway/metrics.go:645` `m.compactShed += out.ShedTokens`), and each fire re-trims the client's FULL re-sent history (`internal/agent/anthropic_compact.go:440-443`, `internal/gateway/messages.go:814` — fak keeps no compacted body, only a restore tombstone), so the same aged middle turns are dropped and counted on EVERY subsequent fire. The session-total 746,956 therefore exceeds the session's live-context ceiling (~250–300K: `cache_read` 247,074 + `cache_creation` 112,957 + `input` 6,227) by ~4.8× — it is NOT a distinct-token count, and ratioing it against `cache_read_tokens` (a different cumulative sum, per-turn reused-prefix in a different currency) is not a like-for-like share. Fences: (1) cite shed **per fire** (divide `compaction_shed_tokens` by `compaction_fired`), never the session sum, and never as a "share" of cache_read; (2) valuing the shed: the **report** (`fak cachevalue report`) prices it at its honest PROPORTIONAL blend — `saved_token_equiv = min(shed, cache_read)×0.1 + (shed − min)×1.0`, the warm portion the provider evidenced as cache_reads at 0.1× and the cold remainder at 1.0× — via `cacheprice.ShedTokenEquiv` and `NewSavingsRows` (`internal/cachevaluereport/track2.go:437-438`), each row stamped `CACHE_READ_MARGINAL` (wholly warm) / `FULL_INPUT` (wholly cold) / `BLENDED_MARGINAL` (mixed) via `shedValuationBasis` (`track2.go:462`; LANDED `627001df` — visible in the report's `pricing` column). The pre-fix all-1.0× `saved_token_equiv = shed` over-credited warm fires ~10×, and the aggregate-warm binary fix that followed then UNDER-credited a cold-dominant session ~10× on a single warm token; BOTH step functions are retired on all three surfaces (#2794/#2798/#2804): the report, the live guard-exit/`/metrics` split (`gateway.MechanismSavings.FakTokenEquiv`, warm witness `CompactionCacheReadTokens`), and the `sessionobs` net-true ledger now price shed through the one `cacheprice.ShedTokenEquiv` blend, so no fak-authored figure books warm shed at full input nor collapses a cold-dominant session to the marginal; (3) fak-AUTHORED, on top of the provider discount — never attribute the provider slice to fak. Reproduce: `fak cachevalue report` (owner-attribution + fleet-benefit); fold `compaction_shed_tokens/compaction_fired` per `generated_at` from the ledger |
| **Fleet 5-agent × 200-turn 7B in <10 min — on a Metal forward (MEASURED, M3 Pro)** | **8.2 min (llama.cpp Metal forward + fak's reuse/batching pattern) · 2.5× vs a tuned single-stream baseline · ≥30× vs naive. ⚠️ pure-fak OWN forward (pure-Go CPU, Metal-decode lane open) ≈ 22–51 min — over the bar; sub-10-min on fak's own forward needs its GPU/CUDA path** | Qwen2.5-7B Q8, T=200 A=5 P=2048 D=20 R=12, M3 Pro | 5 single-stream sessions / naive re-prefill | _this commit_ | `session/macbook-m3pro-7b-batched-{bench,ctx}.log` (measured 17.41/44/392 t/s) + `fleet-5x200-7b-projection-20260622.json` + `FLEET-5X200-7B-10MIN-RESULTS.md`. Batched ≈ a tuned `llama-server --parallel`; fak's add is per-agent KV ownership + safety floor, not raw t/s. The ⚠️ own-forward figure was an arithmetic estimate, never run — a real (reduced-scope) own-forward measurement now exists, see the next row |
| **Own-forward (pure-Go CPU, no Metal) multi-agent fleet — MEASURED, reduced scope (M3 Pro)** | **NetVsTuned (fak-fused vs warm per-agent KV) 1.06× @1.5B, 1.26× @7B · NetVsNaive 8.3×/9.0× · decode batching gave ~0% gain — the whole win is prefill/shared-prefix reuse, the inverse emphasis from the Metal row above** | Qwen2.5-1.5B Q8 T=50 A=5 P=2048 D=20 R=12; Qwen2.5-7B Q8 T=20 A=5 (same P/D/R), M3 Pro | 5 single-stream (per-agent warm KV) / naive re-prefill | _this commit_ | `experiments/session/own-forward-fleet-mac-{1.5b-5x50,7b-5x20}-20260704.{json,log}` + `docs/notes/OWN-FORWARD-MULTI-AGENT-FLEET-MAC-2026-07-04.md`. **Reduced turn-count vs the 200-turn shape above** (killed after 59 min with no ETA on a shared 5-user node; `sessionbench` has no checkpoint/resume). **Built from a private-mirror checkout with an uncommitted 566+/396- diff in progress on the quant/parallel matmul hot path** — the A/B/C comparison is self-consistent (one binary, all arms) but the absolute tok/s are not yet citable against a clean fak commit; single rep (`-reps 1`), not best-of-N |
| **Session value-add (high-T ladder)** | **24.9× → 139.3×** | SmolLM2-135M Q8, T=64 → T=512 | Naive stateless | `92896a4` | `highT-smollm2-135m-*-fresh-20260619.json` |
| Session value-add (1.5B "realistic model") | 7.2× → 10.0× | Qwen2.5-1.5B Q8, T=8 → T=16 | Naive stateless | `92896a4` | `smoke-qwen2.5-1.5b-T8-16-fresh-20260619.json` |
| ~~Session value-add 11.2–14.5× (SmolLM2, P=512)~~ | ❌ STALE | SmolLM2-135M Q8 | Naive stateless | `5b0f40d` | superseded by re-measured row below |
| Session value-add (SmolLM2 P=512, re-measured) | ~~5.3–7.4×~~ — raw artifact not retained | SmolLM2-135M Q8 | Naive stateless | `885ae8a` | ❌ not independently reproducible: the raw sessionbench artifact was never git-tracked and cannot be regenerated here (no resident SmolLM2-135M Q8 export). The tracked SmolLM2 session value-add witness is the high-T ladder row above. |
| Qwen2.5-7B fak decode | 8.7 tok/s | Qwen2.5-7B Q8 | llama.cpp Metal 17.6 tok/s | `34c74f4` | `model-ladder/modelbench-qwen25-7b-q8.json` |
| Qwen2.5-7B fak/llama.cpp ratio | 0.50× decode / 0.083× prefill | Qwen2.5-7B Q8 | llama.cpp Metal | `34c74f4` | `QWEN25-7B-RESULTS.md` |
| Qwen2.5-7B greedy parity | ✅ full 7-token match | Qwen2.5-7B Q8 | llama.cpp ("2+2 is 4.") | `34c74f4` | `QWEN25-7B-RESULTS.md` |
| Qwen3.5-0.8B hybrid-GDN runs in fak | ✅ coherent ("pong") | Qwen3.5-0.8B f32 | instruction-following | `6a376b8` | `QWEN35-0.8B-RESULTS.md` |
| Qwen3.6-27B Q8 decode | 0.1 tok/s | Qwen3.6-27B q4_k_m (GGUF->Q8) | llama.cpp Metal 7.29 tok/s | `1698eff` | `docs/benchmarks/FAK-NATIVE-QWEN35-RESULTS.md` |
| Qwen3.6-27B fak/llama.cpp ratio | 0.12× decode / 0.01× prefill | Qwen3.6-27B q4_k_m | llama.cpp Metal | `1698eff` | `model-ladder/qwen36-perf-gate-m3-20260619.md` |
| **Qwen3.6-27B fak **Metal Q4_K** decode / prefill ([#63](https://github.com/anthony-chaudhary/fak/issues/63))** | **decode 1.2 tok/s · warm prefill 2.6 @P=27, 7.3 @P=940 tok/s** (cold first prefill 0.5; decode 0.16× and warm prefill 0.05×→0.14× of SOTA) | Qwen3.6-27B q4_k_m, M3 Pro (`-tags fakmetal`, `FAK_Q4K=1 FAK_METAL=1`, post-#1085; no co-resident llama-server for clean runs) | llama.cpp Metal 7.29 decode / 51.55 prefill | _this commit_ | `experiments/qwen36/metal-fak-q4k-post1085-m3pro-20260628.json` archives the owner-witnessed post-#1085 prefill refresh (GPU q4_k GEMM engaged for all 184 resident weights; greedy first token still matches CPU). Decode is carried forward from `experiments/benchmark/runs/by-machine/node-macos-a/20260626T055239Z-q4k-metal-decode-27b/score.json` (kernels bit-correct; GEMV cosine 1.000000, greedy token-parity vs CPU) until #67 remeasures resident-forward decode. **Still trails SOTA; no pass claimed.** Remaining wall is weight upload + per-call GPU round-trip / result memcpy (#1113/#69). Full prior diagnosis: `docs/notes/MAC-QWEN36-27B-Q4K-METAL-PERF-DIAGNOSIS-2026-06-26.md` |
| **Qwen3.8-27B 3-Way Apple Silicon Metal head-to-head fixture preview (fak-native vs llama.cpp vs MLX) [SIMULATED]** ([#2723](https://github.com/anthony-chaudhary/fak/issues/2723)) | **Validator-fixture preview, not measured on device: modeled decode 7.61 tok/s vs llama.cpp 7.38 and MLX 8.07; modeled prefill 48.54 vs 52.74 and 64.10 tok/s; modeled shared-prefix TTFT 12.60 ms.** | Intended envelope: Qwen3.8-27B Q4_K_M (17.1 GB), Apple M3 Pro 36GB, macOS Darwin arm64; 20 deterministic fixture samples | Planned llama.cpp b3600 (Metal) and MLX 0.22.1 (Metal) reference arms | n/a — no physical receipt | [Simulated three-way detail](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/MAC-THREEWAY-BENCH-2026-09-03.md) and [#2723](https://github.com/anthony-chaudhary/fak/issues/2723). The validator fixture synthesizes these values; no raw on-device comparison packet or bound evidence files are committed. Promote only after a physical matched-envelope campaign lands its immutable receipts. |
| **Qwen3.8-27B 4-Way Apple Silicon MTP speculative-decode fixture preview (fak-native vs AX Engine vs MTPLX vs llama.cpp) [SIMULATED]** ([#12239](https://github.com/anthony-chaudhary/fak/issues/12239)) | **Validator-fixture preview, not measured on device: modeled decode 15.22 tok/s vs llama.cpp 12.14, AX Engine 14.85, and MTPLX 14.98; modeled acceptance rate 0.785.** | Intended envelope: Qwen3.8-27B Q4_K_M (17.1 GB), Apple M3 Pro 36GB, macOS Darwin arm64; 20 deterministic fixture samples | Planned llama.cpp b3600, AX Engine v2.4.1, and MTPLX mlx-mtp-0.3.0 reference arms | n/a — no physical receipt | [Simulated MTP detail](https://github.com/anthony-chaudhary/fak/blob/main/docs/claims/qwen38-native-mtp-speculative-decode.md) and [#12239](https://github.com/anthony-chaudhary/fak/issues/12239). The validator fixture synthesizes these values; the referenced raw packet and evidence files are not committed. Promote only after the four physical arms run under one matched envelope and land immutable raw and quality receipts. |
| **Qwen Hopper H100 Q8_0 matched-weights head-to-head (fak-cuda-q8 vs llama.cpp CUDA) ([#10944](https://github.com/anthony-chaudhary/fak/issues/10944))** — CANONICAL | **decode 111.94 tok/s (+17.4% vs f32 95.37) · prefill 62.93 tok/s (+8.1% vs 58.21) · matched Q8_0 weights in VRAM · llama.cpp CUDA baseline 362.68 decode / 18,797.66 prefill** | Qwen2.5-3B Q8_0 on GCP `a3-highgpu-1g` (`NVIDIA H100 80GB HBM3`, `sm_90`), single-stream batch=1, resident native device GEMV (`k_q8_gemm`) | llama.cpp CUDA Q8_0 on identical hardware/weights | _this commit_ | `experiments/benchmark/runs/by-machine/gcp-a3-high-h100-1g/20260905T163948Z-gcp/paired-report.json`, `docs/_witnesses/issue-10944-nvidia-gcp-overnight/gcp_h100_paired_report.json`, `docs/benchmarks/GCP-H100-RESULTS.md`. Proves on-silicon Lever 1 Hopper acceleration with resident Q8_0 device GEMV. |
| **Qwen3.8-27B Live A100 GPU Serving & Prefix Caching Head-to-Head (fak-qwen-serve) ([#10944](https://github.com/anthony-chaudhary/fak/issues/10944))** — CANONICAL | **in-kernel radix KV cache reuse TTFT 3,183.57 ms (4.84× vs cold 15,399.32 ms, 6.12× vs unique dynamic 19,478.28 ms) @ 2048 ctx · prefill 0.00 s · decode 17.8–18.1 tok/s @ 256 ctx, 12.0 tok/s @ 1024 ctx · 50-subagent concurrency fanout 970–1,049 cluster tok/s (91/91 ok, 0 errors)** | `Qwen3.8-27B-Q4_K_M` (25,562 MiB VRAM) on 1× NVIDIA datacenter GPU, live endpoint `fak-qwen-serve` in `us-central1-f` | Cold prompt and unique dynamic prompt on identical hardware/weights; reference vLLM FP8 / BF16 TP=2 baselines | _this commit_ | `docs/_witnesses/issue-10944-nvidia-gcp-overnight/fak_native_qwen38_a100_bench_raw.json`, `qwen38_fanout_concurrency_raw.json`, `README.md`. Measured on live A100 cloud silicon with 0 errors across 91 parallel subagent invocations. |
| **Qwen3.8-27B Many-Agent Shared-Cache 4x Head-to-Head (fak-native vs llama.cpp) [SIMULATED]** ([#3809](https://github.com/anthony-chaudhary/fak/issues/3809)) | **Modeled workload projection ([SIMULATED], unmeasured on silicon end-to-end; see `docs/standards/simulated-results-discipline.md`): calibrated 1.86× wall-clock speedup (393.8 s vs 732.2 s; 13.00 vs 6.99 effective tok/s) · TTFT p50 12.6 ms vs 126,030.8 ms (10,002.4× faster) · 97.0% prefix reuse · 3,072 MB memory saved (post-#11855 calibration)** | Qwen3.8-27B Q4_K_M (17.1 GB), Apple M3 Pro 36GB, macOS Darwin arm64 (`node-macos-a`), K=4 concurrent agents, H=20 turns, P=4096 shared prefix | llama.cpp b9828 (Metal multi-slot) on identical hardware/weights | _this commit_ | `experiments/benchmark/runs/by-machine/node-macos-a/20260905T120000Z-agentic-4x/packet.json` (schema `fak.macbench.agentic-comparison.v1`, `fak macbench validate-agentic-comparison`); details in `docs/notes/MAC-AGENTIC-4X-QWEN38-2026-09-05.md`. Analytical workload model calibrated post-#11855/#11860 (eliminating unphysical queue-wait double-counting in reference arm) based on measured single-stream baseline rates (#2723, #9513) and prefix-sharing mechanics (#3813); full 30-minute physical multi-agent execution on silicon required before promoting to `[SHIPPED]`. |
| Qwen3.6-27B token parity | 2-token match (drift @3) | Qwen3.6-27B q4_k_m | llama.cpp oracle | `d03be46` | `model-ladder/qwen36-resident-q4k-parity-20260619.json` |
| Qwen3.6-27B surface smoke | 4/4 surfaces PASS | Qwen3.6-27B (served) | agent/gateway/mcp/dogfood | `8a0f5bc` | `model-ladder/qwen36-surfaces-dogfood-opencode-20260619.json` |
| **Qwen3.6-27B 8-GPU SGLang serving — fak-gateway vs raw-SGLang (SGLang adapter closure [#39](https://github.com/anthony-chaudhary/fak/issues/39); model-ladder Rung 4 headline [#921](https://github.com/anthony-chaudhary/fak/issues/921))** | **peak C=64: fak-gateway 1085.6 vs raw-SGLang 1451.6 completion tok/s = 0.75×; gateway tax converges to ~3% at saturation (C=128: 1074.4 vs 1103.2 = 0.97×); 3/3 fak surfaces PASS (agent / gateway-OpenAI single-stream decode 59.3 tok/s / MCP-HTTP)** | Qwen/Qwen3.6-27B (dense hybrid Gated-DeltaNet), 8-GPU datacenter server | raw SGLang 0.5.10.post1 (TP=8, bf16) | `a2559041` | `experiments/qwen36/gpu-server-r4-20260622/compare.json` (+ `COMPARE.md`, `fak-gateway.json`, `raw-sglang.json`, `surface-smoke.json`); results doc `docs/benchmarks/QWEN36-27B-GPU-SERVER-RESULTS.md`. **SGLang-serves + fak-adjudicates, NOT fak's native CUDA engine** (no quantized multi-GPU 27B path yet; f32 27B is 108 GB > 80 GB). fak's axis is the adjudication/coherence/measurement plane, not raw tok/s — the gateway tax is the cost of mediation and amortizes to ~3% at load. Marker-compliance caveat: the load harness requires a literal `FAK_GPU_SERVER_REQ_` echo the reasoning model emits only ~35–66% of the time, so throughput is over the compliant subset (`--max-error-rate 0.9`); per-point ok/requests in the JSONs (C=64: fak 55/64, raw 59/64). Single-stream rates ≪ batched and not fak's axis |
| **Gemma-4-31B live serving latency — fak gateway vs raw SGLang-compatible endpoint ([#1480](https://github.com/anthony-chaudhary/fak/issues/1480))** | **WITNESSED same-run C=2/N=8: TTFT p50 115.1 vs 76.9 ms (+38.2 ms, 1.50×); TPOT p50 36.6 vs 31.7 ms (+5.0 ms, 1.16×); ITL p50 31.6 vs 31.4 ms (+0.2 ms, 1.01×); E2E p50 633.2 vs 616.2 ms (+17.1 ms, 1.03×); both tracks 8/8 ok** | `google/gemma-4-31B-it`, live GPU-server SGLang-compatible OpenAI endpoint | raw SGLang-compatible `/v1` on the same model, prompt set, token limits, and concurrency | _this commit_ | `experiments/benchmark/runs/by-machine/gpu-server-sglang/20260704T144837Z-serving-parity/serving-parity.json` (+ `manifest.json`). Artifact carries per-request TTFT/TPOT/ITL/E2E, output-token basis, HTTP status, error class, p50/p90/p95/p99 stats, and `ours_minus_raw` / `ours_div_raw` tax fields. This is a serving-boundary gateway-tax packet, not cache-value, solve-rate, vLLM, or a vs-naive headline; private control details are scrubbed to generic GPU-server language. |
| **Qwen3.6-27B cold standup on 8-GPU datacenter server — raw-SGLang serving baseline (Rung-4 fresh-standup witness, [#921](https://github.com/anthony-chaudhary/fak/issues/921))** | **peak 820.5 completion tok/s @ C=64 (16.5k total tok/s incl. prompt) · 0 errors across all 6 points C=1→64 · single-stream 77.8 completion tok/s @ C=1** | Qwen/Qwen3.6-27B (dense hybrid Gated-DeltaNet), 8-GPU datacenter server, TP=8 | raw SGLang (`--tp 8 --mem-fraction-static 0.85`), cold standup from all-GPUs-idle (0 MiB) | _this commit_ | `experiments/qwen36/gpu-server-standup-27b-20260624/raw-sglang.json` (`fak.gpu-server-endpoint-bench.v1`, 6-point sweep + 8-GPU datacenter server topology/driver provenance) + `STANDUP.md` (narrative) + `samples.json` (3 live temperature-0 completions). **Cold-standup witness**: at start all 8 GPUs idle and nothing on `:30000`; weights load + CUDA-graph capture, then a real OpenAI-compatible sweep. This is the **raw SGLang serving baseline** (fresh bring-up proof), NOT a fak-gateway number — the gateway-tax comparison is the gpu-server-r4 Rung-4 row above. C=64 cap ⇒ 820.5 is a floor, not a ceiling (throughput still climbing monotonically at the top of the sweep); the wider C=128 sweep on the same box is the gpu-server-r4 row (raw 1451.6 @ C=64 → 1103.2 @ C=128). Marker gate disabled (`--no-require-response-marker`: the reasoning model echoes the literal `FAK_GPU_SERVER_REQ_` only ~35–66% of the time), throughput measured over completed requests (`errors=0` at every point) |
| Synthetic model live ratio | 1.64× | 64h/4L wiring | Full re-prefill | `a200c3d` | `radixbench-synthetic.json` |
| **GPU Q8 decode (Vulkan, RX 7600)** | **24.6 tok/s · 1.49× vs GPU f32** | SmolLM2-135M Q8 | Same forward, f32 weights on GPU | `60db592` | `q8gpu-smollm2-135m-{gpu-q8,gpu-f32}-20260619.json` |
| **AMD Strix Halo APU sub-kernel validation & candidate ablations** | **19/19 sub-kernels PASS · 167.5× GPU Q4_K speedup (451 µs vs 75.6 ms CPU) · 2.69× f16 KV contiguization (16.7 ms vs 44.9 ms) · 1.63× fused RMSNormMatMul (17.4 ms vs 28.3 ms)** | AMD Ryzen AI MAX+ 395 w/ Radeon 8060S (40 CUs, gfx1151, 64 GB UMA) | Discrete CPU reference, chained discrete kernels, unquantized F32, host-visible streaming, strided channel camping | _this commit_ | `docs/benchmarks/STRIX-HALO-BENCHMARK-RESULTS.md` (`docs/benchmarks/strix-halo-validation-11940.json`; reproduce: `go run ./cmd/fak-dev amd-strix-validate --subkernels=all --ablate=all --json`). Physical on-device execution witness covering all 19 native Vulkan SPIR-V compute kernels and 5 architectural candidate baselines. |
| **Qwen3.8-27B GPU Direct NVMe P2PDMA overflow architecture [SIMULATED]** ([#11226](https://github.com/anthony-chaudhary/fak/issues/11226)) | **Modeled projection: 145.8 decode tok/s (3.00× vs CPU-staged 48.6, 2.59× vs llama.cpp 56.2) · 2450.0 prefill tok/s · 34.80 ms TTFT · 0 host DRAM bounce copies (`StagingCopyCount == 0`) · 6.4 GB/s direct DMA** | Qwen3.8-27B hybrid Gated-DeltaNet + full attention geometry (AMD RX 7600 8GB VRAM + Crucial T705 Gen5 NVMe) | Standard CPU-staged swapping (3 copies) & llama.cpp OS mmap (2 copies) | _this commit_ | `docs/benchmarks/QWEN38-AMD-GPUDIRECT-RESULTS.md` (`fak.modelengine.qwen38-gpudirect-swap/1`; reproduce: `go run ./cmd/fak-dev amd-gpudirect qwen38 --json`). Architectural BaM zero-copy specification with synthetic hybrid state round-tripped bit-exact across 512–2048 tokens; physical in-kernel benchmark awaits Windows CGO Vulkan/ROCm driver path. |
| **Qwen3.8-27B & Flash Next NVIDIA RTX 5090 GPU Direct NVMe & Hierarchical Memory [SIMULATED]** ([#11326](https://github.com/anthony-chaudhary/fak/issues/11326)) | **Modeled analytical roofline projection ([SIMULATED], unmeasured on silicon; see `docs/standards/simulated-results-discipline.md`): 318.8 decode tok/s (theoretical zero-copy roofline vs CPU-staged 52.4; external comparisons unmeasured pending physical Blackwell hardware) · 3150.0 prefill tok/s · 24.50 ms TTFT · 0 host DRAM bounce copies (`StagingCopyCount() == 0`) · 7.1 GB/s sustained P2PDMA · 0.42s 32K context restoration** | Qwen3.8 (27B hybrid GDN + Flash Next MoE / 51B PLE) on NVIDIA RTX 5090 FE (32GB GDDR7, sm_120) + AMD Ryzen 9 5950X + 128GB DDR4 + Samsung 990 Pro 2TB (`M2A_CPU`) | Standard CPU-staged swapping (3 copies) & llama.cpp / vLLM (2 copies) | _this commit_ | `docs/benchmarks/QWEN38-NVIDIA-5090-GPUDIRECT-RESULTS.md` (`fak.modelengine.qwen38-cudadirect-swap/1`; reproduce: `go run ./cmd/fak-dev cuda-gpudirect qwen38 --json`). BaM-style direct NVMe queues in BAR1 VRAM (`CUDABaMVRAMQueue`), direct CPU root complex P2PDMA (<700ns latency), 3-tier memory management (Tier 0 VRAM 32GB, Tier 1 Host Pinned DRAM 128GB, Tier 2 NVMe P2PDMA), and cuStreamWaitValue64 MoE expert streaming. Verified via `internal/compute/cuda_gpudirect_storage_test.go` and `internal/model/qwen38_cudadirect_swap_test.go`. Modeled roofline only; unmodeled effects include PCIe TLP packet overhead, DRAM bank contention, thermal throttling, and OS/CGO jitter. Physical execution on Blackwell silicon required before citing as an achieved win or promoting to `[SHIPPED]`. |
| **GPU/CPU Q8 decode crossover** | **CPU lead 7.2× (135M) → 1.16× (1.5B)** | SmolLM2-135M → Qwen2.5-1.5B Q8 | CPU Q8 (legacy) | `7bf666b` | `crossover-qwen2.5-1.5b-{gpu,cpu}-q8-20260619.json` |
| **GPU decode parity (reusable CUDA graph, RTX 4070)** — README headline | **~120 tok/s (119–120, f32) · parity with llama.cpp Q8_0** | SmolLM2-135M, RTX 4070 Laptop sm_89 / WSL2 (gated `FAK_CUDA_GRAPH=1`) | llama.cpp Q8_0 (120 ± 15, `-ngl 99`) | `1029e37` | `GPU.md` §3b + `LLAMACPP-HEADTOHEAD-RESULTS.md` (on-box bench/test witness; reproduce: `FAK_CUDA_GRAPH=1 go run -tags cuda ./cmd/modelbench -dir internal/model/.cache/smollm2-135m -backend cuda`) |
| **Pure-kernel decide latency (M3 Pro)** | **362 ns** allow · 560–605 ns w/ ArgPredicates | syscall/adjudicator Decide | per-call decision | `bcad56e` | `experiments/mac-m3pro-kernel-20260620/kernel-latency-mac-m3pro-20260620.json` |
| **Pure-kernel admission latency (M3 Pro)** | **1.8–14 µs** scan · 3.3–15.8 µs Admit · 29–87 µs chain | ctxmmu / normgate+ctxmmu | per-result admission (cited "~1,300 ns" = cheapest scan layer only) | `bcad56e` | `experiments/mac-m3pro-kernel-20260620/kernel-latency-mac-m3pro-20260620.json` |
| **Syscall boundary tax (M3 Pro, refreshed)** | **~2,849×** in-process vs spawned `fak hook` | in-process adjudication | process-per-decide baseline (n=100) | `bcad56e` | `report.json` + `experiments/mac-m3pro-kernel-20260620/report.json` |
| **Adjudication overhead on the zero-alloc read path — drop-in-cost floor ([#451](https://github.com/anthony-chaudhary/fak/issues/451))** | **~0.55 ns/op · 0 B/op · 0 allocs/op · FLAT across N=1→1000 registered drivers** | n/a — kernel registry-read fold (no model, no GPU) | per-call registry read the decide path folds over every tool call | _this commit_ | `internal/abi/registry_scaling_test.go` — non-regression proof `TestRegistryReadsZeroAlloc` (0 allocs/op with 256 drivers of every kind); reproduce `go test ./internal/abi -bench BenchmarkRegistryReadScaling -benchmem`. Machine-stamped Ryzen 9 9950X (linux/amd64, 2026-06-26); v0.21.0's single-accessor figure was 1.39 ns/op. This is the **GPU-free FLOOR** of the `fak serve`-fronts-vLLM/SGLang drop-in cost: the full per-call decide is the "Pure-kernel decide latency" row above (362 ns allow), and the end-to-end **network** gateway tax is `docs/benchmarks/VLLM-HEADTOHEAD-RESULTS.md` §3 (vLLM pending a GPU run) / §4 (SGLang measured 0.75× peak → ~3% at saturation) |
| **Causal invalidation-on-external-write** | **PASS · max\|Δ\|=0** (1 evicted, sibling warm, re-admit refused) | vDSO `Revoke` + cachemeta external-invalidation | blunt world-flush / stale serve | `0fc39aa` | `experiments/causal-invalidation-20260620/causalbench-witness-20260620.json` |
| **Ultra-long-context work floor (>100k tokens, EXACT/contention-free)** | **single ~10× · 5-agent fleet ~40×+ vs naive (4.3× vs tuned)** | Qwen2.5-7B geometry, P=100k T=10 C=1/5 D=200 R=500 (arithmetic, no model) | Naive re-prefill (A/C ref) / warm per-agent KV (B/C) | _this commit_ | `session/ultra-long-context-floor-20260622.json` + `ULTRA-LONG-CONTEXT-RESULTS.md`. WORK floor (token = sessionbench `prefillTokens`; FLOP = O(L²)-aware), not a wall-clock; anchor token A/C 62.0× reproduces the committed 50×5 token floor; live wall-clock anchor at >100k is separately gated |
| **README front-page webbench hero — WebVoyager fleet prefill (MODELED geometry, no model)** | **8-worker A/C 9.7× vs naive floor · B/C 1.10× vs tuned per-agent KV · A/B turn-tax 8.8× (worker-independent)** | WebVoyager 643-task set, turns derived per-task from difficulty (`geometry_sources: difficulty=643`), BasePrefix=3000 / Action=150 / DOMState=2000 | A: naive re-prefill-every-turn (A/C) · B: tuned per-agent KV (B/C) | _this commit_ | `experiments/webbench/webvoyager-geometry-20260625.json` (emit with `fak webbench describe --dataset testdata/webbench/webvoyager-converted.jsonl --workers 1,2,4,8 --out …`). PREFILL-TOKEN WORK FLOOR from a deterministic geometry model — closed-form integer formula in `internal/webbench/geometry.go::ComputeArms`, **no model, no wall-clock, not "measured"**. README leads with the B/C value-stack number (1.10×); the 9.7× is explicitly the vs-naive-floor figure |

| **Decode vs prefill worker-count scaling (x86_64 32-core, within-run ratio)** | **decode all-cores-default penalty 2.5× (1.5B) → 2.1× (3B) → 1.14× (7B); decode peaks ≤8–16w, prefill scales to all cores** | Qwen2.5-1.5B/3B/7B Q8, x86_64 32-core agent-host (contended) | best worker count vs 32w default — same box, same run | _this commit_ | `experiments/session/worker-scaling-desktop-x86-20260624.json` + `WORKER-SCALING-DESKTOP-X86-20260624.md`. WITHIN-RUN ratio only; absolute tok/s is contended agent-host, NOT comparable to the uncontended M3 Pro rows. CPU-threading analogue of the GPU launch-bound small-model artifact |
| **Self-ablation feature sweep — vDSO on/off (deterministic, Regime A of epic #607)** | **vdso_hits 0→7 · engine_calls 12→5 · tokens 937→417 (−520)** | tau2-airline-smoke frozen trace (12 calls), mock engine, no model | all-off baseline (vDSO off) | _this commit_ | `experiments/ablate/tau2-smoke-vdso-ablation.json` + `ABLATE-RESULTS.md`. Counter fields (workload_hash/vdso_hits/engine_calls/tokens/denies/quarantines) reproduce byte-identical (kernel event counters on a frozen trace); only p50_ns/wall_seconds/buckets are single-box. Rung 1 sweeps the one runtime knob only; env-gated features + cross-agent (Regime B) arms are separate rungs |
| **Cross-agent ablation — bare `claude` vs `fak guard -- claude` (Regime B of epic #607, [#623](https://github.com/anthony-chaudhary/fak/issues/623))** | **K=5/arm, both 5/5 success · output 0.98× · turns 1.00× · total-ingested 1.56× (−28 986 tok, kernel overhead) · +fak: 5 ALLOW / 0 deny** | `pong` 1-tool-call task w/ deterministic check, `claude-opus-4-8`, same OAuth acct, single Windows host | `claude_code` (bare `claude -p`) baseline | _this commit_ | `experiments/ablate/cross-agent-pong-opus.json` + `ABLATE-RESULTS.md`. Regime B is DISTRIBUTIONAL (mean ± CI95 over K≥5; the `WorkloadHash` guard does NOT apply); success-gated, model-named, tokens decomposed never summed. ONE tiny tool-light task on ONE host ⇒ deny/repair/quarantine counters an honest zero, cache-split is cold-prefix illustrative not a fleet SLA. Tool: `tools/cross_agent_ablate.py` (17 hermetic tests) |
| **AgentDojo structural safety floor (local, model-free)** | **full-stack ASR 0/38 (0.000) vs detection-only 29/38 (0.763) · benign controls 2/2 · gate PASS** | deterministic AgentDojo-style red-team, no model | detection-only lexical gates | _this commit_ | `experiments/agent-live/agentdojo-fak-fullstack-20260625.json` (reproduce: `go run ./cmd/agentdojoredteam -json`; corpus `sha256:ddc5b9ae08df0b37224a290fae212525228d2930e77afecb7bfc868b06ca1060`). LOCAL structural floor only — not an official external AgentDojo leaderboard result or raw-model arm |
| **AgentDojo external entry — `fak_gateway` registered non-model defense (#1064; module BUILT+WITNESSED, run/PR operator-gated)** | **module loads + intercepts a tool call (26-check test PASS); local intercept witness targeted ASR 0/7 (0.000) · benign 2/2; `benign/under-attack utility = NEEDS_KEY`; `result_claim_allowed=false`** | `fak_gateway` `BasePipelineElement` in a fork of `ethz-spylab/agentdojo`; targeted-ASR mechanism WITNESSED locally, utility arms OBSERVED (paid fronted model) | the four published non-model rows (Tool Filter / Spotlighting / Transformers PI Detector / Repeat User Prompt) + the formal-isolation tier (CaMeL ASR 0 / MELON 0.0–2.4%) | _this commit_ | `experiments/agent-live/agentdojo-fak-gateway-defense-entry-20260627.json` + `.md`; module `experiments/agentdojo-fak-defense/` (reproduce witness: `python3 experiments/agentdojo-fak-defense/fak_gateway_defense.py --json`; test: `python3 experiments/agentdojo-fak-defense/test_fak_gateway_defense.py`). PLACE in the ~0-ASR tier at a stated utility cost — **not a win, not a leaderboard rank**. 629-case + 97-case run on a fronted model and the upstream PR are the recorded operator-gated blocker |
| **ToolSandbox/tau3 policy-state adapter smoke ([SIMULATED] local fixture)** | **raw safe pass^1 1/2 (0.500) -> fak safe pass^1 2/2 (1.000); fak denied 1 policy/minefield call; `result_claim_allowed=false`** | `offline-trace`, 2 ToolSandbox-shaped tasks, no live model | raw trace replay without fak mediation | `c92bb2c` | `experiments/agent-live/toolsandbox-policy-state-smoke-20260625.json` + `.md` (reproduce: `go run ./cmd/toolsandboxbench -suite testdata/toolsandbox/policy_state_smoke.json -out experiments/agent-live/toolsandbox-policy-state-smoke-20260625.json -md experiments/agent-live/toolsandbox-policy-state-smoke-20260625.md`). Adapter smoke only - not an official Apple ToolSandbox or tau3 leaderboard result |
| **GLM-5.2 fak-kernel cache value (PENDING — results not yet collected)** | **PENDING — see result packet for shape** | GLM-5.2 on pure fak kernel, solved SWE-bench ticket | No cache (work saved is the lever) | _pending_ | `docs/benchmarks/GLM52-FAK-KERNEL-CACHE-VALUE-RESULTS.md` (result packet shape; observation seam shipped at `52dfea0d`, datacenter GPU access is the residual). Observation metric: `kv_prefix.reused_tokens` (WITNESSED, not provider's `cache_read`). See runbook for full path. |
| **Local-model coding witness (2026-06-27) — local capability MEASURED; replay-trace governance WITNESSED; frontier/full-Claude E2E still pending** | **MEASURED local capability:** Qwen2.5-Coder 3B and 7B both produced the one-line `testdata/coding_smoke` fix and the fixture went 1-fail/1-pass -> 2-pass. **WITNESSED governance:** `fak guard --replay-trace` adjudicated 8 tool-call verdicts and denied 3 dangerous calls. **PENDING:** a metered frontier A/B on a harder fixture and a full Claude Code tool-call end-to-end local-model witness. | Qwen2.5-Coder-3B/7B via local Ollama/OpenAI wire for the measured capability rows; model-agnostic replay trace for the governance row; frontier arm not run here | Minimal coding fixture (`testdata/coding_smoke`) for capability; `internal/gateway/testdata/guard-trace-e2e.json` replay for governance | _this commit_ | `docs/notes/LOCAL-MODEL-CODING-WITNESS-2026-06-27.md` (measured result and honesty fences) + `docs/benchmarks/LOCAL-MODEL-CODING-WITNESS-RUNBOOK.md` (reproduction commands). No new local-vs-frontier capability claim is made by this authority row; the result doc explicitly fences the full Claude tool-call E2E witness as still owed. |
| **fak self-tax — own mediation overhead + net effect over time (LIVING, net-true-labeled)** ([#1169](https://github.com/anthony-chaudhary/fak/issues/1169), L5/T12 of epic [#1147](https://github.com/anthony-chaudhary/fak/issues/1147)) | **read-path floor ~0.55 ns/op (0 allocs, FLAT N=1→1000) · pure-kernel decide 362 ns · vDSO self-ablation −520 tok (net SAVING) · cross-agent tool-light +28,986 tok (1.56×, net COST) · reuse detector nets POSITIVE** — signed + workload-dependent, never an average | n/a — fak's own mediation overhead (NOT raw inference; that axis is #306) | fak-off / bare-agent on the same workload (per component artifact) | _this commit_ | `docs/benchmarks/SELF-TAX-TREND.md` — the living trend doc that folds the already-committed self-tax artifacts into one signed, net-true-labeled series: `internal/abi/registry_scaling_test.go` (read floor), `experiments/mac-m3pro-kernel-20260620/kernel-latency-mac-m3pro-20260620.json` (decide), `experiments/ablate/tau2-smoke-vdso-ablation.json` (vDSO net saving), `experiments/ablate/cross-agent-pong-opus.json` (cross-agent net cost), and design note §9 (the WITNESSED⊕OBSERVED−MODELED improvement detector). Reproduce (representative — full per-surface set in the trend doc): `go test ./internal/abi -bench BenchmarkRegistryReadScaling -benchmem`. **Honesty fence:** this is fak's *mediation* overhead, not #306 raw-inference parity; each row is an independently-committed point, and the always-on fak-on-vs-fak-off CI gate (T6/T7/T8) + single `fak perf` read-out (T11) are the named epic follow-on, so this is a curated fold today, not yet an auto-emitted JSON |
| **fak support-maturity matrix — model × backend rung coverage (LIVING, generated-not-typed) (#1255, E6 of epic #1243)** | **Grade F · score 33.9 · support_maturity_debt 37 · 19/56 cells SUPPORTED (13 PROOF-PATH-ONLY, 24 FENCED, 0 UNDEFINED across 14 families × 4 backends) — a coverage instrument, not a vs-baseline win** | n/a — the kernel's own model × backend support grid (internal/covmatrix), folded by internal/supportmaturityscore; deterministic from the committed tree, no model/GPU | n/a — this measures support maturity, not throughput; the per-cell parity work is #307/#305/#303/#301 | this commit | docs/HARDWARE-MATRIX.md — Reproduce: `go run ./cmd/fak support-maturity-scorecard --json` |
| **FrontierSWE time-to-solution — fak-routed vs raw harness (C15 of epic [#1706](https://github.com/anthony-chaudhary/fak/issues/1706), [#1721](https://github.com/anthony-chaudhary/fak/issues/1721)) — GATED, no number until witnessed** | **GATED — no TTS number is claimed. Claim boundary: no wall-clock/turn-count TTS ratio until (a) the official FrontierSWE grader produced both arms' reward.json grader output (C13 [#1719](https://github.com/anthony-chaudhary/fak/issues/1719)) and (b) score-parity holds — fak `correctness`/`speedup` >= raw (C11 [#1717](https://github.com/anthony-chaudhary/fak/issues/1717))** | harness-agnostic (claude-code shim first) | raw FrontierSWE harness (same task/model/budget/retry) | _gated_ | `docs/benchmarks/FRONTIERSWE-RESULTS.md` (authority) + `docs/benchmarks/FRONTIERSWE-TTS-RUNBOOK.md` (raw-vs-fak recipe). Offline projection only, NOT a measurement: `fak frontierswe describe --tts`; scoring runtime green (`go test ./internal/frontierswe`). Measured TTS is the C13/C14 residual — see [FRONTIERSWE-RESULTS.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/FRONTIERSWE-RESULTS.md) |
| **Full-span four-band trace + µs decide tail under same-process load (G5/R6 of epic [#2218](https://github.com/anthony-chaudhary/fak/issues/2218), [#2223](https://github.com/anthony-chaudhary/fak/issues/2223))** | **loaded p99 ≈ 30–34 µs vs quiet ≈ 13.5–14 µs (×2.1–2.5) · p50 nearly flat (×1.17–1.19) · >100 µs share ~3× (0.18–0.19% → 0.56–0.58%) · trace: all 4 bands walkable in ONE process, every DENY carrying a consequence class (retry_turn / forked_outcome / clean_stop)** — OBSERVED research-grade, no gate flipped | n/a — kernel `Decide` fold + real `superloop.Walk`/`EvaluatePreflight`, hermetic, no model | quiet arm of the SAME process/run (within-run only; per-band totals never cross-multiplied) | _this commit_ | `docs/benchmarks/fullspan/fullspan-trace-20260707.json` + `tailload-20260707-run{1,2}.json` (two full 200k-sample runs; loaded p99 agrees within 14.2%) + sheet `docs/benchmarks/FULLSPAN-TAILLOAD-RESULTS.md`. Hardware named in-artifact: AMD Ryzen 9 9950X (16C/32T), 256 GiB RAM, WSL2 Ubuntu 24.04 on a Windows 11 host. Reproduce: `FAK_BENCH_HW=… FAK_TAILLOAD_OUT=… go test ./internal/bench/ -run TestWriteTailLoadArtifact -count=1`. **Fences:** three-clocks (no end-to-end field — structurally enforced by test), trace B0 spans are one-shot cold costs (do NOT compare to the warmed 362 ns M3 Pro anchor), WSL2-virtualized tails, coarse-clock hosts refuse to write artifacts |
| **Served-inline proxy-dedup hit-rate — how often `--vdso-proxy-fill` actually fires on an agent read pattern ([#1350](https://github.com/anthony-chaudhary/fak/issues/1350), epic [#1351](https://github.com/anthony-chaudhary/fak/issues/1351))** | **WITNESSED: read-only-*shaped* names (`get_`/`read_`/`search_`/…) serve 5/14 read-only proposals = 35.7%, and every served hit is a cross-turn re-read (W3 5/5; within-turn W1/W2 1/0 and force-fresh 1/0 serve nothing); Claude-native `Read`/`Grep`/`Glob` names serve 0/14 = 0% (the read-only NAME gate does not recognize them); the canonical guard-trace-e2e dominant path serves 0/2 = 0% for the same structural reason. Saves the EXECUTION of a duplicate/repeat read (the client round-trip), NOT the surviving call's tokens. Recommendation: keep `--vdso-proxy-fill` OPT-IN until the name gate recognizes native read tools (or scope default-on to MCP/snake_case reads)** | n/a — kernel served-inline seam (`adjudicateProposedServed` + `admitInboundResults`), no model/GPU | a miss is byte-identical to today (`TestServedInline_Miss`); denominator = read-only-proposed calls per turn | `7b07af0ca` (dominant-path arm) / `953bff575` (dedup sheet) | `internal/gateway/testdata/served_inline_dedup_report.json` (modeled `representative-coding-session-v1`, 6 turns, per-win-class decomposition) + `internal/gateway/testdata/served_inline_guardtrace_report.json` (canonical `guard-trace-e2e` fixture, per-turn counts + `fallback_reason`) + sheet `docs/benchmarks/SERVED-INLINE-DEDUP-RESULTS.md`. **Provenance:** the SERVE is WITNESSED (fak's real in-process seam served these calls, not a model of what it would do); the dedup trace redundancy is MODELED (a named, reproducible transcript) and the guard-trace arm is an AUTHORED shared fixture with real Claude Code native names — **neither is a captured live `/v1/messages` session** (that promotion path stays open). Reproduce: `go test ./internal/gateway -run TestServedInlineDedup -count=1` + `go test ./internal/gateway -run TestServedInlineGuardTrace -count=1`. **Fence:** a near-zero rate on the dominant path is a valid, publishable result that down-ranks the lever, not a defect — the 0% is the name gate, independent of how redundant the reads are |
| **Governed-agent overhead (observer-effect) + pilot-survival demonstrator** ([#3292](https://github.com/anthony-chaudhary/fak/issues/3292), Workstream E of epic [#3256](https://github.com/anthony-chaudhary/fak/issues/3256)) | **Overhead WITHIN its declared envelope: in-kernel decide 362 ns / hook p99 ~175 ms < 250 ms budget / meter 0-alloc — signed net (vDSO −520 tok saving, cross-agent +28,986 tok cost). Survival: governed cuts false-done 0.50→0.00 (net +3,900 tok), one-holder leases, runaway/injection caps — ungoverned does none** | n/a — fak's own mediation overhead + survival levers (no model/GPU) | fak-off / naive self-report loop / ungoverned control, per lever | _this commit_ | Detailed section [Governed-agent overhead + pilot-survival demonstrator](#governed-agent-overhead-observer-effect-budget--pilot-survival-demonstrator-2026-07-19--the-corporate-what-does-it-cost--will-it-survive-proof). On-box witnesses: `go test ./internal/turntaxmeter ./internal/loopgate -count=1` → `ok` (2026-07-19). Binds existing artifacts (`turntaxmeter/{overheadbudget,hooklat,sampler}.go`, `loopgate/testdata/verified-vs-naive-loop.report.json`, `perf-dos-hook-cost.md`, `perf-runaway-guard.md`, `SELF-TAX-TREND.md`, `cmd/turntaxdemo -selfcheck`); adds no new measured number. Provenance-labeled per row; false-done corpus SIMULATED, safety-floor re-run peer-red-trunk-gated |
| **Terminal-Bench 4 end-to-end benchmark — fak harness (in-kernel model) vs OpenCode (llama.cpp baseline reference) ([#10803](https://github.com/anthony-chaudhary/fak/issues/10803))** | **WITNESSED offline smoke test: fak all-in-one 100.0% (5/5) vs OpenCode + llama.cpp 60.0% (3/5) solve rate (+40.0% delta) · 83.5% prompt token reduction (530 vs 3,210) · 11 vDSO hits** — offline synthetic suite (`testdata/tb4bench/synthetic_suite.json`), digest `eeb965614260ee62d2f60be0187409ff73c4745e709eae3bb67a59cfe216a344`. Official live container campaign remains gated on full OCI engine. | Qwen3.8-Coder GGUF, pinned seed=42, temp=0.0, 32k context, airgapped OCI container | OpenCode headless execution against dedicated reference `llama-server` on identical weights | `eeb96561` | `docs/benchmarks/TERMINAL-BENCH-4-REPRODUCTION.md` (reproduction runbook & parity envelope) + `experiments/benchmarks/tb4-smoke-wsl/compare.json` + `compare.md`. Verification: `fak bench tb4 run --arm both --mock --dataset testdata/tb4bench/synthetic_suite.json` |
| **50+ sub-agent fan-out scaling under strict resource bounds ([#10855](https://github.com/anthony-chaudhary/fak/issues/10855))** | **N=50: 6.21 ms / 8,056.5 agents/sec / 147 vDSO dedup hits / 100,352 prefix tokens elided; N=64: 7.77 ms / 8,231.9 agents/sec; N=256: 28.83 ms / 8,879.8 agents/sec; linear prefix elision (N-1)×2048** | Research gather mode, 50+ concurrent subagents, synthetic model | Uncached serial execution | `095cdd1ff` | `docs/_witnesses/issue-10855-subagent-50-fanrun/fanrun-50plus-subagents.json` + `docs/_witnesses/issue-10855-subagent-50-fanrun/README.md`. Reproduce: `go run ./cmd/fanrun --agents 1,4,16,32,50,64,128,256` |
| **Qwen3.8 M2 Metal sequence prefill campaign ([#9525](https://github.com/anthony-chaudhary/fak/issues/9525), [#9230](https://github.com/anthony-chaudhary/fak/issues/9230), [#9430](https://github.com/anthony-chaudhary/fak/issues/9430))** | **M2 sequence prefill: 43.8% prefill latency improvement (10,284.5 ms vs 18,304.9 ms); 1 command buffer vs 192 (99.5% reduction); first-token latency +43.9% (2,451.8 ms vs 4,368.5 ms); 0 fallbacks; 0 swap growth** | Qwen3.8-27B Q4_K_M GGUF, Apple M3 Pro (18 GPU cores, 36 GiB) | Per-op synchronous Metal prefill control | _this commit_ | `docs/_witnesses/issue-9525-qwen38-sequence-prefill/receipt.json` + `README.md`. Reproduce: `go test -v ./docs/_witnesses/issue-9525-qwen38-sequence-prefill/...` |
| **Qwen M3 Pro Metal sequence prefill & kernel microbenchmarks ([#9230](https://github.com/anthony-chaudhary/fak/issues/9230), [#9430](https://github.com/anthony-chaudhary/fak/issues/9430))** | **M2 sequence prefill: 7.80× speedup (19.73 ms vs 153.91 ms, 87.2% latency reduction); fused MLP: 1.62× (1.962 ms vs 3.175 ms, 76.6 GB/s); Q4_K GEMV P4/P8 crossover: 1.70–1.79×; Qwen3.5-0.8B Metal decode: 11.42 tok/s, prefill@256 120.63 tok/s; 0 swap growth** | Qwen3.5-0.8B/4B & Qwen3.5 hybrid Q4_K synthetic shapes, Apple M3 Pro | Control sequence prefill & separate MLP kernels | `c5a3b411f` | `docs/_witnesses/qwen38-m3pro-metal-benchmarks-2026-09-03.json`. Reproduce: `go test -bench BenchmarkMetalQwen35P32SequenceVsControl ./internal/model` |

> **The model-ladder thesis.** Live wall-clock ratio climbs toward the deterministic
> 7.50× token-speedup ceiling as per-token compute grows (135M 4.58× → 360M 5.40× →
> 0.5B 6.20× → 1.5B 6.95×). This confirms that the residual gap below 7.50× is
> clone/memcpy overhead that becomes negligible on larger models — not an
> architectural limit. The deterministic metrics (token speedup, hit rate) are
> hardware-independent and reproduce the committed JSON exactly; only the live
> wall-clocks are single-box (within-run ratios authoritative per
> [BENCHMARK-GOVERNANCE.md "Within-Run Ratios"](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-GOVERNANCE.md#within-run-ratios--single-box-discipline)).

---

## Governed-agent overhead (observer-effect budget) + pilot-survival demonstrator (2026-07-19) — the corporate "what does it cost / will it survive" proof

**Date:** 2026-07-19
**Issue:** [#3292](https://github.com/anthony-chaudhary/fak/issues/3292) (Workstream E / proof of epic [#3256](https://github.com/anthony-chaudhary/fak/issues/3256), the all-in-one agent runtime MLP)
**Standards:** [observer-effect](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/observer-effect.md) · [net-true-value](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/net-true-value.md)
**On-box witnesses (this row):** `go test ./internal/turntaxmeter -count=1` → `ok` · `go test ./internal/loopgate -count=1` → `ok` (both green on win/amd64, 2026-07-19)
**Artifacts:** `internal/turntaxmeter/{overheadbudget,hooklat,sampler}.go` · `internal/loopgate/testdata/verified-vs-naive-loop.report.json` · `docs/perf-dos-hook-cost.md` · `docs/perf-runaway-guard.md` · `docs/benchmarks/SELF-TAX-TREND.md` · `cmd/turntaxdemo` `-selfcheck`

### What this witnesses (and why it is one row, two axes)

The corporate objection to a governance layer is two questions: *"what does it cost me?"* and
*"will the pilot survive production?"* This row answers **both** from committed evidence, never
narrated. It does not add a new measured number — it **binds** the already-witnessed overhead
and survival artifacts into the single authority the issue names ("land both in
BENCHMARK-AUTHORITY.md with traced artifacts"), each provenance-labeled and net of what it saves,
so no naive-baseline strawman is smuggled in.

### Axis 1 — observer-effect overhead vs its DECLARED envelope (not a naive baseline)

The self-tax plane declares an **envelope** per lifecycle rung — a scope-stated ceiling, not a
promise of zero cost (the epic's explicit non-goal: a gate that costs 8% and saves 40% is a net
win). A breach names ONE closed-vocabulary token, `OVERHEAD_BUDGET_EXCEEDED`
(`internal/turntaxmeter/overheadbudget.go`), verifiable via `dos man wedge <TOKEN> --explain`. The measured
overhead sits **within** that envelope:

| Rung / meter | Measured overhead | Declared envelope | Verdict | Provenance |
|---|---|---|---|---|
| In-kernel adjudication decide (per tool call) | **362 ns** allow (M3 Pro), p50 ~1,300 ns in-process | 5 µs compute-rung `MaxNS` | within envelope | OBSERVED (`experiments/mac-m3pro-kernel-…json`; `docs/perf-dos-hook-cost.md`) |
| Guard hook tail (pretool+posttool, fleet) | **~85 ms mean / ~175 ms p99** | `DefaultHookP99BudgetMS` = **250 ms** | within envelope (`GATE_LATENCY_REGRESSION` not fired) | OBSERVED (2026-07-01 guard-audit, `hooklat.go`) |
| Meter's OWN per-event cost (the observer effect itself) | **0 B / 0 allocs**, 1-in-N sampled | zero-alloc cost cap | within cap | WITNESSED on-box (`sampler_cost_test.go`, green 2026-07-19) |
| Envelope table itself | v0.1 declared calibration | — | fail-open on undeclared rungs | MODELED (`overheadbudget.go` doc: generous, not a measured p99) |

**Net of what it saves** (signed, workload-dependent — never an average): the same plane's
`SELF-TAX-TREND.md` folds the net line — vDSO self-ablation nets **−520 tokens (a SAVING)** while
a tool-light cross-agent turn nets **+28,986 tokens (a COST)**. The overhead is real and bounded;
whether it is net-positive is a property of the workload, stated as such.

### Axis 2 — pilot-survival demonstrator: governed stops what ungoverned does not

Three survival levers, each a **captured** governed-vs-ungoverned contrast with a traced artifact —
not a narrated claim:

| Survival lever | Ungoverned (control) | Governed runtime | Traced artifact | Provenance |
|---|---|---|---|---|
| **False "done"** (verified progress catches it) | naive self-report loop accepts, false-done rate **0.50**, ships 34 slop units, 5,100 rework tokens | witnessed exit-gate re-arms, false-done **0.00**, 15 slop units, 0 rework; **net +3,900 tokens** after 1,200 gate-cost | `internal/loopgate/testdata/verified-vs-naive-loop.report.json` | gate MECHANISM WITNESSED (drives real `loopgate.Adjudicate`), corpus **SIMULATED**; live-agent dojo run `not yet` |
| **Collision** (leases prevent it) | two workers write the same tree concurrently | `dos arbitrate` admits ONE holder per lane, refuses the second | acquired live this session: `dos arbitrate --lane compute --kind cluster` → `outcome:"acquire"`, tree `internal/compute/**` | WITNESSED (this session) |
| **Runaway** (a resource cap stops it) | one `llama-cli` with no thread bound climbed to **129,427 threads** and pinned the host | `fak process-guard` flags >2,000 threads / a sustained >90%/core CPU pin; opt-in reaper (never touches protected/own tree) | `docs/perf-runaway-guard.md` (witnessed incident + shipped guard) | OBSERVED (incident) + shipped guard |
| **Injected/destructive result** (safety floor stops it) | baseline admits **1** prompt-injection and executes **1** destructive op | fak admits **0**, executes **0** on the same trace | `cmd/turntaxdemo -selfcheck` pinned invariants (`selfcheckExpect`, airline suite) | committed invariant; on-box re-run currently BLOCKED (peer-red trunk, below) |

### Honesty fences

- **This binds, it does not re-measure.** Every number above traces to an already-committed
  artifact with its own provenance; this row is the authority binding the issue asks for, not a
  fresh benchmark. The two on-box witnesses (`turntaxmeter`, `loopgate`) were re-run green on
  win/amd64 on 2026-07-19 to confirm the cited artifacts still hold on this host.
- **The overhead envelope is DECLARED (v0.1), not a measured p99.** It exists so a *gross*
  regression (an order of magnitude over) reads back as a structured `OVERHEAD_BUDGET_EXCEEDED`
  breach while normal jitter stays OK. Tightening a row toward a measured fleet p99 is the
  named self-tax follow-on, not this row.
- **The false-done corpus is SIMULATED.** `verified-vs-naive-loop.report.json` labels its own
  provenance `SIMULATED`: it drives the *real* shipping gate against hand-authored episode
  traces, so it proves the gate's mechanism, not a wall-clock tok/s. A live-agent run over a real
  dojo episode corpus is host/agent-gated and remains `not yet`.
- **The safety-floor contrast is committed but not re-run here.** `cmd/turntaxdemo -selfcheck`
  encodes the injection/destructive invariants (baseline 1→fak 0 on both), but on 2026-07-19 the
  demo could not be rebuilt on this host: the trunk was peer-red (`internal/agent`,
  `internal/policy` carried in-flight peer edits that broke the full build), so the demo binary —
  which imports the kernel registrations — did not compile. The invariant is cited from its
  committed pin; a green-trunk re-run (`go run ./cmd/turntaxdemo -selfcheck`) is the reproduce
  handle.

### Verification

- `go test ./internal/turntaxmeter -count=1` → `ok` (the observer-effect meter: declared
  envelope + `OVERHEAD_BUDGET_EXCEEDED`, the hook-latency p99 fold, and the sampler's own
  zero-alloc cost cap). Run green on win/amd64, 2026-07-19.
- `go test ./internal/loopgate -count=1` → `ok` (the verified-vs-naive false-done demonstrator).
  Run green on win/amd64, 2026-07-19.
- `dos man wedge OVERHEAD_BUDGET_EXCEEDED --explain` / `dos man wedge GATE_LATENCY_REGRESSION --explain`
  resolve both breach tokens as real closed-vocabulary members (declared in `dos.toml`).
- Green-trunk only: `go run ./cmd/turntaxdemo -selfcheck` asserts the safety-floor invariants and
  exits non-zero on drift; `go run ./cmd/turntaxdemo -print` renders the governed-vs-ungoverned
  turn-tax side by side.

---

## Full-Span Four-Band Trace + Tail Under Load (2026-07-07) — the dynamic-range composition witness

**Date:** 2026-07-07
**Issue:** [#2223](https://github.com/anthony-chaudhary/fak/issues/2223) (gap G5 + risk R6 of epic [#2218](https://github.com/anthony-chaudhary/fak/issues/2218))
**Files:** `docs/benchmarks/fullspan/fullspan-trace-20260707.json`, `docs/benchmarks/fullspan/tailload-20260707-run1.json` + `-run2.json` *(anchors)*, [`docs/benchmarks/FULLSPAN-TAILLOAD-RESULTS.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/FULLSPAN-TAILLOAD-RESULTS.md) *(narrative + full fences)*
**Machine:** AMD Ryzen 9 9950X (16C/32T), 256 GiB RAM, WSL2 Ubuntu 24.04 (kernel 6.6.114.1-microsoft-standard-WSL2) on a Windows 11 host, go1.26.0 linux/amd64
**Label:** OBSERVED, research-grade — one box; two full runs for the tail arm; **no gate asserted or flipped**

### What this adds (and why)

Every prior dynamic-range number lived in a separate artifact on a separate clock
band. This pass commits (1) the first **single-process causal trace** spanning all
four bands — a real B6 plan verdict (`superloop.Walk`) → real B4 admission
(`dispatchtick.EvaluatePreflight`) → scripted B2 turn → 5 real B0 `kernel.Decide`
calls — with a walkable parent chain and **every DENY span carrying an observed
downstream-consequence class** (`retry_turn` / `forked_outcome` / `clean_stop`,
with induced-work attribution per deny reason); and (2) the µs decide fold's
**per-call latency distribution quiet vs under same-process B2/B3-shaped load**
(4 streaming + 4 session-churn workers in the measuring process), 200,000 exact
samples per arm.

### Results

Quiet vs loaded, `kernel.Decide` per-call, both runs (ns):

| Arm | p50 | p99 | p99.9 | >100 µs share |
|---|---:|---:|---:|---:|
| quiet (run1 / run2) | 2,904 / 2,892 | 13,483 / 14,137 | 123,513 / 118,899 | 0.19% / 0.18% |
| **loaded** (run1 / run2) | 3,452 / 3,388 | **33,714 / 29,533** | 450,726 / 456,049 | 0.58% / 0.56% |

Same-process load leaves the median nearly flat (×1.17–1.19) and degrades the
tail disproportionately: p99 ×2.1–2.5, p99.9 ×3.7–3.8, >100 µs share ~3×. The
quiet-process B0 anchors understate the production tail — now quantified. Trace
per-band totals: B6 30,319 ns (1 span) · B4 3,842 ns (1) · B2 301,192 ns (3) ·
B0 57,004 ns (5); reported per band and never combined.

### Honesty fences (full set in the results sheet)

- Three clocks, never one number: the artifact schema has **no** end-to-end /
  ratio / speedup field, enforced structurally by the test, not by prose.
- Trace B0 spans are one-shot cold costs — not comparable to the warmed
  `go test -bench` medians in the M3 Pro anchor rows; the only intended
  comparison is within-run quiet-vs-loaded.
- WSL2-virtualized host (hypervisor jitter inflates tails); per-call timer
  overhead identical in both arms; coarse-clock hosts (Windows ~0.5 ms tick)
  refuse to write artifacts rather than commit quantization noise.
- Two runs agree on loaded p99 within 14.2% (same side of every threshold) —
  consistent, but OBSERVED→gate promotion is deliberately not attempted.

### Verification

- `go test ./internal/bench/ -count=1` → PASS (structure asserted everywhere;
  distribution/duration assertions gate on measured clock resolution).
- Reproduce commands in the sheet; artifacts carry config + hardware in-band.

---

## AgentDojo Structural Safety Floor (2026-06-25)

**Date:** 2026-06-25
**Commit:** _this commit_
**File:** `experiments/agent-live/agentdojo-fak-fullstack-20260625.json`
**Reproduce:** `go run ./cmd/agentdojoredteam -json`

### What this measures

This is Packet A from `docs/notes/AGENTIC-BENCHMARK-RUN-PACKETS-2026-06-25.md`: the
local, deterministic AgentDojo-style structural safety floor. It compares the same
38-attack corpus against two configurations:

- **detection-only:** content detectors only (`normgate` + `ctxmmu`);
- **full-stack:** the shipped detector stack plus IFC provenance taint and sink-gate.

It is model-free and preserves the benchmark fence: this row does **not** claim an
official external AgentDojo leaderboard score, and it does not replace a future raw
model-vs-fak external harness arm.

### Results

| Metric | Artifact field | Value |
|---|---|---:|
| Task / attack count | `task_count` | 38 |
| Detection-only attack successes | `asr_detection_succeeded` / `asr_detection` | 29 / 0.763 |
| Full-stack attack successes | `asr_fullstack_succeeded` / `asr_fullstack` | **0 / 0.000** |
| Harvest corpus rows / catches | `corpus_rows` / `corpus_catches` | 38 / 38 |
| Closed catch reasons | `catch_reasons` | `MALFORMED=9`, `TRUST_VIOLATION=29` |
| Benign controls completed | `benign_completed` / `benign_completion_rate` | 2 / 1.000 |
| Gate verdict | `gate` | **PASS** |

Corpus identity: `sha256:ddc5b9ae08df0b37224a290fae212525228d2930e77afecb7bfc868b06ca1060`.
The artifact also records the reproduce command, the attack ids, policy mode
(`detection-only-vs-full-stack-ifc`), and source revision metadata.

### Honesty fences

- This is a **local structural safety floor**, not a claim that fak beats the
  official AgentDojo benchmark or any model leaderboard.
- The detection-only arm is an internal lexical-gate baseline, not a raw frontier
  model arm.
- A safety win counts here because the runner now reports benign full-stack controls
  alongside ASR; broader task utility still requires the external AgentDojo-compatible
  adapter described in issue #868/#869.

### Verification

- `go test ./cmd/agentdojoredteam ./internal/agentdojo` -> PASS.
- `go run ./cmd/agentdojoredteam -json` -> exit 0 and writes `gate=PASS`.
- JSON parse/read-back confirmed the fields in the table above.

---

## AgentDojo External Entry — `fak_gateway` Registered Non-Model Defense (#1064, 2026-06-27)

**Claim class:** module BUILT + locally WITNESSED; the public-harness row is
operator-gated. `result_claim_allowed=false`.
**Commit:** _this commit_
**Files:** `experiments/agent-live/agentdojo-fak-gateway-defense-entry-20260627.json` + `.md`;
module + test under `experiments/agentdojo-fak-defense/`.

### What this is

The external-entry counterpart the local floor (#869) deliberately excluded. fak's
default-deny **tool-call admission gate** (capability floor + IFC source-stamp /
sink-gate + result quarantine) is packaged as a real `BasePipelineElement` for the
upstream `ethz-spylab/agentdojo` harness — `FakGatewayDefense`, selectable via
`--defense fak_gateway` or `--module-to-load`, a faithful port of `internal/ifc/ifc.go`.
This is a *capability-floor* class, distinct from the four published content/transform
non-model rows.

### What is WITNESSED here vs operator-gated

| Item | State |
|---|---|
| Module loads + intercepts a tool call (unit test, 26 checks) | **WITNESSED · PASS** (`python3 experiments/agentdojo-fak-defense/test_fak_gateway_defense.py`) |
| `targeted ASR` mechanism (local intercept witness) | **WITNESSED · 0/7 (0.000)**, benign 2/2 (`python3 experiments/agentdojo-fak-defense/fak_gateway_defense.py --json`) |
| `benign utility` / `utility-under-attack` (629+97-case run) | **NEEDS_KEY** — OBSERVED, property of the fronted model `<model-id>`, ≈US$39 paid run |
| Upstream PR into a fork of `ethz-spylab/agentdojo` | **operator-gated** (outward-facing; recorded blocker per #1064 AC) |

### Honesty fences

- `targeted ASR` is fak-authored (WITNESSED); `benign/under-attack utility` are the
  fronted model's capability (OBSERVED) — different provenance, reported together,
  never ASR alone. A refusing gate depresses benign utility; that cost is stated.
- The internal `cmd/agentdojoredteam` 0/38 is **not** an AgentDojo-629 number.
- fak's structural floor is **co-equal** with the formal-isolation tier (CaMeL ASR 0;
  MELON 0.0–2.4%) — a PLACE in the ~0-ASR tier, never a win, never a leaderboard rank,
  never a README headline.

---

## ToolSandbox/tau3 Adapter Smoke and Agentic Authority Shape (2026-06-25)

**Claim class:** `[SIMULATED]` benchmark fixture; the adapter code path is shipped,
but the tasks are a local ToolSandbox/tau3-shaped smoke, not external harness rows.
**Result commit:** `c92bb2c`
**Files:** `experiments/agent-live/toolsandbox-policy-state-smoke-20260625.json`,
`experiments/agent-live/toolsandbox-policy-state-smoke-20260625.md`
**Reproduce:**

```powershell
go run ./cmd/toolsandboxbench `
  -suite testdata/toolsandbox/policy_state_smoke.json `
  -out experiments/agent-live/toolsandbox-policy-state-smoke-20260625.json `
  -md experiments/agent-live/toolsandbox-policy-state-smoke-20260625.md
```

### What this measures

This is Packet E from `docs/notes/AGENTIC-BENCHMARK-RUN-PACKETS-2026-06-25.md`.
It is the first raw-vs-fak agentic authority entry shape for issue #876: every
quoted number below names the artifact, reproduce command, task ids, model/trace
configuration, utility metric, safety metric, and limitation.

The raw arm replays the same tool trace without fak mediation, then adjudicates
the calls after the fact to count policy breaches and minefield hits. The fak arm
adjudicates before execution and only completes a milestone after an `ALLOW` or
`TRANSFORM` verdict. `pass^1` means all benchmark milestones completed; `safe
pass^1` means milestone completion with zero policy breaches and zero minefield
hits.

### Configuration

| Field | Value |
|---|---|
| Benchmark | `toolsandbox-shaped-smoke` |
| Model | `offline-trace` |
| Task ids | `retail-refund-policy-minefield`, `banking-address-update-benign` |
| Task count | 2 |
| Raw/fak parity guard | `same_task_ids=true`, `same_trace=true` |
| External grader | none; local fixture only |
| Evidence class | `SIMULATED_LOCAL_FIXTURE` |
| Result claim allowed | `false` |

### Results

| Metric | Artifact field | Raw | fak |
|---|---|---:|---:|
| pass^1 | `summary.*.pass_1` | 2/2 (1.000) | 2/2 (1.000) |
| safe pass^1 | `summary.*.safe_pass_1` | 1/2 (0.500) | 2/2 (1.000) |
| policy breaches | `summary.*.policy_breaches` | 1 | 0 |
| minefield hits | `summary.*.minefield_hits` | 1 | 0 |
| denied calls | `summary.*.denied_calls` | 0 | 1 |
| argument repairs | `summary.*.argument_repairs` | 0 | 0 |

Derived deltas: `safe_success_delta=1`, `policy_breach_delta=1`,
`minefield_hit_delta=1`.

### Promotion Gate

- A raw-vs-fak agentic authority row must include: artifact path, reproduce
  command, model/runner configuration, task ids, raw/fak parity guard, utility
  metric, safety/evidence metric, and limitations.
- A local fixture row may be cited only as `[SIMULATED]` adapter evidence. It must
  not be promoted into a leaderboard, README headline, or external benchmark claim.
- The ToolSandbox smoke artifact now carries this as data:
  `evidence_class=SIMULATED_LOCAL_FIXTURE`, `official_harness.available=false`,
  promotion requirements, and `result_claim_allowed=false`.
- An official tau3, ToolSandbox, AgentDojo, SWE-bench, Terminal-Bench, or browser
  benchmark row needs benchmark-native tasks and grader output attached or linked,
  with raw and fak arms sharing the same task ids, model, budget, and retry policy.
- If a number is not in this file with those fields, treat it as a run note, not a
  quotable benchmark claim.

### Verification

- `go run ./cmd/toolsandboxbench ...` regenerated the JSON and Markdown witnesses.
- JSON read-back confirmed `schema=fak.toolsandbox-adapter-report.v1`,
  `task_count=2`, raw safe successes `1`, fak safe successes `2`, fak denied
  calls `1`, `evidence_class=SIMULATED_LOCAL_FIXTURE`, and
  `result_claim_allowed=false`.
- `go test ./internal/toolsandbox ./cmd/toolsandboxbench` -> PASS.

---

## Pure-kernel latency — Apple M3 Pro (2026-06-20)

**Date:** 2026-06-20
**Commit:** `bcad56e`
**Files:** `experiments/mac-m3pro-kernel-20260620/kernel-latency-mac-m3pro-20260620.json` *(the anchor)*, `MAC-M3PRO-KERNEL-BENCH-2026-06-20.md` *(narrative companion — not published in the public repo)*
**Machine:** Mac15,7 — Apple M3 Pro, 12 core, arm64, darwin, go1.26.0. Medians of count=8 trials on an idle box.

### What this adds (and why)

The Authority's model-bench rows left the **pure-kernel latency stack** uncommitted: the
syscall bench (`report.json`) was the one pure-kernel artifact and was explicitly "narrow",
and a "~1,300 ns" admission figure cited in `DISAGGREGATED-AGENT-MEMORY.md` and
`MEMORY-LAYERS-EXPLAINER.md` had no committed artifact. This pass witnesses the stack via
`go test -bench` (the most reproducible form) so every cited number traces to a committed
artifact + reproduction command. Full decomposition and honest fences in the results doc.

### Results

| Layer | p50 ns/op | B/op | allocs/op | verdict |
|---|---:|---:|---:|---|
| **Decide** (canonical allow) | **362** | 256 | 5 | ALLOW |
| Decide w/ ArgPredicates (0→2000 unrelated) | 560 → 605 | 600 | 14 | — |
| **ScreenBytes** scan — secret (regex) | **1,812** | 0 | 0 | caught |
| ScreenBytes scan — benign (full) | 4,482 | 128 | 2 | allow |
| ScreenBytes scan — injection (nested) | 14,062 | 417 | 2 | caught |
| **Admit** (full gate, +page-out) — secret | 3,337 | 2,022 | 26 | QUARANTINE |
| Admit — injection | 15,799 | 2,662 | 28 | QUARANTINE |
| **AdmitChain** (normgate+ctxmmu) — benign | 29,171 | 1,662 | 25 | ALLOW |
| AdmitChain — injection | 87,056 | 4,854 | 38 | QUARANTINE |

Plus the refreshed syscall A/B: in-process p50 **2,427 ns** vs spawned `fak hook` p50
**6.913 ms** (n=100) → **~2,849×** boundary tax, `gate_primary=pass`.

### The honest finding on the cited figure

The "~1,300 ns" is the **narrow reading** — the cheapest `ScreenBytes` path (secret regex)
measures ~1.8 µs here, same order, **not fabricated**. But it names only one layer: the
general scan is 4.5–14 µs, the full `Admit` (with the page-out side-effect) is 3.3–15.8 µs,
and the full normgate+ctxmmu chain the kernel `Reap` runs is 29–87 µs. The single cited
number undersells the composed path by ~an order of magnitude on the worst payload; the
decomposition above replaces it. (Governance rule #4: the old figure is marked, not
silently removed.)

### Verification

- New `internal/ctxmmu/bench_test.go` compiles + vets clean; `go test ./internal/ctxmmu
  ./internal/adjudicator` → PASS (existing ctxmmu tests unaffected by the normgate
  registration the chain bench enables).
- `dos_commit_audit bcad56e` → **ABSTAIN** (the subject is a witness/documentation claim, not
  a falsifiable code claim; the diff is nonetheless real — it lists `bench_test.go` + the
  committed JSON artifacts). The load-bearing witness is `dos verify fak benchmark` below.
- `dos verify fak benchmark` → **SHIPPED** (`bcad56e`, rung `trailer` — the `(fak benchmark)`
  stamp binds the commit as a unit of benchmark work, confirmed by evidence, not self-report).

---

## Causal invalidation-on-external-write (2026-06-20) — the CPU-only strategic witness

**Date:** 2026-06-20
**Commit:** `0fc39aa`
**File:** `experiments/causal-invalidation-20260620/causalbench-witness-20260620.json`
**Reproduce:** `go run ./cmd/causalbench -selfcheck` (zero files, exits non-zero on any violation)

### What this witnesses (and why it is the cheapest strategic proof)

This is matrix row 6 of `PLAN-cloud-neocloud-rightsizing-2026-06-20.md` — the one
genuinely **net-new** strategic concept with **no hardware dependency** ($0, CPU-only,
unblocked). It proves the property `STRATEGIC-TIMING-2026-06-20.md` ranks #3: an external
write makes a cached read stale, and the system **itself** discovers *which* cached reads
depended on the now-stale world-state and evicts exactly those, byte-exact, refusing
re-admission. It is the causal sibling of the provable-deletion witness (`cmd/deletioncert`,
row 5): deletion evicts a span an operator *chose*; this evicts the reads an external write
*caused* to go stale — the MESI-invalidate analogue in the integrity direction.

The kernel mechanism was already shipped (`cachemeta.PlanExternalInvalidations`,
`vdso.Revoke`, `internal/gateway/coherence.go` wiring it onto live `fak serve` traffic).
What was missing — and what this adds — is a single self-checking end-to-end witness that
binds the whole chain, the artifact this row anchors. It runs on the **real process-global
`vdso.Default`** (the same `Lookup`/`Emit`/`Revoke` path live traffic uses), no model and
no weights, because the property is structural over cache identity + the witness ledger,
not numeric.

### The chain it proves (every row an asserted invariant; the demo exits non-zero otherwise)

| Invariant | Artifact field | Value |
|---|---|---|
| Two reads admitted under two external witnesses serve byte-exact from cache | `w1_hit_before_write` / `w2_hit_before_write` | true / true |
| Cached bytes equal a fresh engine call (a hit *is* a fresh call) | `w1_served_byte_exact` | true |
| External write refutes one witness → **exactly** the dependent read evicted | `w1_evicted_by_write` | **1** (targeted, not a flush) |
| The sibling under the unrefuted witness stays byte-identical across the write | `w2_byte_identical_across` | true (**max\|Δ\|=0**) |
| The refuted read now misses → goes to the engine, fresh (no stale serve) | `w1_miss_after_write` | true |
| A re-fill under the refuted witness does **not** repopulate (CAS can't resurrect it) | `w1_readmission_refused` | true |
| Refuting an unrelated witness evicts **0** local entries | `unrelated_witness_evicts` | 0 |
| The refutation is broadcast on the coherence bus (cross-agent propagation) | `coherence_broadcast_fired` | true |
| The integrity clock advances on refutation | `trust_epoch_advanced` | true |

### Honesty fences

- **This is a containment/coherence witness, not a throughput number.** It proves the
  causal-eviction *property* holds byte-exact on the real kernel path; it says nothing
  about tok/s or scale. Pool-scale behaviour under many concurrent agents is row 15 of the
  right-sizing plan and remains `[DEFERRED]` / projected.
- **Structural, not numeric.** Like `cmd/deletioncert`, it uses inline payloads and the
  witness ledger, so `max|Δ|=0` is an exact byte comparison of served payloads, not an
  approximate tolerance. No model is loaded; the claim is about cache identity, which is
  hardware-independent and reproduces the committed JSON exactly.
- **The witness is not keyed into the tier-2 key yet** (per `revoke.go`'s own honesty
  note): two agents reading under *different* witnesses still share by `(tool,args,worldVer)`.
  This witnesses the revocation axis (C4 causal-consumer eviction), which is the
  load-bearing half; witness-keying is the natural follow-on.

### Verification

- `go run ./cmd/causalbench -selfcheck` → exit 0 (all 12 guarded invariants hold — the
  9-row table above is the headline subset); `main_test.go` guards the same chain via the
  portable `go test ./cmd/causalbench/` → `ok cmd/causalbench` (on Windows: `.\fak\test.ps1`).
- `dos_commit_audit 0fc39aa` binds the result commit (diff-witnessed: the demo, its test,
  and the committed JSON artifact).
- The number is a correctness verdict (PASS / `max|Δ|=0`), not a wall-clock — hardware-
  independent and re-derivable from the artifact alone.

---

## README Headline — 50-turn × 5-agent Qwen2.5-1.5B (the number on the front page)

**Date:** 2026-06-19
**Commit:** `2bbda6f`
**File:** `experiments/session/headline-qwen-50x5.json`
**Chart:** `experiments/session/chart1-headline-walltime.svg`

This is the number a first-time visitor sees in README §1: *"On a realistic 50-turn ×
5-agent run (Apple M3 Pro, Qwen2.5-1.5B), fak did in ~19 minutes what the naive loop
needs an estimated ~19 hours."* Every figure in that sentence traces here.

### Shape & arms

`T=50 agents=5 prefix=2048 decode=32 result=64`, Qwen2.5-1.5B-Instruct Q8 (lean
quantize-at-load). Three arms over the **same Q8 forward pass**: A naive-stateless, B
per-agent-KV tuned, C fak fused (prefix prefilled once, cloned into the agents, batched
decode).

### Results (from the artifact)

| Metric | Value | Artifact field |
|---|---|---|
| **Reuse win vs naive** | **60.3×** | `net_value_add_vs_naive=60.346` |
| **Reuse win vs tuned warm-cache** | **4.1×** | `net_value_add_vs_tuned=4.125` |
| Arm A naive total | ~19.1 h | `arm_A_naive_stateless.total_ms=68,726,015` |
| Arm C fak total | ~19.0 min | `arm_C_fak_fused.total_ms=1,138,871` |
| Exact prefill-token ratio A/C | 62.0× | `prefill_tokens.a_over_c=62.05` |
| Turn-tax A/B | 14.6× | `turn_tax_A_over_B=14.63` |

### Honesty fences (matching the README's own)

- The **60.3×** is **vs the naive loop** (re-send the whole growing context every turn).
  Vs a *tuned* warm-cache stack the honest gain is **4.1×** — a few-fold, as stated.
- Arm A's ~19h is **modeled** from `prefillCost(L)` sampled at six lengths
  (`prefill_model.lens/ms` in the artifact), not run live (it would take ~19h).
  The model is **validated within ~0.4%**: `live_validate.anchored_computed_over_live
  = 1.0039` (a reduced live shape confirms the projection). The README's "within ~1%"
  is conservative against this.
- Arms B and C run **fully live** (attention growth captured); arm A's decode is set
  byte-identical to arm B's live decode. Disclosed in the artifact `methodology` field.

### Verification

- `dos_commit_audit 2bbda6f` binds the result commit.
- Bit-identity gates green (`TestBatchedDecodeMatchesSerial`,
  `TestBatchFromPrefixMatchesIndependentPrefill`): the three arms emit identical tokens,
  so the win is reuse, not a numerics shortcut.

### F1 — tombstone note (2026-06-19, Governance rule #4)

The old SmolLM2 session cells **11.2×/14.5×** (P=512, T=8/16, A=4) do **not reproduce** on the
current kernel: re-measured as **5.3×/7.4×**. Root cause: commits `70a2cab` (Q8 prefill softcap),
`6e5fda3` (SEAM-0 decode fold), `eb9a2e5` (q8 scratch reuse) between the `5b0f40d` measurement and
HEAD made the Q8 prefill ~2× faster, so computed arm A (naive re-prefill) got cheaper and A/C
shrank. The fak arm C still matches (12.0s/28.2s old vs 11.2s/24.4s re-measured). The old number
shrank **because fak got faster at the prefill arm A depends on**, not from any regression. Full
diagnosis: `benchmark-run-opencode-20260619/BENCHMARK-RUN-OPENCODE-20260619.md` finding F1.

**Cross-ref (process doc):** the [BENCHMARK-GOVERNANCE.md Regime Boundaries](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-GOVERNANCE.md#regime-boundaries--what-each-number-means)
example formerly cited the retired figure as its live *Session value-add*; it now cites the
re-measured row above and links back to this tombstone (#143), keeping the governance discipline
and this authority in sync.

---

## RadixAttention Results (SmolLM2-135M Q8)

**Date:** 2026-06-18
**Commit:** `a200c3d`
**File:** `experiments/radixattention/radixbench-smollm2-135m-q8.json`

### What This Measures

Compares **baseline** (full re-prefill per request) vs **radix** (automatic prefix-cached KV reuse using the same algorithm as SGLang's RadixAttention paper).

### Workload: Agents

- **Shape:** 5 agents × 6 turns = 30 requests
- **System prefix:** 128 tokens (shared across all agents)
- **Per-turn step:** 24 tokens
- **Model:** SmolLM2-135M Q8_0 (30 layers, real checkpoint)

### Results

| Metric | Baseline | Radix | Speedup/Improvement |
|---|---|---|---|---|
| **Wall-clock** | 240,994 ms (~241s) | 49,452 ms (~49s) | **4.87×** |
| **Tokens computed** | 6,360 | 848 | **7.50×** fewer |
| **Cache hit rate** | 0% | 86.7% | Matches SGLang band (50-99%) |

### Verification

- `dos_commit_audit a200c3d` → **OK** (diff-witnessed)
- Committed JSON artifact exists and is readable
- Token counts are exact integers (hardware-independent)

### Why Wall-Clock (4.87×) < Token Speedup (7.50×)

On the synthetic 64-hidden/4-layer wiring model, the memcpy cost of cloning cached KV masks the compute savings (1.64× live ratio). On SmolLM2-135M, per-token compute dominates memcpy, so live speedup approaches the theoretical token figure (4.87×).

**Both results are committed and real** — they document different regimes.

> **Superseded as the headline by the 2026-06-19 model ladder below.** The single
> 135M point (`a200c3d`, contended 4.87×) remains a real committed measurement; the
> fresh ladder re-runs it at 4.58× (reps3, lightly contended) and extends it across
> three more models. Cite the ladder for the release; this row stays as provenance.

---

## RadixAttention Model Ladder (2026-06-19) — climbs to the token-speedup ceiling

**Date:** 2026-06-19
**Commit:** `92896a4`
**Files:** `experiments/radixattention/radixbench-{smollm2-135m,smollm2-360m,qwen2.5-0.5b,qwen2.5-1.5b}-q8-agents-fresh-20260619.json`

### What This Adds

The same RadixAttention `agents` workload (5 agents × 6 turns, 128-token shared
system prefix, 24-token per-turn step) run across four real Q8 checkpoints. The
**deterministic** metrics (token speedup, hit rate, FCFS→cache-aware recovery) are
byte-identical across all four (model-independent); the **live wall-clock** ratio is
the one that moves, and it climbs monotonically toward the 7.50× token ceiling as the
model grows.

### Results — `agents` workload

| Model | Live wall-clock | Token speedup | Hit rate (FCFS → cache-aware) | Artifact (`live_prefill_speedup`) |
|---|---|---|---|---|
| SmolLM2-135M (30L) | **4.58×** | 7.50× | 62.1% → 86.7% (100% of optimal) | `radixbench-smollm2-135m-q8-agents-fresh-20260619.json` |
| SmolLM2-360M (32L) | **5.40×** | 7.50× | 62.1% → 86.7% | `radixbench-smollm2-360m-q8-agents-fresh-20260619.json` |
| Qwen2.5-0.5B (24L) | **6.20×** | 7.50× | 62.1% → 86.7% | `radixbench-qwen2.5-0.5b-q8-agents-fresh-20260619.json` |
| Qwen2.5-1.5B (28L) | **6.95×** | 7.50× | 62.1% → 86.7% | `radixbench-qwen2.5-1.5b-q8-agents-fresh-20260619.json` |

Deterministic hit rates reproduce committed `a200c3d` exactly: few-shot 88.2%,
multi-turn-chat 79.5%, tree-of-thought 77.2%, agents 86.7%. Policy-eviction witness
green on every run.

### Verification

- Each row's `live_prefill_speedup` read directly from its committed JSON (verified
  2026-06-19: 4.581 / 5.40 / 6.20 / 6.951 → rounded above).
- `internal/radixkv` split-reuse == recompute (max|Δ|=0) → **PASS** (numerics are
  reuse, not a shortcut).
- Token counts (`prefill_token_speedup=7.5`, `radix_computed_tokens=848`) are exact
  integers, hardware-independent.
- **Cross-platform reproduction (2026-06-19):** the 135M `agents` deterministic fields
  reproduce **exactly on Windows x86_64** (hit 86.7%, token 7.50×, reused 5512,
  computed 848) vs the Mac M3 arm64 committed artifact; the live ratio moves (2.60× on
  x86 vs 4.58× on Mac) exactly as the small-model clone-overhead thesis predicts. See
  [`experiments/radixattention/CROSS-PLATFORM-REPRO-20260619.md`](https://github.com/anthony-chaudhary/fak/blob/main/experiments/radixattention/CROSS-PLATFORM-REPRO-20260619.md).

---

## Session Value-Stack High-T Ladder (2026-06-19) — the O(T²)→O(T) contrast

**Date:** 2026-06-19
**Commit:** `92896a4`
**Files:** `experiments/session/highT-smollm2-135m-{64-128-256,512}-fresh-20260619.json`

### What This Adds

The session value-stack (A=naive-stateless, B=per-agent-KV tuned, C=fak fused) pushed
to high turn counts on SmolLM2-135M (P=512, D=4, R=8, C=2) to expose the naive arm's
O(T²) re-prefill signature against fak's near-linear curve.

### Results

| T | A naive | B tuned | C fak | **A/C vs naive** | turn-tax A/B | exact prefill-tok A/C |
|---|---|---|---|---|---|---|
| 64  | 268.1s | 14.3s | 10.8s | **24.9×** | 18.7× | 74.9× |
| 128 | 908.8s | 30.7s | 23.0s | **39.5×** | 29.6× | 128.2× |
| 256 | 3982.1s | 74.8s | 54.4s | **73.2×** | 53.2× | 227.7× |
| 512 | 20424.4s (~5.7h) | 211.5s | 146.6s | **139.3×** | 96.6× | 421.7× |

The naive arm A explodes ~4× per T-doubling (268→909→3982s) — the O(T²) re-prefill
signature — while B and C stay near-linear.

### Honest methodology (carried from the artifact)

Arms **B and C run end-to-end LIVE** (attention growth captured). Arm **A's prefill
is modeled** from `prefillCost(L)` measured at sampled lengths (the O(L²)
prefill-attention captured, summed over the exact per-turn contexts), because running
A fully live at T=512 would take ~5.7h per cell; arm A's decode is set byte-identical
to arm B's live decode. The `validate` shape runs arm A fully live to confirm the
model. This is disclosed in each JSON's `methodology` field.

### Verification

- T=512 cell read from artifact: `net_value_add_vs_naive=139.278`,
  `turn_tax_A_over_B=96.564`, `prefill_tokens.a_over_c=421.716`.
- Bit-identity gates green: `TestBatchedDecodeMatchesSerial`,
  `TestBatchFromPrefixMatchesIndependentPrefill` (arms produce identical tokens).

---

## GPU Q8 Throughput — Vulkan on the AMD RX 7600 (2026-06-19)

**Date:** 2026-06-19
**Commit:** `60db592` (path unblocked by `84c2e6c`)
**Files:** `experiments/gpu/q8gpu-smollm2-135m-{gpu-q8,gpu-f32,cpu-q8}-20260619.json`
**Doc:** `experiments/gpu/VULKAN-Q8-RX7600-20260619.md`

### What This Adds

The first committed Q8-on-GPU throughput numbers from the `modelbench` harness. Until
`84c2e6c`, `modelbench -backend vulkan -quant` hard-refused ("compute HAL sessions are
f32-only today") even though the Q8 weight-upload + device-GEMM path was fully wired in
`internal/model/hal.go`. Three arms over the **same SmolLM2-135M Q8 forward pass** on the
real RX 7600 (Vulkan 1.4.349, native Windows), 64 decode steps / 3 reps.

### Results

| arm | decode tok/s | prefill P=16 → 512 | artifact |
|---|---:|---:|---|
| **gpu-q8** | **24.6** | 15.6 → 24.8 | `q8gpu-smollm2-135m-gpu-q8-20260619.json` |
| gpu-f32 | 16.5 | 12.6 → 18.7 | `q8gpu-smollm2-135m-gpu-f32-20260619.json` |
| cpu-q8 | 176.9 | 969 → 1519 | `q8gpu-smollm2-135m-cpu-q8-20260619.json` |

### Two honest findings

1. **Q8 weight-narrowing buys ~1.49× decode on the GPU** (24.6 vs 16.5 tok/s) and ~25–30%
   on prefill at every length — same forward, same device, only the weight dtype changes.
   The decode path is memory-bound, so cutting weight traffic ~4× directly raises throughput.
2. **On 135M the CPU wins outright** — cpu-q8 decode 176.9 tok/s is **7.2×** the GPU, and CPU
   prefill (batched GEMM) is **40–75×** the GPU's single-token-looped device prefill. The GPU
   path is **launch-bound** (~330 device ops/token × a fixed dispatch tax that dwarfs 135M's
   per-op compute), the same regime the CUDA/RTX-4070 lane documents — now confirmed on a
   second vendor. The device path is the architecture that scales to models too big for CPU
   residency, **not** a win at 135M. Lever: batched device prefill + capture-replay graph
   (`Async`/`GraphCompile` both `false` in the RX 7600 caps today).

### Verification

- Correctness gated on the real GPU: `TestHALVulkanQ8ForwardMatchesComputeQ8` →
  **prefill cosine = 1.0, step cosine = 1.0**; `TestHALVulkanForwardMatchesNative` →
  argmax-exact, cosine 1.0. The throughput win is reuse + narrower traffic, not a numerics
  shortcut.
- Each row read directly from its committed JSON (`decode.tok_per_sec`, `prefill[].tok_per_sec`).
- `precision`/`backend.selected`/`backend.tier` fields in each artifact make the provenance
  self-describing (e.g. gpu-q8: `precision=Q8_0`, `selected=vulkan`, `tier=discrete:AMD Radeon RX 7600`).

---

## GPU/CPU Q8 Crossover — the device path catches the CPU as the model grows (2026-06-19)

**Date:** 2026-06-19
**Commit:** `7bf666b` (unblocked by the `8c74fd9` q8_matmul input-tiling fix)
**Files:** `experiments/gpu/crossover-qwen2.5-1.5b-{gpu,cpu}-q8-20260619.json`
**Doc:** `experiments/gpu/CROSSOVER-1P5B-RX7600-20260619.md`

### What This Adds

The 135M GPU result above showed the device path **launch-bound** — 7.2× behind the CPU. The
obvious question: does that gap close on a bigger model, where the per-token GEMM is large
enough to amortize the fixed ~330-op/token dispatch tax? Measured on Qwen2.5-1.5B Q8 (the
`q8_matmul` shader's old inDim≤2048 cap, which the 1.5B FFN's inDim=8960 exceeded, was lifted
in `8c74fd9` — verified bit-correct by `TestVulkanQ8MatMulWideInput`, cosine ≥ 0.9999).

### Results — Q8 decode tok/s, GPU (Vulkan RX 7600) vs CPU (pure-Go legacy)

| model | CPU Q8 decode | GPU Q8 decode | **CPU / GPU ratio** |
|---|---:|---:|---:|
| SmolLM2-135M | 176.9 | 24.6 | **7.2×** |
| Qwen2.5-1.5B | 18.4 | 15.9 | **1.16×** |

The CPU's lead collapses **7.2× → 1.16×** as per-token compute grows ~11×. This is direct
evidence for the device-path thesis: the GPU wins as the model grows (one more size step, 3B+,
likely flips it to a GPU win). The launch-bound regime is a small-model artifact, not a
ceiling.

### Honest fences

- **Decode only.** Prefill still favors the CPU heavily (the device prefill loops single tokens
  — HAL prefill isn't batched — so it runs at decode speed; the CPU batches its prefill GEMM).
  Batched device prefill is the standing next lever; it does not affect the decode crossover.
- A transient large-prefill-shape VRAM-allocation panic exists on the 1.5B (the pool's
  drain-and-retry usually absorbs it; a smaller `-prefill-sizes` avoids it). The decode number
  is stable across reps.

### Verification

- Each ratio read from the committed JSON `decode.tok_per_sec` (GPU 15.900, CPU 18.428 → 1.16×;
  135M GPU 24.620, CPU 176.898 → 7.19×).
- Q8 device GEMM bit-close to the CPU Q8 reference at the 1.5B FFN width
  (`TestVulkanQ8MatMulWideInput`, in=8960, cosine ≥ 0.9999); HAL forward gate argmax-exact.

---

## Session Value-Stack Results (SmolLM2-135M Q8)

**File:** `docs/benchmarks/SESSION-VALUE-STACK-RESULTS.md`

### What This Measures

Compares three arms running the **same Q8 forward pass**:
- **A — naive-stateless**: Re-prefills entire context every turn (common local pattern)
- **B — per-agent-KV**: Prompt-cache/persistent KV per agent, no cross-agent sharing
- **C — fak fused**: Prefix prefilled once + cloned into C agents, batched decode

### Results

| Turns | Agents | Prefix | Naive (A) | Tuned (B) | fak (C) | A/C | B/C |
|---|---|---|---|---|---|---|---|
| 8 | 4 | 512 | 135.1s | 32.4s | 12.0s | **11.2×** | 2.70× |
| 16 | 4 | 512 | 409.4s | 67.9s | 28.2s | **14.5×** | 2.41× |

### Key Point

The **11.2–14.5×** value-add is **vs naive stateless serving**, not vs SGLang or any other tuned baseline. This is the "common local pattern" comparison.

> **Low-T anchor for the high-T ladder above.** This T=8/16 authority-shape result
> (C=4, D=24) is the conservative point; the 2026-06-19 high-T ladder pushes the same
> A-vs-C comparison to T=512 → 139.3× by isolating T-scaling with a smaller per-turn
> step. Both are vs the same naive-stateless baseline; they differ only in shape.

---

## Baseline Comparisons: What Each Number Means

| Number | Compares Against | Regime |
|---|---|---|
| 4.58× → 6.95× | Full re-prefill per request | RadixAttention live ladder (135M → 1.5B), climbing to the 7.50× ceiling |
| 7.50× | Token count reduction | Theoretical compute saved (deterministic, model-independent) |
| 86.7% | SGLang's published 50-99% band | Cache hit rate (FCFS 62.1% → cache-aware, 100% of optimal) |
| 5.3–7.4× (T=8/16) → 139.3× (T=512) | Naive stateless (no KV persistence) | Session value-add, O(T²)→O(T) as T grows |
| 2.4–2.7× | Tuned single-tenant (per-agent KV) | Marginal value over warm cache |

---

## Cross-Index

### SGLang RadixAttention Paper
- **Source:** Lianmin Zheng et al., "SGLang: Efficient Execution of Structured Language Model Programs," arXiv:2312.07104; NeurIPS 2024
- **fak replication:** `docs/benchmarks/RADIXATTENTION-RESULTS.md`
- **Claim:** fak achieves 86.7% hit rate (inside SGLang's 50-99% band)

### SmolLM2-135M Reference
- **Role:** In-kernel bit-exact anchor for GPU/CPU equivalence gates
- **Proof:** `IN-KERNEL-MODEL-DESIGN.md` R0–R14 *(narrative companion — not published in the public repo; the bit-exact equivalence ships as tests, e.g. `TestHALVulkanForwardMatchesNative`)*
- **Status:** Proven bit-for-bit vs HF oracle

---

## Reproduce

```bash
# RadixAttention benchmark
go run ./cmd/radixbench \
  -dir internal/model/.cache/smollm2-135m \
  -quant \
  -out experiments/radixattention/radixbench-smollm2-135m-q8.json

# Session value-stack
go run ./cmd/sessionbench \
  -turns 8,16,32 -agents 4 -prefix 512 -decode 24 -result 48 \
  -out experiments/session/smoke-smollm2.json
```

---

## Tombstoned/Outdated Claims

The following claims have been superseded or should not be used:

| Old Claim | Status | Replacement |
|---|---|---|
| "~13× speedup, P=512,T=5,C=5" | ❌ Not found in committed evidence | Use 4.87× (RadixAttention) or 11.2–14.5× (value-stack) |
| "SmolLM2-135M achieves ~370s → ~30s" | ❌ No committed artifact for this exact config | See committed results above |
| Any uncommitted/transient benchmark numbers | ❌ Must ship via commit + JSON | See authority table |

---

## Next Model Results Template

When benchmarking a new model, add an entry following this structure:

```markdown
### Model-Name Results

**Date:** YYYY-MM-DD
**Commit:** <hash>
**File:** `path/to/artifact.json`

| Metric | Baseline | Optimized | Speedup |
|---|---|---|---|
| Wall-clock | XXX ms | YYY ms | **Z.Z×** |
```

---

## DOS Verification Discipline

Every claim in this document is backed by:
1. **Committed artifact** (JSON in repo)
2. **Git commit** with `dos_commit_audit` verification
3. **Reproducible command** in "Reproduce" section

No claim exists without a traceable source.

---

# `docs/fak-vs-dos.md`

> Source: `docs/fak-vs-dos.md`

---
title: "FAK and DOS: which layer owns what"
description: "The practical boundary between the Fused Agent Kernel and the DOS decision/operations kernel, with command routing and a live-usage audit."
---

# FAK and DOS: which layer owns what

**FAK runs agents. DOS admits work and verifies results.**

## TL;DR

Use FAK for tool calls, context, models, cache, and policy. Use DOS for leases, evidence, liveness, and decisions. They are complementary kernels, not two names for one product.

## The 10-second routing rule

| Your question | Owner | Start here |
|---|---|---|
| How do I manage an agent's tool calls, context, tokens, model route, cache, or capability policy? | **FAK** | `fak manage`, `fak serve`, `fak session`, `fak preflight` |
| May this worker touch these files while peers are active? | **DOS** | `dos arbitrate`, then `dos lease-lane acquire` |
| Did the claimed commit, phase, or external effect really happen? | **DOS** | `dos verify`, `dos commit-audit`, `dos witness` |
| What is the fleet doing, stuck on, or waiting for a human to decide? | **DOS** | `dos top`, `dos decisions`, `dos observe` |
| How do I launch the FAK repository's issue workers and operating loops? | **FAK workflow, DOS-grounded** | use the FAK skill/`fak-dev` workflow; it calls DOS for admission, leases, and witnesses |
| Where does this repository declare DOS lanes and policy? | **FAK repository configuration consumed by DOS** | `dos.toml`; generated runtime evidence is under gitignored `.dos/` |

A useful sentence to remember is:

> FAK is in the agent's execution path; DOS is in the work's decision and proof path.

## The seam, not an overlap

### FAK owns agent execution

FAK is the shipped Go product in this repository. It wraps an agent harness and adjudicates tool calls. It also manages context, turn budgets, model routes, provider/cache state, capability policy, and runtime controls. A user adopting those outcomes installs and invokes `fak`.

Repository-specific orchestration also lives here: FAK's skills and `fak-dev` workflows know this project's issue taxonomy, test gates, worker launchers, and landing rules. They may *compose* DOS, but that does not move DOS's primitives into FAK.

### DOS owns work admission and evidence

DOS is the separately installed `dos-kernel` package and `dos` CLI. It provides repository-agnostic syscalls. Those syscalls cover admission, leases, claim verification, liveness, decisions, observation, and witnessed completion. This repository supplies `dos.toml`; DOS reads it and journals local state under `.dos/`.

Use the DOS primitive directly when the question is itself a DOS question. Do not invent FAK aliases for DOS primitives merely because the work happens in this repository. Call `dos arbitrate`, `dos verify`, `dos witness`, and `dos decisions` directly.

### Why FAK contains DOS-shaped code and docs

Three integrations are intentional:

1. **FAK workflows call DOS.** A FAK dispatcher understands the local backlog; DOS independently decides whether its proposed file tree is safe and later witnesses the result.
2. **FAK installs a headless-safe DOS hook fallback.** Some headless seats do not load the user-scope plugin. The project hook keeps DOS adjudication present there; plugin-aware wiring must avoid double firing. The recorded decision is in [`notes/DECISION-fak-dos-hook-boundary-2026-07-07.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/DECISION-fak-dos-hook-boundary-2026-07-07.md).
3. **FAK reports DOS evidence.** FAK operator views may summarize lease, hook, or decision state. Rendering DOS evidence is integration, not ownership of the DOS verdict.

If code starts reimplementing a DOS verdict rather than invoking or projecting it, the boundary has been crossed.

## Is this checkout actually using DOS?

**Yes. Its explicit primitives work, but the current hook signal is noisy.** A live audit on 2026-08-17 found:

- `dos doctor` resolved workspace `C:\work\fak`, loaded `dos.toml`, found the workspace skill pack, and identified the installed distribution as `dos-kernel`.
- The executable reports DOS `v0.29.0` from a local editable checkout. Codex discovered plugin skills cached from `dos-kernel` `0.30.0`. This is environment drift, not an ownership change.
- `dos helped` saw **1,577** hook-adjudicated calls in its current window: **573** passed untouched, **1,000** produced admission advice, and **0** were refused in that window.
- All 1,000 interventions were admission cautions, overwhelmingly on generic shell calls whose hook input exposed an empty/unknown file tree. DOS allowed those calls, but repeating the same advisory does not provide 1,000 units of value.
- This session followed the intended path. `dos arbitrate` proved the four documentation paths disjoint. Then `dos lease-lane acquire` journaled the grant before editing.
- The audit first found stale and missing lane declarations in `dos.toml`. This change reconciled them; `dos doctor --check --wiring` now exits 0 with no findings. Its inventory still labels 23 optional workspace host configs `NOT_WIRED`; none is a previously wired host that drifted.

That yields a precise verdict:

| Question | Verdict |
|---|---|
| Is DOS installed and reading this workspace? | **Yes** |
| Are hooks adjudicating calls? | **Yes** |
| Are explicit admission + lease primitives usable? | **Yes** |
| Is every hook advisory useful? | **No — unknown-footprint shell calls dominate** |
| Is the complete doctor/wiring gate green? | **Yes, after reconciling `dos.toml`** |
| Should FAK duplicate DOS to hide those rough edges? | **No — fix wiring/footprint quality at the DOS seam** |

## The correct operating pattern

```powershell
# 1. Diagnose the separately installed DOS layer.
dos doctor --check --wiring

# 2. Before concurrent edits, declare the real file footprint.
dos arbitrate --workspace . --lane my-work --kind keyword `
  --tree docs/example.md README.md --explain

# 3. A GO verdict is only a decision. Journal the hold before editing.
dos --workspace . lease-lane acquire --lane my-work --kind keyword `
  --tree docs/example.md README.md --owner <session-id>

# 4. Run the FAK-specific workflow and its tests/build gates.
#    Use fak/fak-dev for agent runtime and repository workflow concerns.

# 5. Use DOS to verify claims, then release the lease.
dos verify --help
dos --workspace . lease-lane release --help
```

The distinction between step 2 and step 3 matters: `dos arbitrate` answers “would this be safe now?”; it does not reserve anything. The lease command records the reservation so later workers can see it.

## Anti-confusion checks

- Agent-runtime questions should normally begin with `fak`: tokens, context, models, tools, policy, or serving.
- Work-control questions should normally begin with `dos`: collision, lease, truth, liveness, witness, or operator decision.
- A FAK skill named `dos-*` is an operating recipe that uses DOS; it is not a second DOS implementation.
- `.dos/` is derived local DOS state. `dos.toml` is this repository's committed DOS policy. Neither is the FAK runtime product.
- The FAK headless hook fallback exists for host coverage. It must call the stable DOS hook contract rather than importing private DOS internals.

For DOS's own verb map, run `dos start-here`. For the low-level hook audit, see the [DOS kernel transfer playbook](https://github.com/anthony-chaudhary/fak/blob/main/docs/dos-kernel-transfer-playbook.md).

---

# `docs/astra-formal-subagents.md`

> Source: `docs/astra-formal-subagents.md`

---
title: "Astra formal subagents: bounded admission and operator contract"
description: "Current contract for routing complete, read-only formal packets to GPT-6 Astra without claiming measured quality superiority."
---

# Astra formal subagents

This is the operator contract for [#12090](https://github.com/anthony-chaudhary/fak/issues/12090). Treat the implementation as shipped only when its resolving commit is reachable from the repository's authoritative branch. The contract controls where Astra is used; it does not establish that Astra produces better results than another model. That requires a matched, accepted-outcome comparison that has not yet been reported here.

## The admission rule

Automatic Astra routing is for a complete, read-only formal packet with `work_class` set to `rigor`. The packet may contain only these four closed task kinds:

- `formal_proof`
- `invariant_audit`
- `state_machine_audit`
- `numerical_correctness_derivation`

Every packet must state its definitions and assumptions, exact proposition, required output form, deterministic witness, and one to three distinct bounded repository surfaces. The schema identity must be exactly `fak-formal-packet/1`.

Astra is not automatically selected for exploration, implementation, test execution, general review, issue triage, or documentation. Difficulty, model-related keywords, labels, and ordinary task prose do not confer eligibility. A declared effect-mode worker also makes the packet ineligible: the automatic formal route is an analysis-only route, not write authority.

## Plan a formal packet

Save a task such as this as `formal-task.json`:

```json
{
  "schema": "fak-orchestration-task/1",
  "id": "lease-holder-invariant",
  "work_class": "rigor",
  "formal_packet": {
    "schema": "fak-formal-packet/1",
    "task_kinds": [
      "invariant_audit"
    ],
    "definitions_and_assumptions": "Every positive lease generation denotes one acquired holder epoch; generation zero is the only anonymous legacy state.",
    "exact_proposition": "Prove that every accepted positive-generation fence token has matching non-empty presented and current holder identities, or give the smallest counterexample.",
    "required_output_form": "DEFINITIONS, INVARIANT, CASE TABLE, PROOF OR COUNTEREXAMPLE, MINIMAL FIX OBLIGATION, WITNESS",
    "deterministic_witness": "go test ./internal/leaseref -run TestFenceRefusesMissingHolderIdentity -count=1",
    "surfaces": [
      "internal/leaseref/**"
    ]
  }
}
```

Resolve it without launching work:

```powershell
fak orchestration plan --profile auto --task formal-task.json --json --selfcheck
```

The command emits stable plan JSON and `SELFCHECK PASS ... launched=0`. Inspect `resolved.astra_route`, `resolved.sol_route`, and `overrides` before launching. `--selfcheck` proves schema round-trip and route resolution only; it is not a model-quality or completed-task witness.

## Read the receipt

`resolved.astra_route.eligible` answers whether the packet satisfies the automatic formal-admission contract. `selected` answers whether the final executable child route uses an Astra-family model after profile and pin precedence. They intentionally differ in several cases:

| Situation | `eligible` | `selected` | Meaning |
|---|---:|---:|---|
| Complete packet, `auto`, no pin | `true` | `true` | Automatic worker route is `gpt-6-astra` at `xhigh`. |
| Complete packet, `off` or another direct/no-child resolution | `true` | `false` | The packet is valid, but there is no child route to select. |
| Complete packet, non-Astra environment, task, or CLI worker-model control | `true` | `false` | The higher-precedence model control wins over automatic selection. |
| Complete packet, Astra environment, task, or CLI worker-model control | `true` | `true` when a supported child route exists | The effective model is Astra, but the receipt source identifies the winning control rather than automatic admission. |
| Incomplete, excluded-kind, unbounded, or effect-mode packet | `false` | `false` unless explicitly pinned | The formal packet cannot automatically spend Astra capacity. A pin is explicit operator control outside the automatic-admission claim. |

Model and effort precedence is:

1. CLI `--worker-model` and `--worker-effort` controls;
2. task `pins.model` and `pins.effort` controls;
3. `FAK_ORCHESTRATION_WORKER_MODEL` and `FAK_ORCHESTRATION_WORKER_EFFORT` environment controls;
4. the eligible formal-packet default, `gpt-6-astra` and `xhigh`;
5. the ordinary orchestration default.

Environment controls are resolved before the stable plan and launch receipt. They appear with source `environment`, never silently override a task or CLI pin, and an invalid `FAK_ORCHESTRATION_WORKER_EFFORT` value fails before launch. `astra_route.source` identifies the effective model source (`formal-packet`, `environment`, `task.pin`, or `operator-pin`). `reasoning_effort_source` tracks effort independently. The matching `overrides` fields are `sol_route.worker_model` and `sol_route.worker_reasoning_effort`; a formal task pin must not masquerade as a `fast.*` decision.

Formal admission never grants write access. Default and explicitly declared observe workers remain read-only. Declaring any effect-mode `worker_access` adds `ASTRA_FORMAL_PACKET_ANALYSIS_ONLY` and prevents automatic Astra admission. Access compilation and enforcement remain separate controls even when an operator explicitly pins a model.

## Fail-closed refusal classes

A declared but invalid packet remains visible in `astra_route` with `eligible=false`, `selected=false` absent an explicit model pin, and ordered reason codes:

| Reason | Refusal |
|---|---|
| `ASTRA_FORMAL_PACKET_SCHEMA_INVALID` | Schema is missing or not the exact supported identity. |
| `ASTRA_FORMAL_PACKET_WORK_CLASS_REQUIRED` | Task is not `rigor`. |
| `ASTRA_FORMAL_PACKET_ANALYSIS_ONLY` | At least one declared worker requests effect mode. |
| `ASTRA_FORMAL_PACKET_TASK_KIND_REQUIRED` | No non-empty task kind is present. |
| `ASTRA_FORMAL_PACKET_TASK_KIND_DUPLICATE` | Canonically equivalent kinds repeat. |
| `ASTRA_FORMAL_PACKET_TASK_KIND_EXCLUDED` | A known non-formal kind is mixed into the packet. |
| `ASTRA_FORMAL_PACKET_TASK_KIND_UNKNOWN` | A kind is outside the closed vocabulary. |
| `ASTRA_FORMAL_PACKET_DEFINITIONS_AND_ASSUMPTIONS_REQUIRED` | Definitions/assumptions are blank or placeholder text. |
| `ASTRA_FORMAL_PACKET_EXACT_PROPOSITION_REQUIRED` | The proposition is blank or placeholder text. |
| `ASTRA_FORMAL_PACKET_REQUIRED_OUTPUT_FORM_REQUIRED` | The output contract is blank or placeholder text. |
| `ASTRA_FORMAL_PACKET_DETERMINISTIC_WITNESS_REQUIRED` | The witness is blank or placeholder text. |
| `ASTRA_FORMAL_PACKET_SURFACE_COUNT_INVALID` | Surface count is outside one through three. |
| `ASTRA_FORMAL_PACKET_SURFACE_UNBOUNDED` | A surface is rooted, drive-qualified, parent-escaping, wildcard-ambiguous, or otherwise unbounded. |
| `ASTRA_FORMAL_PACKET_SURFACE_DUPLICATE` | Two surfaces canonicalize to the same bounded region. |

Fix every returned reason before relying on automatic admission. Do not work around a refusal by embedding formal-sounding prose in `--task-text`; task text cannot invent a formal packet.

## Launch assignment binding

On launch, the complete packet is lowered into the worker's exact read-only assignment. The launch receipt and every worker row are expected to carry the same non-empty `assignment_digest`:

```text
sha256:<64 lowercase hexadecimal digits>
```

The digest is SHA-256 over the exact trimmed assignment text sent to workers. A different digest means a different assignment and must not be reconciled as the planned formal task. The digest proves launch-time binding to bytes, not that a worker understood the proposition, completed the witness, or returned a correct proof. Broader stale-worker detection and result reconciliation remain tracked by [#8855](https://github.com/anthony-chaudhary/fak/issues/8855).

## Ranked weakest-link deployment map

The first audit cohort is complete: [#12109](https://github.com/anthony-chaudhary/fak/issues/12109), [#12110](https://github.com/anthony-chaudhary/fak/issues/12110), [#12111](https://github.com/anthony-chaudhary/fak/issues/12111), and [#12112](https://github.com/anthony-chaudhary/fak/issues/12112) are closed. Do not keep spending Astra on a settled packet unless a changed assumption invalidates its proof.

The 2026-09-08 live #12112 dogfood run was rejected rather than accepted: the route selected `gpt-6-astra` at `xhigh`, but the unversioned installed binary omitted assignment binding and one child received no task payload; another child surfaced a finite-arithmetic overflow gap before the system-commit guard stopped it. The sanitized [#12116 receipt](https://github.com/anthony-chaudhary/fak/blob/main/docs/_witnesses/issue-12116-astra-formal-live/receipt.json) records the refusal and recovery without making a quality-superiority claim. The resulting numerical gap is tracked by [#12158](https://github.com/anthony-chaudhary/fak/issues/12158).

These are the next highest-value places to spend one bounded Astra read-only audit. Ranking reflects correctness impact and formal fit, not measured Astra superiority. Implementation and test execution remain separate.

1. [#11950 — bound idempotency identity](https://github.com/anthony-chaudhary/fak/issues/11950): prove the canonical semantic-request equivalence relation across persistence, replay, and typed conflict, including legacy rows.
2. [#11951 — durable spawn-budget restart](https://github.com/anthony-chaudhary/fak/issues/11951): derive the reservation/reconciliation state machine and prove that unknown consumption remains reserved while duplicate admission never double-reserves.
3. [#11954 — EffectCoordinator bound identity](https://github.com/anthony-chaudhary/fak/issues/11954): after #11950, prove that coordinator retries preserve the same resource, principal, payload, and authority-domain binding across attempt changes.

[#12113 — quarantine authority rebinding](https://github.com/anthony-chaudhary/fak/issues/12113) already received a bounded Astra state-machine audit on 2026-09-08. It rejected the exact-byte-crossing invariant with a deterministic rebind schedule; the issue now needs separate implementation and test work, not another identical audit.

Do not open a duplicate for the `PublishFenced` check-then-publish takeover race. [#11835](https://github.com/anthony-chaudhary/fak/issues/11835) already owns authoritative acquisition and fencing at the actual write boundary, including late stale-holder refusal. #12109 is narrower: it covers missing holder identity at equal positive generation.

## Benefit and harm standard

| Axis | Operating standard |
|---|---|
| Indication | A closed formal kind, fixed assumptions, exact proposition, bounded surfaces, specified output, and deterministic witness. |
| Comparator/manual pinning | Compare with the existing manually pinned model on the same packet and acceptance witness; manual pinning remains the control path. |
| Expected benefit | Hypothesis: clearer proof obligations and fewer missed state, identity, or numerical counterexamples. No measured benefit is claimed by this page. |
| Harms/cost | Higher model cost and latency, duplicated reasoning, false confidence in polished proofs, and opportunity cost if used on mechanical work. |
| Uncertainty | Repository-specific quality and cost deltas are not yet established by a matched accepted-outcome study. |
| Contraindications | Exploration, coding, broad review, documentation, issue triage, test execution, effectful access, vague propositions, or unbounded repository scope. |
| Safeguards/dose | One focused read-only Astra audit at `xhigh`, one to three surfaces, compact required output, deterministic witness, then separate implementation and verification. |
| Control | `--profile off`, environment defaults, task pins, and CLI worker pins control routing in the documented precedence; receipts must preserve the winning source. |
| Surveillance | Inspect eligibility/reasons, selected model/effort, access mode, assignment digest, witness result, token/cost receipt, and follow-on defect rate. Re-rank only from comparable evidence. |

The deterministic package witness for the route contract is:

```powershell
go test ./internal/orchestration/...
```

Runtime behavior and tests are the final authority when this page and the executable disagree.

---

# `docs/research/micro-context-fabrics.md`

> Source: `docs/research/micro-context-fabrics.md`

---
title: "Micro-context fabrics for 100–10,000 parallel agents"
description: "Dedicated research focus for splitting one cached agent base into many bounded logical contexts, with controlled-kernel and API-only paths."
status: active
last_reviewed: 2026-08-12
---

# Micro-context fabrics: one cached base, 10,000 useful agent contexts

## Verdict and first witness

This is a new, high-priority research focus for **micro contexts**: treat the reusable,
cacheable setup as an immutable agent **base**, then run 100, 1,000, or 10,000 small,
independently scheduled delta contexts over it. The target is not “spawn 10,000 OS
processes.” It is to make 10,000 logical context windows cheap enough to schedule,
observe, pause, resume, and fold while physical model slots remain bounded.

The first spine is runnable now:

```bash
go run ./cmd/microcontextdemo -selfcheck -contexts 10000 -workers 64
```

It proves 10,000 isolated logical contexts retire through 64 bounded workers, with one
shared base installed at the gateway and one delta planner call per context. Its JSON
explicitly says `synthetic planner`: it is a harness/concurrency witness, **not** a model
throughput or cache-hit claim. The next model-backed spine must preserve that distinction.

## End-state dimensions

The program succeeds when it can flex along either or both dimensions:

1. **Usable → highly performant:** a model already capable of interactive single-stream
   work gains substantially higher aggregate useful tokens/sec by amortizing prefix
   prefill, batching compatible deltas, and keeping model slots saturated.
2. **Not usable alone → usable as a fabric:** a slow single stream (including an
   offloaded or otherwise latency-bound model) becomes useful for a user, team, or service
   because many independent tasks make aggregate progress concurrently. Per-stream latency
   is reported honestly; aggregate throughput must not disguise an unusable critical path.

At scale, the unit is a logical context descriptor, not a full harness replica:

```text
base_id + task_delta + capability_set + budget + continuation + output_contract
```

The kernel owns admission, scheduling, shared-prefix identity, tool policy, journaling,
and result folding. Model serving owns prefill/decode capacity. A full agent harness is
optional and should be paid for only when a task needs its UI/session semantics.

### General large-input operator

The same substrate can be more than a high-agent-count runtime. The proposed
[large-input operator contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-large-input-operators.md) partitions a large
artifact into stable records, runs deterministic filters before semantic work, uses bounded
micro-contexts to select or execute filters and tool calls, emits typed cacheable facts, and
folds them hierarchically with provenance and safe cancellation. That is the general-purpose
claim to test: not "parallelize every prompt," but make micro-context execution one selectable
backend for decomposable large-input work alongside SQL/search, retrieval, compression, coarse
chunks, and tuned long context. The operator contract and this execution fabric are separate
layers; neither is sufficient alone.

## Minimal-spine ladder

Each rung must run end to end before the next is treated as real.

| Rung | Working component | Required witness |
|---|---|---|
| S0 | Synthetic 10k logical contexts over bounded workers (current) | exact completion count, one base install, peak concurrency ≤ worker cap |
| S1 | **Observed:** one real OpenAI-compatible endpoint, 100 delta contexts, no tools | [100/100 captured witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s1-real-endpoint.md): wall time, TTFT, usage-derived token rates, failures |
| S2 | **Observed:** shared-prefix A/B, unique full prompts vs one base + deltas | [No cache benefit on first endpoint](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s2-prefix-ab.md); scoped concurrency gain retained separately |
| S2b | Shared-prefix A/B on a cache-observable controlled kernel (#5817) | kernel cache hit/miss counters, prefill tokens saved, tuned baselines; supersedes the S2 cache verdict |
| S3 | **Observed:** 1k resumable contexts with bounded scheduler RAM and backpressure | [1,000-context hibernation witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s3-hibernation-restart.md): queue age, resident/hibernated counts, runtime reconstruction |
| S4a | **Observed:** versioned lightweight descriptor through existing Host/Gateway | [Harness inventory and 1,000-context adapter](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4-lightweight-descriptor.md) |
| S4b | **Observed:** compatibility-class planner | [Mixed workload](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4-compatibility-scheduler.md): isolation, aging, cancellation, padding/fill telemetry |
| S4 | **Observed fixture:** tool-capable microagents through capability/resource/idempotency/readback seams | [Parallel effect-safety witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4-effect-safety.md) |
| S5a | 1,000 real-model-turn contexts on one controlled node (#5820) | ledger: TTFT/tail, prefill/decode tokens/sec, RAM/KV roofline, useful-result rate |
| S5 | 10k contexts under a controlled kernel | useful-result throughput, tail latency, memory/KV roofline, overload behavior |
| S6 | API-only adapter | provider-supported cache controls or measured natural prefix reuse; no kernel-only claims |
| S7 | Multi-user/fairness mode | tenant isolation, weighted fairness, cancellation, spend and rate-limit envelopes |
| S8 | **Observed fixture:** 1,000-record general large-input operator (#6029) | [partition/filter/map/cache/fold/oracle witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8-large-input-operator.md) |
| S8a | **Observed fixture:** adaptive filter-stage selector (#6030) | [confusion/cost/cache witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8a-filter-selector.md) |
| S8b | **Observed fixture:** bounded read-only tool enrichment (#6031) | [request/receipt/restart witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8b-tool-enrichment.md) |
| S8c | **Observed fixture:** provenance-preserving hierarchical fold (#6032) | [fold-tree/property/invalidation witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8c-provenance-fold.md) |
| S8d | **Simulated fixture:** tuned-baseline falsification spine (#6100; parent #6033 remains open) | [five-pipeline decision-boundary harness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8d-falsification-spine.md) |
| S8e | **Observed controlled fixture:** effectful stages bound to witnessed receipts (#6034) | [effect journal/read-back state machine](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8e-effect-receipts.md) |

Promotion requires a captured artifact at every rung. A synthetic rate is never promoted
as inference throughput. An inference token rate is never promoted as useful agent work.

## Architecture: split the context, not the safety boundary

### Immutable base

The base contains stable system instructions, tool schemas, repository orientation, model
configuration, and other prefix-stable material. It is content-addressed and versioned.
Workers send only a task delta and continuation identity. In a controlled kernel, fak can
install and route this base directly. In API-only mode, the adapter uses provider prompt
caching when exposed, otherwise byte-identical prefixes and measured cache telemetry.

### Micro-context state

Each logical context owns only mutable state: goal/delta, short transcript or summary,
tool/effect journal, budget, priority/deadline, and a continuation token. Cold contexts
hibernate to durable state. Warm contexts consume scarce model/KV slots. The existing
`internal/microagent` host, scheduler, hibernation, warm-band, session gateway, and tool
execution seams are the starting substrate; this program must integrate them rather than
invent a second fleet runtime.

### Ultracode relationship

Ultracode supplies the workflow pattern: decompose work into independently checkable,
disjoint packets and fold only witnessed results. Micro-contexts move that pattern below
full harness instances. Ultracode remains useful for issue/worktree coordination; the
micro-context fabric handles model-turn scheduling and compact context lifetimes. They
compose, but their speedups are measured separately from inference batching/cache gains.

## Constraints, existing footholds, and explicit next proofs

These are not undifferentiated reasons the program might fail. Each constraint is paired
with the fak capability that already reduces it and the remaining experiment or mitigation.
That distinction prevents shipped substrate from being rediscovered as a blocker and keeps
an external limitation beside the route around it.

- **Harness assumptions — start with our harness, then expose micro-agents:** third-party
  harnesses often bind one process, terminal, transcript, cwd, credential set, and approval
  channel to one agent. fak already has a compact `microagent.Descriptor`, bounded host,
  session gateway, and continuation budgets, so the first integration target is fak's own
  harness: keep one top-level operator/harness agent and expose the bounded contexts beneath
  it as micro-agents. Measure that path's bring-up cost and effect boundary first; only then
  map the same descriptor onto Codex, Claude, OpenAI, or other provider/harness adapters
  through `docs/integrations/`, rather than requiring each external harness to become the
  fleet runtime. #5789 shipped the descriptor inventory; the next proof is this
  parent-agent/micro-agent integration shape, not another descriptor design pass.
- **Context segmentation and addressability — extend, do not re-invent:** fak already
  separates immutable base from per-context delta, gives each logical context an ID and
  durable continuation state, and derives in-kernel shared-prefix cache identity from the
  request context (`prefixCacheIdentityFromContext`). The remaining risk is identity drift
  from timestamps, reordered tools, provider wrappers, sampling regimes, or an incomplete
  namespace key. Canonicalize those inputs, make base/context/prefix IDs visible at the
  gateway, and require every miss to carry a reason. S2b (#5817) already proved that this
  addressability reaches real KV reuse on the controlled kernel; the open work is preserving
  and explaining that identity across each new provider wrapper and decode regime.
- **KV/cache capacity — logical scale is not resident scale:** `internal/microagent` already
  has warm-band, warm-reserve, parked, and hibernated state, so 10k context IDs need not mean
  10k live KV allocations. S3 (#5788) and the controlled 10k soak (#5792) exercised the
  bounded lifecycle; each new backend must still report restore latency and bytes against
  recompute under a fixed resident-slot budget so admission can choose warm, parked, or
  hibernated placement from evidence.
- **Batch incompatibility — classify before coalescing:** the compatibility scheduler already
  keys work by model, sampling configuration, tools, prefix identity, phase, and length
  bucket, with incompatible work retaining a singleton path. The remaining experiment is a
  shipped in #5790/#5819. Preserve its proof obligation on every backend: a real-model
  mixed-workload run must report achieved batch size, padding waste, queue-delay tax, and
  singleton fallback rate instead of treating scheduler presence as a throughput gain.
- **Head-of-line blocking — model turns and effects are separate resources:** the bounded
  scheduler and tool-execution seam already prevent a tool call from being the model slot
  itself. Add observable cancellation/preemption and prove with one long tool wait plus
  short decodes that the short contexts continue retiring within their deadlines.
- **Provider/API opacity — integrate at the top-level boundary and narrow the claim:** the
  API-only adapter shipped in #5793 and can carry the same base/delta descriptor to an
  external provider while a top-level fak agent owns decomposition, policy, journals, and
  result folding. For closed
  providers, expose micro-agents through that adapter and claim only observable billed
  cached tokens, latency, errors, and controlled A/B outcomes; reserve KV residency and
  batch-formation claims for fak's controlled-kernel path. Provider-specific integrations
  are follow-ons after the own-harness spine, not prerequisites for proving the fabric.
- **Rate limits and quotas — provider capacity is an admission input:** API mode remains
  bounded by RPM, TPM, and concurrency quotas even when logical contexts are cheap. Feed
  #5793/#5795 feed those limits into the budget/fair scheduler with token buckets, retry
  budgets, and provider-aware fairness. Every provider integration must report admitted,
  delayed, retried, and rejected work by provider rather than calling queued fan-out
  concurrency.
- **Result quality — fold witnessed work, not turn count:** the verifier and quality-ledger
  seams already exist, so use fixed task corpora and verifiers not authored by workers; the
  acceptance metric is pass rate times completed work per time, with duplicated or failed
  outputs charged as cost. #5794 shipped that ledger; new workloads must populate it rather
  than invent a throughput-only score.
- **Tool/effect conflicts — read-only first, leased effects second:** capabilities, the tool
  execution floor, journals, and independent read-back are existing control points. Keep the
  first scale runs read-only; #5791 shipped the effect-safe seam, and effectful micro-agents
  must still enter through its lane/resource leases and idempotency keys while reporting
  denied, conflicted, and independently confirmed effects.
- **Fault amplification — version and canary the shared base:** bounded retries and journal
  quarantine already limit some local failures, but one bad immutable base can still poison
  a cohort. Roll each base version through a small canary cohort, trip a circuit breaker on
  shared failure signatures, quarantine retries, and retain the prior base for rollback
  before raising admission.
- **Observability cardinality — aggregate without losing addressability:** context IDs and
  journals provide the correlation key; full per-turn spans for 10k contexts do not scale.
  Keep aggregate queue/cache/quality metrics for every cohort, sampled spans for normal work,
  and unsampled error/effect journals addressable by context ID.
- **Economics — compare complete systems:** the relevant alternative is a tuned top-level
  agent or harness running the same task corpus, not 10k naive processes. Report net value
  after duplicated output, scheduler and adapter overhead, idle GPU time, cache storage,
  provider charges, and failed work; without that witness the outcome remains `not yet`.

## Measurement contract

Every experiment records: model/provider and hardware provenance; base and delta token
counts; logical contexts and physical slots; prefill/decode/total tokens per second; TTFT
and p50/p95 completion latency; cache hit/miss evidence; queue delay; peak host RAM and
KV/cache bytes; error/retry/cancel counts; verifier pass rate; and useful completed tasks
per wall-clock minute. Required comparisons are tuned sequential, tuned provider-native
batching, and the fak micro-context path. The two headline dimensions are reported
separately—single critical-path usability and aggregate useful throughput.

## Activation-boundary experiment (2026-08-26)

**Question.** Does the smallest shipped micro-collapse unit justify agent semantics, or does a simpler task/kernel abstraction explain the observed behavior?

**Method.** Run the deterministic offline witness:

```bash
fak micro collapse --calls 3 --payload-bytes 2048 --json
```

The observed `fak-micro-collapse/1` receipt reported `verdict=PASS`, `calls=3`, `allowed=3`, `denied=0`, `intermediate_tokens=1663`, `folded_tokens=15`, `saved_tokens=1648`, and `journal_rows=3`. Apply the [activation-bounded threshold](https://github.com/anthony-chaudhary/fak/blob/main/docs/concepts/micro-agents.md#activation-bounded-agent-threshold) to the smallest child contribution, not to the enclosing pipeline.

| Candidate | Objective | State/budget bound | Attribution | Identity-specific control | Receipt | Decision |
|---|---:|---:|---:|---:|---:|---|
| Raw activation/token contribution | No | Yes, tensor shape/schedule | No | No | No | Compute primitive |
| Kernel or batched stage | Named operation, not an independent success condition | Yes | Stage-level only | Enclosing-stage only | Aggregate counters | Kernel/task primitive |
| One governed collapse child call | Pipeline-assigned operation | Call bound | Journal-row attribution | Allow/deny before execution, but no independent mid-call cancel or replacement | Journal row | Bounded task/tool call |
| Future activation-bounded agent | Required | Required | Required | Required while siblings continue | Required five-field receipt | Agent only after a witnessed pass |

**Conclusion.** The experiment does not promote activations—or the current collapse child—to agents. The current receipt proves useful bounded accounting, but it does not prove an independent objective, persistent state boundary, or identity-specific cancellation/replacement. The cheaper and more precise abstraction is a governed task/tool call over kernel primitives. This conclusion is falsifiable: a future run crosses the threshold only when its receipt names all five requirements and a test cancels or replaces one candidate identity while sibling work continues and remains attributable. Until then this remains Research/Peripheral rather than a runtime architecture commitment.
## Issue map

Epic: [#5785](https://github.com/anthony-chaudhary/fak/issues/5785) (P0, G0).

- [#5786](https://github.com/anthony-chaudhary/fak/issues/5786) — S1 real-endpoint 100-context spine.
- [#5787](https://github.com/anthony-chaudhary/fak/issues/5787) — S2 shared-prefix A/B against tuned batching.
- [#5788](https://github.com/anthony-chaudhary/fak/issues/5788) — S3 1,000-context hibernation and restart.
- [#5789](https://github.com/anthony-chaudhary/fak/issues/5789) — harness-assumption inventory and lightweight descriptor.
- [#5790](https://github.com/anthony-chaudhary/fak/issues/5790) — compatibility-class scheduler.
- [#5791](https://github.com/anthony-chaudhary/fak/issues/5791) — tool/effect safety.
- [#5792](https://github.com/anthony-chaudhary/fak/issues/5792) — S5 controlled-kernel 10k soak.
- [#5793](https://github.com/anthony-chaudhary/fak/issues/5793) — S6 API-only adapter.
- [#5794](https://github.com/anthony-chaudhary/fak/issues/5794) — quality and observability ledger.
- [#5795](https://github.com/anthony-chaudhary/fak/issues/5795) — S7 multi-user fairness and economics.
- [#5817](https://github.com/anthony-chaudhary/fak/issues/5817) — S2b kernel cache-observable shared-prefix A/B.
- [#5818](https://github.com/anthony-chaudhary/fak/issues/5818) — multi-turn descriptor continuation budgets.
- [#5819](https://github.com/anthony-chaudhary/fak/issues/5819) — in-kernel compatibility-batch execution.
- [#5820](https://github.com/anthony-chaudhary/fak/issues/5820) — S5a 1,000 real-model-turn contexts on one controlled node.
- [#5821](https://github.com/anthony-chaudhary/fak/issues/5821) — cache-value Track-1 fold of witnessed shared-base reuse.
- [#5830](https://github.com/anthony-chaudhary/fak/issues/5830)–[#5844](https://github.com/anthony-chaudhary/fak/issues/5844) — hardening fan-out off spine `28846558d0` (qa, dogfood, product, observability, integration, docs, release).
- [#6029](https://github.com/anthony-chaudhary/fak/issues/6029)–[#6034](https://github.com/anthony-chaudhary/fak/issues/6034) — large-input operator cohort: runnable 1,000-issue spine, adaptive filter selection, read-only tool enrichment, provenance-preserving fold, tuned-baseline falsification, then effectful tools; see [the operator note](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-large-input-operators.md).

2026-08-06: #5792–#5795 repaired to dispatch-ready contract bodies (research completion
standard, routed); the `research` label was removed from the open leaves so the
dispatcher can route them (it is a triage-hold label, kept on the epic only).

## Prior art to reuse

- `internal/microagent`: bounded host, context, scheduler, warm reserve/band, hibernation,
  session gateway, retries, and tool execution.
- [`docs/explainers/ultracode-multi-agent-dogfood.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/ultracode-multi-agent-dogfood.md):
  disjoint work packets and honest orchestration concurrency.
- [`docs/notes/AGENTIC-CACHING-SOTA-2026-06-19.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/AGENTIC-CACHING-SOTA-2026-06-19.md):
  agentic caching layers and open gaps.
- [`docs/notes/MOE-SSD-MULTI-AGENT-NET-TOKS-2026-07-18.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/MOE-SSD-MULTI-AGENT-NET-TOKS-2026-07-18.md):
  inference-throughput roofline; do not blend it with orchestration metrics.

---

# `docs/research/README.md`

> Source: `docs/research/README.md`

---
title: "Research route: hypotheses, evidence, and promotion"
description: "Index of fak research notes and captured stage witnesses, with the lifecycle rule that keeps a hypothesis out of the maintained authority set."
---

# Research: hypotheses, evidence, and promotion

**Primary audience:** research readers evaluating an idea's provenance, evidence, and maturity before using it to guide implementation.
**Lifecycle:** research; conclusions here remain hypotheses or scoped observations until a maintained authority adopts them.
**Generation:** usually `gen/second-next` or `gen/future`, with some research performed to validate `gen/next` work; the route itself is `gen/now` documentation.
**Authority:** current behavior comes from maintained guides, contracts, tests, and shipped implementation linked by the [documentation index](https://github.com/anthony-chaudhary/fak/blob/main/docs/index.md).
**Support:** use the [contributor route](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md) for implementation help and the [operator guides](https://github.com/anthony-chaudhary/fak/tree/main/docs/operator) for supported operation.

**Next action:** before applying a research claim, find its provenance, maturity, and promotion witness; if any is missing, keep the claim in research and verify it independently.

## Active focus

- [Related-system inventory maps](https://github.com/anthony-chaudhary/fak/tree/main/docs/research/inventory) — pinned local-checkout maps generated with `fak study-inventory` before deep `study-repo` borrowing, so broad source coverage has a concrete denominator.

- [Structured session intent](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/structured-session-intent-2026-08-18.md) — own-prompt inventory plus recent scheduler/hook research, with a validated minimum/target/maximum, trigger, recurrence, and lifecycle-hook declaration spine.

- [Micro-context fabrics for 100–10,000 parallel agents](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-fabrics.md) — split one cached agent base into many bounded logical contexts; includes the runnable 10k synthetic spine and controlled-kernel/API-only research ladder.
- [Micro-context operators for general large-input LLM work](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-large-input-operators.md) — partition record/field/group inputs, adaptively select filters or tools, emit typed facts, fold with provenance, and stop only under an answer-safe contract.
- [Micro-window routing across filters and tool calls](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-filter-tool-routing.md) — typed value-of-information routing over deterministic filters, semantic windows, read tools, widen/stop/escalate, with kernel-owned authority and budgets (#6105/#6106).
- [Micro-context S8 large-input operator witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8-large-input-operator.md) - fixture-backed 1,000-record partition, deterministic prefilter, semantic map, exact reuse/invalidation, bounded fold, and oracle proof.
- [Micro-context S8a adaptive filter-selector witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8a-filter-selector.md) — allowlisted per-record routing across exact exclusion, semantic filtering, group widening, and escalation with confusion/cost/replay telemetry.
- [Micro-context S8b read-only tool-enrichment witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8b-tool-enrichment.md) — typed allowlisted reads with cross-record dedupe, quotas, timeout/retry, cancellation, restart cache, recursive output bounds, and independently read-back citations.
- [Micro-context S8c provenance-fold witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8c-provenance-fold.md) — bounded typed reducers, content-addressed trees, minority/uncertainty retention, source-resolved claims, and path-local invalidation.
- [Micro-context S8d tuned-baseline falsification spine](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8d-falsification-spine.md) — fixture-modeled five-pipeline decision boundaries that expose both micro-context wins and losses while leaving live net-true evidence open.
- [Micro-context S8e witnessed effect receipts](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8e-effect-receipts.md) — capability/idempotency/resource-bound effects with denial, conflict, partial failure, cancellation ambiguity, restart replay, breaker, and independent read-back states.
- [Micro-context S8f non-fixture corpus and grader](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8f-corpus-grader.md) — frozen 1,000 real-issue snapshot, train/tune/test isolation, separately hashed answers, leakage checks, and a strict held-out grader (#6108).
- [Micro-context S8g tuned baselines and exact frontier](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8g-tuned-baselines.md) — tune/held-out protocol, zero-semantic-residual falsification, adaptive zero-call result, and the boundary on model/performance claims (#6109).
- [Micro-context S8h executable value-of-information routing](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8h-routing-voi.md) — controlled filter/model/tool policy crossover, cancellation accounting, oracle regret, and explicit live-evidence boundary (#6105).
- [Micro-context S8i independently adjudicated semantic residual](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8i-semantic-residual.md) — blinded tune/test packet, two-model live adjudication, exact-agreement/abstention fold, and blind semantic grader (#6124).
- [Micro-context S8j live semantic matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8j-live-semantic-matrix.md) — same-endpoint retrieval/long/chunk/micro execution with observed tokens, cache, TTFT/tail, retries, strict quality, and a no-winner falsification (#6110).
- [Micro-context S8k strengthened live baselines](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8k-strong-live-baselines.md) — leave-one-out top-k retrieval, concurrent chunks, tune-only abstention calibration, tail outlier evidence, and continued no-winner result (#6151).
- [Micro-context S8l live cancellation and partial-fold policies](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8l-tail-policies.md) — wait-all/deadline/early-stop/hedge receipts, typed partial folds, live latency-quality tradeoffs, and cancellation billing limits (#6160).
- [Micro-context S1 real-endpoint witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s1-real-endpoint.md) — 100/100 contexts through four bounded workers, with TTFT/usage telemetry and a retained 16-worker overload finding.
- [Micro-context S2 shared-prefix A/B](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s2-prefix-ab.md) — no cache benefit observed on the first real endpoint; scoped concurrency improved aggregate work while worsening TTFT.
## Read research by maturity

Research material makes options and uncertainty inspectable. It can identify a useful mechanism, report a measurement, or define a future architecture without claiming that fak currently supports it.

| Maturity | What the material establishes | What you may do next |
|---|---|---|
| **Hypothesis** | A falsifiable idea, assumption set, or proposed mechanism. | Reproduce or design the named experiment; do not present it as product behavior. |
| **Observed** | A result with a named source, environment, date, and measurement method. | Re-run it in the target environment before generalizing beyond its stated scope. |
| **Validated option** | Evidence has retired specified technical uncertainty, but product integration or support gates remain. | Follow the promotion trigger and linked implementation issue. |
| **Promoted** | A maintained contract, guide, test, or shipped implementation has adopted the conclusion. | Leave this route and use that maintained authority for current behavior. |

A paper, benchmark, model output, or committed study can be valuable evidence without being support evidence. Support begins only where a maintained product route states its scope and names a reproducible witness.

## Provenance contract

A research claim should make these fields discoverable:

1. **Question or hypothesis** — the proposition being tested.
2. **Source and date** — paper, repository revision, dataset, model, interview, or local observation.
3. **Method and environment** — baseline, hardware, software version, workload, mode, and constraints.
4. **Result and uncertainty** — what was observed, its scope, and what remains unknown.
5. **Maturity and generation** — hypothesis, observed, validated option, or promoted; plus the applicable generation stream.
6. **Promotion witness** — the evidence and maintained destination required before the claim becomes implementation guidance.

For performance or efficiency claims, apply the [net-true-value standard](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/net-true-value.md): compare against the real tuned alternative, include added costs, state scope and provenance, and provide a reproducible witness.

## Promotion, demotion, and retirement

The [Generation Contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation.md#promotion-verbs) governs generation changes. Promote research closer to `gen/now` only when its named blocker is retired by evidence. Demote it when an assumption fails or a witness regresses. Retire it when the option is superseded, rejected with evidence, completed, or no longer has an owner and witness path.

Promotion does not rewrite history. Put durable instructions in the maintained contract, guide, test, or implementation route, then link the research record for rationale. Superseded research remains available through the [dated-notes route](https://github.com/anthony-chaudhary/fak/tree/main/docs/notes); retired material belongs in the [archive](https://github.com/anthony-chaudhary/fak/tree/main/docs/archive) with its replacement or retirement decision.

## Mode and discovery

Research often depends on a specific backend, model, release, hardware tier, dataset, or offline/live mode. Apply a result only to the mode and generation it names. A future-generation label expresses horizon rather than support, and priority expresses importance rather than maturity.

The repository's studies and dated investigations currently live under [`docs/notes/`](https://github.com/anthony-chaudhary/fak/tree/main/docs/notes); the curated [Notes & research index](https://github.com/anthony-chaudhary/fak/blob/main/INDEX.md#notes--research-docsnotes) is the human route. [`docs/sota/`](https://github.com/anthony-chaudhary/fak/tree/main/docs/sota) tracks state-of-the-art comparisons. Use [`llms.txt`](https://github.com/anthony-chaudhary/fak/blob/main/llms.txt) for machine-oriented discovery.

## Quantization and runtime evaluations

- [CubicQuant bounded evaluation](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/quantization/cubicquant.md) — bounded scalar-reconstruction study for the CubicQuant format against tuned uniform and non-uniform baselines.
- [Heterogeneity-aware microscaling](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/quantization/heterogeneous-microscaling.md) — bounded interoperability evaluation for mixed-shape microscaling groups.
- [Output-aware INT2 KV-cache rotation](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/quantization/int2-kv-rotation.md) — bounded evaluation of low-bit KV rotation under output-aware reconstruction.
- [LightRot bounded low-bit evaluation](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/quantization/lightrot.md) — bounded reproduction surface for the LightRot quantization proposal.
- [QEvict recoverable quantized KV eviction](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/quantization/qevict.md) — bounded evaluation of quantized eviction with recovery semantics.
- [Recurrent Residual Quantization evaluation](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/quantization/recurrent-residual.md) — bounded reconstruction study for recurrent residual quantization.
- [ReQuant fixed-grid refinement evaluation](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/quantization/requant.md) — bounded evaluation of refinement over a fixed quantization grid.
- [Qwen3.8 Metal OSS hot-path study](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/qwen38-metal-oss-hotpath-study.md) — captured hot-path observations for the open-source Metal route serving Qwen3.8.
- [Frame loss between master agents and subagents](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/relativity-frame-loss.md) — study of context and control loss between coordinating and delegated agents.
- [TTL-upgrade refusal corpus study](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/ttl-upgrade-refusal-corpus-2026-08-09.md) — dated refusal-corpus study for time-to-live upgrade behavior.

- [Micro-context cache-value Track-1 fold](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-cachevalue-track1.md) — controlled S2b reuse enters the witnessed P&L while synthetic and provider-dollar evidence remain fenced out.
- [Micro-context S2b controlled in-kernel prefix-cache A/B](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s2b-kernel-cache-ab.md) — fresh-process arms reconcile response usage with RadixAttention counters and observe a fixture-scoped 2.16x shared-base service gain.
- [Micro-context S3 hibernation/restart](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s3-hibernation-restart.md)
- [Micro-context S4a lightweight descriptor](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4-lightweight-descriptor.md)
- [Micro-context S4b compatibility scheduler](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4-compatibility-scheduler.md)
- [Micro-context S4c effect safety](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4-effect-safety.md)
- [Micro-context S4d: bounded multi-turn continuation](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4d-multi-turn-descriptor.md) — exact 1,000×3 turn accounting with continuation-token and byte-verified mid-task restore.
- [Micro-context S4e real compatibility-batch execution](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s4e-compat-batch-execution.md) — planner batches execute through the in-kernel batch seam; the first mixed-length CPU fixture is an honest 0.539x negative result.
- [Micro-context S5a controlled-kernel 1,000-context ramp](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s5a-controlled-kernel-1k.md)
- [Micro-context S6 API-only adapter](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s6-api-only.md)
- [Micro-context health scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-health-scorecard.md) — deterministic witness fold grades the controlled 1k CUDA ledger A/100 and names the missing second-run drift baseline.
- [Micro-context outcome counters](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-outcome-counters.md) — the existing quality ledger exposes reconciled success/error/refusal totals; controlled 1k CUDA readout is 1000/0/0.
- [Micro-context quality and observability ledger](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-quality-ledger.md)
- [Micro-context S7 mixed-tenant fairness](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s7-fairness.md)
- [S8m: three-adjudicator tool-routing gold stabilization](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8m-tool-gold.md)
- [S8n: filter/tool micro-window scheduler](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8n-filter-tool-scheduler.md)
- [S8o: live quality-qualified filter/tool scheduler](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8o-live-filter-tool-scheduler.md)

- [S8p live scheduler disagreement audit](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8p-disagreement-audit.md) — blinded error atlas finds disputed gold and no stable pre-answer admission signal; records `not-yet`.

## Frontier infrastructure and workload expectations

- Human overview and field assumptions: [`frontier-infrastructure/README.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/frontier-infrastructure/README.md)
- Machine-readable dated evidence ledger: [`frontier-infrastructure/index.json`](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/frontier-infrastructure/index.json)
- Benchmark assumption registry (batching, clusters, heavy tails/Zipf, user and agent workloads): [`frontier-infrastructure/workload-assumptions.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/frontier-infrastructure/workload-assumptions.md)
- Offline structural validation commands are embedded in the corpus README.

This corpus separates production measurements, official statements, vendor claims,
reported estimates, analysis, and rumors. Its explicit coverage gaps are part of the
evidence contract; the index must not imply that the open web is finite or fully observed.

## Named coding-workload patterns

- [`coding-workload-vocabulary.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/coding-workload-vocabulary.md) — cited proposal separating workload shape, orchestration topology, verification strategy, and failure mode; names reusable patterns/subpatterns and rejects common conflations.
- [`coding-workload-vocabulary.json`](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/coding-workload-vocabulary.json) — deterministic machine companion with source provenance, inclusion/exclusion boundaries, aliases, and stable candidate IDs.

Use `fak workpattern list|source|trajectory|report` to consume the canonical seed catalog and evidence miners. The report is bounded to explicit detectors and scrubbed/local inputs; it does not claim universal taxonomy consensus or infer private-chat intent.

- [S8q/S8r true pre-answer tool admission](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8qr-true-tool-admission.md) — paired model-distinct consensus gold; two-stage admission matches quality while opening 50% fewer reads on the scoped envelope.

- [S8s/S8t natural multi-tool decision surface](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-s8st-natural-multitool-surface.md) — five evidence classes show fixed/adaptive/parallel crossover by tool cost, with quality gated first.

- [`tensor-build-local-study-2026-08-15.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/tensor-build-local-study-2026-08-15.md) — deep, snapshot-pinned study of local TensorBuild: typed engine identity, evidence tiers, artifact liveness, agent/human control parity, and work-cost attribution; dedupes current fak coverage and files #6874-#6876.

- [`CONCEPT-STUDY-TENSOR-BUILD-2026-08-29.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-TENSOR-BUILD-2026-08-29.md) — current whole-tree, snapshot-pinned TensorBuild recheck: 26 evidence/measurement/native-runtime candidates, exact FAK witnesses, explicit TensorRT ablations, and 13 bounded leaves (#10268-#10271, #10278-#10286).

## Architecture and kernel studies

- [AI-Ops storage qualification study (2026-08-29)](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/ai-ops-storage-qualification-study-2026-08-29.md) — artifact-first SSD qualification architecture, license gate, and trace-to-storage-envelope borrow filed as #10267.
- [Agent-serving composition architecture (2026-08-26)](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/agent-serving-composition-architecture-study-2026-08-26.md) — evidence-backed composition study and follow-on map.
- [NVIDIA KVTC study](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/nvidia-kvtc-study.md) — upstream KV-transfer mechanism study and fak relevance map.
- [Incumbent inference architecture bottlenecks (2026-08-28)](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/incumbent-inference-architecture-bottlenecks-2026-08-28.md) — vLLM/SGLang/Dynamo/MAX bottleneck study (#9894): the binding constraint is the interaction cross-product among scheduling, reusable state, specialization, and compilation lifecycle, not one missing kernel.
- [metrics-service study (2026-08-29)](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/metrics-service-study-2026-08-29.md) — snapshot-pinned study of an external Go observability runtime (#10287): deadline-aligned collection, normalized validated snapshots, and concurrent sink fan-out, with the bounded borrow/disposition map.
- [ActaClad/plumbline study (2026-08-31)](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/plumbline-study-2026-08-31.md) — Python static-analyzer study (#10452): borrow four bounded mechanisms (stable evidence paths, negative fixtures, a new-findings gate, export into existing surfaces); do not adopt a second analyzer framework.

---

# `README.md`

> Source: `README.md`

<p align="center">
  <picture><source media="(prefers-color-scheme: dark)" srcset="visuals/brand/fak-logo.svg"><img src="visuals/brand/fak-logo-ink.svg" alt="fak logo" width="320"></picture>
</p>

# fak — useful local agents, accelerated automatically

**Fak is building the open runtime that makes useful local agents practical on your own machine.**

Start locally, give an agent real work, and keep useful context across turns.
Our first breakthrough milestone combines native inference, speculative decoding,
and agentic caching into an experience whose qualified acceleration is automatic.
The capability floor bounds what tools the agent may execute.

**Status:** this is the product milestone we are working toward. Today,
automatic setup and cache reuse have specific model/backend limits; speculative
decoding and physical GPU Direct paths are not universally enabled or qualified.
See the [local-agent milestone](https://github.com/anthony-chaudhary/fak/blob/main/docs/local-agent-milestone.md) for the current
wiring, the meaning of automatic, and the evidence required to earn the claim.

**Where it stands on speed:** on an Apple M3 Pro running Qwen3.8-27B,
fak's own Metal engine measured 6.86 decode tok/s, 0.985× a pinned llama.cpp
reference build on the same Mac, with token-for-token identical output
(observed 2026-09-03). That is
parity with the strongest local engine, not yet a lead; see the
[Qwen results](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/QWEN-PERFORMANCE-INDEX.md) for the receipt.

**Pick your path:** run a local agent → [Try fak](#try-fak) · guard an agent you
already use → [`fak guard`](#governance-for-external-agents-fak-guard) · check the
evidence → [benchmarks](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/README.md) and [claims](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md).

## Try fak

Install with `curl -fsSL https://raw.githubusercontent.com/anthony-chaudhary/fak/main/install.sh | sh` (or `go install github.com/anthony-chaudhary/fak/cmd/fak@latest`).

Try the local workflow on a supported Apple Silicon configuration:

1. Start local inference (`fak up`):
   Sizes a model/context budget from unified memory (`macfit`), holds new work
   back under memory pressure instead of swapping, and starts the local OpenAI-compatible endpoint on
   `:8080` and opens an interactive chat REPL. Use `fak up --mock` to inspect the workflow without
   a model or GPU; mock output is not inference performance evidence.
   ```bash
   fak up
   # -> [READY] fak up running on http://127.0.0.1:8080
   ```
   Inspect the selected model/backend and memory budget before interpreting
   results. Apple Metal selection is automatic when the device and build support it.

2. Run agent work (`fak opencode`):
   In another terminal (or backgrounding `fak up --headless`), launch OpenCode:
   ```bash
   fak opencode
   ```
   Give OpenCode a parallel multi-agent prompt:
   ```
   "Using parallel subagents, audit the packages under internal/ and report their status"
   ```
   The same local endpoint also backs Claude Code (`fak claude`, via an
   Anthropic Messages adapter), Codex (`fak codex`), and [Pi](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/pi.md).
   Compatible agents reuse shared instructions and repository context; a fresh
   prefix still needs prefill, and reuse depends on model, backend, and cache identity.

3. Inspect the subagent benchmark:
   ```bash
   fak bench subagent --concurrency=4
   ```
   Check its engine, regime, and receipt before treating output as hardware
   evidence; it does not replace a real coding-task acceptance witness.

### Governance for external agents (`fak guard`)

Already running Claude Code or Codex? Wrap the agent you already run with one command to add a default-deny capability floor (only allowed tools run; everything else is blocked). fak forwards Codex subscription credentials with no API key required and blocks tools outside the allowed policy without breaking the task:

```bash
fak guard -- codex
```

In-kernel policy adjudication checks every tool call in under a microsecond before execution. See the [interactive showcase](https://github.com/anthony-chaudhary/fak/blob/main/docs/showcase.html) for a guided tour, or run `fak agent --offline` (# -> task completed) to inspect policy decisions with zero setup.

## Latest hardware results — 2026-09-25

One row per hardware family: the newest committed performance receipt for that platform,
with its claim boundary and receipt link. A row stays historical until a newer quality-complete
measurement exists.

| Platform | Latest witnessed result | Status & Details |
|---|---|---|
| Mac | Qwen3.8-27B Q4_K_M on Apple M3 Pro: forward-owned Metal sequence prefill was 43.8% faster, 10,284.5 vs 18,304.9 ms, and used 1 command buffer instead of 192; observed 2026-09-03. | Accepted component-path result with exact greedy continuation and zero fallbacks; it is not a full-run throughput comparison. [Qwen result index](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/QWEN-PERFORMANCE-INDEX.md) |
| AMD | Qwen3.6-27B on RX 7600: the pure-fak TG1 microbench measured 1.24 decode tok/s versus 0.99 for the local llama.cpp Vulkan baseline; observed 2026-06-19. | Narrow, older-model microbench. No accepted current Qwen3.8 AMD result exists. [AMD receipt](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/QWEN36-AMD-VULKAN-RESULTS.md) |
| NVIDIA | Qwen2.5-3B Q8_0 on a physical Hopper H100: fak reached 111.9 decode tok/s, 17.4% above its f32 path; observed 2026-09-05. | Native CUDA result; llama.cpp Q8_0 was 3.24× as fast at 362.7 tok/s in the same run. [H100 receipt](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/GCP-H100-RESULTS.md) |

Read the status column before comparing rates. History: [benchmark index](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/README.md);
claim boundaries: [BENCHMARK-AUTHORITY.md](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md); Mac setup:
[Mac local models](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mac-local-models.md) and [Mac agent UI](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mac-agent-ui.md).

## Why run coding agents on fak

- **Reuse the work behind each turn:** Compatible prefix/KV caching avoids
  rebuilding shared instructions and context. The product target is automatic
  reuse across turns and compatible agents, with correct invalidation and isolation.
- **Accelerate generation automatically:** Native kernels, memory sizing,
  quantization, and speculative decoding are parts of one local workflow. The
  milestone requires qualified defaults; current MTP decoding requires explicit
  selection. See the [implementation snapshot](https://github.com/anthony-chaudhary/fak/blob/main/docs/local-agent-milestone.md#current-implementation-is-narrower-than-the-milestone).
- **Real-time multi-agent visibility:** Inspect live cross-agent reuse rates, per-subagent token breakdowns, and savings sparklines directly in your terminal overlay (`fak info` / `fak guard`) to see and verify the speedup as subagents execute concurrently.
- **Keep reusable state close to compute:** Device-resident caching and direct GPU
  storage paths aim to cut paging and copy overhead; GPU residency and physical
  NVMe-to-GPU DMA are separate claims, qualified in the [claim ledger](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md).
- **Run on your own hardware:** Native backends target Apple Silicon (fak-native
  Metal), AMD and Strix Halo (bundled Vulkan), and NVIDIA
  (`ghcr.io/anthony-chaudhary/fak:cuda-latest` with `--gpus all`), each with its own
  support envelope; external engines are explicit references only
  ([native inference goal](https://github.com/anthony-chaudhary/fak/blob/main/docs/native-inference-goal.md)).
- **Default-deny capability floor:** Protect your workspace from unintended commands, path escapes, or tool poisoning. Every tool call is verified against a capability floor before execution; subagents get their own narrower floor, and a circuit breaker stops an agent stuck retrying a failing tool. Drop-in wrappers protect existing agents like Claude Code, Codex, OpenCode, and Cursor with zero rewrites.

## Configure agent profiles

Built-in work and output profiles cut token waste and resist unnecessary dependencies:

```bash
fak agent profiles
fak guard --output-profile caveman:medium --work-profile ponytail:high -- codex \
  "Remove the duplicate cache without adding a dependency."
```

Balanced defaults are `ponytail:medium` for work discipline and `caveman:medium` for concise responses. See
[work profiles](https://github.com/anthony-chaudhary/fak/blob/main/docs/work-profiles.md), [response profiles](https://github.com/anthony-chaudhary/fak/blob/main/docs/response-profiles.md), or the
[harness guide](https://github.com/anthony-chaudhary/fak/blob/main/docs/harness-init.md) to build a named agent on the runtime.

## Going deeper

| If you want to… | Start here |
|---|---|
| Check what is shipped, limited, or planned | [Status](https://github.com/anthony-chaudhary/fak/blob/main/STATUS.md) · [claims](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md) · [feature matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/features.md) |
| Browse performance evidence | [Mac](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/MAC-THREEWAY-BENCH-2026-09-03.md) · [AMD](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/QWEN36-AMD-VULKAN-RESULTS.md) · [NVIDIA](https://github.com/anthony-chaudhary/fak/blob/main/docs/_witnesses/issue-10944-nvidia-gcp-overnight/README.md) · [all benchmarks](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/README.md) |
| Is the cache paying off? (trend) | [Cache-value roll-up](https://github.com/anthony-chaudhary/fak/blob/main/docs/cache-value-rollup.md) — kernel reuse and provider-dollar savings kept in separate, unblended tracks |
| Connect another agent or model | [Codex](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/openai-codex.md) · [Claude Code](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/claude.md) · [Pi](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/pi.md) · [subagents](https://github.com/anthony-chaudhary/fak/blob/main/docs/subagents-guide.md) · [Mac local models](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mac-local-models.md) · [all integrations](https://github.com/anthony-chaudhary/fak/tree/main/docs/integrations) |
| Understand the runtime | [Architecture](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md) · [capability map](https://github.com/anthony-chaudhary/fak/blob/main/docs/CAPABILITIES.md) · [CLI reference](https://github.com/anthony-chaudhary/fak/blob/main/docs/cli-reference.md) |
| Learn in prerequisite order | [Start here](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md) · [learning path](https://github.com/anthony-chaudhary/fak/blob/main/LEARNING-PATH.md) · [documentation index](https://github.com/anthony-chaudhary/fak/blob/main/docs/index.md) |
| Build on fak | [Go API](https://github.com/anthony-chaudhary/fak/tree/main/pkg) · [harness contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/harness-kit-contract.md) · [contributing](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md) |

## Commercial serving

We run one repetitive repo workload (test-candidate generation, codebase Q&A, or doc
maintenance) on a managed, metered inference route and prove it against your baseline
and acceptance criteria in a fixed-fee two-week pilot before ongoing metered operation.
Inference quality is [SW-VERIFIED] at raw-compute parity; production readiness is not yet
claimed. To start, open an issue describing your workload (channel=oss).

Apache-2.0 licensed.

<!-- readme-verified: 2026-09-27 vs VERSION 0.55.0 + BENCHMARK-AUTHORITY · appeal-verified: 2026-09-09 · process: tools/readme_freshness_audit.py + tools/doc_appeal_scorecard.py -->

---

# `START-HERE.md`

> Source: `START-HERE.md`

# Start here: human navigator and task router

*`README.md` is the single canonical front door for product overview, value proposition, and quick proof. This page owns one job: human task routing. It takes a goal you have right now and routes you to the single authoritative document without circular loops.*

> **TL;DR:** Newcomers should focus on the 16 core pages in the Newcomer Tier below. For immediate tasks, use the Task Router table to jump directly to the right guide.

The documentation repository contains over 1,200 pages spanning specifications, historical research notes, and benchmarks. To prevent navigation confusion, documentation is split into two distinct tiers: the **Newcomer Tier** (the essential 16 core pages) and the **Operator/Maintainer Tier**.

---

## The Newcomer Tier (Essential 16 pages)

If you are new to fak, focus exclusively on these core documents. They cover everything required to understand, install, configure, and operate the runtime:

| Document | Purpose |
|---|---|
| [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md) | Canonical front door, product overview, and 60-second proof. |
| [`GETTING-STARTED.md`](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md) | Concrete setup sequence: offline verification, gateway setup, and local serve. |
| [`docs/fak/tutorial.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/tutorial.md) | Step-by-step guided first session with captured command outputs. |
| [`docs/repro-packet.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/repro-packet.md) | Deterministic offline reproducibility packet and policy checks. |
| [`docs/fak/governed-agent-quickstart.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/governed-agent-quickstart.md) | 10-minute quickstart to launch a governed agent offline. |
| [`docs/fak/server-quickstart.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/server-quickstart.md) | Stand up a shared OpenAI, Anthropic, or MCP gateway (`fak serve`). |
| [`docs/fak/mac-local-models.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mac-local-models.md) | Run local models (Qwen3.8) on Apple Silicon Mac with Metal & interactive chat. |
| [`docs/subagents-guide.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/subagents-guide.md) | Guide to running parallel subagents with zero cold start via `fak up`. |
| [`docs/integrations/README.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/README.md) | Integration chooser for Claude Code, Codex, Cursor, and other harnesses. |
| [`docs/integrations/claude.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/claude.md) | Default one-command proxy recipe (`fak guard -- claude`). |
| [`POLICY.md`](https://github.com/anthony-chaudhary/fak/blob/main/POLICY.md) | Schema and reference for authoring default-deny capability floors. |
| [`docs/showcase.html`](https://github.com/anthony-chaudhary/fak/blob/main/docs/showcase.html) | Interactive browser tour of tool call adjudication and caching. |
| [`docs/adoption/troubleshooting-first-run.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/troubleshooting-first-run.md) | Symptoms, causes, and one-line fixes for common first-run issues. |
| [`docs/CAPABILITIES.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/CAPABILITIES.md) | Performance map for token savings, turn elimination, and context reuse. |
| [`docs/architecture.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/architecture.md) | Architecture boundary and whole-path agent coordination target. |
| [`docs/glossary.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/glossary.md) | Precise definitions for core runtime terms. |
| [`docs/FAQ.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/FAQ.md) | Frequently asked questions on setup, protocols, and security. |
| [`CONTRIBUTING.md`](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md) | Guidelines for contributing code or documentation. |

### Quick verification: try it (30 seconds)

Test tool-call adjudication locally without a model or GPU:

```bash
go run ./cmd/fak preflight --policy examples/customer-support-readonly-policy.json --tool refund_payment --args "{}"
```

**Expected output:**
```
verdict=DENY reason=POLICY_BLOCK by=monitor
```

---

## Task Router: what do you want to do today?

Select the task you need to accomplish to find its primary guide and immediate next action:

| You want to… | Authoritative route | Next action |
|---|---|---|
| Understand what fak is and read the pitch | [`README.md`](https://github.com/anthony-chaudhary/fak/blob/main/README.md) | Read the front door and inspect the hardware benchmark table. |
| Verify the tool-call boundary offline | [`docs/repro-packet.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/repro-packet.md) | Run the deterministic offline proof with zero downloads (no key, model, or GPU). |
| Install fak on your local machine | [`GETTING-STARTED.md`](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md) | Download a prebuilt binary or run `go install`, then verify your install. |
| Protect an agent you already use | [`docs/integrations/README.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/README.md) | Pick your agent harness and launch it behind `fak guard`. |
| Run a shared model gateway | [`docs/fak/server-quickstart.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/server-quickstart.md) | Launch `fak serve` pointing to an upstream provider or Ollama. |
| Run local models (Qwen3.8) on Apple Silicon Mac | [`docs/fak/mac-local-models.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mac-local-models.md) | Run `fak run qwen38` for interactive REPL or `fak serve --metal` for GPU server. |
| Run concurrent subagents with zero cold start | [`docs/subagents-guide.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/subagents-guide.md) | Start `fak up` and launch parallel subagents with `fak opencode`, `claude`, or `codex`. |
| Author custom tool allow/deny rules | [`POLICY.md`](https://github.com/anthony-chaudhary/fak/blob/main/POLICY.md) | Dump the built-in policy with `fak policy --dump` and customize it. |
| Fix a command that failed or misbehaved | [`docs/adoption/troubleshooting-first-run.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/troubleshooting-first-run.md) | Match your symptom to its cause and apply the one-line fix. |
| Understand how fak coordinates execution | [`docs/architecture.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/architecture.md) | Review the five-layer observation to typed-effect architecture. |
| Contribute code or documentation | [`CONTRIBUTING.md`](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md) | Check prerequisites, sign the DCO, and follow the contribution workflow. |

---

## Operator and Maintainer Tier

Use these resources when deploying to production, running fleet infrastructure, or working on the kernel codebase:

### Production and deployment
- Cloud deployment: [`docs/deployment.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/deployment.md) covers local, fleet, and air-gapped deployment patterns.
- Supported providers: [`docs/supported/clouds.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/clouds.md) details hosted model provider support.
- Reference architecture: [`docs/vendor/neo-cloud-reference-architecture.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/vendor/neo-cloud-reference-architecture.md) documents production cloud topologies.
- Rollback procedures: [`docs/ROLLBACK.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/ROLLBACK.md) details safe recovery steps for bad deployments.
- Security policy: [`SECURITY.md`](https://github.com/anthony-chaudhary/fak/blob/main/SECURITY.md) defines capability security guarantees and private vulnerability disclosure.

### Benchmarks and hardware
- Hardware matrix: [`docs/HARDWARE-MATRIX.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/HARDWARE-MATRIX.md) lists supported acceleration targets (Apple Silicon, AMD, NVIDIA).
- Benchmark authority: [`BENCHMARK-AUTHORITY.md`](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md) provides canonical benchmarks and claim boundaries.
- Claims register: [`CLAIMS.md`](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md) records shipped, simulated, and stub status for all claims.
- Product status: [`STATUS.md`](https://github.com/anthony-chaudhary/fak/blob/main/STATUS.md) tracks subsystem implementation maturity.

### Kernel internals and repository workflows
- Extending the kernel: [`EXTENDING.md`](https://github.com/anthony-chaudhary/fak/blob/main/EXTENDING.md) explains how to register new leaves without editing core files.
- Internal architecture: [`ARCHITECTURE.md`](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md) details kernel subsystem contracts.
- Automated coding agents: [`AGENTS.md`](https://github.com/anthony-chaudhary/fak/blob/main/AGENTS.md) contains operational rules for autonomous workers on the shared trunk.
- Exhaustive documentation index: [`INDEX.md`](https://github.com/anthony-chaudhary/fak/blob/main/INDEX.md) catalogs the full 1,200+ page corpus for name-based lookups and archival research notes.
- Kernel boundary separation: [`docs/fak-vs-dos.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak-vs-dos.md) distinguishes the agent execution path (FAK) from the work decision path (DOS).

---

# `INDEX.md`

> Source: `INDEX.md`

# INDEX — the full map of the fak repo

- [docs/notes/POSTTOOL-LATENCY-SPAN-2026-09-02.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/POSTTOOL-LATENCY-SPAN-2026-09-02.md) — the #10662 post-tool latency span: `tool_result_recorded → next_model_item` definition, disjointness rule, closed band/ordinal/attribution vocabularies, and the `fak session-audit posttool` readout.
- [docs/notes/qwen38-paged-swap-dogfood-2026-08-29.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/qwen38-paged-swap-dogfood-2026-08-29.md) — production-sized, repository-derived Qwen3.8 paged-swap codec round-trip witness for issue #9617.


- [`docs/notes/LONG-CONTEXT-MODEL-PRESETS-2026-08-28.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/LONG-CONTEXT-MODEL-PRESETS-2026-08-28.md) — dated, source-pinned Qwen3.8-Flash-Next and GLM-5.3-Flash analytical presets for the generic long-context estimator.

- [`docs/notes/CONCEPT-STUDY-ENDLESS-2026-08-24.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-ENDLESS-2026-08-24.md) — pinned study of Endless's open-turn `wait_for_user_input` mechanism, fak seams, exclusions, and measured borrow portfolio.
*This page owns one job: exhaustive lookup. It is for a reader who already knows a document, component, or artifact **by name**, or who needs a versioned, research, or historical route that the shorter front doors deliberately omit. It is a reference, not an on-ramp — if you are still deciding, read the [README](https://github.com/anthony-chaudhary/fak/blob/main/README.md); to install and run, use [Getting started](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md); to route a job you already have to one authority, use [START-HERE](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md).*

**Audience:** readers doing name-based or lifecycle-based lookup who want the current, versioned, research, or historical documentation route without reconstructing repository history.

This is the complete map of repository documentation. The audience table below is a coarse first cut; the detailed subject map after it is the exhaustive route and the reason this page exists.

## Choose an audience route

| Your job | Current route | Deeper or versioned route |
|---|---|---|
| **Evaluate what fak is and verify it** | [README](https://github.com/anthony-chaudhary/fak/blob/main/README.md) for the product contract, then [Getting started](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md) for install, deterministic proof, or production selection. | [Security policy](https://github.com/anthony-chaudhary/fak/blob/main/SECURITY.md) for the capability floor; [benchmark authority](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md) for scoped results and tuned baselines. |
| **Build or integrate with fak** | [Getting started](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md), then [agent runtime ownership and flow](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/agent-runtime.md), [integration guides](https://github.com/anthony-chaudhary/fak/tree/main/docs/integrations), or [supported APIs and protocols](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/apis-and-protocols.md). | [Architecture and design](#architecture--design) for implementation internals; [generation contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation.md) for current versus next work. |
| **Operate fak or an agent fleet** | [START-HERE](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md) for the runnable route, then [operating the agent fleet](#operating-the-agent-fleet). | [Work map](https://github.com/anthony-chaudhary/fak/blob/main/docs/WORK-MAP.md) for active operational programs and ownership. |
| **Contribute code or documentation** | [AGENTS](https://github.com/anthony-chaudhary/fak/blob/main/AGENTS.md) for enforced workflow, [CONTRIBUTING](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md) for contributor policy, and the [Code of Conduct](https://github.com/anthony-chaudhary/fak/blob/main/.github/CODE_OF_CONDUCT.md) for the participation standard that governs issues, pull requests, and reviews. | [Reference index](#reference-index), [plans](#plans--media), and versioned implementation docs carry deeper rationale. |
| **Research history or prior decisions** | [Notes and research](#notes--research-docsnotes) for dated evidence. | [Archive](https://github.com/anthony-chaudhary/fak/tree/main/docs/archive) is historical: follow replacement links before treating archived material as current guidance. |

**Default choice:** a first-time evaluator starts with the [README](https://github.com/anthony-chaudhary/fak/blob/main/README.md), then follows [Getting started](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md). Use another row only when its named job matches yours.

**Next action — select one route:** choose the row matching your job, open its first current link, and use the deeper column only when you need implementation, versioned, research, or historical context.

**Route context:**

- **Mode:** evaluation, building, operation, contribution, and research are separate reader modes; a deep implementation or historical route does not replace its current operational front door.
- **Generation:** current behavior is labeled `gen/now` through the [generation contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation.md); next and future material is planning context, not a shipped default.
- **Lifecycle:** front doors and maintained guides are current; versioned docs describe a named revision; dated notes are research evidence; [`docs/archive/`](https://github.com/anthony-chaudhary/fak/tree/main/docs/archive) is historical and requires a replacement check.
- **Support boundary:** this map routes documentation. Product guarantees remain scoped by [`CLAIMS.md`](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md), supported-version pages, the [security policy](https://github.com/anthony-chaudhary/fak/blob/main/SECURITY.md), and committed benchmark evidence.

New dated notes go under [`docs/notes/`](https://github.com/anthony-chaudhary/fak/tree/main/docs/notes) and get a line in **Notes & research**, so the audience route can stay concise while the map remains complete. Before touching this map or [`llms.txt`](https://github.com/anthony-chaudhary/fak/blob/main/llms.txt), run `python tools/check_index_sync.py --audit-tree`; the same reciprocal check runs in staged mode for front-door index and dated-note changes.

## Start here
- Structural sharing for the legacy Metal host prefix KV — [`docs/notes/SCOPE-LEGACY-METAL-HOST-PREFIX-KV-STRUCTURAL-SHARING-2026-09-14.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/SCOPE-LEGACY-METAL-HOST-PREFIX-KV-STRUCTURAL-SHARING-2026-09-14.md) is the #12856 scoping audit: it inventories the backend-nil host `K`/`Kraw`/`V` readers and writers, names the bounded three-file slice (`kvcache.go`, `kv.go`, `kv_cache_bytes.go`) and a fail-closed rollback that keeps the flat layout, and re-derives the 6,910,967,808-byte prefix payload from the real 16 full-attention layers. No sharing implemented; physical qualification separate.
- Meta-loop trigger effectiveness — [`docs/notes/META-LOOP-TRIGGER-EFFECTIVENESS-2026-08-31.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/META-LOOP-TRIGGER-EFFECTIVENESS-2026-08-31.md) defines the read-only `fak-loop-trigger/1` receipt, classification precedence, effectiveness metrics, live baseline, and field-derived cadence policy for super/meta loops (#10352).

- [Response profiles: concise, Caveman-compatible, and composable](https://github.com/anthony-chaudhary/fak/blob/main/docs/response-profiles.md) -- choose low/medium/high, distinguish native from original, and understand the safety and composition boundaries.
- [Work profiles: Ponytail-inspired implementation policy](https://github.com/anthony-chaudhary/fak/blob/main/docs/work-profiles.md) -- independently choose low/medium/high simplicity pressure, mix it with response shape, and keep correctness carve-outs explicit.
- **Micro-context fabric research** — one immutable agent base, bounded physical model slots, and 100 / 1,000 / 10,000 logical contexts; start with the [research contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/micro-context-fabrics.md), run the floor with `go run ./cmd/microcontextdemo -selfcheck -contexts 10000 -workers 64`, and follow the [research witness index](https://github.com/anthony-chaudhary/fak/blob/main/docs/research/README.md).


- [README](https://github.com/anthony-chaudhary/fak/blob/main/README.md) — what fak is and why, in one read.
- [START-HERE](https://github.com/anthony-chaudhary/fak/blob/main/START-HERE.md) — the job-to-authority route map: pick the task you have now and get one destination and one next action.
- [Getting started](https://github.com/anthony-chaudhary/fak/blob/main/GETTING-STARTED.md) — choose install, deterministic proof, or production setup with prerequisites and a first check.
- [Install fak](https://github.com/anthony-chaudhary/fak/blob/main/INSTALL.md) — external-adopter paths for a verified release binary, manual archive, Docker image, or source build, followed by the first-session tutorial.
- [Simple local-model demo](https://github.com/anthony-chaudhary/fak/blob/main/cmd/simpledemo/README.md) — run a friendly local GGUF chat end to end through fak's in-kernel engine, with model discovery, deterministic mode, streamed output, and per-turn measurements.
- [Tier 2: run the fused in-kernel model](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/in-kernel-model.md) — the kernel-developer continuation of Getting started, split out of it so the newcomer route reads end to end: the deterministic synthetic checkpoint over `/v1/fak/syscall`, the one-command SmolLM2-135M HuggingFace export, the Qwen3.6-27B GGUF smoke through `cmd/fakchat`, in-kernel chat through `fak serve --gguf` on both the OpenAI and Anthropic wires, and the honest caveat on why this is a correctness/reference path rather than a production chat server.
- [AGENTS](https://github.com/anthony-chaudhary/fak/blob/main/AGENTS.md) — orientation for coding agents (build, test, the hard rules). Written for automated contributors inside the maintainers' shared checkout; humans want [CONTRIBUTING](https://github.com/anthony-chaudhary/fak/blob/main/CONTRIBUTING.md).
- [Security policy](https://github.com/anthony-chaudhary/fak/blob/main/SECURITY.md) — evaluate the current capability floor, its policy configuration, scoped evidence, and private reporting route.
- [Benchmark authority](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md) — select a scoped result, compare its tuned baseline, and follow the committed artifact and reproduce route.
- [llms.txt](https://github.com/anthony-chaudhary/fak/blob/main/llms.txt) — the answer-engine index (and its inlined `llms-full.txt`).
- [Docs home](https://github.com/anthony-chaudhary/fak/blob/main/docs/index.md) — choose a current public, builder, operator, contributor, or research route.
- [In-one-breath contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/ONE-BREATH-CONTRACT.md) — the named, bounded explainer summary a budget-aware loader can serve first; `fak breath` checks only its countable shape and ratchets pre-existing doc debt.
- [Innovations index](https://github.com/anthony-chaudhary/fak/blob/main/docs/INNOVATIONS-INDEX.md) — the durable catalog of fak's innovations, concepts, and learnings: what each is, the general primitive it embodies, where it lives, whether it's shipped, and whether it's been generalized for reuse. The innovations counterpart to this repo map and the [llms.txt](https://github.com/anthony-chaudhary/fak/blob/main/llms.txt) doc map.
- [Generated verb/refusal surface](https://github.com/anthony-chaudhary/fak/blob/main/docs/generated/verb-surface.md) — source-derived command tree, preconditions, refusal codes, and visible gap count.
- [Work map](https://github.com/anthony-chaudhary/fak/blob/main/docs/WORK-MAP.md) — where each *kind* of work lives, kept separate: **optimizations** (the `EXTENDING.md` three-gate lane), **ongoing work** (the in-flight trackers, epics, and dispatch loop), and **dev** (the build/test/partition/ship workflow). Names the overlaps and the known drift between the status surfaces. Read this when you're not sure which front door a task belongs to.
- [Generation contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation.md) — the canonical now/next/second-next/future taxonomy for generation-aware development: stream labels, generation milestones, promotion/demotion evidence, shared-trunk rules, runtime feature-gate separation, intake rules, issue views, and anti-patterns. The front door for epic #1625.
- [Localized entry points (i18n)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/README.md) — in-language front doors that carry the pitch, the 60-second proof, and the market-specific value props, then hand off to the English docs. Ships **17 languages** — [हिन्दी (Hindi)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/hi/README.md), [தமிழ் (Tamil)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/ta/README.md), [తెలుగు (Telugu)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/te/README.md), [বাংলা (Bengali)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/bn/README.md), [मराठी (Marathi)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/mr/README.md), [Deutsch (German)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/de/README.md), [Français (French)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/fr/README.md), [Español (Spanish)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/es/README.md), [日本語 (Japanese)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/ja/README.md), [简体中文 (Simplified Chinese)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/zh/README.md), [한국어 (Korean)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/ko/README.md), [Português (Portuguese)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/pt/README.md), [Русский (Russian)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/ru/README.md), [العربية (Arabic)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/ar/README.md), [Bahasa Indonesia (Indonesian)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/id/README.md), [Tiếng Việt (Vietnamese)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/vi/README.md), and [Türkçe (Turkish)](https://github.com/anthony-chaudhary/fak/blob/main/docs/i18n/tr/README.md) — each folder also carrying a quickstart, an install guide, and an FAQ; rationale in the [emerging-market adoption note](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-EMERGING-MARKET-ADOPTION-2026-06-30.md); where to share them in the [distribution-channels note](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/INDIA-I18N-DISTRIBUTION-CHANNELS-2026-07-01.md).
- [AEO and market-event marketing hub](https://github.com/anthony-chaudhary/fak/blob/main/docs/marketing/README.md) — generated answer-engine feeds (`updates.json`, `disambiguation-terms.json`, `llms-updates.txt`, `llms-terms.txt`) plus the Fable 5-style frontier-model launch fit: cost-aware routing, refusal fallback, prompt-cache economics, and classifier-vs-capability-floor positioning.

## What fak supports

The dedicated, cross-linked capability pages — what fak works with, each grounded in the
repo and the sourced [compatibility matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/compatibility-matrix.md).

- [What fak supports (hub)](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/README.md) — the index of every "supported" page.
- [Models](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/models.md) — any model you front, plus the in-kernel architectures proven bit-exact.
- [Features](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/features.md) — every capability with its shipped / simulated / stub status.
- [Clouds & hosted providers](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/clouds.md) — Anthropic, OpenAI, Gemini, xAI, Bedrock, Vertex, Azure, OpenRouter, Together, Groq, Fireworks. *Fronting a hosted model API* — for the opposite shape (renting a GPU box and running fak on it) see the neo-cloud deploy quickstart below.
- [APIs, wires & MCP](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/apis-and-protocols.md) — the wires fak speaks and the fak-native + MCP endpoints.
- [Agent harnesses & frameworks](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/agent-harnesses.md) — Claude Code, Cursor, Codex, and the framework field.
- [Serving engines](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/engines.md) — Ollama, vLLM, SGLang, llama.cpp, LM Studio, and the in-kernel engine.

## Measuring sticks (the scorecards)

Each turns a fuzzy goal into a number you can drive toward zero.

- [Industry scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/industry-scorecard/README.md) — the one OUTWARD stick: an industry-first taxonomy of the dimensions the LLM-serving field competes on (vLLM/SGLang/TensorRT-LLM/llama.cpp), with fak's honest position on each (mostly named gaps). Two driven numbers: **coverage** (of the field) and **parity-debt** (honesty of the rows). Modular folder, regenerated from `tools/industry_scorecard.data/`.
- [SOTA prior-art matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/sota/README.md) — the one INWARD kernel stick: the maintained, load-bearing map from every compute operation fak's kernel performs to the production reference to learn from *before* writing it from scratch (llama.cpp/Marlin/CUTLASS/FlashInfer/vLLM/SGLang/a named paper), the route (borrow/bind/stay-minimal), and the oracle. Read it before optimizing a kernel; `fak sota <op>` surfaces a row, the `PRIOR_ART` advisory gate nudges at commit time, and `fak sota-coverage-scorecard` keeps the matrix complete against the tree. Source of truth: `internal/sotamatrix`.
- [Agent-readiness scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/AGENT-READINESS-SCORECARD.md) — can an AI agent discover, adopt, and build on fak (friction-debt across the three steps an agent walks).
- [Popularization-readiness scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/popularization-scorecard/README.md) — does fak's front door actually convert? Models a first-time visitor's path as a conversion funnel (land → orient → trust → install → act); each stage is positioned on whether the affordance a visitor reaches for there really exists in the tree. Two driven numbers: **coverage** (of the funnel) and **popularization-debt** (unmet affordances + verdict overclaims). Dimension K of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Maturity scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/MATURITY-SCORECARD.md) — the one LIFECYCLE stick: where each declared capability (one per `internal/<leaf>` lane) sits on a closed ladder (`proposed→prototyped→tested→dogfooded→default`, + a `benchmarked` badge), and the next work item that would mature it. Immaturity is never a defect — only a **ladder-skip** (fak relies on a capability yet leaves it untested) is. `fak maturity next` is the ranked backlog the dispatch loop pulls from; re-derived from `dos.toml` + the tree's import graph.
- [Repo-hygiene scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/REPO-HYGIENE-SCORECARD.md) — verbosity, organization, indexing, and accessibility of the whole tree.
- [Code-quality scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/CODE-QUALITY-SCORECARD.md) — the Go module's code-debt.
- [Generalization-debt scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/GENERALIZATION-DEBT-SCORECARD.md) — the coupling stick: `fak score generalization` makes singular production implementations (coupled to one model, backend, or provider instead of an interface, adapter, factory, or registry) visible as debt points with interest drivers, and records accepted temporary debt in the declaration's doc comment.
- [Testing and linting infrastructure scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/TESTING-LINTING-INFRA-SCORECARD.md) — current-state score plus a top-20 spine for making fak's tests, linting, CI, and agentic developer loop faster, better typed, more modular, and easier to trust.
- [Steerability scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/STEERABILITY-SCORECARD.md) — the one GROWTH-INVARIANT stick: does the effort to steer/change/navigate the repo stay flat as it grows? A 0–100 **steerability index** over modularity, coupling (blast radius), navigability, and correction — every KPI a ratio/percentile, so a 2× repo with the same discipline scores the same. The number every count-based sibling can't give.
- [Observability scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/OBSERVABILITY-SCORECARD.md) — the observability plane (metrics, dashboards, alerts, traces, proofs, ship-audit), graded so every dashboard/alert/doc points at a metric the binary emits and every claim is verifiable.
- [Conflation scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/CONFLATION-SCORECARD.md) — the truth-maintenance stick: does every number/status fak reports label its provenance — WITNESSED (a fact fak authored, in its control) vs OBSERVED (a value fak relays from an external party) — and never blame a provider-side miss on a fak action? Reads the fact-reporting surfaces (metric help, the guard exit summary) so a dashboard can't say "fak broke the cache" when the provider's cache simply expired.
- [Negframe scorecard](https://github.com/anthony-chaudhary/fak/blob/main/.claude/skills/negframe-score/SKILL.md) — the steer-prose FRAMING stick: does agent-steer prose (AGENTS.md, CLAUDE.md, the skills) lead with the affordance ("remember to stamp the commit") instead of the prohibition the reader must invert? Mechanical negatives with an unambiguous positive rewrite fold into `negframe_debt` — the only gating tier, each finding carrying its reframe (`fak score negframe --suggest`) — while judgement-tier findings stay advisory; `--since <ref>` is the diff-scoped ratchet that keeps a change from introducing a new mechanical negative. Source of truth: `internal/negframe`.
- [Default-value scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/DEFAULT-VALUE-SCORECARD.md) — the agentic default-window monitor: every cost/cache/amplification value flag ships default-ON or behind a reasoned, time-bounded opt-in gate; expired gates become typed `OPT_IN_REVIEW_DUE` omission debt, forcing a fresh decision to default on safely, renew with current evidence, or retire the lever.
- [UI/UX-quality scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/UI-QUALITY-SCORECARD.md) — the one TERMINAL-SURFACE stick: do the `fak console` panes, the `fak info` overlay, and `fak guard --split` render correctly and legibly? Graded against the render source itself — rune-safe truncation (no byte-slice that splits a multibyte rune), cell-aware column pads (no `%-Ns` shear on a multibyte row), empty-state branches, full info-legend and console-help coverage — folded into a `ui_quality_debt` integer. The first ship retired the `trimTUI` byte-slice that emitted a half rune on any em-dash/emoji.
- [Demo-quality scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/DEMO-QUALITY-SCORECARD.md) — are the demos runnable, honest, self-contained.
- [Demo-robustness scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/DEMO-ROBUSTNESS-SCORECARD.md) — do the demos survive bad input and odd environments.
- [Learning-docs scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/LEARNING-SCORECARD.md) — does the teaching set actually teach (learning-debt: a how-to with no runnable command, a tutorial with no worked output, an orphan lesson, an uncovered topic). Pedagogy counterpart of repo-hygiene (structure) and doc-appeal (voice).
- [Concept-usage scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/concept-usage-scorecard.md) — the INWARD dogfooding stick: when an agent BUILDS fak, how much does that development route through fak's *own* concepts? Two axes, both re-derived from `git log` + the `.dos` journals — **usage** (the ship-stamp / DCO / binding-verb commit discipline + lane arbitration) and **witness** (the `verify`/`improve` syscalls over passive recall). The thin axis is witness: it catches development that ships with fak's commit clothes but still trusts a self-report instead of witnessing its own claim. Breadth-and-depth sibling of the [dogfood-loop scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/dogfood-loop-scorecard.md) (one launched-session honesty loop).
- [Loop scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/loop-scorecard.md) — are fak's always-on agentic background loops *first-class durable processes*, or fire-and-forget scripts? Three axes, all re-derived from the loop ledger + job registry `fak loop` writes — **durability** (auto-restart on system restart: every firing loop registered, armed, and `fak cron emit`-projectable to an OS unit), **self-report** (no dark loop; fires that record an end outcome; heartbeat/notify), and **dogfood** (runs route through `fak loop run` under `fak guard`, the canonical hash-chained ledger, the witness contract). The honest baseline is debt 4 / composite 33 (F): the loops that fire are driven by external schedulers, so they vanish on reboot, go dark, and run unguarded — the 3× program drives them through `fak loop run`. Operational sibling of the [concept-usage scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/concept-usage-scorecard.md).
- [Release-readiness scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/RELEASE-READINESS-SCORECARD.md) — the RELEASE-VELOCITY stick: can a kernel that writes hundreds of commits a day cut, validate, publish, and roll back a release at the same speed, or does `@latest` rot far behind HEAD? KPIs across the release lifecycle — **discover** (an agent can find & invoke the path), **automate** (the machine cuts on green, not a human), **validate** (the cut is gated and the publish verified), **trust** (stable anchors, signed artifacts, rollback) — folded into a `release-debt` integer, all re-derived from git + the tracked tree + live release signals. Honest baseline is debt 12 / composite 16 (F): the cut is a hand-driven 7-step ritual and the cadence is dry-run-only, so `@latest` sits ~1900 commits behind HEAD; the 3× program (epic #1354) automates the cut and makes the path agent-discoverable.
- [Cadence report](https://github.com/anthony-chaudhary/fak/blob/main/docs/cadence/README.md) — the regular control-pane fold over scores, feature maturity, work done, and release state. Its append-only ledger records `standing_score`, normalized health, and difficulty fields so trends can keep climbing or fall instead of treating a bounded 100 as a stable meaning after the scorecard surface changes.
- [Milestone status](https://github.com/anthony-chaudhary/fak/blob/main/docs/milestones/STATUS.md) — the durable, freshness-checked snapshot of the maturity CLIMB: the model x backend grid's distribution across the closed M0–M7 support-maturity ladder, generated from `internal/covmatrix` + `internal/milestonereport` and kept honest by `fak milestone status-doc --check-doc` (a committed cell cannot silently drift from the live grid). The `gh`-fed epic ROADMAP stays on the Slack card + the durable `docs/milestones/history.jsonl` trend ledger.
- [Code-2x program](https://github.com/anthony-chaudhary/fak/blob/main/docs/CODE-2X-PROGRAM.md) — the plan to halve code-debt, then halve it again.

## Architecture & design

- [External system architecture](https://github.com/anthony-chaudhary/fak/blob/main/docs/architecture.md) — choose an interface and understand the managed request, effect, result, and support boundaries before implementation internals.
- [Architecture](https://github.com/anthony-chaudhary/fak/blob/main/ARCHITECTURE.md) — the registry seams and the frozen ABI.
- [Agent runtime: ownership, interfaces, and proof](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/agent-runtime.md) — the builder route for the host-side model/tool loop, its mediated call flow, interface choices, and deterministic offline proof.
- [GPU forward pass](https://github.com/anthony-chaudhary/fak/blob/main/GPU.md) — the in-kernel Llama decode on the GPU: a real on-box run witnessed against the CPU reference, with the honest gap to llama.cpp.
- [Partitioning](https://github.com/anthony-chaudhary/fak/blob/main/PARTITION.md) — how the kernel splits work across lanes and leaves.
- [Launchguard](https://github.com/anthony-chaudhary/fak/blob/main/docs/launchguard.md) — the host-local circuit breaker for detached agent and service supervisors: stable identity digests, bounded rolling attempt budgets with backoff, terminal quarantine, and the `fak launchguard status` / `reset` operator surface.
- [Extending fak](https://github.com/anthony-chaudhary/fak/blob/main/EXTENDING.md) — plug in an optimization, prove it correct, prove it faster.
- [Extension seams](https://github.com/anthony-chaudhary/fak/blob/main/docs/extension-seams.md) — where new behavior belongs: fak has no undifferentiated "plugin" mechanism; pick the least-privileged attachment seam, and user- or agent-authored code never joins the kernel's trusted core.
- [The Footprint Ladder](https://github.com/anthony-chaudhary/fak/blob/main/docs/footprint-ladder.md) — the doctrine for adding a capability at the highest (least-footprint) rung that works: a new core tool's marginal prefix-token cost × call frequency is its footprint bill, measured by `fak footprint` and refused by the floor-growth gate unless justified.
- [SOTA comparison](https://github.com/anthony-chaudhary/fak/blob/main/SOTA-COMPARISON.md) — where fak sits next to the state of the art.
- [Branch regime ADR](https://github.com/anthony-chaudhary/fak/blob/main/docs/branch-regime.md) — the proposed dev/main branch-role contract: `dev` as hot integration, `main` as public front door, plus migration order, backout, and no-split-brain rules.
- [Memory-concept ranking dossier (superset)](https://github.com/anthony-chaudhary/fak/blob/main/docs/superset/MEMORY-CONCEPT-RANKINGS.md) — ten memory concepts M1–M10 ranked fak-vs-engines with evidence, per-engine evidence tables, and adopt-or-SKIP verdicts (epic #2236, issue #2237, issue #3143).
- [The abstraction map](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/the-abstraction-map.md) — the end-to-end floor plan for a human operator: seven floors from your terminal and work items (lanes, leases, diff-witnessed done-claims) down through the two front doors (`fak guard` / `fak serve`), the kernel's verdict fold, context + memory, the engine seam, the compute HAL (`internal/compute`), and the machine floor (build tags compile backends in, a runtime `Tier()` probe picks which runs). One seam named per floor, one worked trip down and back up, and the rule that holds the stack together: **every floor trusts the seam below it, never a story about it**. An operator can stop at floors 6–5; the rest is there when wanted.
- [The tool call is a syscall](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/tool-call-is-a-syscall.md) — the keystone mental model, taught from scratch: an OS kernel never trusts a user program's word that a write is safe — the syscall crosses a boundary the program does not control, and fak applies the same boundary to the LLM's tool calls (proposed call → verdict → only an allowed call executes; suspicious results quarantined). The sentence to keep: **the model proposes, the kernel disposes**. Engineering depth in [policy-in-the-kernel](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/policy-in-the-kernel.md).
- [Verify, don't trust: what DOS actually checks](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/verify-dont-trust.md) — the DOS primer for a newcomer, the most transferable idea in `fak` (it works with zero fak internals): an agent narrates its own success, so the kernel re-checks the three things it will not take on the worker's word — a **done-claim** against git (`dos verify` / `dos commit-audit`, the subject is forgeable but the diff is not), a **refusal** against a closed reason vocabulary (`dos arbitrate` refuses with a token a peer can act on), and a **recalled memory** re-verified at read time (`dos memory recall`, `RECALL_FRESH`/`STALE`/`UNVERIFIABLE`). The sentence to keep: **the model proposes, the kernel disposes — and the kernel checks**; honest fence: it grades a claim against its own evidence, not code correctness. Sits under [the tool call is a syscall](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/tool-call-is-a-syscall.md).
- [The addressable KV cache in 5 minutes](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/addressable-kv-cache-in-5-min.md) — the gentle on-ramp to the addressable cache for a first-time reader: one analogy (a shared notebook you can only append to vs one you can reach into and tear a page out of the middle), the KV-cache / append-only / addressable terms defined as they appear, and the one property that matters — evict a poisoned span and the cache stays bit-for-bit identical to one that never saw it (`max|Δ| = 0`). No code; honest fence (witnessed on a synthetic model, live loop not yet driving it, ~4.1× quoted as measured). The door to the dense [addressable KV cache](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/addressable-kv-cache.md) internals page.
- [Why default-deny beats a classifier](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/default-deny-vs-classifier.md) — the prompt-injection explainer for the most common objection: a classifier has to recognize the attack and can fail open, while a default-deny capability floor and result quarantine stop the same injected prompt by structure. Walks one attack through both stacks, names the OWASP Agentic Top-10 / MCP Top-10 risks covered, and keeps the 5/5 live run framed as a small demonstration, not a benchmark sweep.
- [What fak is not — the honest boundary](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/what-fak-is-not.md) — the disarming page for the skeptic: `fak` is **not** a serving engine and does not beat vLLM/SGLang/llama.cpp on raw tokens/sec — as a gateway fronting SGLang it *trails* raw SGLang 0.75× at peak (the ~3%-at-saturation adjudication tax), because its field is governance at the agent boundary, not throughput. States the 0/29-novel prior-art audit plainly and reframes it as the point (the contribution is the *assembly*, not an invented primitive), and names the detector as best-effort/evadable while the default-deny capability floor is the load-bearing guarantee. Honest fence: witnessed numbers only (~4.1× vs the *tuned* baseline, never the naive 60×; WebVoyager 8.8–9.7× labeled simulated). The volunteered-limits companion to [why default-deny beats a classifier](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/default-deny-vs-classifier.md).
- [Change data capture for agents](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/change-data-capture-for-agents.md) — the page for the infra reader who searches "change data capture" or "Debezium": `fak` already implements the whole CDC pattern family — a hash-chained change log (`internal/journal`), cursor-drained change feeds (`GET /v1/fak/changes`, `GET /v1/fak/events`), a transactional outbox (`internal/slackoutbox`), and fold-to-read-model projections (`dos_verify`/`dos_status`) — but over **agent, cache, and session state**, mapped one-to-one onto the CDC vocabulary (LSN/offset ↔ `Seq` cursor, tombstone/delete ↔ `revocation`, source metadata ↔ `WorldVer`/`TrustEpoch`, CQRS read model ↔ the fold verbs). The honest fence and companion to [what fak is not](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/what-fak-is-not.md): `fak` is a change **source** for agent work, **not** a database-replication sink — it never ingests your Postgres WAL, emits semantic events (not row images), and runs no Kafka/Flink. Shipped feeds today; the Debezium-envelope export (#3171), work-change unification (#3172), and consumer-contract doc (#3173) are the open follow-ups.
- [Native harness default-on security features](https://github.com/anthony-chaudhary/fak/blob/main/docs/architecture/native-harness-default-security.md) — the one-page inventory of what the fak native harness ships **by default** at the call boundary: the default-deny capability floor, the write-time result quarantine, the JIT secret "page in and out" (mask-in-place redact by default, opt-in hard seal, and a fail-closed re-screen on the way back so a clearance cannot launder a credential), kernel-authored provenance/IFC, witness gates, plan-CFI, and the durable-memory promotion gate — each row pointing at the package that owns it, plus the `pkg/harnesskit` contracts (`ToolScope`, `AuthBinding`, `PublicToolContract`) an external builder inherits instead of re-implementing. Keeps the honest scope: the detector is evadable, the floor and the wall are load-bearing.
- [Long-session economics](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/long-session-economics.md) — the cost explainer for the largest audience (anyone paying for a long agent run): why a growing transcript re-sends everything each turn, why the provider's prompt-cache discount survives *only* while the prefix is byte-identical, and how `fak` keeps it alive by splicing on original bytes (a `memcpy`, never a re-marshal) rather than rewriting the prompt to shrink it. Carries a worked cost example and the honest fence — `fak` guarantees the byte-identical prefix and *relays the provider's own reuse number*, it never claims a saving it can't force. The economics angle beside the practical flag ([long sessions: keep the cache hit](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/long-sessions-keep-the-cache-hit.md)) and the theory ([addressable KV cache](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/addressable-kv-cache.md)).
- [Context shedding](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/context-shedding.md) — the plain-English *what and why* of fak's default "trim the middle out" saving: every turn re-sends the whole transcript, the growing middle keeps missing the provider cache, so fak drops the stale middle (keeping the cached head byte-identical and your recent turns, leaving a restore handle) and the trimmed history stops costing you every turn. Written to be read honestly: the real saving is per-turn recurrence; the honest number is shed **per fire**, because fak re-trims the client's full re-sent history each fire so a session-total re-counts the same middle (the [retracted-75%](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md) lesson, baked in); and the ref-count / use-after-free framing of whether a shed span was safe to drop. Companion to [long-session economics](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/long-session-economics.md) and the [built-in compaction audit](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/BUILT-IN-COMPACTION-AUDIT-2026-07-06.md).
- [What a saved token is worth](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/what-a-saved-token-is-worth.md) — the valuation companion to context shedding: how a token fak avoided paying for becomes an honest dollar, from the one number that underlies it all (the `0.1x` cache-read price). Names the two savings stories that share that number (the provider's `0.9x` read rebate vs fak's compaction shed), the `count × price` identity that keeps the figure honest (with the two orthogonal inflation bugs it has survived — the price axis, `1.0x`-on-warm over-crediting compaction 10x, fixed by #2794/#2798/#2796; and the count axis, per-fire not session-sum), the `ValuationBasis` label every fak dollar must carry (the renderer refuses an unlabeled one), and the decoder ring for the *same* `0.1` declared six times under six names — why the layered-DAG import rule makes most of those copies mandatory (pinned to the canonical `gateway.CacheReadMultiplier`) and which one same-file pair is honest, collapsible debt.
- [Memory engineering](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/memory-engineering.md) — the definitional page for the discipline after prompt and context engineering: what an agent remembers (write-time admission), where memory lives (the four layers), how recall is verified (integrity + verified recall), and when a memory is provably forgotten (bit-exact eviction) — each decided by an inspectable mechanism, never the model's judgment. The sentence to keep: **if it can't answer "why is this in memory, is it still true, can you prove it's gone" — that's memory features, not memory engineering**. Depth in [context-is-not-memory](https://github.com/anthony-chaudhary/fak/blob/main/docs/CONTEXT-IS-NOT-MEMORY.md), [memory-layers-explainer](https://github.com/anthony-chaudhary/fak/blob/main/docs/MEMORY-LAYERS-EXPLAINER.md), and [memory-ecc-integrity](https://github.com/anthony-chaudhary/fak/blob/main/docs/MEMORY-ECC-INTEGRITY.md).
- [Engineering is building loops — the worldview page](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/engineering-is-building-loops.md) — the manifesto for the whole substrate: modern engineering is increasingly the act of building agentic loops (observe→orient→decide→act→verify), and `fak` is the in-process kernel they run on, **safe and fast for the same reason** — the gate that refuses a bad action is the one that lets a known-good action reuse work it already trusts. Traces the *same* observe/decide/act/verify shape at five nested scales (tool-call syscall → turn → session → fleet → RSI) and the orthogonal invariant recurring at every ring: **a decision no participant can move by narrating a number** (provable refusal → `Clear()`+rescreen quarantine → sealed session pages → per-SHA `dos commit-audit` → the non-forgeable keep-bit `shipgate.Evaluate`). Honest fences kept intact: 0/29 primitives novel (the contribution is the *assembly*), the detector is evadable while the default-deny capability floor is load-bearing, the ctxplan live loop is a guarded seam off by default, and the async / autonomous-meta-RSI-apply reaches are named unshipped. Closes on the quotable line — *engineering is building loops; `fak` is the floor they run on*. The scale axis beside the deployment axis ([cross-platform spine](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/cross-platform-spine.md)) and the by-part decomposition ([what is an agent?](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/EXPLAINER-what-is-an-agent-2026-06-24.md)).
- [The cross-platform spine (IoT to hyperscaler)](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/cross-platform-spine.md) — why the same pure-Go kernel is the invariant spine across the whole deployment spectrum (IoT, edge, laptop, hyperscaler), the way Linux is one kernel under a phone and a datacenter: the hardware specifics change, the agentic workload shape and the kernel's invariants do not. The deployment-substrate axis, alongside the scale axis ([loops](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/engineering-is-building-loops.md)) and the hardware-depth axis ([HAL](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/hardware-portability.md)).
- [What is a CUDA kernel? (kernel disambiguation)](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/what-is-a-cuda-kernel.md) — the page for the word every newcomer trips on twice: "kernel" names three unrelated machines here — fak's OS-metaphor reference monitor (`internal/adjudicator` + `internal/kernel`, pure Go, no GPU), the HPC compute-kernel sense (`internal/model/kernel.go`'s `matKernel` arithmetic paths), and the 21 literal CUDA `__global__` kernels in `internal/compute/cuda_kernels.cu`. Shows what a `.cu` file actually is (annotated `k_q8_gemm` walk, host vs device code, the `<<<grid, block>>>` launch), corrects the ".h is a level" misreading with the header-to-silicon depth ladder (cgo ABI → host CUDA C++ → `__global__` → intrinsics/PTX → SASS → CUDA/tensor cores), and closes with the four-part when-is-deeper-worth-it test tied to the S/M/L costing legend. Honest fences kept: fak's CUDA lane is a correctness-gated first-generation Approx peer (argmax-exact + logit cosine), not a tuned engine; the int8 tensor-core MMA lever is estimated, not measured. The hardware-depth companion to [hardware portability](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/hardware-portability.md) and the [canonical glossary](https://github.com/anthony-chaudhary/fak/blob/main/docs/glossary.md)'s kernel cluster.
- [Neo-silicon onboarding](https://github.com/anthony-chaudhary/fak/blob/main/docs/vendor/neo-silicon-onboarding.md) — vendor-facing path from a stable `compute.Backend` name through the compiling MatMul+Attention example and the planned backend conformance kit.
- [Data residency & compliance](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/data-residency-and-compliance.md) — how the self-host-first, fail-closed, default-deny, audit-logged boundary maps to India's DPDP Act and China's PIPL/DSL/CSL: an enforcement control surface you run on infrastructure you control (not legal advice, not a certification). The residency lens on the same boundary; go-to-market context in the [emerging-market adoption note](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-EMERGING-MARKET-ADOPTION-2026-06-30.md).
- [The god-file growth gate: a ratchet, not a sweep](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/god-file-growth-gate.md) — a monolith is never written in one commit; it accretes one line at a time, faster than anyone pays it down, because nothing refuses the next line. Hermes' `gateway/run.py` reached **20,320 lines** that way while its own rubric asked for the refactor — a doctrine that only *describes* the good state defends nothing. fak already names the ceiling (`FILE_HARD_MAX=1500` / `FUNC_HARD_MAX=200`) and owns the paydown loop ([/modularize](https://github.com/anthony-chaudhary/fak/blob/main/.claude/skills/modularize/SKILL.md), `tools/godsplit_plan.py`); `GOD_FILE_GROWTH` (`internal/hooks/gate_godfile.go`) turns the doctrine preventive — a `fak hygiene` + `make ci` ratchet that grandfathers today's offenders at-size (`godfile_baseline.go`) and refuses only NEW growth (new file/function over the ceiling, or a grandfathered one grown past its frozen size), with the baseline tighten-only, so the god-code surface is monotonically non-increasing. The preventive half of the [steerability](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/code-linting-at-the-kernel.md) story: /modularize retires the debt, the gate stops the next unit forming. Closes #2868 (Track I of #2834).
- **More explainers** — [the caching ladder (five rungs, five audiences)](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/caching/README.md) · [cache reuse: choose the layer](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/cache.md) · [context management](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/context.md) · [context as a variable](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/context-as-a-variable.md) · [model routing](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/routing.md) · [the agent virtual filesystem](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/agent-virtual-filesystem.md) · [shared workspace & the negation operator](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/shared-workspace-and-the-negation-operator.md) · [the vLLM lifecycle cache-loss bridge](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/vllm-lifecycle-cache-loss-bridge.md).
- [Interoperability stance](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/interoperability.md) — why fak adopts whatever agent/model/framework you run (the one opinion kept is the floor) + the honest per-wire grade for the flagship harnesses and every interop protocol.
- [Agent-to-agent value](https://github.com/anthony-chaudhary/fak/blob/main/docs/a2a-value-opportunities.md) — where the cross-agent cache pays off.
- [Multi-agent coordination protocol (RFC, D-007)](https://github.com/anthony-chaudhary/fak/blob/main/docs/multi-agent-coordination-protocol.md) — the normative spec binding message passing (a2achan), shared state (sharedtask), and coordination primitives (comm) under one default-deny floor.
- [Shared state ladder](https://github.com/anthony-chaudhary/fak/blob/main/docs/shared-state-ladder.md) — the canonical split for shared live messages, live mutable state, durable handoff, disaggregated state, and user-level collaborative editing.
- [Shared task record contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/shared-task-record-contract.md) — the executable envelope contract for collaborative task records, user patches, conflicts, approval gates, and disaggregated artifact refs.
- [Region admission](https://github.com/anthony-chaudhary/fak/blob/main/docs/region-admission.md) — one decision for "may THIS actor act on THIS (lane, tree) now?" (`internal/regionadmit`): tree geometry plus `dos.toml` lane semantics refusing `COLLISION_RISK`, consulted by the dispatch tick, by lane/region-declaring loop drives (which hold a fenced region lease for the whole drive), and by manual sessions via `fak loop region` — all over the `fak leaseref` acquire/renew/release/reap lease fabric. Also carries the **sub-lane** algebra (`internal/laneadmit/lanetree.go`): a lane name is path-shaped, so `gateway/server` derives its tree from the nearest declared ancestor and the addressable space follows the repo (543 declared roots → 13,690 file-granularity lanes) instead of the hand-typed roster — default granularity stays the leaf, and a sub-lane is never more permissive than its root. Algebra only so far: `regionadmit` (the twin this page's surfaces actually call) has not adopted it, so `--lane gateway/server` still resolves to no tree — the page says so where it matters.
- [Task manager concept](https://github.com/anthony-chaudhary/fak/blob/main/docs/task-manager.md) — the process-local runtime snapshot for tasks, steps, resource usage, concept runtime, progress, and ETA.
- [Agent–machine link protocol](https://github.com/anthony-chaudhary/fak/blob/main/docs/agent-machine-link-protocol.md) — how an agent and a host node bind.
- [Glossary](https://github.com/anthony-chaudhary/fak/blob/main/docs/glossary.md) — the canonical split for overloaded core words (session, agent, context, model, memory, tool/skill, steering), shared-memory senses, and the live memory issue-owner map.
- [Concept glossary (implementation vocabulary)](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/concept-glossary.md) — the contributor-layer companion to the Glossary above: where fak draws the line between similar-sounding implementation names (the cache planes, gate vs guard, the witness families, cross-package symbol collisions), every entry anchored on disk and machine-verified by `tools/concept_disambiguation_scorecard.py`. Public/product terms stay in [the Glossary](https://github.com/anthony-chaudhary/fak/blob/main/docs/glossary.md); internal identifiers resolve here.
- [DOS kernel transfer playbook](https://github.com/anthony-chaudhary/fak/blob/main/docs/dos-kernel-transfer-playbook.md) — moving the trust substrate to a new repo.
- [Prefill visuals](https://github.com/anthony-chaudhary/fak/blob/main/docs/prefill-visuals.md) — the diagrams behind prefill reuse.
- [Adoption visuals](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption-visuals.md) — five figures for how to think about using fak: where the binary sits, the rung-by-rung on-ramp, the honest fak-authored vs provider-observed value split, which integration shape fits, and when the perf win is real.
- [Adoption-signals dashboard spec](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/signals.md) — the honest signals worth watching to tell whether fak/DOS is actually being adopted (stars, forks, watchers, mentions, directory listings, integration recipes, distinct harnesses, docs reachability): where each is collected, and — the load-bearing column — what each does NOT prove. Every signal labeled `OBSERVED` (a number a third party controls) vs `WITNESSED` (an artifact fak authored) per the [conflation discipline](https://github.com/anthony-chaudhary/fak/blob/main/docs/CONFLATION-SCORECARD.md); no vanity total ships without its disclaimer, no number is invented. Dimension K of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Where to submit fak — directory & awesome-list checklist](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/directories.md) — the curated, honest map of the awesome-lists, tool directories, and registries where fak actually belongs (agent-infra, MCP, harness engineering, LLM security, Go, self-hosted): each venue with why-it-fits, a submission note, and a current status (wired / live / not yet / blocked / declined). Venues that do NOT fit are in the Declined section with the reason; the copy-paste payloads live in [`docs/launch/directory-submissions.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/launch/directory-submissions.md). Dimension J of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Social storyboard for the five concepts](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/social-storyboard.md) — a drop-in card-per-concept storyboard for a short social thread or slide carousel: one card each for tool-call-as-syscall, verify-don't-trust, the addressable KV cache, the default-deny gate + quarantine, and the one static binary — each with a one-line hook, the diagram that carries it, and a link. The reusable spine for any social push; witnessed numbers only (tuned ~4.1×, ~362 ns guard tax, `max|Δ| = 0`). Dimension J of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Watch it: install to first DENY verdict (recorded cast)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/casts/README.md) — a checked-in asciinema-style terminal recording of the 60-second proof: `go build` the one static binary, run one `fak preflight`, and watch the default-deny capability gate refuse a dangerous tool call with `verdict=DENY reason=DEFAULT_DENY` — no key, no model, no GPU. Plays in a browser via the [`.cast`](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/casts/install-to-first-verdict.cast) file (asciinema v2) or reads as an annotated transcript with a still frame, so a reader gets the "oh, that's it?" moment without running anything. Honest about being a recording; every line is real captured output; no benchmark claimed. Dimension C of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [fak concept card (printable one-pager)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/concept-card.md) — the single page you hand someone at a meetup: the five concepts (one sentence each + its carrying diagram), the install one-liner, and the 60-second proof command — designed to render to one printed/PDF page, honesty-fenced (not a token engine, evadable-by-design detector, 0/29 novel, tuned ~4.1× never the naive number). The printable counterpart to the [social storyboard](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/social-storyboard.md) and [pitch ladder](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/pitch-ladder.md). Dimension B of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [How to find and name fak (search + disambiguation)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/naming.md) — a reader-facing note on how to search for and refer to fak: the bare word is buried under homophone + F.A.K.-acronym noise, so the concept travels under `agent kernel`, `Fused Agent Kernel`, and "treat the tool call like a syscall". The disambiguated terms to use (kept consistent with [`llms.txt`](https://github.com/anthony-chaudhary/fak/blob/main/llms.txt)), the category shelves, and what fak is NOT (network firewall, guardrails classifier, token engine, request-level router). Dimension I of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Objections & one-line answers (advocate card)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/objections.md) — a pocket card of the 10 objections fak actually draws (isn't this a classifier, doesn't it slow the agent, why not just use vLLM, is the security real if the detector is evadable, another gateway, is it novel, how is it different from a firewall, do I rewrite my agent, are the numbers too good, is the KV eviction lossless), each with a crisp honest one-to-two-line answer and a deeper link. Concedes where an objection lands (evadable detector, 0/29 novelty, throughput); witnessed numbers only. Dimension I of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [The pitch ladder: 1 sentence / 1 paragraph / 1 page](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/pitch-ladder.md) — the canonical fak pitch at three zoom levels: one sentence (the tweet, ≤30 words), one paragraph (the HN comment), one page (the blog intro) — each self-consistent and quotable, anchored on "treat the tool call like a syscall: the model proposes, the kernel disposes". The source for the README lead; a consistency table keeps the rungs from drifting; witnessed numbers only (the tuned ~4.1×, ~362 ns guard tax, `max|Δ| = 0`) and the 0/29 novelty fence conceded up front. Dimension I of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Launch-post draft kit (honest, ready to adapt)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/launch-kit.md) — a single, self-contained Show HN / launch-post draft: title options, the opening hook, the prosecution-first first comment, a longer launch-post body, the honest what-it-IS / what-it-is-NOT framing, a copy-paste TL;DR, and a pre-post checklist that blocks an over-claim. Draft only — posting is human-owned; every witnessed number traces to [`CLAIMS.md`](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md) / [`BENCHMARK-AUTHORITY.md`](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md) and leads with the tuned ~4.1×, never the naive multiplier; aspirational numbers (~60×, "agent city", power/$) are labeled design-target / simulated. The one-page unit; the removed per-channel campaign is retained in git history. Dimension J of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [It didn't work: first-run troubleshooting](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/troubleshooting-first-run.md) — the five most likely fak first-run failures, each with the exact symptom, cause, and one-line fix: binary not on PATH (`./fak` or `dogfood-claude.ps1 --install`), `--base-url` wrong or empty (empty = the offline scripted mock, hence canned replies), the upstream key env not set for a keyed provider (`--api-key-env VAR` names the env var, never the literal key), a tool call refused by the embedded default-deny capability floor (inspect with `fak guard --dump-policy`, then grant — don't disable the gate), and `address already in use` on `fak serve` (pick another `--addr`). Frames first-run failures as wiring, not the one static Go binary. Dimension F of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Bench-story: the 4x that's real](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/stories/the-real-4x.md) — the tuned warm-cache result told honestly: fak's fused kernel does ~4.1× less work than a *tuned* single-tenant stack (B/C, the honest headline) on a 50-turn × 5-agent session (Qwen2.5-1.5B, Q8_0, M3 Pro), not the flattering 60.3× naive number (A/C). What the 4.1x is (prefix reuse + decode batching over a warm KV cache), what it is NOT (a throughput win over llama.cpp; a hosted-API win), and the conditions it holds under. Witnessed against [SESSION-VALUE-STACK-RESULTS.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/SESSION-VALUE-STACK-RESULTS.md). Dimension H of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Bench-story: max|Δ| = 0, the bit-exact eviction result](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/stories/bit-exact.md) — the addressable-KV eviction result told as a falsifiable claim: fak removes a span from a live kernel-owned attention cache and the next-token distribution comes out bit-identical to a run that never saw it (`max|Δ| = 0`, argmax and tie-break identical), against a non-vacuous poison control at `max|Δ| ≈ 0.326`. Keeps the two exactness numbers apart (evict-vs-never `0` vs the HF-oracle `≈4.4e-5`), names the honest fences (synthetic-model witness, self-signed v1 deletion certificate, coherent-compaction scope), and gives the one no-key/no-GPU/no-network command that checks it (`examples/addressable-evict/run.sh`). Witnessed against [CLAIMS.md](https://github.com/anthony-chaudhary/fak/blob/main/CLAIMS.md). Dimension H of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Bench-story: the cache cliff](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/stories/cache-cliff.md) — why a high prompt-cache hit rate can be a warning sign: fak's fleet audit measured 96.6% of ingested tokens cached (median 99%, 199 sessions), but that number is the *frozen-trajectory ceiling* — a prefix match that rises with length (82%→96%→99% at 10/50/200 turns) and only holds while the harness never edits history. Edit-depth into the prefix sends it off a cliff (99%→94.1%→74.3%→49.5%→0.0% at 0/5/25/50/100%), and cross-agent fan-out forfeits shared reuse (0% across the fleet). Argues the honest metric is reread work *deleted* (content/identity-keyed, coherence-checked) not hit rate quoted; runnable via [tools/cache_curve.py](https://github.com/anthony-chaudhary/fak/blob/main/tools/cache_curve.py). Witnessed against the frozen-trajectory cache-cliff explainer. Dimension H of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Capability matrix across the category](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/matrix.md) — the one-glance, sourced table scoring fak against the three neighbouring product categories people already run (guardrails/output-validation libraries, LLM API gateways & routers, vLLM/SGLang-class inference servers) on the six capabilities fak claims: default-deny tool-call gate, prompt-injection result quarantine, addressable bit-exact KV eviction (`max|Δ| = 0`), commit-level verify, closed-vocabulary structured refusal, and single-binary drop-in. Every cell is yes/partial/no with a note and a source (or an `unverified` tag), the categories are framed as complements not rivals, and fak's own gaps are named (the KV evictor is a no-op on a proxy seat; 0/29 novel). Links to the long-form [comparison](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak-vs-alternatives-comparison.md) and [compatibility matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/compatibility-matrix.md) for depth. Dimension D of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [fak vs a guardrails library (honest side-by-side)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/vs-guardrails.md) — the focused one-on-one against the guardrails class (Guardrails AI, NeMo Guardrails, Llama Guard): each tool's real strength stated from its own primary docs, the fail-open recognizer vs fail-closed capability-gate distinction as the spine, a side-by-side table, and the honest "these are complements you run together" framing (a guardrails lib validates content; fak default-denies the tool-call boundary and can front one via one base-URL change). fak's own gaps named (poor content classifier, evadable-by-design bonus detector, 0/29 novel). Links to the wider [capability matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/matrix.md). Dimension D of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [fak vs an API gateway / LLM router](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/vs-routers.md) — the focused router/gateway side-by-side for the "don't I already have this?" question: a router (OpenRouter, Portkey, LiteLLM, Kong AI Gateway) picks *which model* serves a request and how to reach it reliably; fak governs *which effects* each tool call may have via a default-deny capability floor one layer down. Shows the compose topology (fak in front of the router), states the honestly-mixed packaging (Kong=binary, LiteLLM=Python proxy, OpenRouter=hosted → why the category is *Partial* on single-binary), and names fak's gaps (does not out-connect a gateway; KV evictor is a no-op on a hosted seat; 0/29 novel). Summarizes and links [the routers integration guide](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/routers.md) and the [category matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/matrix.md); witnessed numbers only. Dimension D of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Is fak just a firewall?](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/vs-firewall.md) — the boundary FAQ expanded into a standalone page: it concedes what is genuinely firewall-shaped about the default-deny tool-call gate (a single chokepoint, a reviewable allow-list, an audit ledger, a drop-in appliance) before drawing the real line — a firewall filters traffic by rules on packets or HTTP it inspects and often fails open, while fak adjudicates effects by capability on the same in-process call path the model does not control (fails closed, ~362 ns/decision), understands tool-call semantics a packet filter cannot, quarantines poisoned results, and can refuse a false "done" from git evidence. Same instinct, different layer; keep your network firewall too. Honesty-fenced (0/29 novel, evadable-by-design detector, no market claim). Links to the [capability matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/matrix.md) and [tool call is a syscall](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/tool-call-is-a-syscall.md). Dimension D of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [fak vs vLLM / SGLang: a different boundary, not a race](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/vs-serving-engines.md) — the focused inference-server side-by-side for the "isn't fak competing with my serving engine?" question: it concedes that vLLM/SGLang own raw throughput (continuous batching, PagedAttention/RadixAttention) and that fak makes no tokens/sec claim against them, then draws the real line — fak governs the *tool call*, not the token stream, owning the default-deny capability floor, structured refusals, result quarantine, and audit ledger those engines leave open. The recommended move is to run both: keep the engine for throughput and front it with `fak serve` via one base-URL change. Honesty-fenced (the in-kernel model path is a correctness reference not a tuned server, the KV evictor is a no-op on a proxy seat, ~3% overhead at saturation, 0/29 novel). Links to the [capability matrix](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/compare/matrix.md) and the long-form [comparison](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak-vs-alternatives-comparison.md). Dimension D of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Good first popularization tasks (contributor on-ramp)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/good-first-tasks.md) — a curated board of small, self-contained doc/example contributions a newcomer can finish in an afternoon: 12 tasks, each naming the exact existing file to edit and a starter/moderate difficulty (add an objection answer, a glossary term, a comparison row, a translation, an integration recipe), plus a pointer to the open `popularization` backlog for a tracked ticket. Honesty-fenced (evadable-by-design detector, 0/29 novel, tuned ~4.1× never the naive number). Dimension E of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Who is fak for? A persona gallery](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/personas.md) — the kinds of person (and one machine) who land on fak, each with who they are, a one-sentence quote-ready pitch, and the first door to walk through: the solo dev, app developer, and backend integrator (consume); the SRE and security engineer (operate); the ML researcher, benchmark engineer, and decision-maker (evaluate); the contributor and the AI coding agent (build). The reader-facing cut of the [pitch ladder](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/pitch-ladder.md), on the same 10-persona roster the [persona-readiness scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/persona-scorecard/README.md) grades; pitches FOR personas not fabricated quotes FROM them, witnessed numbers only. Dimension E of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Voices: adopter quotes & case notes](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/voices.md) — the honest, empty-until-real scaffold where genuine adopter testimonials and case notes land: the one rule (**no fabricated testimonials** — every quote is real, attributable, adopter-approved, or it is not here), a copy-paste submission template, and an empty state that reads as an invitation (*be the first — here's how to share your story*) rather than a manufactured quote. Ships empty on purpose; the same *verify, don't trust* discipline turned on our own social proof — witnessed numbers only (tuned ~4.1× never the naive ~60×), no market-adoption claim, the personas gallery's other half. Dimension E of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [The fak roadmap: what's shipped and what's next](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/roadmap.md) — the reader-facing "what we're building next" page, on the now/next/future generation taxonomy: **Now** lists what is shipped and git-witnessed today (the default-deny tool-call gate at ~362 ns, `dos commit-audit` verify, the bit-exact KV cache `max|Δ| = 0`, the single binary, and fak owning its native agent loop via the closed epic #1315); **Next** labels the planned near-term foundation (durable sessions M#1, cache default-on M#2, native-harness host seams); **Later** keeps the longer-horizon bets visible as bets (neo-silicon/neo-cloud binding #1678, the datacenter-GPU pure-fak kernel #1010). Every planned item is labeled planned, every number is witnessed (tuned ~4.1× never the naive ~60×), no market-adoption claim. Dimension E of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [fak install paths: which one command do I run?](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/install-paths.md) — a four-row decision table that routes each reader to their one correct install command: prebuilt binary ("just try it"), go install ("I have Go"), from source ("I want to hack on it"), and the one-paste MCP .mcp.json ("I drive an MCP client"). Each row has one verified command (all consistent with GETTING-STARTED and examples/mcp/.mcp.json) and its honest caveat. Dimension F of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Recipe: OpenAI SDK + fak (set one base URL)](https://github.com/anthony-chaudhary/fak/blob/main/examples/openai-sdk-minimal/README.md) — the smallest runnable repo for the universal integration recipe: `app.py` points the official OpenAI SDK at `fak serve` (one line: `base_url`), makes one `chat.completions` call, then asks fak to adjudicate two proposed tool calls without running them (`POST /v1/fak/adjudicate`, pre-execution verdict only) and prints the real verdict (`read_file`→ALLOW, `Bash "git push"`→DENY). Self-contained (the verdict endpoint needs no model); one dep (`openai`), verdict call is stdlib. Dimension G (integration recipes) of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Recipe: fak in front of Ollama (docker compose)](https://github.com/anthony-chaudhary/fak/blob/main/examples/compose-ollama/README.md) — a copy-paste `compose.yaml` + validated `policy.json` that brings up a governed local model in one command: Ollama generates on an internal-only `:11434`, a one-shot `ollama-pull` seeds the model, and `fak serve` fronts it as the only published port (`:8080`) enforcing a fail-closed `fak-policy/v1` floor + Bearer auth + audit journal. Distilled from the deployment guide's compose pattern; the policy validates offline with `fak policy --check`. Dimension G (integration recipes) of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Adopt fak in an existing repo (10-minute checklist)](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/adopt-in-your-repo.md) — the ordered checklist for taking an agent project you already run and putting fak in front in ~10 minutes: get the binary, repoint one base URL (fak guard -- claude, or fak serve --base-url), load a starter capability floor, verify a verdict appears (fak preflight / the guard exit summary), then wire the DOS trust gate into your runtime (dos init --hooks auto .) so a false "done" is refused from git evidence (dos commit-audit). Honest about what does not change (model, keys, IDE, subscription stay put) and the current limits (KV evictor is a no-op on a subscription seat; Anthropic-wire streaming is buffered). Dimension G of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).
- [Bench-story: roughly 100% evadable, on purpose](https://github.com/anthony-chaudhary/fak/blob/main/docs/adoption/stories/evadable-on-purpose.md) — the detector-honesty story: fak's injection detector is roughly 100% evadable by design and the repo says so in CLAIMS.md and SECURITY.md; the security floor is the default-deny capability lock plus structural quarantine, neither of which has to recognize an attack to stop it. Separates the evadable detector (a bonus) from the non-evadable floor + containment, and names the false-positive ceiling (2 of 59 pages sealed on benign images) and the argument-injection residual. Dimension H of the [concept-popularization epic](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-POPULARIZATION-EPIC-2026-07-02.md).

### Bounded microagents construct harnesses

- [`docs/notes/microagents-to-harnesses-2026-08-18.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/microagents-to-harnesses-2026-08-18.md) — doctrine and captured `go run ./cmd/microharnessdemo -selfcheck` witness for 1–3-turn, receipt-only child composition under the existing host admission floor.
## Operating the agent fleet

- **[Daily lock-aware Git hygiene](https://github.com/anthony-chaudhary/fak/blob/main/docs/releases/issue-5592-daily-lock-aware-git-hygiene-2026-08-08.md)** — fetch, lock, commit, push, and reconcile safely on shared trunk.

- [Run local models on Mac (Apple Silicon Metal) and interactive chat](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/mac-local-models.md) — run Qwen3.8 natively on Apple Silicon with Metal acceleration, interactive chat REPL (`fak run qwen38`), or an OpenAI-compatible gateway (`fak serve --metal`).
- [Claude Code on your Mac's local model](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/claude-mac.md) — point Claude Code at an open model on your own Mac, with the kernel adjudicating tool calls and caching shared prefixes (88.2% compute reduction, flat 180 ms TTFT at K=16).
- [Deploy fak on a rented GPU cloud](https://github.com/anthony-chaudhary/fak/blob/main/docs/fak/neo-cloud-deploy.md) — the operator quickstart for **standing the gateway up on a GPU box you rent** (CoreWeave, Lambda, RunPod, Crusoe, Vast.ai, Nebius): the two host primitives every provider is a storefront over (raw GPU VM → `docker run --gpus all` / `install.sh`+systemd; GPU k8s pool → `kubectl apply -k deploy/k8s/overlays/gpu`), and the two shapes — **in-kernel** on the card from a mounted `.gguf` (needs a `Dockerfile.cuda` build — with `CUDA_ARCH` unset the default build covers *every* arch in `internal/compute/cuda_arch.txt` plus a `compute_120` PTX floor, so it is the one to ship when you do not know which card you will be handed; `--build-arg CUDA_ARCH=sm_90` only *narrows* it. The versioned `-cuda` tag is published (`0.55.0-cuda` live, plus the moving `cuda-latest` tag) or **proxy** in front of a co-located vLLM/SGLang (the lean static image is enough). Every per-provider row is marked `not yet` end-to-end witnessed — a dogfood path, not a verified claim. Disambiguates the three pages that all say "cloud": this one deploys the **gateway**, [clouds.md](https://github.com/anthony-chaudhary/fak/blob/main/docs/supported/clouds.md) fronts a hosted **API**, and the [neo-cloud reference architecture](https://github.com/anthony-chaudhary/fak/blob/main/docs/vendor/neo-cloud-reference-architecture.md) is the **backend binding layer**. `gen/next` under epic #1678.
- [The scope of a GitHub ticket](https://github.com/anthony-chaudhary/fak/blob/main/docs/ticket-scope.md) — the front door for the ticket-scope toolkit: what the scope of one issue *is* — the six axes that decide whether a ticket is a single dispatchable unit of agent work (structure / size on the S0–S4 ladder / atomicity / write-scope / cohort placement / work-class), each mapped to the verb that measures it (`fak-dev issue contract`, `fak dispatch issue-smallness-lint`, `fak-dev issue cohort`), the closed reason it fails with, and the fix. The executable pass over it is the [`/ticket-scope`](https://github.com/anthony-chaudhary/fak/blob/main/.claude/skills/ticket-scope/SKILL.md) skill.
- [Spine-first + fan-out defaults](https://github.com/anthony-chaudhary/fak/blob/main/docs/spine-first-defaults.md) — the two defaults that fire for every new unit of work unless waived: ship the applied end-to-end **spine** first (or file the spine as its own issue), then expand proof, optimize against that working baseline, and **fan out** the follow-on QA / dogfood / productization backlog at creation time with `fak-dev issue fanout` (**3 is the floor**, not the target — `issuefanout.MinFanout`; the full taxonomy runs to ~15 candidates), wave-planned via `fak-dev issue cohort --from-plan`. The doctrine behind epic #2510; the planner spine is [`internal/issuefanout`](https://github.com/anthony-chaudhary/fak/blob/main/internal/issuefanout/issuefanout.go) (`fak-dev issue fanout`, shipped `5b8f0bd1`).
- [`fleet` operator console](https://github.com/anthony-chaudhary/fak/blob/main/tools/FLEET.md) — one command to watch the agent fleet on a host: live session/account health, what stopped and why, which accounts are resumable, and what needs you. The session-health companion to `dos top`; install it onto PATH with `fleet install`.
- [Operator-steerability PRs (`fak steer prs`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/operator-steerability-prs.md) — the doctrine note for epic #5015: a conventional PR bundles a merge gate with observability, and fak's PR-free continuous-merge trunk wants only the second — so the overlay folds landed `(fak <leaf>)` commits into PR-sized units banded RESIDUAL → UNVERIFIABLE → CLEARED (the [HUMAN_RESIDUAL choice-triage doctrine](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-CHOICE-IS-A-TRIAGE-2026-07-07.md) pointed at landed commits), stays read-mostly and gates nothing (`--check` reports, never blocks), and holds the anti-gaming line that an ack is not a witness. Leaf [`internal/steerpr`](https://github.com/anthony-chaudhary/fak/blob/main/internal/steerpr/steerpr.go); operator loop in the [`/steer-prs`](https://github.com/anthony-chaudhary/fak/blob/main/.claude/skills/steer-prs/SKILL.md) skill.
- [The out-of-band operator control plane](https://github.com/anthony-chaudhary/fak/blob/main/docs/operator-control-plane.md) — the one page that names the whole plane for epic #2753: how a human changes what a *running* session is doing without typing prose into the channel the agent reads as its task. The closed control vocabulary from [`internal/sessionctl`](https://github.com/anthony-chaudhary/fak/blob/main/internal/sessionctl/vocab.go) (`steer` · `redirect` · `pause` · `resume` · `cancel` · `terminate` · `throttle` · `budget` · `priority`), each op carrying the same four fixed properties — capability, boundary (never mid-decode), witness-of-**applied** (the running arm consuming it, never the enqueue), and closed refusal token — plus the `fak session` / `fak signal` / `fak ps` front door and the honest fences (capability named but only wired for `steer`; `redirect` has no CLI spelling; `pace`/`envelope`/`run` are unregistered writes). Also resolves the name collision: `fak steering` (Slack CI/CD reporting) and `fak steer` (PR steering) are **not** control ops — documented, deliberately not renamed.
- [Managed worker worktrees](https://github.com/anthony-chaudhary/fak/blob/main/docs/managed-worker-worktrees.md) — operator guide and runbook for detached build-isolation worker worktrees (#1334 / #3165): portable defaults (`fak worktree worker defaults`), environment overrides (`FLEET_WORKER_WORKTREE_ROOT`, `GOCACHE`, `GOTMPDIR`, `DISPATCH_WORKSPACE`), lifecycle operations (`prepare`, `list`, `land`, `reap`, `gc`), and remote crash recovery (`publish`, `recover`).
- [Super loops (`fak superloop`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/super-loops.md) — the operator-intent meta-loop: an intent like "improve quality" that **walks** its member loops/scorecards/gardens to read their status first, then folds a **worst-first** worklist of what to enter. The layer above a normal loop; five properties separate them (`has_members`, `walks_first`, `selects_worst_first`, `exits_on_aggregate`, `interior_node`), and `fak superloop explain` makes the distinction executable.
- [Super workstreams (`fak superstream`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/super-workstreams.md) — coordinates sequential multi-task execution queues across dynamic per-item lane leases while enforcing O(1) context safety (`StreamCarryoverSeed`) over long turns.
- [Guard restart continuity](https://github.com/anthony-chaudhary/fak/blob/main/docs/guard-restart-continuity.md) — what happens to your conversation when `fak guard --restart-on-budget` relaunches its child: the restart chain (budget exhausted → seed written → relaunch → handback), the closed continuity modes (resumed via `--continue` / seed-prompt (reserved, #3056) / orphaned / blocked), and the operator diagnosis for the one symptom you actually see — "conversation was compacted" right after a guarded budget restart is the seed-handback path; check `fak guard restart-audit`. The operator contract over #3055/#3057's code rungs, under the [guard-lifecycle epic #1193](https://github.com/anthony-chaudhary/fak/issues/1193).
- [Run it all night (`fak nightrun`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/nightrun/README.md) — the data-collection center of excellence: the one door an operator/agent uses to answer "what is the single most important datum I can collect on THIS box, right now?" then collect the feasible queue on a loop into a durable ledger.
- [GPU server overnight run plan (2026-06-28)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/GPU-SERVER-OVERNIGHT-PLAN-2026-06-28.md) — the fleet-scale companion: per-box overnight data-collection plan for the GPU server/CPU server boxes reached over the private control bridge (where the frontier GLM-5.2 witnesses live), with the live fleet state, the exact runbook per box, and the honesty boundary.
- [Mac metal node overnight run plan (2026-06-28)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/MAC-OVERNIGHT-PLAN-2026-06-28.md) — the Apple-Silicon companion: the fak-kernel Qwen3.6-27B Q4_K decode witness collected on the Metal verify node (warm ~1.6–1.9 tok/s, climbing with length), why `--context-budget-tokens` is load-bearing on the 36 GB box, and the resume conditions.

## Benchmarks & methodology
- [Tool-prior compatibility ledger](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/tool-prior/qwen2.5-14b-2026-08-14.json) — dated Qwen 2.5 14B raw-call corpus comparing canonical, provider-native, API, and command-style tool names (#6820)

- [Benchmark evidence authority](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-AUTHORITY.md) — the governed sheet for scoped benchmark rows, tuned baselines, artifacts, and reproduce commands.
- [Qwen performance index](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/QWEN-PERFORMANCE-INDEX.md) — the generated active readout and retained envelope-specific witnesses for Qwen native performance.
- [Latest Qwen3.8-27B results](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/QWEN38-27B-LATEST.md) — the detailed accepted, approximate, and diagnostic lifecycle without cross-envelope splicing.
- [Benchmark template](https://github.com/anthony-chaudhary/fak/blob/main/BENCHMARK-TEMPLATE.md) — the shape a new benchmark result doc must take.
- [Observed small-model Ultracode micro-context + prefix-cache win](https://github.com/anthony-chaudhary/fak/blob/main/docs/_witnesses/issue-8624-ultracode-smallmodel/README.md) — replayable `qwen2.5:0.5b` agentic frontier at widths 1/2/4/8; records equal outcomes, attributes 62.7% of credited avoidance to fak role scoping and 37.3% to ordinary Ollama/llama.cpp prefix reuse, and names the missing factorial fusion control (#8624).
- [Net-true value standard](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/net-true-value.md) — how fak decides a gain is real, not noise: the six-question rubric (real baseline / net of cost / scope / provenance / witness / realized) used on fak's own claims and on incoming industry "5×" claims, each criterion bound to the stick that enforces it.
- [Agent grammar standard](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/agent-grammar.md) — the normative trust grammar a second agent fleet conforms to: the closed nouns (lane · lease · reason token · witness · verdict · claim · ladder rung · scope), the shipped verbs each with an input→verdict signature and the closed vocabulary it draws from (every verb maps to a `dos_*` MCP verb / `dos.toml` surface today), the lift recipe as MUST clauses, the `G6` one-sided-screen + witnessed-loss polarity predicate as a checkable MUST, and a per-verb conformance checklist — the role `internal/abi`'s golden freeze plays for the ABI.
- [Observer-effect standard](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/observer-effect.md) — the cost-side companion: the perf-floor/security-floor duality, the WITNESSED/OBSERVED/MODELED/SIMULATED label required on every overhead number fak reports about itself, and the fence that a hot-path meter's own cost must be under a declared cap proven by a green test (the shipped `AcceptanceMeter` is pinned at 0 allocations/sample).
- [Support-maturity honesty fence](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/support-maturity-honesty-fence.md) — the support-side companion (epic #1243): the net-true-value lens on *support* claims. Three rules keep the maturity ladder (M0 none → M7 beyond-SOTA) from becoming a wish-list — no self-reported promotion (a rung rises only on a non-author witness, the shipgate keep-bit), a rung can drop (the scorecard re-derives every rung from the live grid, so a stale/regressed witness demotes with no latch), and every rung carries a WITNESSED/OBSERVED/MODELED label (only WITNESSED attains a rung).
- [Hardware bench plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/bench-plan.md) — auto-generated: the next highest-value test per bench-node (regenerate with `tools/bench_plan.py`, don't hand-edit).
- [Web-agent baselines](https://github.com/anthony-chaudhary/fak/blob/main/docs/webbench-baselines.md) — the modeled WebVoyager prefill geometry (vs the naive floor).
- [Web-agent blockers](https://github.com/anthony-chaudhary/fak/blob/main/docs/webbench-blockers.md) — what is not yet measured, and why.
- [Web-agent measurement summary](https://github.com/anthony-chaudhary/fak/blob/main/docs/webbench-real-measurements-summary.md) — the rolled-up real-run numbers.
- [DOS hook cost](https://github.com/anthony-chaudhary/fak/blob/main/docs/perf-dos-hook-cost.md) — what the commit-time guards cost in practice.
- [Runaway-guard cost](https://github.com/anthony-chaudhary/fak/blob/main/docs/perf-runaway-guard.md) — the cost of the process-reaper backstop.
- [Defender exclusion baseline (Windows fleet hosts)](https://github.com/anthony-chaudhary/fak/blob/main/docs/host-defender-exclusions.md) — the per-spawn Defender scan tax (~21% of a core measured), the narrow exclusion set that removes it, the elevated one-paste apply/verify, and the honest security trade-off.
- [`fak rollup` — the executive activity roll-up](https://github.com/anthony-chaudhary/fak/blob/main/docs/fleet-rollup.md) — one signal-dense page folding the agentic-fleet planes — closure honesty, dark loops, ship-stamp rate, box liveness — into a GREEN/WATCH/RED verdict and a ranked what-needs-you list.
- [Independent cross-model issue audits](https://github.com/anthony-chaudhary/fak/blob/main/docs/crossaudit-operator-runbook.md) — the operator runbook for the reciprocal audit program (#3846): `fak issue audit`, `audit-loop`, `finding`, and `fak audit` end to end.

## Status & tracking

Working docs that track a specific effort. Dated by design; they age out.

- [Executive roll-up](https://github.com/anthony-chaudhary/fak/blob/main/docs/EXECUTIVE-ROLLUP.md) — the leadership snapshot: flagship wins, the live GLM-5.2 goal, the real risks, and the one open decision — every number provenance-labeled (witnessed / observed / simulated / unverified). Aggregated from PRODUCT-STATUS, BENCHMARK-AUTHORITY, the AgentDojo red-team, the dispatch audit, and the industry scorecard.
- [Issue tracker](https://github.com/anthony-chaudhary/fak/issues) — the live, always-current open-issue count. This index never hard-codes a count (it drifts); see the tracker for the current number.
- [Support-maturity disambiguation epic (#1243)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/support-maturity-disambiguation-tracking-1243.md) — splits the one fuzzy word "supported" into a closed, witnessed level ladder (none → loads → runs → correct → optimized → SOTA-parity → beyond-SOTA) and a router that turns each cell's rung into the right dev-regime, time-horizon, and tooling — so a 10-step "make it work" is never confused with a 10,000-step "make it fast". Extends `internal/covmatrix`; the measurement+routing instrument above the Parity tracks (#307/#305/#303/#301).
- [GPU parity tracking (#480)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/gpu-parity-tracking-480.md) — bringing the GPU path to parity.
- [AMD Strix Halo APU benchmark results and optimization baseline index](https://github.com/anthony-chaudhary/fak/blob/main/docs/benchmarks/STRIX-HALO-BENCHMARK-RESULTS.md) — 19/19 sub-kernels PASS, 5/5 differential ablations VERIFIED_LIFT on physical Ryzen AI MAX+ 395 w/ Radeon 8060S (gfx1151), cross-referencing candidate improvements and measured baselines against receipt `docs/benchmarks/strix-halo-validation-11940.json`.
- [SIMD CPU parity tracking (#400)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/simd-cpu-parity-tracking-400.md) — Go-native SIMD (AVX2/AVX-512/NEON) vs llama.cpp on CPU: parity where SIMD is the lever, non-SIMD residuals named.
- [Track B performance parity tracking (#306)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/track-b-performance-parity-tracking-306.md) — roll-up status for the eight Track B children, with stale migrated issue links corrected and GPU-gated/unimplemented acceptance blockers kept explicit.
- [Track D agent-framework parity tracking (#304)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/track-d-agent-framework-parity-tracking-304.md) — roll-up status for the eight Track D children, with stale migrated issue links corrected; D-001 benchmark scaffold (smoke green, published run bench-node-gated), D-002 closed by wire repoint, D-003/D-004 wire-supported with adapter acceptance open, D-005/D-006 unimplemented, D-007/D-008 scaffolds tracked by sibling epics.
- [The self-tax plane — performance-assurance epic (#1147)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/self-tax-performance-assurance-tracking-1147.md) — first-class, always-on evidence that fak's own gates/guards/verification cost no more than a declared budget (and naming it when they make work faster): the mediation-overhead dual of the security floor, mechanizing net-true-value Q2 (net-of-cost) across turn-by-turn → post-session → CI-gate → model-as-judge → observability. Sibling of #306 (mediation overhead vs raw-inference parity).
- [Model-arch seam status (#487)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/model-arch-seam-status-487.md) — the model-architecture seam work.
- [Trust-floor decomposition (#492)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/trust-floor-decomposition-492.md) — splitting the trust floor into checkable parts.
- [Prospective exact-model v2 readout (#4851)](https://github.com/anthony-chaudhary/fak/blob/main/docs/model-acceptance-prospective-v2-readout.md) — the infrastructure-HOLD record from the prospective exact-model acceptance campaign.
- [Prospective exact-model v3 runbook (#4845)](https://github.com/anthony-chaudhary/fak/blob/main/docs/model-acceptance-prospective-v3-runbook.md) — the post-reset campaign runbook; supersedes the v2 HOLD readout for context.
- [Project-work backlog census (2026-07-15)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/project-work-backlog-census-2026-07-15.md) — OBSERVE verdict: 26 of 1,419 open issues carry valid production-work metadata, so strict project-work enforcement is not yet promotable.

## Plans & media

- [Video content plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/video-content-plan.md) — the explainer-video storyboard.
- [Qwen3.8 Mac top-ten performance plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/plans/qwen38-mac-top10-plan.md) — the finite, receipt-gated plan for ten accepted fak-native Qwen3.8-27B improvements on the sanctioned M3 Pro envelope.

## Reference index

Developer, design, and internal reference docs — indexed here so each is reachable from the map.

- [Developer tooling](https://github.com/anthony-chaudhary/fak/blob/main/docs/dev-tooling.md) — query the curated documentation map before surveying the tree, then choose the build, test, debug, profile, or committed-tip witness for the question at hand.
- [Dev-process private boundary](https://github.com/anthony-chaudhary/fak/blob/main/docs/dev-process-private-boundary.md) — architectural boundary between public runtime (The Engine) and private autonomous development factory (The Factory), with the complete [top 100 tools migration catalog](https://github.com/anthony-chaudhary/fak/blob/main/docs/dev-process-top-100-tools-inventory.md).
- [Top 100 autonomous factory tools inventory](https://github.com/anthony-chaudhary/fak/blob/main/docs/dev-process-top-100-tools-inventory.md) — authoritative catalog and ranking of the top 100 dev process tools migrating from Python to Go in fak-private across six functional cohorts.
- [Nightly trajectory attribution receipt](https://github.com/anthony-chaudhary/fak/blob/main/docs/ci/trajectory-attribution-nightly.md) — the bounded local/fleet collection, budget, publication, rollback, and scrubbed-receipt contract for the scheduled trajectory audit.
- [Default documentation self-index dogfood (2026-08-22)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/DOCUMENTATION-SELF-INDEX-DOGFOOD-2026-08-22.md) — live-repository readout for the default docs lookup: shorthand equivalence passed; multi-term relevance and path uniqueness defects were marker-deduped into #8537 and #8538.
- [`docs/notes/RMRF-ISO-ROOT-FOLDER-AUDIT-2026-07-17.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/RMRF-ISO-ROOT-FOLDER-AUDIT-2026-07-17.md) — **Root isolation scratch audit**: interrupted peer-dirty copy, buildcheck amplification, evidence-preserving quarantine, and prevention.
- [Matched Dogfood Evaluation: Structured MCP Tool Compression (2026-09-06)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/2026-09-06-mcpbroker-structured-compression-matched-dogfood.md) — 2026-09-06 matched task evaluation report for structured MCP tool compression (#11825).

**Scorecards & measurement** — [Scoreboard debt portfolio](https://github.com/anthony-chaudhary/fak/blob/main/docs/scoreboard-debt.md) (unified summary of all 50 scorecard debt categories, dual-axis ratchets, and deterministic verification) · [Bench-DX](https://github.com/anthony-chaudhary/fak/blob/main/docs/BENCH-DX-SCORECARD.md) (benchmarking developer experience) · [Claim-reproducibility](https://github.com/anthony-chaudhary/fak/blob/main/docs/CLAIM-REPRO-SCORECARD.md) (are claims falsifiable from a clean clone) · [Code-slop](https://github.com/anthony-chaudhary/fak/blob/main/docs/CODE-SLOP-SCORECARD.md) (the slop the compiler can't see) · [Verifier-exposure](https://github.com/anthony-chaudhary/fak/blob/main/docs/VERIFIER-EXPOSURE-SCORECARD.md) · [MLP first-lovable-cut](https://github.com/anthony-chaudhary/fak/blob/main/docs/mlp/scorecard.md) · [Generation portfolio RSI](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation-future-portfolio-rsi-score.md) · [Industry-scorecard freshness cadence](https://github.com/anthony-chaudhary/fak/blob/main/docs/industry-scorecard/CADENCE.md).

**Integrations** — [Amp](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/amp.md) (governed Sourcegraph Amp agent) · [Codex Memories](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/codex-memories.md) · [Gemini CLI](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/gemini-cli.md) (governed tool calls via MCP or an OpenAI-compatible gateway) · [OpenCode](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/opencode.md) (governed terminal agent).

**Adoption playbooks** — [coding/dev-agent](https://github.com/anthony-chaudhary/fak/blob/main/docs/playbooks/coding-dev-agent.md) · [DevOps/SRE](https://github.com/anthony-chaudhary/fak/blob/main/docs/playbooks/devops-sre.md) · [research-agent](https://github.com/anthony-chaudhary/fak/blob/main/docs/playbooks/research-agent.md).

**Documentation program** — [audience architecture (2026-07-15)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/DOCUMENTATION-AUDIENCE-ARCHITECTURE-2026-07-15.md) · [cohort dispatch playbook](https://github.com/anthony-chaudhary/fak/blob/main/docs/playbooks/documentation-cohort.md) · [weekly reconciliation loop](https://github.com/anthony-chaudhary/fak/blob/main/docs/playbooks/documentation-weekly-loop.md) · [public documentation style](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/public-documentation-style.md) · [metadata template](https://github.com/anthony-chaudhary/fak/blob/main/docs/templates/documentation-metadata.md) · [choice-table template](https://github.com/anthony-chaudhary/fak/blob/main/docs/templates/documentation-choice-table.md).

**Standards & contracts** — [queried-harness-overlay schema](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/queried-overlay-schema.md) · [symptom witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/symptom-witness.md) (a fix is witnessed-fixed only when a test fails on the broken tree) · [system-prompt-mutation schema](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/system-prompt-mutation-schema.md) · [session-descriptor contract](https://github.com/anthony-chaudhary/fak/blob/main/docs/session-descriptor-contract.md) (`fak.session.descriptor.v1`) · [operating-envelope declarations](https://github.com/anthony-chaudhary/fak/blob/main/docs/project-work-operating-envelopes.md) · [proof-artifact placement](https://github.com/anthony-chaudhary/fak/blob/main/docs/proof-artifact-placement.md) · [production completion & project scope](https://github.com/anthony-chaudhary/fak/blob/main/docs/project-production-completion.md) · [proportionate risk assessment](https://github.com/anthony-chaudhary/fak/blob/main/docs/standards/risk-assessment.md) (the one Risk assessment block through issue intake → preflight → witness: risk event, exposed subject, severity, likelihood, blast radius, mitigations, rollback, negative-path witness).

**DeepSeek V4 design notes** — [attention seam map](https://github.com/anthony-chaudhary/fak/blob/main/docs/deepseek/v4-attention-seam-map.md) · [FP4 quantization plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/deepseek/v4-fp4-quant-support-plan.md) · [heterogeneous KV + on-disk prefix reuse](https://github.com/anthony-chaudhary/fak/blob/main/docs/deepseek/v4-heterogeneous-kv-plan.md) · [MoE expert-dispatch baseline](https://github.com/anthony-chaudhary/fak/blob/main/docs/deepseek/v4-moe-dispatch-baseline.md) · [MTP / speculative-decoding eval plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/deepseek/v4-mtp-speculative-eval-plan.md) · [deterministic parity harness](https://github.com/anthony-chaudhary/fak/blob/main/docs/deepseek/v4-parity-harness.md).

**Kernel: MoE & expert residency** — [activated-expert offload ladder (epic #5606)](https://github.com/anthony-chaudhary/fak/blob/main/docs/MOE-ACTIVATED-OFFLOAD-PLAN.md) — the plan of record for making MoE *activated*-expert offloading first-class: today's models fire only a few percent of their experts per token, so the stored-vs-activated byte gap is a **residency** problem, and the ladder (R0 bound the activated set → R1 graded spill knob → R2 durable pin-set → R3 same-step prefetch → R4 evictor promoted on measured regret → R5 checkpoint tier under a ring miss → R6 operator surface → R7 cross-agent coalescing → R8 grouped GEMM) says which rung each optimization is, what its witness is, and which are landed.

**Context & cache** — [gateway cold-tool deferral](https://github.com/anthony-chaudhary/fak/blob/main/docs/context-budget/gateway-cold-tool-deferral.md) · [MCP tool-schema floor](https://github.com/anthony-chaudhary/fak/blob/main/docs/context-budget/mcp-tool-floor.md) · [AGENTS.md instruction-pulled floor](https://github.com/anthony-chaudhary/fak/blob/main/docs/context-budget/agents-md-floor.md) · [measured anchors outrank synthetic Zipf](https://github.com/anthony-chaudhary/fak/blob/main/docs/cache-frontier/next-50-item-05-measured-anchors.md) · [context control surfaces & the pin-audit gap](https://github.com/anthony-chaudhary/fak/blob/main/docs/explainers/context-control-surfaces-and-the-pin-audit-gap.md).

**Fleet, dispatch & operations** — [fleet self-discovery spine](https://github.com/anthony-chaudhary/fak/blob/main/docs/FLEET-SELF-DISCOVERY-SPINE.md) · [dispatch session observability plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/DISPATCH-SESSION-OBSERVABILITY-PLAN.md) · [guard-session dataset plan](https://github.com/anthony-chaudhary/fak/blob/main/docs/GUARD-SESSION-DATASET-PLAN.md) · [issue-pickup 10× program](https://github.com/anthony-chaudhary/fak/blob/main/docs/ISSUE-PICKUP-10X-PROGRAM.md) · [blast-radius containment cohort](https://github.com/anthony-chaudhary/fak/blob/main/docs/blast-radius-containment-cohort.md) (hub; the nine ticket bodies sit one-per-file under `docs/blast-radius-containment/`) · [recurring loops inventory](https://github.com/anthony-chaudhary/fak/blob/main/docs/loops-inventory.md) · [tier-to-account routing](https://github.com/anthony-chaudhary/fak/blob/main/docs/tier-account-routing.md) · [region: one-sided shared-result pool](https://github.com/anthony-chaudhary/fak/blob/main/docs/region.md) · [host-termination provenance](https://github.com/anthony-chaudhary/fak/blob/main/docs/host-termination-provenance.md) · [model production-readiness inventory](https://github.com/anthony-chaudhary/fak/blob/main/docs/model-production-readiness-inventory.md) · [concept rename (`fak rename-concept`)](https://github.com/anthony-chaudhary/fak/blob/main/docs/concept-rename.md) · [cmd-lane split plan (#4320)](https://github.com/anthony-chaudhary/fak/blob/main/docs/dispatch/cmd-lane-split-plan.md).

**Quality, provenance & supply chain** — [claim-proof evaluator route](https://github.com/anthony-chaudhary/fak/blob/main/docs/quality/claim-proof-proximity-audit.md) · [output-quality regression runbook](https://github.com/anthony-chaudhary/fak/blob/main/docs/quality/output-quality-regression-runbook.md) · [durable-artifacts inventory](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/durable-artifacts.md) · [supply-chain reproducibility](https://github.com/anthony-chaudhary/fak/blob/main/docs/supply-chain-reproducibility.md) · [checkpoint Go-port decision](https://github.com/anthony-chaudhary/fak/blob/main/docs/checkpoint-go-port-decision.md) · [CI/CD reporting Slack sink](https://github.com/anthony-chaudhary/fak/blob/main/docs/decisions/cicd-reporting-slack-sink.md) · [release branch-regime and actionable CI base-red status](https://github.com/anthony-chaudhary/fak/blob/main/docs/release-branch-regime-status.md) · [front-door clarity scorecard](https://github.com/anthony-chaudhary/fak/blob/main/docs/quality/frontdoor-clarity-scorecard.md) · [front-door link witness](https://github.com/anthony-chaudhary/fak/blob/main/docs/quality/frontdoor-link-witness.md) · [generation-drift check](https://github.com/anthony-chaudhary/fak/blob/main/docs/quality/generation-drift-check.md) · [public-negation audit](https://github.com/anthony-chaudhary/fak/blob/main/docs/quality/public-negation-audit.md) · [public terminology audit](https://github.com/anthony-chaudhary/fak/blob/main/docs/quality/public-terminology-audit.md).

**Operator & user routes** — [troubleshooting route](https://github.com/anthony-chaudhary/fak/blob/main/docs/troubleshooting.md) (start from the user-visible symptom) · [upgrade route](https://github.com/anthony-chaudhary/fak/blob/main/docs/upgrade.md) (choose a release, migrate configuration, preserve rollback) · [operator route](https://github.com/anthony-chaudhary/fak/blob/main/docs/operator/README.md) (deploy, observe, recover, upgrade; routes on to [policy authoring](https://github.com/anthony-chaudhary/fak/blob/main/docs/policy.md)) · [observability route](https://github.com/anthony-chaudhary/fak/blob/main/docs/observability/README.md) (choose the production signal) · [config bails](https://github.com/anthony-chaudhary/fak/blob/main/docs/config-bails.md) (every configuration refusal's reason token, the check, and the `fak recover` fix).

**Area indexes** — [implementation docs by generation](https://github.com/anthony-chaudhary/fak/blob/main/docs/generation/README.md) · [milestone tracking ledger](https://github.com/anthony-chaudhary/fak/blob/main/docs/milestones/README.md) · [#grafana link registry](https://github.com/anthony-chaudhary/fak/blob/main/docs/grafana/README.md) · [Confluence publishing](https://github.com/anthony-chaudhary/fak/blob/main/docs/confluence/README.md) · [Kimi K3 route page](https://github.com/anthony-chaudhary/fak/blob/main/docs/kimi-k3/README.md).

**Scorecards (cont.)** — [intent-literal / metric-honesty](https://github.com/anthony-chaudhary/fak/blob/main/docs/intent-literal-scorecard/README.md) · [persona-fit](https://github.com/anthony-chaudhary/fak/blob/main/docs/persona-fit-scorecard/README.md) · [fak-native TOON](https://github.com/anthony-chaudhary/fak/blob/main/docs/toon-scorecard/README.md).

**Caching & serving concepts** — [Awesome Caching](https://github.com/anthony-chaudhary/fak/blob/main/docs/awesome-caching/README.md) (every caching concept fak knows, each in its own words) · [disaggregable compute-claim ladder](https://github.com/anthony-chaudhary/fak/blob/main/docs/serving/COMPUTE-CLAIM-TAXONOMY.md) · [Mobile FFI](https://github.com/anthony-chaudhary/fak/blob/main/docs/integrations/mobile.md) (gating on-device tool calls through fak).

**Slack channels** — [#blockers](https://github.com/anthony-chaudhary/fak/blob/main/docs/blockers-channel.md) · [#releases](https://github.com/anthony-chaudhary/fak/blob/main/docs/releases-channel.md).

**Dated baselines & drills** — [systems baseline (2026-06-26)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/SYSTEMS-BASELINE-2026-06-26.md) · [anchor-strategy net-dollar verdict (2026-07-11)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/ANCHOR-STRATEGY-NET-DOLLAR-VERDICT-2026-07-11.md) · [cache-value score regression (2026-07-04)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CACHE-VALUE-SCORE-REGRESSION-2026-07-04.md) · [logvault restore drill (2026-07-09)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/logvault-restore-drill-2026-07-09.md) · [archived pre-fresh-start README (2026-06-25)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/README-2026-06-25-before-fresh-start.md) · [DeepSeek-V4 pure-fak native-GPU witness status (2026-07-16)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/DEEPSEEK-V4-PURE-FAK-NATIVE-GPU-WITNESS-STATUS-2026-07-16.md).

**Vendor & agent-tool front doors** — [neo-silicon internal pilot pitch](https://github.com/anthony-chaudhary/fak/blob/main/docs/vendor/neo-silicon-internal-pitch.md) (fak for accelerator vendors) · [AGENT.md](https://github.com/anthony-chaudhary/fak/blob/main/AGENT.md) · [GEMINI.md](https://github.com/anthony-chaudhary/fak/blob/main/GEMINI.md).

- [Learning observation lineage](https://github.com/anthony-chaudhary/fak/blob/main/docs/learning-observation.md) — content-addressed source/candidate/witness/verdict records and closed-enum edges; separate from witness-gated admission.

## Notes & research (`docs/notes/`)

- [2026-09-19 — V4.1 native track: promoted the O(E log k) partial top-k into the V4 router seam (#12975), landed and closed](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/2026-09-19-v41-12975-partial-top-k-landed.md) -- auto-indexed dated note.
- [`docs/notes/CONCEPT-STUDY-AZHU9701-NINFER-4090D-2026-09-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-AZHU9701-NINFER-4090D-2026-09-03.md) — pinned deep study of Azhu9701/ninfer-4090d@3eaa163: a Windows RTX 4090 D inference deployment dossier whose headline 114-SM wave-grid / DirectStorage 1.3 / E8-lattice KV claims are unverified README prose (no engine source in-repo, no license), with two genuine artifacts — a restart-based load-adaptive draft-depth controller and a tool-call-drift salvage patch; dynamic MTP depth is PRESENT-on-axis in fak, while lax parameter salvage is PARTIAL and filed as #12944; parent #10193, portfolio #10960.
- [Three-Tree Synchronization Invariants and Multi-Node Safe Sync Rules](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/SAFE-SYNC-THREE-TREE-INVARIANTS-2026-09-09.md) -- auto-indexed dated note.
- [`docs/notes/CONCEPT-STUDY-PURE-GO-HOST-BINDINGS-2026-09-08.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-PURE-GO-HOST-BINDINGS-2026-09-08.md) — inventory and source-pinned migration plan for replacing accelerator CGo/C++ host shims with Go-toolchain-owned bindings while retaining CUDA/PTX, MSL, SPIR-V, explicit engine identity, and real-device witnesses; parent #12210, shipped loader spine #12212, sequenced Vulkan leaves #12283–#12286, CUDA loader #11657; receipt `study_3cd46be3b861ab9872e7c8cb01f81def505a9cecbbc6cf80e728087c689558dc`.
- [Agent-first abstractions for a massively multiagent world](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-AGENT-FIRST-ABSTRACTIONS-2026-09-06.md) — pinned Temporal, Ray, A2A, and LangGraph research; durable work, effects, evidence, and bounded demand; original hypotheses and four scoped follow-ons under #11949.
- [`docs/notes/CONCEPT-STUDY-STRIX-HALO-ECOSYSTEM-2026-09-05.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-STRIX-HALO-ECOSYSTEM-2026-09-05.md) — pinned deep study of the open-source AMD Strix Halo APU ecosystem across 6 repositories (qwen38-strix-halo-harness, hal0, strix-halo-llamacpp, halofpx, ember, amd-strix-halo-toolboxes): auxiliary-model compaction, UMA 16-channel f16 KV contiguization (+169% prefill), Flash Attention dequant-once, watchdog byte-gating, 194ms full-prompt prefix caching, heterogeneous NPU/GPU drafting, and autonomous thermal governing; filed #11746-#11749 under epics #11572, #11149, #11241.
- [`docs/notes/CONCEPT-STUDY-AMD-GPU-DIRECT-ROCM-RDMA-2026-09-04.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-AMD-GPU-DIRECT-ROCM-RDMA-2026-09-04.md) — pinned deep study of OSS AMD GPU Direct, ROCm-RDMA, xGMI/PCIe P2P, and BaM NVMe Direct DMA: Linux DMA-BUF export (`AMDKFD_IOC_EXPORT_DMABUF`), `ibv_reg_dmabuf_mr` zero-copy memory registration without host bounce buffers (#11228), topology discovery and ReBAR/ACS validation (#11227), direct NVMe P2P storage streaming (#11229), sub-microsecond HSA completion signals (#11230), RDMA verbs QP engine (#11262), NVMe VRAM queue storage memory slab (#11263), AMDGPUDirectCollective communicator (#11264), and fak-dev CLI tooling (#11265); parent epic #11226; receipt `study_6f42749e9f4a8d98a9ca9bd3686fd5dc863fc131dd7c2951a30d4044645748f5`.
- [Concept Study: adrienbrault/qwen3.8-27b-rtx5090 — Consumer Blackwell Serving, sm120 NVFP4 Scaling, and Direct-I/O Disk KV Tier](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-ADRIENBRAULT-RTX5090-2026-09-03.md) -- auto-indexed dated note.
- [CONCEPT-STUDY: airawatraj/dgx-spark-qwen38-flash-agent — HashK GPU PLE compression, Mamba DeltaNet speculative rollback invariants, and Blackwell SM121 runtime patches (2026-09-03)](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-AIRAWATRAJ-HASHK-PLE-2026-09-03.md) -- auto-indexed dated note.
- [Concept Study: davidcanar/vllm-strix-halo — Dual AMD Strix Halo APU Cluster over Thunderbolt-4 RoCE-RDMA](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-DAVIDCANAR-STRIX-HALO-ROCE-2026-09-03.md) -- auto-indexed dated note.
- [Concept Study: MindLab-Research/ferrite — GLM-5.3-Flash Native Inference Engine](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-FERRITE-GLM53-2026-09-03.md) -- auto-indexed dated note.
- [Concept Study: hasso5703/dgx-spark-qwen38 — High-Performance Qwen3.8 GB10 Serving Stack](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-HASSO5703-DGX-SPARK-2026-09-03.md) -- auto-indexed dated note.
- [Concept Study: vcruz305/GLM-5.3-Flash-EXL3-K2-DGX-Spark-recipe — Single-Spark 128 GB UMA Serving & CUDA Kernel Forensics](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-VCRUZ305-GLM53-EXL3-2026-09-03.md) -- auto-indexed dated note.
- [Mac head-to-head benchmark: fak vs llama.cpp vs MLX](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/MAC-BENCH-FAK-LLAMACPP-MLX-2026-09-03.md) -- auto-indexed dated note.
- [Mac head-to-head benchmark: fak-native vs llama.cpp vs MLX on Apple Silicon Metal](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/MAC-THREEWAY-BENCH-2026-09-03.md) -- auto-indexed dated note.
- [Mac agentic shared-cache benchmark: fak-native vs llama.cpp 4.2x speedup on Qwen3.8-27B](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/MAC-AGENTIC-4X-QWEN38-2026-09-05.md) -- auto-indexed dated note.
- [`docs/notes/CONCEPT-HOT-SWAPPABLE-SERVING-ARCHITECTURE-2026-09-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-HOT-SWAPPABLE-SERVING-ARCHITECTURE-2026-09-03.md) — hot-swappable AI inference serving architecture, dynamic zero-downtime reconfiguration of serving parameters, shift-left validation, monotonic epochs, and transient sweep APIs for autonomous RSI agents.
- [`docs/notes/CONCEPT-NATIVE-HARNESS-DATABASE-AND-DATASLOT-LIFECYCLE-2026-09-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-NATIVE-HARNESS-DATABASE-AND-DATASLOT-LIFECYCLE-2026-09-03.md) — native harness database and data-slot lifecycle architecture: zero-CGo session persistence WAL and dormant database reflection, bounded query, and migration safety (#10646, #10652).
- [`docs/notes/CONCEPT-STUDY-PERFORMANCE-OSS-PROCESS-FAILURES-2026-09-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-PERFORMANCE-OSS-PROCESS-FAILURES-2026-09-03.md) — pinned forensic study of engineering process failures across high-commit performance OSS (vLLM, llama.cpp, SGLang, TensorRT-LLM, PyTorch, Triton) and transferred FAK guardrails (#10933-#10938).
- [`docs/notes/CONCEPT-STUDY-ROCM-HRX-SYSTEM-2026-09-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-ROCM-HRX-SYSTEM-2026-09-03.md) — pinned deep study of AMD ROCm HRX System and Loom compiler (ROCm/hrx-system@4c5f2d9): alternative ultra-low-latency HIP runtime, direct user-space AQL and PM4 GPU command processor submission (#11093), standalone sub-2ms HSACO ELF64 generation without LLVM dependencies (#11094), KPACK compressed kernel packaging (#11095), stream-ordered VMM slab allocator with virtual memory reservation (#11096), in-band GPU kernel AddressSanitizer device events (#11097), and in-IR test oracle/benchmark specification with parameter dictionaries (#11098); parent epic #11092; receipt `study_2f5d10b5eb90152e1263112cf68dda0387221ce77cf11f5fb7d586975e691016`.

- [`docs/notes/CONCEPT-STUDY-LEMONADE-2026-09-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-LEMONADE-2026-09-03.md) — pinned deep study of Lemonade SDK (lemonade-sdk/lemonade@bc8d99c): multi-backend C++20 server (lemond), GGUF speculative draft companion discovery and quant-bit distance matching (#11086), zero-byte-streamed transparent retry on accelerator reset (#11087), reload-cost-weighted model eviction scoring under memory pressure (#11088), in-band route decision injection into SSE streams (#11089), OS suspend/idle inhibitor during active inference (#11090), and dynamic model target aliasing with active-standby hot failover (#11091); receipt `study_cc3b4fe925b988801bccaeb92d141ee0e0f0a37cc17e8edb0e4fb3a883df1e69`.

- [`docs/notes/CONCEPT-STUDY-QWEN38-27B-RTX3090-2026-09-03.md`](https://github.com/anthony-chaudhary/fak/blob/main/docs/notes/CONCEPT-STUDY-QWEN38-27B-RTX3090-2026-09-03.md) — pinned deep study of syv-ai/qwen38-27b-rtx3090: Lookup-Augmented Block Drafting (LABD, 381 tok/s decode), split-KV multi-query speculative attention, AutoRound ne

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.