agentleFS
Sign inSign up

contextweaver

dgenio/contextweaver/llms-full.txt

Dynamic context management for tool-using AI agents. Source files (concatenated in order): README.md docs/architecture.md docs/concepts.md docs/quickstart.md docs/dailydriver.md docs/securitymodel.md docs/securitymcpgateway.md docs/sensitivity.md docs/recipes/index.md docs/recipes/claudecode.md docs/recipes/okfbundle.md docs/integrationmcp.md docs/integrationa2a.md docs/errors.md docs/agent-context/architecture.md docs/agent-context/invariants.md docs/agent-context/workflows.md docs/agent-context/lessons-learned.md docs/agent-context/review-checklist.md docs/guideagentloop.md To regenerate: make llms (or python scripts/gen_llms.py). --> [](https://github.com/dgenio/contextweaver/actions/workflows/ci.yml) [](https://pypi.org/project/contextweaver/) [](https://pypi.org/project/contextweaver/) [](LICENSE) [](https://scorecard.dev/viewer/?uri=github.com/dgenio/contextweaver) [](https://dgenio.github.io/contextweaver) [](https://github.com/dgenio/contextweaver/discussions) Capture an agent's effective capability surface, commit it, and see semantically meaningful changes before deployment. ContextWeaver is currently testing a deliberately narrow product hypothesis: capability snapshot + semantic drift. Given an OpenAPI document,…

llms.txt9 starsChanged 3 months ago
  • Reads credentials
  • Installs packages
# contextweaver — Full Documentation

> Dynamic context management for tool-using AI agents.

---

<!--
  GENERATED FILE — do not edit by hand.

  Source files (concatenated in order):
    README.md
    docs/architecture.md
    docs/concepts.md
    docs/quickstart.md
    docs/daily_driver.md
    docs/security_model.md
    docs/security_mcp_gateway.md
    docs/sensitivity.md
    docs/recipes/index.md
    docs/recipes/claude_code.md
    docs/recipes/okf_bundle.md
    docs/integration_mcp.md
    docs/integration_a2a.md
    docs/errors.md
    docs/agent-context/architecture.md
    docs/agent-context/invariants.md
    docs/agent-context/workflows.md
    docs/agent-context/lessons-learned.md
    docs/agent-context/review-checklist.md
    docs/guide_agent_loop.md

  To regenerate: `make llms` (or `python scripts/gen_llms.py`).
-->

<!-- FILE: README.md -->

# contextweaver

<!-- mcp-name: io.github.dgenio/contextweaver -->

[![CI](https://github.com/dgenio/contextweaver/actions/workflows/ci.yml/badge.svg)](https://github.com/dgenio/contextweaver/actions/workflows/ci.yml)
[![PyPI version](https://img.shields.io/pypi/v/contextweaver.svg)](https://pypi.org/project/contextweaver/)
[![Python versions](https://img.shields.io/pypi/pyversions/contextweaver.svg)](https://pypi.org/project/contextweaver/)
[![License: Apache-2.0](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)
[![OpenSSF Scorecard](https://api.securityscorecards.dev/projects/github.com/dgenio/contextweaver/badge)](https://scorecard.dev/viewer/?uri=github.com/dgenio/contextweaver)
[![Docs](https://img.shields.io/badge/docs-mkdocs--material-blue.svg)](https://dgenio.github.io/contextweaver)
[![GitHub Discussions](https://img.shields.io/github/discussions/dgenio/contextweaver)](https://github.com/dgenio/contextweaver/discussions)

> **Capture an agent's effective capability surface, commit it, and see semantically meaningful changes before deployment.**

ContextWeaver is currently testing a deliberately narrow product hypothesis:
**capability snapshot + semantic drift**.

Given an OpenAPI document, a captured MCP `tools/list` response, or a native
ContextWeaver catalog, the D1 experiment produces a deterministic normalized
snapshot that you can inspect, verify, and compare with a later candidate.
It does not require a model account, a gateway, a tool executor, or the Weaver
Stack.

**Status:** alpha, and specifically a **product experiment**. The implementation
works and is tested; the user-value hypothesis is not yet proven. The project is
actively measuring whether independent users keep this workflow after trying it
on real projects.

## Try the capability-drift experiment

Clone the repository and install that checkout so the maintained example
fixtures and the code under evaluation are guaranteed to match:

Commands below use `python3`, matching the repository's documented default (`docs/agent-context/workflows.md`). On Windows installations that expose Python through the launcher, replace `python3` with `py -3`; if your environment already exposes the intended interpreter as `python`, use that consistently.

```bash
git clone --depth 1 https://github.com/dgenio/contextweaver.git
cd contextweaver
python3 -m pip install .
```

Run the maintained OpenAPI example:

```bash
python3 -m contextweaver.d1 snapshot examples/d1/openapi_before.json --source-type openapi --output ./cw-before.json
python3 -m contextweaver.d1 snapshot examples/d1/openapi_after.json --source-type openapi --output ./cw-after.json
python3 -m contextweaver.d1 inspect ./cw-after.json
python3 -m contextweaver.d1 verify ./cw-after.json
python3 -m contextweaver.d1 diff ./cw-before.json ./cw-after.json
```

The candidate fixture intentionally:

- makes `customer_id` required on the existing `listInvoices` capability;
- changes its description;
- adds a new `getInvoice` capability.

The diff separates capability additions/removals from changes to an existing
logical capability and reports the structured paths that changed. Contract
changes are separated from documentation-only changes. Changes involving fields
such as `required`, `type`, or `enum` are flagged as **potentially breaking** for
review.

That flag is intentionally conservative: ContextWeaver does **not** claim to be
a complete JSON-Schema compatibility checker.

Full walkthrough: [Capability drift experiment](docs/d1_capability_drift.md).

## Use it on your own source

### OpenAPI

```bash
python3 -m contextweaver.d1 snapshot ./openapi.yaml --source-type openapi --output ./capabilities.json
python3 -m contextweaver.d1 verify ./capabilities.json
```

After the API changes:

```bash
python3 -m contextweaver.d1 snapshot ./openapi.yaml --source-type openapi --output ./capabilities-candidate.json
python3 -m contextweaver.d1 diff ./capabilities.json ./capabilities-candidate.json
```

### Captured MCP tools

If you already have an MCP `tools/list` response saved as JSON:

```bash
python3 -m contextweaver.d1 snapshot ./tools-list.json \
  --source-type mcp \
  --output ./capabilities.json
```

For MCP, D1 compares tools by their upstream logical name so an input-schema
edit appears as a change to the same capability rather than an unexplained
remove/add pair. The historical schema-sensitive routing ID is retained
separately as `normalized_id` for inspection.

Capturing a live MCP server is a separate operation. `snapshot`, `inspect`,
`diff`, and `verify` do not execute discovered capabilities.

### Native ContextWeaver catalog

```bash
python3 -m contextweaver.d1 snapshot ./catalog.json \
  --source-type native \
  --output ./capabilities.json
```

## What `verify` means

`verify` checks the D1 snapshot contract: structure, deterministic ordering,
logical-ID uniqueness, and the canonical capability digest.

It is **not**:

- deployment approval;
- security certification;
- authentication or authorization;
- a guarantee that a tool implementation is correct;
- routing-quality evaluation;
- production runtime attestation.

## When not to use ContextWeaver D1

A negative answer is useful evidence for this project. Do **not** add
ContextWeaver just because capability snapshots sound tidy.

Use something simpler when:

- **ordinary Git diff, config review, and tests already make your capability
  changes obvious;**
- your tool/API surface is tiny and rarely changes;
- provider-native tool search is the only problem you are trying to solve;
- you need an agent loop, tool executor, IAM layer, or production orchestrator;
- maintaining another committed artifact costs more than the review/debugging
  problem it removes.

If you try D1 and conclude that Git/tests are cheaper, that is a valid product
result — please say so.

## What is being tested

The current survival experiment asks a stronger question than whether the code
works:

> Do capability snapshots and semantic drift reports improve a real
> review/manual/risk process enough that independent users keep them?

The project distinguishes:

```text
qualified exposure
  -> understood the problem
  -> chose to evaluate
  -> attempted setup
  -> reached first useful output
  -> used on a real project
  -> retained independently / removed
```

Stars, forks, downloads, a successful demo, and maintainer-created integrations
are not treated as retained adoption.

The controlling product decision is tracked in
[#758](https://github.com/dgenio/contextweaver/issues/758), and the distribution
quality gate is [#855](https://github.com/dgenio/contextweaver/issues/855).
Unassisted first success and retention are tracked in
[#658](https://github.com/dgenio/contextweaver/issues/658) and genuine adoption
in [#551](https://github.com/dgenio/contextweaver/issues/551).

## What about routing, context compilation, and the MCP gateway?

ContextWeaver already contains substantial historical runtime functionality.
That code still exists and currently shipped behavior should remain truthful and
safe, but **existing implementation is not evidence that the project should
keep expanding it**.

Two broader hypotheses are explicitly evidence-first:

- **D2 — bounded / phase-aware context compilation:** conditional. It must show
  consequential value beyond contemporary provider/runtime-native mechanisms.
- **D3 — custom deterministic tool selection:** a falsification track. It must
  beat modern provider-native tool search/deferred loading or a simple retrieval
  baseline on something target users actually care about.

During the D1 experiment, the project is not expanding routing sophistication,
runtime bundle machinery, memory/session surfaces, framework breadth, gateway
scope, vector stores, or model-assisted enrichment without a concrete external
blocker or approved falsification experiment.

If you are maintaining an existing integration that uses those historical
surfaces, the relevant documentation remains available:

- [Which historical pattern fits?](docs/which_pattern.md)
- [Context firewall](docs/context_firewall.md)
- [MCP Context Gateway architecture](docs/architectures/mcp_context_gateway.md)
- [MCP gateway security model](docs/security_model.md)
- [Comparison / alternatives](docs/comparison.md)
- [Ecosystem map](docs/ecosystem.md)

## Evidence and claims

The D1 implementation supports scoped engineering claims such as deterministic
snapshot construction under the documented source/adapter contract and
structured semantic-diff output. It does **not** yet support the stronger claim
that users need or retain the product.

The historical token-reduction headline is intentionally not used to sell D1.
The current evidence-integrity work for those older benchmark claims is tracked
in [#841](https://github.com/dgenio/contextweaver/issues/841).

See [Claims & evidence](docs/claims.md) for the claim registry and
[Capability drift experiment](docs/d1_capability_drift.md) for the exact D1
contract and limitations.

## Python API stability

D1 is intentionally exposed through:

```bash
python3 -m contextweaver.d1 ...
```

rather than being promoted immediately into the historical top-level CLI or a
large new public Python API. That is deliberate. The experiment should earn a
permanent surface through real retained use before the project takes on another
compatibility obligation.

## Part of the Weaver Stack — optionally

ContextWeaver can be used standalone. It has no hard dependency on the sibling
Weaver projects.

The wider Weaver Stack contains adjacent experiments/components for planning,
execution boundaries, guardrails, lessons, and evaluation. That ecosystem is
**not required** to evaluate D1, and Stack coherence is not a reason to preserve
a ContextWeaver feature that does not justify itself independently.

See the [Ecosystem map](docs/ecosystem.md) only if you actually need those
adjacent responsibilities.

## Install and compatibility

```bash
pip install contextweaver
```

Python 3.10–3.14 are covered by the repository CI matrix.

Current package version: **0.18.2**

| Project | Release |
|---|---|
| ContextWeaver (this repo, [v0.18.2](https://github.com/dgenio/contextweaver/releases/tag/v0.18.2)) | current package release |

The repository is pre-1.0. Prefer the latest supported patch release for bug and
security fixes, and check the changelog before relying on historical runtime
APIs.

## Current roadmap

The roadmap is intentionally a product-decision sequence, not a feature queue.

| Milestone | Status | Meaning |
|---|---|---|
| **v0.18.1 — D1 survival experiment baseline** | ✅ complete | Offline snapshot/inspect/diff/verify exists; user value remains unverified. |
| **v0.18.2** | ✅ current (v0.18.2) | D1 survival experiment and release-path recovery |
| **D1 distribution gate** | 🔬 evidence first | Make the front door understandable, recruit qualified evaluators, measure first success and retention. |
| **D1 decision** | ⏸ next decision | Continue, shrink further, or kill based on retained value after competent distribution. |
| **D2 / D3** | 🧪 conditional | Run only if D1 evidence or independent problem discovery justifies bounded falsification experiments. |

A green CI run does not advance this roadmap by itself.

## Contributing

The most valuable contributions during the survival experiment are narrow and
evidence-linked:

- a real D1 evaluator blocker;
- a semantic-diff case that is currently misleading or silently lost;
- deterministic normalization correctness;
- security/release maintenance for behavior the package still ships;
- negative evidence showing a simpler alternative wins.

Please do not add a framework adapter, routing policy, storage backend, runtime
phase, or ecosystem integration solely for completeness.

See [CONTRIBUTING.md](CONTRIBUTING.md) and [AGENTS.md](AGENTS.md) for repository
engineering conventions.

## Security

See [SECURITY.md](SECURITY.md) for supported-version and vulnerability-reporting
guidance. Do not include credentials, customer data, proprietary schemas, or
private prompts in public adoption/evaluation reports.

## Documentation

- [Documentation site](https://dgenio.github.io/contextweaver)
- [Capability drift experiment](docs/d1_capability_drift.md)
- [Claims & evidence](docs/claims.md)
- [Daily Driver guide](docs/daily_driver.md) — historical/runtime users
- [Cookbook](docs/cookbook.md) — broader shipped surfaces
- [FAQ](docs/faq.md)
- [Changelog](CHANGELOG.md)

## License

Apache-2.0. See [LICENSE](LICENSE).

---

<!-- FILE: docs/architecture.md -->

# Architecture

> **New here?** [Which pattern fits my use case?](which_pattern.md) routes you
> to the smallest piece that fixes your symptom — most callers only need one
> of the two engines, not the whole pipeline.

contextweaver is structured around two cooperating engines that together solve
the "context window problem" for tool-using AI agents.

## Why context engineering matters

The discipline of **context engineering** — deciding *what* goes into a model's
context window, *when*, and *at what cost* — has emerged as the lever that
moves quality and latency once tool-use agents reach production scale. Even
with 200K-token windows, dumping every tool schema and conversation turn into
the prompt is expensive, slows latency, and degrades output quality as the
model's effective attention thins. The lever is selective compilation
(per-phase budgets, tool shortlisting, oversized-output firewalling), not raw
window size.

contextweaver implements that lever as two cooperating engines — the Context
Engine (eight-stage pipeline) and the Routing Engine (bounded DAG + beam
search) — and treats determinism, dependency closure, and sensitivity filters
as load-bearing invariants rather than nice-to-haves. For background on the
term itself, see Atlan's
[*What Is Context Engineering*](https://atlan.com/know/what-is-context-engineering/).

## High-level overview

```
               ┌────────────────────────────┐
  Events ─────>│      Context Engine         │──> ContextPack (prompt)
               │  candidates → closure →     │
               │  sensitivity → firewall →   │
               │  score → dedup → select →   │
               │  render                     │
               └────────────────────────────┘
                          ▲ facts / episodes
               ┌──────────┴─────────────────┐
  Tools ──────>│      Routing Engine         │──> ChoiceCards
               │  Catalog → TreeBuilder →    │
               │  ChoiceGraph → Router       │
               └────────────────────────────┘
```

## Package layout

| Path | Responsibility |
|---|---|
| `types.py` | Core dataclasses and enums (`SelectableItem`, `ContextItem`, `Phase`, `ItemKind`) |
| `envelope.py` | Result types (`ResultEnvelope`, `BuildStats`, `ContextPack`, `ChoiceCard`, `HydrationResult`) |
| `diagnostics.py` | Versioned gateway event schema, JSONL/in-memory sinks, aggregate reports |
| `inspection.py` | Payload-safe offline context/routing/artifact reports |
| `config.py` | Configuration dataclasses (`ContextBudget`, `ContextPolicy`, `ScoringConfig`) |
| `protocols.py` | Protocol interfaces (`TokenEstimator`, `EventHook`, `Summarizer`, …) |
| `exceptions.py` | Custom exception hierarchy |
| `_utils.py` | Text similarity primitives (`tokenize`, `jaccard`, `TfIdfScorer`) |
| `serde.py` | Serialisation helpers for `to_dict` / `from_dict` patterns |
| `store/` | In-memory data stores (`EventLog`, `ArtifactStore`, `EpisodicStore`, `FactStore`) |
| `summarize/` | Rule engine and structured fact extraction |
| `context/` | Full context compilation pipeline |
| `routing/` | Catalog, DAG builder, beam-search router, card renderer |
| `adapters/` | MCP, FastMCP, and A2A protocol adapters |
| `__main__.py` | CLI entry point (`inspect` includes context/routing/artifact diagnostics) |

## Context Engine pipeline

The Context Engine compiles a phase-aware, budget-constrained prompt from
the event log. The pipeline has eight stages:

1. **generate_candidates** — pull phase-relevant events from the event log
   into the initial candidate pool.
2. **dependency_closure** — if a selected item has a `parent_id`, bring
   the parent along even if it scored lower.
3. **sensitivity_filter** — drop or redact items whose `sensitivity`
   level meets or exceeds `ContextPolicy.sensitivity_floor`.
4. **apply_firewall** — tool results are stored out-of-band in the
   ArtifactStore and replaced with summarized/truncated text for prompt
   assembly.
5. **score_candidates** — rank candidates by recency, tag match, kind
   priority, and token cost.
6. **deduplicate_candidates** — remove near-duplicate items using Jaccard
   similarity over tokenised text.
7. **select_and_pack** — greedily pack the highest-scoring candidates
   into the token budget for the current phase.
8. **render_context** — assemble the final prompt string, grouped by
   section (facts, history, tool results), with `BuildStats` metadata.

The pipeline owns `BuildStats` construction after selection. Candidate totals
are captured before sensitivity filtering, while sensitivity, deduplication,
kind-limit, and budget exclusions are attributed per item. This preserves the
invariant `included_count + dropped_count == total_candidates` and keeps
lifecycle hooks aligned with the returned statistics.

## Routing Engine pipeline

The Routing Engine efficiently navigates large tool catalogs so the LLM
never sees all tools at once:

1. **Catalog** — register and manage `SelectableItem` objects.
2. **TreeBuilder** — convert a flat item list into a bounded
   `ChoiceGraph` DAG using namespace grouping, Jaccard clustering, or
   alphabetical fallback.
3. **Router** — beam-search over the graph to find the top-k items most
   relevant to a user query.
4. **ChoiceCards** — render compact, LLM-friendly cards for the selected
   items (never includes full schemas).

## Data stores

All stores are protocol-based with in-memory defaults:

- **EventLog** — append-only log of `ContextItem` events.
- **ArtifactStore** — blob storage for raw tool outputs intercepted by
  the firewall.
- **EpisodicStore** — short episodic memory entries (keyed by episode ID).
- **FactStore** — key-value fact entries persisted across turns.
- **StoreBundle** — convenience wrapper grouping all four stores.

## Progressive disclosure

`context/views.py` provides a `ViewRegistry` that maps content-type patterns
to view generators. When the firewall stores a large tool output as an artifact,
the view system generates alternative representations (JSON subset, CSV summary,
etc.) the agent can drilldown into without retrieving the full blob.
`drilldown_tool_spec()` exposes drilldown as an agent-callable tool.

## Design principles

- **Minimal core dependencies** — a small, audited set (`tiktoken`, `PyYAML`, `rank-bm25`, `mcp`, `jsonschema`, `typer`, `rich`); Python ≥ 3.10.
- **Deterministic** — tie-break by ID, sorted keys.
- **Protocol-based** — all store and estimator interfaces are
  `typing.Protocol`, allowing custom implementations.
- **Async-first** — the Context Engine exposes `build()` (async) with a
  `build_sync()` wrapper for synchronous callers.
- **Budget-aware** — every build is constrained by the phase-specific
  token budget; `BuildStats` explains what was kept and what was dropped.

---

<!-- FILE: docs/concepts.md -->

# Concepts

This document explains the core concepts in contextweaver.

## Context Item

A `ContextItem` is the atomic unit of the event log. Every user turn,
agent message, tool call, tool result, documentation snippet, memory
fact, plan state, or policy rule is represented as a `ContextItem`.

Key fields:

| Field | Description |
|---|---|
| `id` | Unique identifier |
| `kind` | One of the `ItemKind` enum values |
| `text` | The textual content |
| `parent_id` | Optional link to a parent item (e.g. tool_result → tool_call) |
| `token_estimate` | Pre-computed token count (optional) |
| `sensitivity` | Data sensitivity level (`public`, `internal`, `confidential`, `restricted`) |
| `metadata` | Arbitrary key-value metadata |

## Phases

contextweaver organises agent execution into four phases, each with its
own token budget:

- **route** — selecting which tool(s) to call.
- **call** — preparing tool call arguments.
- **interpret** — understanding tool results.
- **answer** — composing the final response to the user.

The `ContextBudget` dataclass defines the token limit for each phase.
Different phases emphasise different item kinds — for example, the
`answer` phase prioritises user turns and tool results, while the
`route` phase prioritises tool descriptions.

### User Query vs Routing Query
Production agents may transform raw user input into a routing query prior to retrieval or context selection. Routing queries remove conversational context so retrievers can match tools and context more precisely.

`Router.route(query)` takes a query that is formatted like a routing query, not the actual user input query.
ContextWeaver does not mandate what happens during the query transformation step. Different applications might adopt different strategies, such as LLM rewriting, classification, templates, or custom middleware.
Check out the MCP context gateway example [MCP gateway example](../examples/mcp_gateway_demo.py) for a worked example.

## Selectable Item (ToolCard)

A `SelectableItem` is the unified representation of anything the Routing
Engine can select — a tool, agent, skill, or internal function. The type
alias `ToolCard` (historically used when emphasising the LLM-facing card
framing) is **deprecated** — use `SelectableItem`; it is scheduled for removal
in 1.0 (see [Upgrading](upgrading.md)).

Key fields: `id`, `kind`, `name`, `description`, `tags`, `namespace`,
`side_effects`, `cost_hint`.

## Context Firewall

The context firewall intercepts `tool_result` items before raw output
reaches the prompt. It stores the raw output in the `ArtifactStore`,
replaces the prompt-facing text with a compact summary, and prevents
large tool outputs from consuming the entire token budget. In practice:

1. Stores the raw output in the `ArtifactStore`.
2. Generates a compact summary using the `Summarizer`.
3. Extracts structured facts into the `ResultEnvelope`.
4. Replaces the original item text with a summary + artifact reference.

## Result Envelope

A `ResultEnvelope` captures the processed output of a tool call:

- `summary` — compact text summary of the result.
- `facts` — list of extracted factual statements.
- `artifacts` — list of `ArtifactRef` handles for raw data.
- `views` — optional alternative representations.
- `status` — success / error / partial.

## Sensitivity Enforcement

Each `ContextItem` has a `sensitivity` field (default: `public`) that
classifies its data sensitivity level. The `ContextPolicy.sensitivity_floor`
setting (default: `confidential`) determines which items are subject to
filtering during context compilation.

Items whose sensitivity level meets or exceeds the floor are either:

- **Dropped** (`sensitivity_action="drop"`, the default) — removed from
  the candidate list before scoring or rendering.
- **Redacted** (`sensitivity_action="redact"`) — text replaced with
  `[REDACTED: {sensitivity}]` via the `MaskRedactionHook`, while
  preserving all item metadata.

Dropped items are recorded in `BuildStats.dropped_reasons["sensitivity"]`.

## Build Stats

Every context build produces a `BuildStats` object that explains exactly
what happened:

- How many candidates were generated.
- How many were included, dropped, or deduplicated.
- Token usage per section.
- Which items were dropped and why (`dropped_items` carries item ID + reason).
- Dependency closures applied.

`total_candidates` is measured after dependency closure and before sensitivity
filtering. Every later exclusion is counted, so completed builds satisfy
`included_count + dropped_count == total_candidates`.

## Choice Graph

The `ChoiceGraph` is a bounded DAG used by the Routing Engine. Interior
nodes are labelled navigation points; leaf nodes are items from the
catalog. The `TreeBuilder` constructs the graph using one of three
strategies:

1. **Namespace grouping** — items sharing a namespace prefix are grouped.
2. **Jaccard clustering** — farthest-first seeding + nearest assignment
   based on text similarity.
3. **Alphabetical fallback** — sorted by name, split into labelled
   chunks.

The `Router` performs beam search over this graph, scoring each path
to find the top-k most relevant items for a given query.

## Choice Cards

A `ChoiceCard` is the LLM-friendly representation of a routing result.
It contains the item name, description, relevance score, and optional
side-effect warning — but **never** the full argument schema. This keeps
the LLM's context focused on *which* tool to use, not *how* to call it.

## Episodic Memory & Facts

contextweaver supports two forms of persistent memory:

- **EpisodicStore** — stores short summaries of past interactions,
  keyed by episode ID. These are injected into the prompt header.
- **FactStore** — stores key-value pairs (e.g. `user_timezone=UTC`).
  Facts are injected into the prompt alongside episodic memory.

Both are capped in the prompt to prevent memory from crowding out the
current conversation.

## View Registry

The `ViewRegistry` (in `context/views.py`) maps content-type patterns to view
generators. When the context firewall stores a large tool output as an artifact,
the view system can generate alternative representations — a JSON subset, a CSV
summary, or a column listing — that the agent can request via drilldown without
retrieving the full blob. This progressive-disclosure mechanism keeps the
context window focused while preserving access to the raw data.

## Hydration Result

A `HydrationResult` (in `envelope.py`) captures the output of hydrating a
tool call with context. Hydration enriches a tool call's arguments or
description with context-aware information before execution, and the
`HydrationResult` carries both the enriched payload and metadata about
what context was used.

---

<!-- FILE: docs/quickstart.md -->

# 10-Minute Quickstart

This guide gets you to a working context build, a firewall-protected tool result,
and a routed tool shortlist in under 10 minutes.

> **Not sure which path to take?** Run `contextweaver start` after installation,
> or try `uvx contextweaver start` without installing. The deterministic wizard
> asks one deployment-intent question and only prints guidance.
>
> **In a hurry?** Run the built-in demo in an isolated environment with no
> persistent install:
>
> ```bash
> uvx contextweaver demo --scenario killer
> ```
>
> After `pip install contextweaver`, the installed CLI exposes the same
> scenarios:
>
> ```bash
> contextweaver demo                                  # friendly walkthrough
> contextweaver demo --scenario large-catalog         # 1,000 tools → compact cards
> contextweaver demo --scenario huge-tool-output      # context firewall on a big tool result
> contextweaver demo --scenario mcp-gateway           # MCP gateway meta-tools end-to-end
> ```
>
> Each scenario is deterministic and network-free. Run `contextweaver demo
> --help` to see the full list.

Time budget:

- Prerequisites: 30 seconds
- Install: 30 seconds
- Your first context build: 3 minutes
- Try the context firewall: 4 minutes
- Try tool routing: 2 minutes
- What to try next: 1 minute

## Choose a starting profile

`contextweaver start` asks one local question and prints an exact command sequence,
a configuration hint, and a verification checklist. It never executes commands,
writes files, or contacts the network.

| Profile | Start here when… | Non-interactive form |
|---|---|---|
| `gateway` | You want a bounded MCP surface in front of a tool catalog. | `contextweaver start --profile gateway` |
| `library` | You want budgeted context inside a custom Python loop. | `contextweaver start --profile library` |
| `routing` | You only need to shortlist a large tool catalog. | `contextweaver start --profile routing` |
| `integration` | You already have a provider or framework agent. | `contextweaver start --profile integration` |

The gateway profile covers deterministic static-catalog evaluation and points to the
[MCP client recipes](recipes/index.md) for real client wiring. For symptom-based
feature selection, use [Which pattern fits?](which_pattern.md).

## Adopting from an existing chat history (5-line drop-in)

> **Already have an OpenAI, Anthropic, or Gemini agent?** You don't need to
> walk the full quickstart — drop contextweaver in front of your existing
> message history with one call. The full walkthrough below is for new agents.

If you have an OpenAI Chat Completions session saved as JSON, you can build
a context pack in five lines (plus imports):

<!-- snippet: skip (illustrative; reads the reader's own session.json) -->
```python
import json
from contextweaver.adapters.openai_messages import from_openai_messages
from contextweaver.context.manager import ContextManager
from contextweaver.types import Phase

mgr = ContextManager()
from_openai_messages(json.load(open("session.json")), into=mgr)
pack = mgr.build_sync(phase=Phase.answer, query="What did we decide?")
print(pack.prompt)
```

The adapter handles every OpenAI Chat Completions role — `system`, `user`,
`assistant` (with optional `tool_calls`), `tool` — and threads
`tool_call_id` ↔ `ContextItem.id` so `to_openai_messages(...)` is the exact
inverse for round-tripping back into the OpenAI SDK.

Anthropic and Google Gemini have sibling adapters with the same shape:

<!-- snippet: skip (illustrative fragment; mirrors the OpenAI block above) -->
```python
from contextweaver.adapters.anthropic_messages import from_anthropic_messages
from contextweaver.adapters.gemini_contents import from_gemini_contents

# Anthropic Messages API: content blocks (text / tool_use / tool_result)
from_anthropic_messages(anthropic_messages, into=mgr)

# Google Gemini: contents[].parts[] (text / functionCall / functionResponse)
from_gemini_contents(gemini_contents, into=mgr)
```

All three adapters are pure stateless converters — they accept plain `dict`s
and never import a provider SDK at module load time. Session payloads may
contain sensitive prompt content; the adapters do not log message bodies
above `DEBUG` level.

> **Want to keep working with the provider SDK after building a pack?** Each
> adapter ships an inverse: `to_openai_messages`, `to_anthropic_messages`,
> `to_gemini_contents`. Round-trip equality holds for the representative
> fixtures in `tests/test_adapters_*.py`.

## 1. Prerequisites (30 seconds)

`contextweaver` requires Python 3.10 or newer.

Commands below use `python3`, matching the repository's documented default
(`docs/agent-context/workflows.md`). On Windows installations that expose Python
through the launcher, replace `python3` with `py -3`; if your environment
already exposes the intended interpreter as `python`, use that consistently.

Check your Python version:

```bash
python3 --version
```

Create and activate a virtual environment:

```bash
python3 -m venv .venv
```

Linux and macOS:

```bash
source .venv/bin/activate
```

Windows PowerShell:

```powershell
.venv\Scripts\Activate.ps1
```

If you see an error like `running scripts is disabled on this system`, either:

- run the activation script from **Command Prompt (cmd.exe)** instead:

  ```cmd
  .venv\Scripts\activate.bat
  ```

- or relax the execution policy for your current user in **PowerShell** (recommended only on machines you control):

  ```powershell
  Set-ExecutionPolicy -Scope CurrentUser -ExecutionPolicy RemoteSigned
  ```

## 2. Install (30 seconds)

For a zero-install trial:

```bash
uvx contextweaver demo --scenario killer
```

`uvx` creates an isolated temporary environment. Its first run may be slower
while dependencies resolve. Pin a release with
`uvx contextweaver@0.14.0 demo --scenario killer`.

Install from PyPI:

```bash
pip install contextweaver
```

If you are working from a repository checkout instead, install the package in editable mode:

```bash
pip install -e ".[dev]"
```

Verify that the package imports from the Python environment you just activated:

```bash
python3 -c "import contextweaver; print(contextweaver.__version__)"
```

Expected output is the installed version, for example `0.14.0`. Then run
the network-free demo as an end-to-end smoke test:

```bash
contextweaver demo
```

You should see the friendly walkthrough start without `ModuleNotFoundError` or
`contextweaver: command not found`.

For a quieter, faster confidence check (especially useful in CI or after
upgrading), run the built-in verify command:

```bash
contextweaver verify
```

This checks import path, ContextManager instantiation, a minimal context
build, token counting, and routing — all without network dependencies.
For machine-readable output: `contextweaver verify --json`.

If your network blocks `openaipublic.blob.core.windows.net`, the demo or
token-budget helpers may print `tiktoken cl100k_base encoding unavailable`.
That warning is harmless: contextweaver falls back to the deterministic
chars/4 estimator and keeps enforcing budgets. To suppress it while keeping
exact `tiktoken` counts, pre-warm a `TIKTOKEN_CACHE_DIR` on a connected
machine and copy that cache into the offline environment; the
[troubleshooting guide](troubleshooting.md#offline-air-gapped-tiktoken-warning)
has the full workflow.

## 3. Your First Context Build (3 minutes)

Scenario: an agent receives a question, decides to query a database, and builds
an answer-phase prompt from the conversation history.

Save this as `first_agent.py`:

```python
"""Your first contextweaver context build."""

from contextweaver.context.manager import ContextManager
from contextweaver.types import ContextItem, ItemKind, Phase

mgr = ContextManager()
mgr.ingest(ContextItem(id="u1", kind=ItemKind.user_turn, text="How many active users do we have?"))
mgr.ingest(ContextItem(id="a1", kind=ItemKind.agent_msg, text="I'll check the database for you."))
mgr.ingest(
    ContextItem(
        id="tc1",
        kind=ItemKind.tool_call,
        text='db_query(sql="SELECT COUNT(*) FROM users WHERE active=true")',
        parent_id="u1",
    )
)
mgr.ingest(ContextItem(id="tr1", kind=ItemKind.tool_result, text="count: 1042", parent_id="tc1"))

pack = mgr.build_sync(phase=Phase.answer, query="active user count")

print("=== Compiled Context ===")
print(pack.prompt)
print("\n=== Build Stats ===")
print(f"Total candidates: {pack.stats.total_candidates}")
print(f"Included in prompt: {pack.stats.included_count}")
print(f"Dropped: {pack.stats.dropped_count}")
print(f"Deduplicated: {pack.stats.dedup_removed}")
```

Run it:

```bash
python3 first_agent.py
```

Expected output excerpt:

```text
=== Compiled Context ===
[TOOL RESULT [artifact:tr1]]
count: 1042

[TOOL CALL]
db_query(sql="SELECT COUNT(*) FROM users WHERE active=true")

[USER]
How many active users do we have?

[ASSISTANT]
I'll check the database for you.

=== Build Stats ===
Total candidates: 4
Included in prompt: 4
Dropped: 0
Deduplicated: 0
```

What just happened:

- You ingested four events into the event log.
- `build_sync()` ran the context pipeline for the `answer` phase.
- The prompt was compiled from the most relevant items and returned with build stats.

### Sensitivity & Default Drops

Every `ContextItem` has a `sensitivity` level. It defaults to `public`, and
the default policy is conservative: items at or above
`Sensitivity.confidential` are dropped before scoring and rendering.

```python
from contextweaver.config import ContextPolicy
from contextweaver.context.manager import ContextManager
from contextweaver.types import ContextItem, ItemKind, Phase, Sensitivity

mgr = ContextManager()
mgr.ingest(ContextItem(id="u1", kind=ItemKind.user_turn, text="public"))
mgr.ingest(
    ContextItem(
        id="c1",
        kind=ItemKind.user_turn,
        text="confidential",
        sensitivity=Sensitivity.confidential,
    )
)

pack = mgr.build_sync(phase=Phase.answer, query="any")
print(pack.stats.dropped_reasons.get("sensitivity", 0))  # 1
```

If you want to keep `confidential` items but still drop `restricted` items,
raise the floor:

```python
policy = ContextPolicy(sensitivity_floor=Sensitivity.restricted)
mgr = ContextManager(policy=policy)
```

If you want sensitive items to remain visible only as masks, use redact mode:

```python
policy = ContextPolicy(sensitivity_action="redact")
mgr = ContextManager(policy=policy)
```

## 4. Try the Context Firewall (4 minutes)

Problem: a large tool result can dominate the prompt if you include it verbatim.

Save this as `firewall_demo.py`:

```python
"""Show how the context firewall keeps prompts compact."""

from contextweaver.context.manager import ContextManager
from contextweaver.types import ContextItem, ItemKind, Phase

large_result = '{"users": [' + ', '.join(
    [
        f'{{"id": {i}, "name": "User{i}", "email": "user{i}@example.com"}}'
        for i in range(1, 101)
    ]
) + ']}'

mgr = ContextManager()
mgr.ingest(ContextItem(id="u1", kind=ItemKind.user_turn, text="List all users"))
mgr.ingest(ContextItem(id="tc1", kind=ItemKind.tool_call, text="list_users()", parent_id="u1"))
mgr.ingest(ContextItem(id="tr1", kind=ItemKind.tool_result, text=large_result, parent_id="tc1"))

pack = mgr.build_sync(phase=Phase.answer, query="user list")

print(f"Raw tool result size: {len(large_result)} chars")
print("\n=== Compiled Context ===")
print(pack.prompt)
print("\n=== Firewall Impact ===")
print(f"Prompt size after firewall: {len(pack.prompt)} chars")
print(f"Artifacts stored: {len(mgr.artifact_store.list_refs())}")
```

Run it:

```bash
python3 firewall_demo.py
```

Expected output excerpt:

```text
Raw tool result size: 6087 chars

=== Compiled Context ===
[USER]
List all users

[TOOL RESULT [artifact:tr1]]
{"users": [{"id": 1, "name": "User1", "email": "user1@example.com"}, ...

[TOOL CALL]
list_users()

=== Firewall Impact ===
Prompt size after firewall: ... chars
Artifacts stored: 1
```

What just happened:

- The tool result was processed by the firewall during context build (all `tool_result` items go through it by default).
- `contextweaver` stored the raw result in the artifact store.
- The prompt kept only a compact summary plus an artifact reference instead of the full payload.

## 5. Try Tool Routing (2 minutes)

Problem: when a catalog grows, the model should only see the most relevant tools.

Save this as `routing_demo.py`:

```python
"""Route a natural-language request to a focused tool shortlist."""

from contextweaver.routing.catalog import Catalog
from contextweaver.routing.router import Router
from contextweaver.routing.tree import TreeBuilder
from contextweaver.types import SelectableItem

catalog = Catalog()
catalog.register(SelectableItem(id="t1", kind="tool", name="send_email", description="Send email to a recipient", tags=["notify", "team", "message"]))
catalog.register(SelectableItem(id="t2", kind="tool", name="db_query", description="Query the database", tags=["data"]))
catalog.register(SelectableItem(id="t3", kind="tool", name="create_ticket", description="Create support ticket", tags=["support"]))
catalog.register(SelectableItem(id="t4", kind="tool", name="send_sms", description="Send SMS message", tags=["notify", "team", "message"]))
catalog.register(SelectableItem(id="t5", kind="tool", name="schedule_meeting", description="Schedule a calendar meeting", tags=["calendar"]))

graph = TreeBuilder(max_children=3).build(catalog.all())
router = Router(graph, items=catalog.all(), beam_width=2, top_k=2)
result = router.route("notify the team about the deadline")

print("=== Query ===")
print("notify the team about the deadline")
print("\n=== Top Tools ===")
for item_id in result.candidate_ids:
    item = catalog.get(item_id)
    print(f"- {item.name}: {item.description}")
```

Run it:

```bash
python3 routing_demo.py
```

Expected output:

```text
=== Query ===
notify the team about the deadline

=== Top Tools ===
- send_sms: Send SMS message
- send_email: Send email to a recipient
```

What just happened:

- `TreeBuilder` organized the catalog into a bounded routing graph.
- `Router` scored the query against that graph and returned the top two candidates.
- Your model would now see a focused shortlist instead of the full catalog.

## 6. What to Try Next (1 minute)

Available now:

- [README](../README.md) for the top-level package overview
- [Concepts](concepts.md) for phases, the context firewall, and routing terms
- [Architecture](architecture.md) for the pipeline stages and module layout
- [MCP Integration](integration_mcp.md) for MCP adapters and session ingestion
- [A2A Integration](integration_a2a.md) for multi-agent adapter flows
- [Examples directory](../examples/) for larger end-to-end demos

Planned separately:

- Framework-specific integration guides are tracked in separate issues and are not part of this quickstart.

If you want a deeper local smoke test after this guide, run:

```bash
python3 -m contextweaver demo
```

---

<!-- FILE: docs/daily_driver.md -->

# Daily Driver Guide

Use contextweaver as a pressure-relief layer for tool-heavy sessions, not as
the default path for every chat message.

```text
User / IDE chat
  |
  +-- trivial question, tiny tool set, small result --> normal host-agent path
  |
  +-- large catalog, large result, long history ----> contextweaver gateway
                                                        |
                                                        +-- tool_browse
                                                        +-- tool_execute
                                                        +-- tool_view (only as needed)
```

contextweaver prepares bounded tool choices and compact result summaries. The
host application still owns the model call, authorization, user approval, and
execution policy. Upstream MCP servers remain the executors of record.

## Recommended daily loop

1. Start with the host client's normal chat path.
2. Use the gateway when the catalog is difficult to navigate, a result is too
   large for the prompt, or the active history needs deterministic budgeting.
3. Ask the client to call `tool_browse` with a routing-oriented query.
4. Execute only the selected `tool_id` through `tool_execute`.
5. Use `tool_view` for a narrow slice only when the summary is insufficient.
6. Inspect the route explanation, build statistics, and artifact reference
   before increasing budgets or exposing more data.

The gateway should usually replace duplicate direct registrations of the same
upstream tools. Advertising both the raw servers and the gateway gives the
model two competing paths and defeats the bounded-tool benefit.

## Start the gateway

The fastest trial requires no persistent installation:

```bash
uvx contextweaver mcp serve \
  --config examples/recipes/gateway_config.yaml \
  --dry-run
```

For regular use, install the CLI once:

```bash
pip install contextweaver
contextweaver mcp serve --config /path/to/gateway.yaml --dry-run
```

Enable local, payload-safe diagnostics in a directory that already exists:

```bash
contextweaver mcp serve \
  --config /path/to/gateway.yaml \
  --diagnostics /path/to/logs/contextweaver.jsonl \
  --quiet
```

Inspect the static catalog before launch and aggregate the event stream later:

```bash
contextweaver mcp inspect --catalog /path/to/catalog.yaml
contextweaver mcp stats --events /path/to/logs/contextweaver.jsonl
```

For support or incident triage, create a bounded local bundle:

```bash
contextweaver mcp incident-pack \
  --config /path/to/gateway.yaml \
  --diagnostics /path/to/logs/contextweaver.jsonl \
  --out /path/to/contextweaver-incident.zip
```

The pack includes a machine-readable manifest, environment summary, redacted
config/catalog excerpts, diagnostics summaries, redacted diagnostics, and a
reproduction checklist. It never reads shell history automatically; pass
`--command-log /path/to/commands.txt` only when you captured a command log
explicitly for that incident.

With a `catalog:` config the CLI represents one static catalog source. Its
catalog report groups tools by namespace; it does not claim live health for
multiple upstream MCP processes.

`pipx run contextweaver ...` is another isolated option. The first `uvx` or
`pipx run` launch resolves a temporary environment and is slower than later
runs. Pin a deployment when reproducibility matters:

```bash
uvx contextweaver@0.14.0 mcp serve --config /path/to/gateway.yaml
pipx run --spec contextweaver==0.14.0 contextweaver mcp serve \
  --config /path/to/gateway.yaml
```

`mcp serve` runs in two modes: a `catalog:` config loads a static catalog with
the deterministic stub upstream (local exercise, CI, offline routing work),
while an `upstreams:` config launches real upstream MCP servers behind the
gateway and executes selected calls live. An existing multi-server client
config migrates with `contextweaver mcp import-vscode <config> --apply`; see
[MCP Integration](integration_mcp.md#connecting-to-real-upstream-mcp-servers)
for the `upstreams:` / `startup:` reference. Live upstream serving covers
tools only — resources and prompts still use the static-catalog path.

## Client instruction

Give the host agent a short operational rule. The same rule works in Cursor,
Claude Desktop, Claude Code, VS Code Copilot agent mode, and generic MCP
clients:

```text
Use the contextweaver MCP gateway when you need to browse or call tools from a
large catalog. Call tool_browse first with a routing-oriented query, execute
only the selected tool_id through tool_execute, and use tool_view only when the
summary is insufficient. Prefer narrow tool_view selectors. The gateway does
not grant authorization; follow the host application's approval and execution
policy.
```

Client-specific placement:

| Client | Where to put the rule |
|---|---|
| Cursor | Project rules or the task prompt |
| Claude Desktop | Project/custom instructions |
| Claude Code | `CLAUDE.md` or the current task prompt |
| GitHub Copilot | `.github/copilot-instructions.md` or repository instructions |
| Generic MCP client | System/developer prompt owned by the host application |

## Use contextweaver when

- The client sees dozens or hundreds of MCP, FastMCP, or Python tools.
- Tool results include large JSON objects, logs, tables, CSV, resources, or
  binary content.
- Multi-turn tool sessions accumulate more history than should reach every
  phase.
- You need deterministic prompt budgets and an inspectable record of what was
  included, dropped, or deduplicated.
- You want schemas hidden until a tool has been selected and hydrated.

For catalogs above roughly 300 tools, treat metadata quality as part of the
deployment: capture/import the upstream `tools/list`, normalize names and
descriptions, then validate routing against representative queries. The
current static-catalog workflow is documented in the
[MCP Context Gateway architecture](architectures/mcp_context_gateway.md).

## Do not use it when

- The agent has only three to five small tools.
- The interaction is one-shot Q&A with no tool or history pressure.
- Tool outputs are already small and the prompt comfortably fits its budget.
- The actual problem is pure retrieval, long-term memory, or observability.
- The host application has not defined who may invoke tools or approve side
  effects.
- You expect contextweaver to be an agent supervisor, model runtime, sandbox,
  or authorization service.

## Debug loop

When a route or prompt looks wrong, inspect these in order:

1. **Gateway configuration.** Confirm `mode`, `top_k`, `beam_width`,
   `cache_stable`, and the catalog path with `mcp serve --dry-run`.
2. **Route result.** Use `RouteResult.explanation()` for ranked candidates,
   score gaps, filters, and ambiguity. Use `debug=True` when you need the
   expansion trace.
3. **Build statistics.** Check `included_count`, `dropped_count`,
   `dropped_reasons`, per-item `dropped_items`, `dedup_removed`, and token
   usage in `BuildStats`. For an ingested session, run
   `contextweaver inspect --session session.json`.
4. **Artifact reference.** Confirm the handle exists before calling
   `tool_view`, then request a bounded `head`, `lines`, `rows`, or `json_keys`
   selector.
5. **Embedded runtime settings.** If you use `ContextManager` directly,
   inspect phase budgets, the firewall threshold, sensitivity policy, and
   scoring/retrieval backend. These are Python runtime settings, not fields in
   the current `mcp serve` YAML.
6. **Telemetry.** Use `mcp serve --diagnostics FILE` for local JSONL counts,
   savings, failures, artifact-view usage, and latency. The built-in stream
   records IDs, sizes, argument key names, and error codes, but not queries,
   argument values, result text, prompt text, or artifact bytes. When the
   `[otel]` extra is enabled, inspect context-build, firewall, and routing spans.

Do not respond to a poor route by immediately increasing every budget or
returning whole artifacts. Better descriptions, a sharper browse query, and a
narrow view usually preserve more of the gateway's benefit.

## Next steps

- [Claude Desktop recipe](recipes/claude_desktop.md)
- [Claude Code recipe](recipes/claude_code.md)
- [Cursor recipe](recipes/cursor.md)
- [GitHub Copilot recipe](recipes/github_copilot.md)
- [MCP Integration](integration_mcp.md)
- [Adopter Benchmark Report](benchmark_report.md)
- [Troubleshooting](troubleshooting.md)
- [Security Model](security_model.md)

---

<!-- FILE: docs/security_model.md -->

# MCP Gateway Security Model

contextweaver is a local context-compilation and MCP gateway layer. It does
not call an LLM or implement upstream tool side effects, but it can process
tool schemas, tool results, session history, and raw artifacts that contain
sensitive data. Treat its artifact store and diagnostics as sensitive as the
upstream outputs they summarize.

This page describes deployment boundaries. Vulnerability reporting remains in
the repository's
[`SECURITY.md`](https://github.com/dgenio/contextweaver/blob/main/SECURITY.md).

For a task-oriented walkthrough of running the gateway in front of powerful
MCP servers (secrets, destructive tools, least privilege), see
[MCP Gateway Security Guide](security_mcp_gateway.md). For configuring the
sensitivity/redaction subsystem, see the
[Sensitivity & Redaction guide](sensitivity.md).

## Default security posture

The serving entrypoints are **secure by default** (issue #744). `contextweaver
mcp serve` runs with the deterministic
[`HeuristicSensitivityClassifier`](sensitivity.md) and secret scrubbing
(`redact_secrets`) enabled, so unlabelled tool output that carries
credential-shaped or PII-shaped content is classified and scrubbed before it
reaches a prompt-bound summary or a `ChoiceCard`. This is a deliberate
divergence from the library-level `ContextManager`, whose defaults stay
permissive for embedding in existing pipelines — the hardening is applied at
the gateway boundary, where an operator reasonably expects the firewall's
headline protections to be on.

Turn the protections off with `contextweaver mcp serve --no-redact` (or
`redact: false` in the config file). Doing so prints a one-line startup
**warning** to stderr, so an unprotected posture is always a visible choice,
never a silent default.

The HTTP sidecar (`contextweaver serve-api`) is a thinner surface: its
`/v1/compact` endpoint **can** scrub secrets end-to-end (issue #745) — enable it
with `redact_secrets: true` in the sidecar config (server-side default) or
`"redact_secrets": true` per request; it is off by default (posture owned by
#744). Binding it without `--api-key` prints a startup warning; do not send
secret-bearing payloads to an unauthenticated sidecar over an untrusted network.

This decision is recorded here so it is not "simplified" away: the serving
defaults must stay secure-by-default, and any change that weakens them is a
security-relevant change requiring review.

## Runtime authorization: the policy gate

Reducing a large MCP surface to `tool_browse` / `tool_execute` / `tool_view`
moves the practical safety boundary into the gateway. contextweaver therefore
provides an explicit, deterministic **policy gate** evaluated before any
upstream dispatch and before any raw artifact egress (issue #373).

A `ToolPolicy` is an ordered list of match → action rules with a default
action. Each rule may match on the upstream `namespace`, a case-sensitive glob
over the tool id/name (`tool`), catalog `tags`, `read_only`, and the surface
(`meta_tool`: `tool_execute` or `tool_view`). The first matching rule wins;
otherwise the `default` applies. Actions are:

- `allow` — proceed normally.
- `deny` — return a `POLICY_DENIED` `GatewayError`; the upstream tool is
  **never** called and no raw content is returned.
- `require_approval` — return an `AUTH_REQUIRED` `GatewayError` (with
  `details.approval = "required"`) so a host or custom loop can surface it for
  human sign-off.

The default (`ToolPolicy()` / no policy configured) allows everything, so
existing deployments are unchanged; opt into `deny` / `require_approval` rules,
or set `default: deny` for an allowlist posture. Configure it under the
`policy` key of `mcp serve --config`:

```yaml
policy:
  default: allow
  rules:
    - { namespace: github, tool: "issues.*", action: allow }
    - { tags: [destructive], action: require_approval }
    - { tool: "*delete*", action: deny }
    - { meta_tool: tool_view, namespace: secrets, action: deny }  # block raw egress
```

MCP annotations such as `readOnlyHint` / `destructiveHint` are untrusted hints
and are **not** the enforcement mechanism — the policy is. Annotations may
inform a rule's authoring, but the gate decides.

## Data flow

```text
                         schemas / calls
MCP client  <------>  contextweaver gateway  <------>  upstream MCP server
    |                    |          |
    |                    |          +--> tool execution and authorization
    |                    |               remain upstream / host concerns
    |                    |
    |                    +--> artifact store
    |                         raw tool bytes, local and out-of-band by default
    |
    +--> model provider
         bounded ChoiceCards, summaries, selected artifact slices
```

Prompt-visible by default:

- Compact `ChoiceCard` fields, without full input schemas.
- Firewalled summaries and extracted facts.
- Artifact handles and metadata needed for progressive disclosure.
- A selected slice returned by `tool_view` after the client requests it.

Out-of-band by default:

- Raw text, resource, image, and audio bytes stored as artifacts.
- Full schemas until the selected tool is hydrated for execution.
- Context items excluded by budget or sensitivity policy.

Out-of-band does not mean harmless or encrypted. The default gateway uses an
in-memory artifact store in the gateway process. Other adapters can use
filesystem or caller-supplied stores.

## Data contextweaver can touch

- MCP tool names, descriptions, annotations, and input schemas.
- MCP text, structured content, resources, images, and audio handled by the
  adapter.
- Session messages and tool history ingested through provider adapters.
- Artifact bytes, labels, media types, sizes, and handles.
- Catalog and gateway configuration files.
- Routing queries, candidate identifiers, scores, and build statistics.
- OpenTelemetry attributes and metrics when the optional integration is
  enabled.
- Local gateway JSONL diagnostics when `mcp serve --diagnostics FILE` is
  enabled.

MCP annotations such as `readOnlyHint` are untrusted metadata. They may improve
presentation or routing, but must not grant permission or bypass approval.

## Network and egress

The context and routing algorithms do not make model calls and do not require
network access. Data can still leave the machine through surrounding systems:

- The MCP client sends the compiled prompt and tool responses to its configured
  model provider.
- Upstream MCP servers may call remote APIs or databases.
- An enabled OpenTelemetry exporter sends spans and metrics to its configured
  endpoint.
- Package installation requires a package index unless artifacts are already
  cached or installed.
- An explicitly selected token estimator may fetch tokenizer data on first use
  when its cache is cold; the documented fallback remains deterministic.

The default OpenTelemetry emission excludes raw queries, full tool
descriptions, schemas, and prompt content. Enabling
`otel_emit_experimental=True` can add sensitive content and should be limited
to a trusted, access-controlled backend.

The built-in gateway JSONL stream excludes query text, argument values, result
text, prompt text, and artifact bytes. It does include canonical tool and
artifact handles, namespaces, argument key names, sizes, timings, and error
codes. Treat those identifiers as operationally sensitive and restrict file
permissions and retention accordingly.

## Trust boundaries

### Host MCP client

The host decides which model receives context, which MCP server is available,
and whether a user must approve a call. contextweaver does not replace those
controls.

### contextweaver gateway

The gateway narrows discovery, validates selected arguments against the
hydrated schema, dispatches calls to an `UpstreamCall`, and firewalls returned
content. Routing is relevance selection, not authorization.

### Upstream MCP server

The upstream server implements the operation and its side effects. It must
authenticate callers, authorize access, validate business rules, and protect
its own credentials. contextweaver does not verify that a tool described as
read-only is actually read-only.

### Artifact store

The artifact store contains the bytes deliberately kept out of the prompt.
Anyone with process or storage access may be able to read them. A handle is an
address, not a capability token.

## Context firewall limits

The context firewall reduces prompt exposure and token use. It is not a data
loss prevention system or security sandbox.

- Raw bytes can remain in the artifact store after a summary is rendered.
- Summaries and extracted facts can still contain sensitive values unless a
  sensitivity or redaction policy removes them.
- `tool_view` deliberately re-exposes selected artifact content to the MCP
  client and therefore potentially to the model provider. It is the intentional
  raw-recovery surface, governed by the same `ToolPolicy` as `tool_execute`
  (issue #746): a `meta_tool: tool_view` rule can `deny` or `require_approval`
  raw egress per namespace/tool. Note that gateway artifacts are stored
  unredacted **at rest** (the scrubbing applies to prompt-bound summaries and
  cards, not the raw bytes), so the policy gate — not the firewall — is what
  bounds raw egress; attribution is best-effort for handles that do not encode
  a tool id.
- The current in-memory gateway store has no TTL, total-size quota, or
  per-handle authorization policy.
- Current selectors accept caller-provided ranges; deployments should use
  narrow selectors and should not assume a built-in maximum response size.
- The gateway does not neutralize prompt injection contained in a tool result.
  The host prompt and execution policy must treat tool content as untrusted.

A policy gate for `tool_view` egress now exists (issue #746, above). Artifact
TTLs, size limits, bounded-view selectors, store-time redaction, and
provenance remain tracked in
[#375](https://github.com/dgenio/contextweaver/issues/375). Until those ship,
high-sensitivity deployments should deny raw `tool_view` via policy (or in
their host integration) rather than relying on artifacts being scrubbed at
rest — they are not.

## Non-goals

contextweaver does not:

- Authenticate users or authorize tool execution.
- Call the model or control how it follows instructions.
- Sandbox or attest upstream MCP server processes.
- Guarantee that MCP annotations are truthful.
- Replace secret scanning, DLP, endpoint security, or storage encryption.
- Prevent a malicious or compromised upstream from returning prompt-injection
  content.
- Make an unsafe tool safe merely because it was selected by the router.

## Hardening checklist

- Use upstream MCP servers you trust and keep them patched.
- Register a tool either directly or behind the gateway, not both.
- Enforce user identity, authorization, and side-effect approval in the host or
  upstream server.
- Keep gateway configs, local state, and artifact directories out of version
  control when they contain machine paths or credentials.
- Put secrets in the client's environment/secret facility, not in committed
  JSON or YAML.
- Use sensitivity labels and redaction before prompt rendering; add
  store-before-view redaction when handling regulated data.
- Prefer `json_keys`, short line ranges, or a small `head` selector over whole
  artifact retrieval.
- Restrict filesystem permissions and process access around persistent
  artifact stores.
- Leave experimental OTel content emission disabled unless the exporter is
  trusted for the data class involved.
- Store gateway JSONL diagnostics in an access-controlled path and define a
  retention policy; use `--quiet` only to suppress lifecycle stderr, not as a
  substitute for diagnostics access control.
- Review tool descriptions and results as untrusted input; do not rely on
  `readOnlyHint`, `destructiveHint`, or similar annotations as policy.
- Pin the contextweaver version in managed deployments and review release notes
  before upgrading.

## Deployment questions

Before exposing a gateway to users, answer:

1. Which identities may use each upstream tool?
2. Where are raw artifacts stored, and who can read that location or process?
3. How long may artifacts remain available?
4. Which content is allowed to reach the model provider?
5. Who approves destructive or externally visible calls?
6. Where do OTel spans go, and can that backend hold the same data class?
7. What is the incident path if a summary or view exposes a secret?

If these answers are undefined, the deployment is not made safe by adding a
context firewall.

---

<!-- FILE: docs/security_mcp_gateway.md -->

# MCP Gateway Security Guide

A local MCP gateway launches upstream servers that may receive secrets through
environment variables and expose tools that read files, modify repositories,
call APIs, or delete resources. This guide is the task-oriented companion to the
[Security Model](security_model.md): how to run contextweaver in front of
powerful MCP servers with least privilege.

For the conceptual trust boundaries and non-goals, read the
[Security Model](security_model.md) first. For tuning the sensitivity/redaction
subsystem, see the [Sensitivity & Redaction guide](sensitivity.md).

## Threats this addresses

- Secrets passed to upstream servers via environment variables.
- Destructive tools exposed to the model by default.
- Workspace filesystem access broader than intended.
- Confusion between read-only and write-capable tools.
- Raw tool output (with embedded secrets) reachable via `tool_view`.
- Secrets leaking into diagnostics or error messages.

## Start secure by default

`contextweaver mcp serve` is secure by default (issue #744): it classifies and
scrubs secret/PII-shaped content in tool output before it reaches the prompt.
Keep it that way — only pass `--no-redact` when you have a specific reason, and
note it prints a startup warning when you do.

```bash
contextweaver mcp serve --config gateway.yaml   # secure by default
```

## Least privilege: deny destructive tools before they can be called

The runtime [policy gate](security_model.md#runtime-authorization-the-policy-gate)
is the enforcement point. Prefer a **default-deny allowlist** for untrusted or
high-blast-radius estates, and require approval for destructive operations:

```yaml
# gateway.yaml
catalog: ./catalog.json
redact: true          # explicit; this is also the default
policy:
  default: deny       # allowlist posture — nothing runs unless matched
  rules:
    - { namespace: github, tool: "issues.*", action: allow }
    - { namespace: github, tool: "pull_requests.*", action: allow }
    - { namespace: filesystem, tool: "read_*", action: allow }
    - { tags: [destructive], action: require_approval }
    - { tool: "*delete*", action: deny }
    - { meta_tool: tool_view, namespace: secrets, action: deny }
```

`deny` means the upstream tool is never invoked; `require_approval` returns an
`AUTH_REQUIRED` error a host can surface for human sign-off. This is enforced by
contextweaver regardless of whether the upstream server implements its own
controls — do not rely on MCP `readOnlyHint`/`destructiveHint` annotations as
policy; they are untrusted hints.

## Policy presets: a faster starting point

Hand-writing a `policy` / `retry` / `rate_limits` / `cache` block from scratch
is a lot of trial-and-error for a first deployment. `mcp serve --policy-preset
<name>` (or the `policy_preset` config key) selects a named `GatewayPreset`
(issue #664) bundling all four:

| Preset       | Authorization                                                              | Retry              | Rate limit (`tool_execute`) | Cache                          |
| ------------ | --------------------------------------------------------------------------- | ------------------ | ---------------------------- | ------------------------------- |
| `safe`       | **Every** `tool_execute` call requires approval — does not rely on the (unverified) `read_only` hint | 2 attempts          | 30/min                       | off                              |
| `balanced`   | Allow-all                                                                  | 3 attempts          | 120/min                      | off                              |
| `throughput` | Allow-all                                                                  | 5 attempts, jittered | none                        | read-only, no allow-list         |

Selecting no preset is inert — behaviour is unchanged. An explicit
`policy` / `retry` / `rate_limits` / `cache` config block still wins over the
preset **for that block** (block-level override, not a field-by-field merge):
start from `safe` and override just `retry` while keeping the preset's
approval-gated policy and quota:

```yaml
# gateway.yaml
catalog: ./catalog.json
policy_preset: safe
retry:
  max_attempts: 5   # overrides the preset's retry block only
```

⚠️ `throughput`'s cache trusts every upstream's self-declared `read_only`
hint with no `allow` list (same caveat as [Bound raw egress](#bound-raw-egress)
below) — do not point it at untrusted or partially-trusted upstreams. Pair
caching with an explicit `cache.allow` list, or use `safe`/`balanced` instead,
for safety-critical estates.

Preview the resolved policy before serving — no catalog required:

```bash
contextweaver mcp serve --policy-preset safe --print-effective-policy
```

## Keep secrets out of committed config

Upstream servers often need tokens (e.g. a GitHub PAT). Never commit them.
Reference environment variables from the client/launcher and keep the values in
your shell or a secret manager:

```jsonc
// .vscode/mcp.json — the token lives in the environment, not the file
{
  "servers": {
    "github": {
      "type": "stdio",
      "command": "docker",
      "args": ["run", "-i", "--rm", "-e", "GITHUB_PERSONAL_ACCESS_TOKEN", "ghcr.io/github/github-mcp-server"]
    }
  }
}
```

Add gateway configs, persistent state directories (`--state-dir`), and artifact
directories to `.gitignore` when they can contain machine paths or credentials.

## Bound raw egress

`tool_view` re-exposes raw artifact bytes and is the intentional
raw-recovery surface. Gateway artifacts are stored **unredacted at rest** — the
scrubbing applies to prompt-bound summaries and cards, not the raw bytes. For
sensitive estates, deny or approval-gate raw egress with a `meta_tool: tool_view`
policy rule (see the example above) rather than assuming artifacts are scrubbed.

## Diagnostics do not print secret values

The built-in JSONL diagnostics stream (`mcp serve --diagnostics FILE`) records
canonical tool/artifact handles, namespaces, argument *key names*, sizes,
timings, and error codes — **not** query text, argument values, result text, or
artifact bytes. Upstream exception text (which can carry hostnames, paths, or
tokens) is control-character-stripped and length-capped before it reaches
model-visible context; the full detail stays operator-side in logs. Still,
treat the diagnostics file as operationally sensitive: restrict its permissions
and define a retention policy. Use `--quiet` only to suppress lifecycle chatter,
not as an access-control substitute.

## The HTTP sidecar

`contextweaver serve-api` is a thinner surface. Its `/v1/compact` endpoint can
scrub secrets end-to-end (issue #745): set `redact_secrets: true` in the sidecar
config to force it on for every request, or let a client opt in per request with
`"redact_secrets": true` in the `/v1/compact` body. It is off by default (posture
owned by #744). An unauthenticated bind still exposes the surface to any local
caller — set `--api-key` (or `CONTEXTWEAVER_SIDECAR_API_KEY`), bind to a trusted
interface, and avoid sending secret-bearing payloads over an untrusted network.
Binding without a key prints a startup warning.

## Checklist

- [ ] Serve with redaction on (default); justify any `--no-redact`.
- [ ] Use a `default: deny` policy for untrusted estates; `require_approval` for
      destructive tools; `deny` `*delete*`-style tools.
- [ ] For a first deployment, start from a named preset (`--policy-preset`)
      and review its `--print-effective-policy` output before going live.
- [ ] Deny or approval-gate `tool_view` for sensitive namespaces.
- [ ] Keep tokens in the environment/secret manager, never in committed config.
- [ ] `.gitignore` gateway config, `--state-dir`, and artifact directories.
- [ ] Restrict permissions/retention on the diagnostics file.
- [ ] Authenticate the sidecar (`--api-key`) and keep it off untrusted networks.
- [ ] Remember: authorization and side-effect execution still ultimately rest
      with the upstream server and host — the gateway narrows and gates, it does
      not replace them.

---

<!-- FILE: docs/sensitivity.md -->

# Sensitivity & Redaction

Sensitivity enforcement is contextweaver's security-grade subsystem. It decides
which context items may reach a prompt and which are dropped or redacted first.
This page is the operator's configuration manual: the levels, the floor/action
knobs, how to write a redaction hook, how enforcement interacts with the
firewall and drilldown, how to verify it, and its limits.

> Misconfiguration here is a data-exposure risk. Every behavioural claim below
> is checked against the source; do not assume defaults are more permissive than
> stated.

## The four levels

`Sensitivity` is an ordered enum (lowest → highest):

| Level | Meaning (suggested) |
|---|---|
| `public` | Safe to send anywhere. **Default for unlabelled items.** |
| `internal` | Team-internal; not for external model providers without review. |
| `confidential` | Sensitive business/PII data. |
| `restricted` | Credential-shaped / regulated content; the tightest class. |

Map these onto your organisation's classification scheme; the ordering is what
enforcement relies on, not the names.

## Floor and action

Enforcement is driven by two `ContextPolicy` fields:

- `sensitivity_floor` (default **`confidential`**): items **at or above** the
  floor are enforced.
- `sensitivity_action` (default **`drop`**): what enforcement does — `"drop"`
  removes the item entirely; `"redact"` replaces its text via the configured
  redaction hooks and keeps a scrubbed placeholder.

```python
from contextweaver import ContextManager
from contextweaver.config import ContextPolicy
from contextweaver.types import Sensitivity

# Redact (not drop) anything confidential-or-higher, using the built-in mask hook.
policy = ContextPolicy(
    sensitivity_floor=Sensitivity.confidential,
    sensitivity_action="redact",
    redaction_hooks=["mask"],
)
manager = ContextManager(policy=policy)
```

| Goal | floor | action |
|---|---|---|
| Drop anything sensitive silently (default) | `confidential` | `drop` |
| Keep structure but scrub sensitive text | `confidential` | `redact` |
| Only ever drop credential-shaped content | `restricted` | `drop` |

The defaults are deliberately conservative. **Do not weaken them** without
review — see the sensitivity rule in the repo's agent guidance.

## Labelling: don't rely on defaults

Unlabelled items default to `public`, so enforcement never sees content the
caller forgot to classify. Two mechanisms raise labels before enforcement:

- **`HeuristicSensitivityClassifier`** (opt-in, deterministic) inspects item
  text and *raises* the label to `restricted` for credential-shaped content or
  `confidential` for PII-shaped markers (email/SSN/card). It can only raise,
  never lower. `contextweaver mcp serve` enables it by default (secure-by-default,
  issue #744); pass it to a library `ContextManager` explicitly:

  ```python
  from contextweaver.context.classify import HeuristicSensitivityClassifier
  manager = ContextManager(sensitivity_classifier=HeuristicSensitivityClassifier())
  ```

- **`redact_secrets=True`** runs a deterministic secret-scrubbing pass over
  firewall summaries and extracted facts before they reach the prompt.

## Redaction hooks

`sensitivity_action="redact"` applies the hooks named in
`ContextPolicy.redaction_hooks`, in order. Two are built in and registered at
import:

- `"mask"` — `MaskRedactionHook`: replaces the item's text with a masked
  placeholder and drops its `artifact_ref` so the rendered prompt cannot
  advertise a handle that drilldown could dereference back to the original.
- `"secret"` — `SecretRedactor`: substring-scrubs secret shapes from the text
  (complements, does not replace, the mask hook).

Register your own hook (it must implement the `RedactionHook` protocol —
`redact(item) -> ContextItem`):

```python
from dataclasses import replace

from contextweaver.context.sensitivity import register_redaction_hook
from contextweaver.types import ContextItem


class BlankRedactor:
    """Mirror MaskRedactionHook's contract: replace text, clear the artifact ref,
    and stamp metadata["redacted"] so the handle can't be dereferenced back."""

    def redact(self, item: ContextItem) -> ContextItem:
        metadata = dict(item.metadata)
        metadata["redacted"] = True
        return replace(item, text="[REDACTED]", artifact_ref=None, metadata=metadata)


register_redaction_hook("blank", BlankRedactor())
# then reference it: ContextPolicy(sensitivity_action="redact", redaction_hooks=["blank"])
```

## How it interacts with the rest of the pipeline

- **Filter runs before the firewall.** In the context pipeline, the sensitivity
  filter (stage 3) runs *before* `apply_firewall` (stage 4), so sensitive
  payloads are dropped/redacted before any summariser or extractor sees them.
- **Drilldown cannot launder content back in.** `ContextManager.drilldown`
  enforces the floor against the artifact's source item: recovering the raw
  bytes of a dropped/redacted item raises `PolicyViolationError` unless
  `ContextPolicy.allow_redacted_drilldown=True` (issue #451).
- **Gateway `tool_view` is governed by the policy gate, not this filter.**
  Gateway artifacts are stored unredacted at rest; bound raw egress with a
  `meta_tool: tool_view` rule in the [security model](security_model.md), not by
  assuming they are scrubbed.

## Verifying your configuration

The repo ships classification fixtures under `tests/fixtures/sensitivity/`
(`public`, `internal`, `confidential`, `restricted`, `pii_like`, `secret_like`).
Use them (or your own) to assert that your floor/action behave as intended, and
add a test that a known-sensitive payload never appears in a built pack's
rendered prompt.

## Limitations

- **Enforcement is label-dependent.** Without a classifier, an unlabelled
  sensitive item is treated as `public`. Enable
  `HeuristicSensitivityClassifier` (or label at ingest) for defence in depth.
- **No content inspection by default.** The classifier is pattern-based and
  deterministic — it is not a full DLP engine and will miss novel secret shapes.
- **`redact` keeps structure.** Redaction replaces text but the item still
  occupies a slot; use `drop` when the item must not appear at all.
- Sensitivity routing does not authenticate users or authorize tool execution —
  those remain host/upstream responsibilities (see
  [Security Model](security_model.md)).

---

<!-- FILE: docs/recipes/index.md -->

# MCP Client Recipes

These recipes put the installed `contextweaver mcp serve` command in front of
an MCP client. The default examples use `uvx`, so the client receives an
isolated current release without requiring a persistent Python environment.

| Client | Recipe | Shipped config |
|---|---|---|
| Claude Desktop | [Claude Desktop](claude_desktop.md) | [`claude_desktop_config.json`](https://github.com/dgenio/contextweaver/blob/main/examples/recipes/claude_desktop_config.json) |
| Claude Code | [Claude Code](claude_code.md) | [`claude_code_mcp.json`](https://github.com/dgenio/contextweaver/blob/main/examples/recipes/claude_code_mcp.json) |
| GitHub Copilot in VS Code | [GitHub Copilot](github_copilot.md) | [`copilot_mcp.json`](https://github.com/dgenio/contextweaver/blob/main/examples/recipes/copilot_mcp.json) |
| Cursor | [Cursor](cursor.md) | [`cursor_mcp.json`](https://github.com/dgenio/contextweaver/blob/main/examples/recipes/cursor_mcp.json) |

## Choose an invocation

Zero-install trial:

```bash
uvx contextweaver mcp serve --config /path/to/gateway.yaml --dry-run
```

Persistent installation:

```bash
pip install contextweaver
contextweaver mcp serve --config /path/to/gateway.yaml --dry-run
```

Isolated pipx run:

```bash
pipx run contextweaver mcp serve --config /path/to/gateway.yaml --dry-run
```

The first `uvx` or `pipx run` invocation resolves an environment and may take
longer. Pin the package in managed environments:

```bash
uvx contextweaver@0.14.0 mcp serve --config /path/to/gateway.yaml
pipx run --spec contextweaver==0.14.0 contextweaver mcp serve \
  --config /path/to/gateway.yaml
```

## Shared gateway config

The shipped [`gateway_config.yaml`](https://github.com/dgenio/contextweaver/blob/main/examples/recipes/gateway_config.yaml)
loads the committed 11-tool filesystem snapshot:

```yaml
catalog: ../architectures/mcp_context_gateway/real_catalogs/filesystem.json
mode: gateway
top_k: 10
beam_width: 3
cache_stable: false
name: contextweaver
```

Relative catalog paths are resolved from the gateway config file's directory.
This keeps project-scoped client configs portable even when the client starts
the server from a different working directory.

## Generate Multi-Client Config Packs

Generate every supported client config from one gateway source file:

```bash
contextweaver mcp generate-configs \
  --config examples/recipes/gateway_config.yaml \
  --out-dir ./generated-recipes
```

By default this emits:

- `copilot_mcp.json`
- `cursor_mcp.json`
- `claude_desktop_config.json`
- `claude_code_mcp.json`

Use `--target` repeatedly to generate only selected clients. The command
validates the gateway config before writing files, fails if outputs already
exist (unless `--force`), and prints target-specific compatibility warnings.

## What the client sees

```text
MCP client
    |
    +-- tool_browse  -> bounded ChoiceCards
    +-- tool_execute -> hydrated, validated selected call
    +-- tool_view    -> selected artifact slice
```

The client sees three meta-tools instead of every full upstream schema. Large
results become summaries plus artifact handles.

## Current runtime boundary

The packaged CLI loads a static JSON/YAML catalog and uses a deterministic
stub upstream handler. It is suitable for client wiring, tool shortlisting,
argument validation, and firewall-shape checks. Live upstream execution
requires a Python composition using `McpClientUpstream` or
`MultiplexUpstream`; see
[MCP Integration](../integration_mcp.md#connecting-to-real-upstream-mcp-servers).

[`examples/recipes/serve_gateway.py`](https://github.com/dgenio/contextweaver/blob/main/examples/recipes/serve_gateway.py)
remains a legacy/development example for custom `ProxyRuntime` wiring. It is
no longer the default client entry point.

## Large catalogs

For 300+ tools, capture/import the upstream `tools/list`, normalize weak names
and descriptions, and test representative `tool_browse` queries before
deployment. The static snapshot workflow and real catalog fixtures are in the
[MCP Context Gateway architecture](../architectures/mcp_context_gateway.md).

## Next reading

- [Daily Driver Guide](../daily_driver.md)
- [MCP Gateway Security Model](../security_model.md)
- [MCP Integration](../integration_mcp.md)
- [Troubleshooting](../troubleshooting.md)

---

<!-- FILE: docs/recipes/claude_code.md -->

# Claude Code + contextweaver gateway

Use contextweaver as one project-scoped MCP server so Claude Code sees three
gateway meta-tools instead of a large set of full upstream schemas.

This recipe's registration syntax and project config were verified on
**Claude Code 2.1.165 on June 10, 2026**. The deterministic gateway surface is
covered by the repository's MCP tests; live model-driven tool selection
remains a manual client check.

## Prerequisites

1. Claude Code installed and signed in.
2. Python 3.10 or newer.
3. `uv` for the zero-install command, or an installed `contextweaver` CLI.
4. A JSON/YAML tool catalog or MCP `tools/list` snapshot.

The worked example uses the committed 11-tool filesystem snapshot and the
gateway config at `examples/recipes/gateway_config.yaml`.

## Validate before registration

From the contextweaver repository root:

```bash
uvx contextweaver mcp serve \
  --config examples/recipes/gateway_config.yaml \
  --dry-run
```

Expected stderr includes:

```text
mode=gateway ... tools=11 top_k=10 beam_width=3 ...
dry-run: catalog validated; not binding stdio.
```

The first `uvx` invocation may take longer while it resolves an isolated
environment. For a persistent installation, use:

```bash
pip install contextweaver
contextweaver mcp serve --config examples/recipes/gateway_config.yaml --dry-run
```

## Option A: commit `.mcp.json`

The shipped
[`examples/recipes/claude_code_mcp.json`](https://github.com/dgenio/contextweaver/blob/main/examples/recipes/claude_code_mcp.json)
contains:

```json
{
  "mcpServers": {
    "contextweaver-gateway": {
      "type": "stdio",
      "command": "uvx",
      "args": [
        "contextweaver",
        "mcp",
        "serve",
        "--config",
        "${CLAUDE_PROJECT_DIR:-.}/examples/recipes/gateway_config.yaml"
      ]
    }
  }
}
```

Copy that structure to `.mcp.json` at the project root and adjust the config
path. Claude Code supports `${VAR}` and `${VAR:-default}` expansion in project
MCP configuration. Keep credentials in environment variables rather than the
committed file.

To pin the package:

```json
"args": [
  "contextweaver@0.14.0",
  "mcp",
  "serve",
  "--config",
  "${CLAUDE_PROJECT_DIR:-.}/examples/recipes/gateway_config.yaml"
]
```

Claude Code asks each user to approve a project-scoped server before
connecting.

## Option B: register the JSON with the Claude CLI

`claude mcp add-json` accepts the server object and writes it to the selected
scope:

PowerShell:

```powershell
claude mcp add-json --scope project contextweaver-gateway `
  '{"type":"stdio","command":"uvx","args":["contextweaver","mcp","serve","--config","${CLAUDE_PROJECT_DIR:-.}/examples/recipes/gateway_config.yaml"]}'
```

macOS/Linux:

```bash
claude mcp add-json --scope project contextweaver-gateway \
  '{"type":"stdio","command":"uvx","args":["contextweaver","mcp","serve","--config","${CLAUDE_PROJECT_DIR:-.}/examples/recipes/gateway_config.yaml"]}'
```

Use `--scope local` for a private entry in the current project or
`--scope user` with an absolute config path for all projects.

Claude Code also documents `claude mcp add <name> -- <command> [args...]`.
On the verified 2.1.165 Windows PowerShell build, a nested server flag such as
`--config` was still parsed as a Claude option. `add-json` and a committed
`.mcp.json` both preserved the arguments correctly, so they are the verified
paths in this recipe.

## Confirm the connection

Run:

```bash
claude mcp list
claude mcp get contextweaver-gateway
```

Then open Claude Code and use `/mcp`. A project-scoped entry may initially
show `Pending approval`; approve it in the interactive session. Once
connected, the server should advertise:

- `tool_browse`
- `tool_execute`
- `tool_view`

It should not advertise the 11 raw filesystem tools from the snapshot.

## Give Claude an operating rule

Add this to the project's `CLAUDE.md` or provide it in the task:

```text
Use contextweaver-gateway for large tool catalogs. Call tool_browse first with
a routing-oriented query, execute only the selected tool_id through
tool_execute, and call tool_view with a narrow selector only when the summary
is insufficient. Do not treat routing as authorization; follow normal approval
rules for tool side effects.
```

Do not keep the same upstream MCP servers registered directly under other
names. Otherwise Claude sees both the raw tools and the gateway.

## Guided first session

1. Open `/mcp` and confirm the three meta-tools.
2. Ask: `Use contextweaver-gateway to find the filesystem tool for listing a directory.`
3. Confirm Claude calls `tool_browse` before `tool_execute`.
4. For a large result, confirm the response is a summary plus an artifact
   handle.
5. Ask for one small slice and confirm Claude uses `tool_view` with `head`,
   `lines`, `rows`, or `json_keys`.

With a `catalog:` config the CLI uses a deterministic stub upstream, so this
checks client wiring, routing, validation, and firewall shape. For real
upstream execution, switch the config to an `upstreams:` block — `mcp serve`
launches the listed MCP servers behind the gateway and executes selected
calls live:

```yaml
upstreams:
  filesystem:
    type: stdio
    command: npx
    args: ["-y", "@modelcontextprotocol/server-filesystem", "/workspace"]
    namespace: fs
startup:
  mode: degraded
```

An existing multi-server client config migrates with
`contextweaver mcp import-vscode <config> --apply` (dry-run by default). See
[MCP Integration](../integration_mcp.md#connecting-to-real-upstream-mcp-servers)
for the full `upstreams:` / `startup:` reference. Live upstream serving covers
tools only; resources and prompts over live upstreams are not yet bridged.

## Troubleshooting

### Gateway is not listed

- Run the dry-run command outside Claude Code first.
- Check `claude mcp get contextweaver-gateway`.
- Approve a project-scoped server in `/mcp`.
- Use an absolute config path for user scope.
- Increase Claude Code's MCP startup timeout if the first cold `uvx` resolve
  exceeds the default.

### Tool names are rejected or missing

The client should see only the gateway's underscore-separated meta-tool names.
If raw upstream names appear, the direct upstream server is still registered.
Remove the duplicate registration and restart the session.

### Server fails after the first call

Stdio servers are not automatically reconnected by Claude Code. Fix the
startup or catalog error, then reconnect from `/mcp` or restart the session.

### Where diagnostics go

`contextweaver mcp serve` writes diagnostics to stderr. Stdout is reserved for
the MCP wire protocol; redirecting application logs to stdout can corrupt the
connection.

### `uvx` is slow or unavailable

Install contextweaver persistently and change the entry to:

```json
"command": "contextweaver",
"args": ["mcp", "serve", "--config", "/absolute/path/to/gateway.yaml"]
```

`pipx run contextweaver mcp serve ...` is also supported for an isolated
invocation.

## Security note

Raw outputs can remain in the artifact store and `tool_view` re-exposes
selected content. Review the [MCP Gateway Security Model](../security_model.md)
before connecting sensitive upstreams.

## See also

- [Daily Driver Guide](../daily_driver.md)
- [Recipes overview](index.md)
- [MCP Integration](../integration_mcp.md)
- [Claude Code MCP documentation](https://code.claude.com/docs/en/mcp)

---

<!-- FILE: docs/recipes/okf_bundle.md -->

# Knowledge-bundle context sources (OKF, repo knowledge, lessons, expertise packs)

Four related adapters let contextweaver ingest external knowledge stored as
Markdown files with YAML frontmatter — the OKF convention — and expose it as
bounded, selectable context candidates that flow through the existing
candidate selection, budget, dedup, sensitivity, and rendering pipeline. None
of them require network access or a runtime dependency beyond PyYAML, which
is already a core dependency.

| Adapter | Module | Use it for |
|---|---|---|
| OKF bundle loader | `contextweaver.adapters.okf` | Generic OKF-format knowledge bundles |
| Repository knowledge | `contextweaver.adapters.repo_knowledge` | Generated repo wikis, `docs/agent-context/`, `AGENTS.md`-style docs |
| Lessons | `contextweaver.adapters.lessons` | LessonWeaver-exported lessons, with lifecycle filtering |
| Expertise packs | `contextweaver.adapters.expertise_pack` | Structured constraints/assumptions with conflict detection |

All four share one permissive parsing core (`contextweaver.adapters._okf_io`):
a missing frontmatter fence, invalid YAML, or a non-mapping frontmatter value
degrades to a diagnostic plus a best-effort node — it never raises, unless
you opt in with `on_invalid="raise"`.

## OKF bundle loader

An OKF bundle is a directory of `.md` files with YAML frontmatter, plus two
optional bundle-level files: `index.md` (overview metadata) and `log.md`
(bundle history) — neither is loaded as ordinary concept content.

```python
from contextweaver.adapters.okf import load_okf_bundle, select_knowledge

bundle = load_okf_bundle("path/to/okf-bundle")
items = select_knowledge(bundle.nodes, "context firewall", budget_tokens=2000)
```

**When to prefer OKF over normal event-log ingestion:** use OKF when the
knowledge is *external, versioned, and reusable across sessions* — a shared
concept library, not this session's conversation. Use the normal event log
(`ContextManager.ingest`) for anything session-specific: tool calls, tool
results, user turns. OKF nodes and event-log items can both be selected into
the same context build; they are independent candidate sources.

Unknown frontmatter fields are preserved verbatim under each node's
`frontmatter` attribute, and every field surfaces in the materialised
`ContextItem.metadata["frontmatter"]` — nothing is silently dropped.

## Repository knowledge

Narrows the OKF loader to repo documentation: generated wikis, architecture
notes, module summaries, `AGENTS.md`/`CLAUDE.md`-style instruction files.
Unlike the OKF loader, plain Markdown files with no frontmatter at all are
still valid candidates (their title falls back to the filename), and
`index.md`/`log.md` carry no special meaning — this is a documentation tree,
not an OKF bundle proper.

```python
from contextweaver.adapters.repo_knowledge import load_repo_knowledge, select_repo_knowledge

bundle = load_repo_knowledge("docs/agent-context", max_files=200)
debugging_docs = select_repo_knowledge(
    bundle.nodes, "why is routing dropping candidates", budget_tokens=2000, usage_tag="debugging"
)
```

References inside a document (`AGENTS.md` links, a node's own `links` field)
are never auto-followed — the loader only reads files under the directory
you point it at, so a documentation tree cannot force-load content beyond
its own root.

Relation to the plain OKF loader: `repo_knowledge` is a thin, purpose-specific
layer over the same core — it adds the plain-Markdown fallback, size
guardrails (`max_files`/`max_total_bytes`), and deterministic usage-tag
classification (`classify_usage`, e.g. `"debugging"`, `"onboarding"`). These
tags are plain metadata strings, not contextweaver `Phase` values.

## Lessons (LessonWeaver-exported bundles)

Lessons differ from repository-knowledge nodes in one key way: **lifecycle
status governs eligibility, not just relevance.** A lesson's `status`
(`candidate`/`reviewed`/`active`/`deprecated`/`rejected`), `scope`, and
`expires_at` decide whether it is even a candidate for selection — repository
documentation nodes have no such gate.

```python
from contextweaver.adapters.lessons import (
    LessonSelectionPolicy,
    load_lesson_bundle,
    select_lessons,
)

nodes, _diagnostics = load_lesson_bundle("path/to/lessonweaver-export")
items, excluded = select_lessons(
    nodes,
    "api design",
    budget_tokens=1500,
    policy=LessonSelectionPolicy(preferred_scope="project"),
)
```

By default, `rejected` and `deprecated` lessons are excluded, and unreviewed
`candidate` lessons are excluded unless you opt in with
`LessonSelectionPolicy(include_candidates=True)`. Every exclusion is reported
back with a reason (`"status:rejected"`, `"expired"`, ...) so you can surface
lifecycle diagnostics rather than silently dropping content.

**End-to-end sketch:** LessonWeaver reviews traces and exports reviewed
lessons as OKF-style Markdown nodes → contextweaver's `select_lessons` picks
the subset relevant to the current task, honoring lifecycle status → a
downstream ChainWeaver flow step can reference the selected lesson IDs (via
each item's `metadata["_contextweaver"]["knowledge_source"]["id"]`) as
provenance for why a particular constraint was applied.

## Expertise packs

An ExpertisePack is a directory bundle of constraint/assumption/verification/
failure-mode nodes. Each node's frontmatter `key` groups related constraints
(e.g. `"api-style"`, `"verification-command"`); an `index.md` declares the
pack's `version`.

```python
from contextweaver.adapters.expertise_pack import (
    detect_conflicts,
    expertise_pack_to_context_items,
    load_expertise_pack,
)

pack = load_expertise_pack("path/to/expertise-pack")
findings = detect_conflicts(pack.nodes, task_tags={"python-library"})
items = expertise_pack_to_context_items(pack, task_tags={"python-library"})
```

Conflict detection is deterministic and literal: it flags constraints that
share a `key` but disagree on text, restricted to nodes that are live
(not expired) and applicable to the given `task_tags`. It does not perform
natural-language contradiction inference — that would require a model call,
which core knowledge-source loading deliberately does not make. Pack
sections only enter bounded context "when relevant" — expired or
inapplicable nodes are excluded by `expertise_pack_to_context_items`, never
injected unconditionally.

**Consuming packs generated by LessonWeaver:** LessonWeaver can export
distilled expertise (goals, constraints, known failure modes) in the same
OKF-style Markdown-plus-frontmatter shape `load_expertise_pack` expects —
point the loader at LessonWeaver's export directory the same way you would
any other ExpertisePack. The pack's `key` field is what LessonWeaver should
use to group related constraints so `detect_conflicts` can catch
contradictions across export runs.

The canonical ExpertisePack schema is tracked externally at
`dgenio/weaver-spec#184`. This adapter validates pack **structure** (an
`index.md` declaring a version, every node carrying a `key`) rather than
that full external schema; see the module docstring in
`contextweaver.adapters.expertise_pack` for the seam to bind it later.

## See also

- [`examples/knowledge_bundles_demo.py`](https://github.com/dgenio/contextweaver/blob/main/examples/knowledge_bundles_demo.py) — a runnable, self-contained walkthrough of all four adapters.
- [Core Concepts](../concepts.md) for `ContextItem`, `Sensitivity`, and the candidate-selection pipeline these adapters feed into.

---

<!-- FILE: docs/integration_mcp.md -->

# MCP Integration

contextweaver provides an adapter for the
[Model Context Protocol (MCP)](https://modelcontextprotocol.io/) that
converts MCP tool definitions and results into contextweaver's native
types.

## Adapter functions

### `mcp_tool_to_selectable(tool_dict)`

Converts an MCP tool definition dict into a `SelectableItem`:

```python
from contextweaver.adapters.mcp import mcp_tool_to_selectable

mcp_tool = {
    "name": "search_database",
    "description": "Search records in the database",
    "inputSchema": {
        "type": "object",
        "properties": {
            "query": {"type": "string"},
            "limit": {"type": "integer", "default": 10}
        }
    },
    "outputSchema": {
        "type": "object",
        "properties": {
            "results": {"type": "array"},
            "total": {"type": "integer"}
        }
    }
}

item = mcp_tool_to_selectable(mcp_tool)
# item.id            == "mcp:search_database"
# item.kind          == "tool"
# item.name          == "search_database"
# item.output_schema == {"type": "object", ...}
```

If the tool definition includes an `outputSchema`, it is preserved in
`item.output_schema`.  When absent the field is `None`.

The namespace is inferred automatically from the tool name prefix:

| Tool name              | Inferred namespace |
|------------------------|--------------------|
| `github.create_issue`  | `github`           |
| `filesystem/read`      | `filesystem`       |
| `slack_send_message`   | `slack`            |
| `search_database`      | `mcp` (fallback)   |

Use `infer_namespace(tool_name)` directly if you need the logic outside of
`mcp_tool_to_selectable()`.

### `mcp_result_to_envelope(result_dict, tool_name)`

Converts an MCP tool result dict into a `ResultEnvelope`:

```python
from contextweaver.adapters.mcp import mcp_result_to_envelope

mcp_result = {
    "content": [{"type": "text", "text": "Found 42 records matching query"}],
    "isError": False
}

envelope, binaries, full_text = mcp_result_to_envelope(mcp_result, "search_database")
# envelope.summary contains truncated text (max 500 chars)
# full_text contains the complete untruncated text
# envelope.status  == "ok"
# binaries maps handle → (raw_bytes, media_type, label)
```

#### Supported content types

| Content type    | Handling                                                                                      |
|-----------------|-----------------------------------------------------------------------------------------------|
| `text`          | Concatenated into `full_text` and `summary`                                                   |
| `image`         | Base64-decoded; stored as binary artifact                                                     |
| `audio`         | Base64-decoded; stored as binary artifact (e.g. `audio/wav`)                                  |
| `resource`      | Text extracted into `full_text`; raw bytes stored as artifact                                 |
| `resource_link` | URI stored as `ArtifactRef`; URI string in `binaries` for caller resolution                |

#### Structured content

If the result contains a top-level `structuredContent` dict, it is
serialized as a JSON artifact and its top-level keys are extracted as
facts:

```python
mcp_result = {
    "content": [{"type": "text", "text": "query done"}],
    "structuredContent": {"count": 42, "status": "done"},
}
envelope, binaries, _ = mcp_result_to_envelope(mcp_result, "query")
# binaries["mcp:query:structured_content"] → JSON bytes
# envelope.facts includes "count: 42", "status: done"
```

#### Content-part annotations

Per-part `annotations` (with `audience` and `priority` fields) are
collected into `envelope.provenance["content_annotations"]`:

```python
mcp_result = {
    "content": [
        {"type": "text", "text": "...", "annotations": {"audience": ["human"], "priority": 0.9}},
    ],
}
envelope, _, _ = mcp_result_to_envelope(mcp_result, "tool")
# envelope.provenance["content_annotations"] == [{"part_index": 0, "audience": ["human"], ...}]
```

### `load_mcp_session_jsonl(path)`

Loads a JSONL session file containing MCP-style events and returns a
list of `ContextItem` objects:

```python
from contextweaver.adapters.mcp import load_mcp_session_jsonl

items = load_mcp_session_jsonl("examples/data/mcp_session.jsonl")
for item in items:
    print(f"{item.kind.value}: {item.text[:60]}...")
```

## Session JSONL format

Each line is a JSON object with at minimum `id`, `type`, and either
`text` or `content`:

```json
{"id": "u1", "type": "user_turn", "text": "Search for open invoices"}
{"id": "tc1", "type": "tool_call", "text": "invoices.search(status='open')", "parent_id": "u1"}
{"id": "tr1", "type": "tool_result", "content": "...", "parent_id": "tc1"}
```

See `examples/data/mcp_session.jsonl` for a complete example.

## End-to-end example

```python
from contextweaver.adapters.mcp import (
    load_mcp_session_jsonl,
    mcp_tool_to_selectable,
)
from contextweaver.context.manager import ContextManager
from contextweaver.types import ItemKind, Phase

# Load session events
items = load_mcp_session_jsonl("examples/data/mcp_session.jsonl")

# Build context with firewall
mgr = ContextManager()
for item in items:
    if item.kind == ItemKind.tool_result and len(item.text) > 2000:
        mgr.ingest_tool_result(
            tool_call_id=item.parent_id or item.id,
            raw_output=item.text,
            tool_name="mcp_tool",
        )
    else:
        mgr.ingest(item)

pack = mgr.build_sync(phase=Phase.answer, query="invoice status")
print(pack.prompt)
```

See `examples/mcp_adapter_demo.py` for the full runnable demo.

## Prompt-caching compatibility

Anthropic (90%), OpenAI (50%), and Google (75%) all discount the prompt-token
cost of tool definitions when the same prefix is reused across requests.
contextweaver's
[`make_choice_cards`](../src/contextweaver/routing/cards.py) function is
**deterministic and byte-stable** for identical inputs (sorted descending by
score, ascending by `id` for ties — see issue #218 for the regression test
that locks this guarantee), so the cards array your downstream prompt
assembler renders is suitable for placement *before* a cache breakpoint.

The repo guarantees this via `tests/test_cards.py::test_make_choice_cards_byte_identical_stable_order`,
which asserts `bytes(card1) == bytes(card2)` across two consecutive calls
on identical inputs. The invariant survives across the full
`SelectableItem → ChoiceCard → cache prefix` chain.

### Worked example: Anthropic `cache_control`

> **Illustrative — requires the Anthropic SDK.** This snippet imports
> `anthropic` to show how the byte-stable cards array slots into the
> provider's cache-control API. contextweaver itself does not depend on
> the Anthropic SDK; install it separately with `pip install anthropic`
> to run the example as-is, or read it as a pattern reference.

```python
import anthropic  # pip install anthropic
from contextweaver.routing.cards import make_choice_cards
from contextweaver.routing.catalog import Catalog

catalog = Catalog()  # populated elsewhere with stable IDs
cards = make_choice_cards(
    catalog.all(),
    scores={item.id: 0.5 for item in catalog.all()},   # deterministic scoring
    max_cards=20,
)

# Render cards into Anthropic's `tools` array (cacheable prefix).
tools = [
    {
        "name": c.name,
        "description": c.description,
        "input_schema": {"type": "object"},  # hydrate per-call when selected
    }
    for c in cards
]

# Place the cache breakpoint on the LAST tool definition. As long as the
# `cards` array is stable, every request reuses the cache prefix and only
# the trailing user turn varies.
if tools:
    tools[-1]["cache_control"] = {"type": "ephemeral"}

client = anthropic.Anthropic()
client.messages.create(
    model="claude-3-5-sonnet-latest",
    max_tokens=1024,
    tools=tools,
    messages=[{"role": "user", "content": "..."}],
)
```

> **Practical guidance for multi-turn navigation.** When the cards array
> *naturally* changes between turns (e.g., user navigated into a sub-tree),
> the cache prefix invalidates — that's expected. To keep the prefix stable
> across navigation, sort hydrated cards by ID once and append newly-discovered
> cards after the breakpoint. The
> [Webfuse MCP cheat sheet](https://www.webfuse.com/mcp-cheat-sheet)
> documents the canonical "append after cache breakpoint" pattern.
>
> **First-class flag:** `ProxyRuntime(cache_stable=True)` implements this
> pattern automatically — see [gateway spec §5](gateway_spec.md#5-cache-stable-tool-browsing-cache_stabletrue).
> Browsed/hydrated tool ids are tracked per session; on each
> `tool_browse` call, previously-seen cards are emitted first in
> ascending-`id` order, followed by a `__cache_breakpoint__` marker
> card, followed by newly-discovered cards (also `id`-ascending).
> First-sighting card content is frozen, so the prefix bytes are
> stable across browses with different queries. **Caveat:** the
> first emitted card is not the highest-ranked when this flag is on
> — read rank from `ChoiceCard.score`.

## Security Considerations

For the full gateway data flow, artifact-store boundary, egress model, and
deployment checklist, read the
[MCP Gateway Security Model](security_model.md).

### MCP annotations are untrusted hints

MCP tool annotations — `readOnlyHint`, `destructiveHint`, `costHint` — are
**server-declared metadata**, not verified security properties.  The
[MCP specification](https://modelcontextprotocol.io/legacy/concepts/tools)
explicitly states:

> _"Clients SHOULD NOT make security-critical decisions based solely on tool
> annotations. Annotations are informational metadata, not security controls."_

contextweaver maps these hints to informational fields on `SelectableItem`:

| Annotation       | Field mapped to             | Purpose               |
|------------------|-----------------------------|-----------------------|
| `readOnlyHint`   | `side_effects=False`, tag `"read-only"` | Routing UX display |
| `destructiveHint`| tag `"destructive"`         | Routing UX display    |
| `costHint`       | `cost_hint` (float)         | Routing cost scoring  |

### `side_effects` is informational only

`item.side_effects = False` (derived from `readOnlyHint=True`) means the
**server advertised** the tool as read-only.  It does **not** guarantee the
tool has no side effects.  A malicious or misconfigured MCP server could
declare `readOnlyHint: True` on a destructive tool; contextweaver would
faithfully tag it `"read-only"` with `side_effects=False`.

**Do not build access-control or safety-gate logic on these fields.**

### Authorization status

contextweaver does **not currently provide an authorization mechanism** for
MCP tools. Do not rely on server-declared annotation hints for access
control.

`CapabilityToken` (see
[issue #20](https://github.com/dgenio/contextweaver/issues/20)) is a
proposed/future feature, not a type that is implemented in the library
today. For actual access control, enforce authorization in your own
application or policy layer.

---

## Runtime modes: transparent proxy and two-tool gateway

The MCP adapter ships two runtime modes for fronting one or more
upstream MCP servers.  Both share the
[`ProxyRuntime`](../src/contextweaver/adapters/proxy_runtime.py) core and
satisfy the contracts in [`docs/gateway_spec.md`](gateway_spec.md):

Production MCP gateway deployments commonly transform raw
user input into routing-oriented queries before calling
`Router.route(query)`. ContextWeaver does not require a
specific rewriting strategy and accepts whichever
routing-shaped query your gateway produces.

| Mode | Discovery channel | Invocation channel | Schema exposure |
|------|-------------------|--------------------|-----------------|
| `ExposureMode.TRANSPARENT` (#13) | Stripped `tools/list` — one entry per upstream tool with sentinel `inputSchema: {"type": "object"}` | `tool_hydrate(tool_id)` + `tool_execute(tool_id, args)` | On demand via `tool_hydrate` |
| `ExposureMode.GATEWAY` (#28 + #34) | None — the agent never sees a `tools/list` | `tool_browse(query|path)` + `tool_execute(tool_id, args)` + `tool_view(handle, selector)` | Internal: `tool_execute` hydrates and validates before upstream dispatch |

Both modes share the same invocation contract: arguments to
`tool_execute` are validated against the hydrated schema via
`jsonschema` before any upstream call, per
[`gateway_spec.md` §4.4](gateway_spec.md#4-schema-exposure-strategy).

### Wiring a gateway over stdio

```python
import asyncio
from contextweaver.adapters import ProxyRuntime, StubUpstream
from contextweaver.adapters.mcp_gateway_server import McpGatewayServer

runtime = ProxyRuntime(StubUpstream([...]))
await runtime.refresh_catalog()
server = McpGatewayServer(runtime, name="example-gateway")
asyncio.run(server.run_stdio())
```

### Wiring a transparent proxy over stdio

```python
from contextweaver.adapters import ExposureMode, ProxyRuntime, StubUpstream
from contextweaver.adapters.mcp_proxy_server import McpProxyServer

runtime = ProxyRuntime(StubUpstream([...]), mode=ExposureMode.TRANSPARENT)
await runtime.refresh_catalog()
server = McpProxyServer(runtime, name="example-proxy")
asyncio.run(server.run_stdio())
```

### Connecting to real upstream MCP servers

Swap [`StubUpstream`](../src/contextweaver/adapters/mcp_upstream.py) for
`McpClientUpstream(session)` (one upstream) or
`MultiplexUpstream([a, b, ...])` (multi-server fan-out).  The runtime
itself is transport-agnostic; the upstream adapter handles the wire
protocol.

### Error shape

Every gateway / proxy meta-tool returns either a `ResultEnvelope` or a
typed
[`GatewayError`](../src/contextweaver/adapters/gateway_error.py)
matching `gateway_spec.md` §3.4:

```json
{
  "error": "PATH_INVALID" | "PATH_NOT_FOUND" | "ARGS_INVALID" | "UPSTREAM_ERROR" | "HYDRATE_FAILED" | "VIEW_FAILED",
  "message": "<human-readable>",
  "path": "<offending path or tool_id>",
  "details": { "...": "..." }
}
```

The meta-tools never raise across the MCP boundary — failures are
delivered as `isError=true` `CallToolResult` payloads.

### See also

- [`docs/gateway_spec.md`](gateway_spec.md) — the normative
  surface specification.
- [`examples/mcp_gateway_demo.py`](../examples/mcp_gateway_demo.py) —
  end-to-end gateway flow using `StubUpstream`.
- [`examples/mcp_proxy_demo.py`](../examples/mcp_proxy_demo.py) —
  end-to-end proxy flow.
- [Recipes > Claude Desktop](recipes/claude_desktop.md) — put a
  contextweaver gateway in front of Claude Desktop's MCP client.
- [Recipes > GitHub Copilot](recipes/github_copilot.md) — put a
  contextweaver gateway in front of VS Code Copilot Chat (agent mode).
- [Recipes > Claude Code](recipes/claude_code.md) — project-scoped gateway
  registration for Claude Code.
- [MCP Gateway Security Model](security_model.md) — data flow, trust
  boundaries, artifact exposure, and hardening.
- [`examples/architectures/mcp_context_gateway/main_real.py`](../examples/architectures/mcp_context_gateway/main_real.py)
  — the reference architecture run against verbatim `tools/list`
  snapshots of MIT-licensed reference MCP servers.

### Transport & compatibility matrix

contextweaver's MCP server supports two transport protocols. The default
is **stdio** (used by every client recipe in the repo). **SSE** is
available for clients that prefer an HTTP-based connection.

| Transport | CLI flag | Status | First shipped |
|---|---|---|---|
| stdio | `--transport stdio` (default) | Stable | v0.4 |
| SSE | `--transport sse --host 127.0.0.1 --port 8000` | Available | v0.15.0 |

Both transports share the same runtime, catalog, and gateway surface —
only the wire binding differs.

#### Client compatibility (tested versions)

| Client | Transport | Required capability | Verified version | Notes |
|---|---|---|---|---|
| Claude Desktop | stdio | `tools` | 0.8.x | macOS / Windows; no SSE support |
| Claude Code | stdio | `tools` | 2.1.x | Project-scoped via `.mcp.json` |
| VS Code + Copilot Chat | stdio | `tools` | 1.96+ | Workspace `.vscode/mcp.json` |
| Cursor | stdio | `tools` | 0.45+ | `.cursor/mcp.json` |
| Generic SSE client | SSE | `tools` | SDK 1.27+ | Any MCP client that speaks HTTP/SSE |

#### Required capabilities

contextweaver advertises only the `tools` capability on initialization.
It does **not** expose `resources`, `prompts`, or `logging` through the
MCP server surface. The gateway's own `tool_browse` / `tool_execute` /
`tool_view` meta-tools are sufficient for every tested client.

#### Choosing a transport

- **stdio** — Use when the client launches the server as a subprocess
  (Claude Desktop, VS Code Copilot, Claude Code, Cursor). This is the
  default and requires no extra flags.
- **SSE** — Use when the client connects to a separate HTTP endpoint
  (e.g. a remote contextweaver instance, a browser-based client, or a
  container sidecar). Enable with:

```bash
contextweaver mcp serve --config gateway.yaml --transport sse --port 8080
```

#### Config-file equivalent

```yaml
catalog: my_catalog.json
mode: gateway
transport: sse
host: 127.0.0.1
port: 8080
```

#### Security note for SSE

SSE binds an HTTP server. By default it listens on `127.0.0.1` only.
If you change `--host` to `0.0.0.0`, ensure a firewall or reverse proxy
sits in front.

The MCP SDK ships DNS-rebinding protection but leaves it **disabled by
default**. contextweaver enables it and scopes the `Host`/`Origin`
allowlist to the address you bind (`--host`/`--port`, plus the usual
`localhost` aliases for loopback binds). A request whose `Host` header is
not in that allowlist is rejected with `421`. Practical consequence: when
binding `0.0.0.0` (or a routable host) and reaching the server under a
different hostname, put a reverse proxy in front that forwards a `Host`
the server accepts. CORS and authentication remain the caller's
responsibility. See [MCP Gateway Security Model](security_model.md).

---

<!-- FILE: docs/integration_a2a.md -->

# A2A Integration

contextweaver provides an adapter for the
[Agent-to-Agent (A2A) protocol](https://google.github.io/A2A/) that
converts A2A agent cards and task results into contextweaver's native
types.

## Adapter functions

### `a2a_agent_to_selectable(agent_card)`

Converts an A2A agent card dict into a `SelectableItem`:

```python
from contextweaver.adapters.a2a import a2a_agent_to_selectable

agent_card = {
    "name": "DataAgent",
    "description": "Retrieves and aggregates data from warehouses",
    "url": "https://agents.example.com/data",
    "skills": [
        {"id": "sql_query", "name": "SQL Query", "description": "Run SQL queries"},
        {"id": "aggregate", "name": "Aggregate", "description": "Aggregate results"},
    ]
}

item = a2a_agent_to_selectable(agent_card)
# item.id    == "a2a:DataAgent"
# item.kind  == "agent"
# item.name  == "DataAgent"
# item.tags  includes skill names
```

### `a2a_result_to_envelope(task_result, agent_name)`

Converts an A2A task result dict into a `ResultEnvelope`:

```python
from contextweaver.adapters.a2a import a2a_result_to_envelope

task_result = {
    "status": {"state": "completed"},
    "artifacts": [
        {"parts": [{"type": "text", "text": "Q4 revenue: $2.1M, +15% YoY"}]}
    ]
}

envelope = a2a_result_to_envelope(task_result, "DataAgent")
# envelope.summary contains the artifact text
# envelope.status  == "ok"
```

### `load_a2a_session_jsonl(path)`

Loads a JSONL session file containing A2A-style multi-agent events:

```python
from contextweaver.adapters.a2a import load_a2a_session_jsonl

items = load_a2a_session_jsonl("examples/data/a2a_session.jsonl")
```

## Session JSONL format

Each line is a JSON object. A2A sessions typically involve multi-agent
handoffs where an orchestrator delegates to specialised agents:

```json
{"id": "u1", "type": "user_turn", "text": "Generate the Q4 report"}
{"id": "tc1", "type": "tool_call", "text": "delegate_to(DataAgent, 'fetch Q4 data')", "parent_id": "u1"}
{"id": "tr1", "type": "tool_result", "content": "...", "parent_id": "tc1"}
```

See `examples/data/a2a_session.jsonl` for a complete multi-agent
session.

## End-to-end example

```python
from contextweaver.adapters.a2a import (
    a2a_agent_to_selectable,
    a2a_result_to_envelope,
    load_a2a_session_jsonl,
)
from contextweaver.context.manager import ContextManager
from contextweaver.types import ItemKind, Phase

# Load multi-agent session
items = load_a2a_session_jsonl("examples/data/a2a_session.jsonl")

# Build context
mgr = ContextManager()
for item in items:
    if item.kind == ItemKind.tool_result and len(item.text) > 2000:
        mgr.ingest_tool_result(
            tool_call_id=item.parent_id or item.id,
            raw_output=item.text,
            tool_name="a2a_agent",
        )
    else:
        mgr.ingest(item)

pack = mgr.build_sync(phase=Phase.answer, query="Q4 report")
print(pack.prompt)
```

See `examples/a2a_adapter_demo.py` for the full runnable demo.

---

<!-- FILE: docs/errors.md -->

# Error Reference

Every error contextweaver raises inherits from `ContextWeaverError`
(in `contextweaver.exceptions`), so you can catch the whole family with one
`except` clause. Each class also carries:

- a stable, machine-readable **`code`** (e.g. `CW_CONFIG`) — branch on this
  instead of string-matching the message; it is safe to log, alert on, and pass
  across the gateway boundary or to non-Python clients, and
- an optional one-line **`hint`** — a remediation pointer, often a link back to
  the relevant section on this page.

Codes are part of the public compatibility surface: they are frozen against a
golden list in the test suite, so a rename or a missing code fails CI.

```python
from contextweaver.exceptions import ContextWeaverError

try:
    pack = manager.build_sync(phase, query)
except ContextWeaverError as exc:
    # exc.code is stable; exc.hint may point at the fix.
    logger.error("contextweaver failed: %s", exc, extra={"cw_code": exc.code})
    raise
```

`str(exc)` renders as `[<code>] <message> (hint: <hint>)` — for example
`[CW_CONFIG] unknown preset 'fast' (hint: check the configuration value or preset name; ...)`.
The message text itself is **not** a stable API; the code and any structured
attributes are.

## Code index

| Code | Exception | Raised when |
| --- | --- | --- |
| `CW_ERROR` | `ContextWeaverError` | Base class; not raised directly. |
| `CW_BUDGET_EXCEEDED` | `BudgetExceededError` | A build would exceed the configured token budget. |
| `CW_BUDGET_OVERFLOW` | `BudgetOverflowError` | Budget pressure dropped candidates under a fail-loud policy. |
| `CW_ARTIFACT_NOT_FOUND` | `ArtifactNotFoundError` | A requested artifact handle is absent from the store. |
| `CW_ARTIFACT_STORE_QUOTA` | `ArtifactStoreQuotaError` | A write would breach an artifact store's size/count quota. |
| `CW_POLICY_VIOLATION` | `PolicyViolationError` | An item violates the active `ContextPolicy`. |
| `CW_ITEM_NOT_FOUND` | `ItemNotFoundError` | A tool/agent/skill ID is not in the catalog or store. |
| `CW_GRAPH_BUILD` | `GraphBuildError` | The routing DAG cannot be constructed (e.g. a cycle). |
| `CW_ROUTE` | `RouteError` | The router cannot produce a valid route. |
| `CW_CATALOG` | `CatalogError` | An invalid catalog operation (duplicate IDs, schema). |
| `CW_CATALOG_VALIDATION` | `CatalogValidationError` | A catalog fails cross-item referential validation. |
| `CW_DUPLICATE_ITEM` | `DuplicateItemError` | A duplicate ID is appended to an append-only store. |
| `CW_CONFIG` | `ConfigError` | A configuration value or preset name is invalid. |
| `CW_VALIDATION` | `ValidationError` | A core data type fails construction-time validation. |
| `CW_DETERMINISM` | `DeterminismError` | A `deterministic=True` firewall path would invoke an LLM. |
| `CW_PATH_INVALID` | `PathInvalidError` | A `tool_browse` path violates the §3.2 grammar. |
| `CW_PATH_NOT_FOUND` | `PathNotFoundError` | A well-formed `tool_browse` path resolves to no node. |
| `CW_UPSTREAM` | `UpstreamError` | An upstream MCP tool call fails for transport/protocol reasons. |
| `CW_STORE_CLOSED` | `StoreClosedError` | An operation is attempted on a closed store. |
| `CW_STORE_TIMEOUT` | `StoreTimeoutError` | An async store operation driven through the sync bridge exceeds its timeout. |
| `CW_UPSTREAM_STARTUP` | `UpstreamStartupError` | Live multi-upstream startup fails under the configured `StartupPolicy`. |

---

## ContextWeaverError

**Code:** `CW_ERROR`

The base of the hierarchy. It is not raised directly; catch it to handle every
contextweaver error in one place. Subclass it (not `Exception`) if you extend
the library so your error stays inside the family.

## BudgetExceededError

**Code:** `CW_BUDGET_EXCEEDED`

The public signal for a hard token-budget violation. The built-in fail-loud
path raises the more specific [`BudgetOverflowError`](#budgetoverflowerror)
(opt-in via `overflow_action="raise"`, issue #510), which attaches the would-be
`BuildStats`. Catch `BudgetExceededError` if you raise budget violations from
your own enforcement code.

**Fix:** raise the per-phase token budget, or trim the candidate set before the
build (see the [budget sizing guidance](troubleshooting.md)).

## BudgetOverflowError

**Code:** `CW_BUDGET_OVERFLOW`

Raised by `context/build_policy.py` when `ContextPolicy.overflow_action="raise"`
and budget pressure would drop candidates. Instead of silently shipping a
subtly-wrong prompt (e.g. a missing mandatory policy item), the build fails
loud. The would-be `BuildStats` is attached as `exc.stats`, and the distinct
dropped kinds as `exc.dropped_kinds`.

**Fix:** raise the phase token budget, relax the policy that marked the dropped
item mandatory, or set `overflow_action="drop"` to accept silent trimming.
Inspect `exc.stats` to see exactly what was kept and dropped.

## ArtifactNotFoundError

**Code:** `CW_ARTIFACT_NOT_FOUND`

Raised by the artifact-store backends (`store/artifacts.py`,
`store/json_file_artifacts.py`, `store/redis_artifacts.py`,
`store/s3_artifacts.py`) when a handle cannot be resolved.

**Fix:** verify the artifact ref came from the same store and has not expired or
been evicted; re-run the build that produced it if the store is ephemeral.

## ArtifactStoreQuotaError

**Code:** `CW_ARTIFACT_STORE_QUOTA`

Raised when a persistent `ArtifactStore` constructed with `max_bytes` /
`max_artifacts` limits (issue #497) would breach a limit on write.

**Fix:** raise the store's quota, prune old artifacts, or shorten artifact
lifetimes so long-running gateways stay within budget.

## PolicyViolationError

**Code:** `CW_POLICY_VIOLATION`

Raised during ingest (`context/ingest.py`) when an item violates the active
`ContextPolicy`.

**Fix:** adjust the item to satisfy the policy, or relax the policy if the
constraint is too strict for your workload.

## ItemNotFoundError

**Code:** `CW_ITEM_NOT_FOUND`

Raised when a requested tool/agent/skill ID is missing from the catalog
(`routing/catalog.py`), from a store (`store/*`), or from an external-memory
backend (`extras/memory/*`).

**Fix:** confirm the ID exists in the catalog/store and matches exactly
(IDs are case-sensitive); rebuild the catalog if it is stale.

## GraphBuildError

**Code:** `CW_GRAPH_BUILD`

Raised by `routing/tree.py`, `routing/graph.py`, and `routing/graph_io.py` when
the routing DAG cannot be built — for example a dependency cycle, a dangling
edge, or a missing root. Structured detail is attached so you can act without
parsing the message: `exc.cycle`, `exc.edge`, `exc.missing_root` (issue #523).

**Fix:** break the reported cycle, remove the dangling `depends_on`/`requires`
reference, or supply the missing root node.

## RouteError

**Code:** `CW_ROUTE`

Raised by `routing/router.py` and `routing/selection.py` when the router cannot
produce a valid route through the choice graph (e.g. no candidate survives the
beam search, or the graph has no reachable selectable items).

**Fix:** widen the routing budget/beam, check that the query matches indexed
items, and confirm the graph contains reachable selectable leaves.

## CatalogError

**Code:** `CW_CATALOG`

The base for catalog problems — duplicate IDs, schema violations, and invalid
catalog operations — raised across `routing/catalog.py`,
`routing/normalizer.py`, `routing/cards.py`, `routing/tool_id.py`, and the
protocol adapters under `adapters/` when they build catalogs from external
sources.

**Fix:** validate the catalog source for duplicate IDs and schema conformance
before loading; run `contextweaver catalog` validation on the file.

## CatalogValidationError

**Code:** `CW_CATALOG_VALIDATION`

A `CatalogError` subclass raised by the loaders' `on_invalid="raise"` path
(`routing/catalog.py`, issue #519) when cross-item referential validation fails.
The full `CatalogValidationReport` is attached as `exc.report` so you can
enumerate every dangling reference at once.

**Fix:** resolve the dangling `depends_on`/`requires` references listed in
`exc.report`, or load with `on_invalid="warn"` to triage incrementally.

## DuplicateItemError

**Code:** `CW_DUPLICATE_ITEM`

Raised when an item with an ID that already exists is appended to an append-only
store (`store/event_log.py`, `store/sqlite_event_log.py`,
`store/redis_event_log.py`, `context/_manager_ingest.py`).

**Fix:** use a unique ID per appended item, or check existence before appending
if duplicates are expected.

## ConfigError

**Code:** `CW_CONFIG`

Raised across the configuration surface (`config.py`, `profiles.py`,
`_scoring_config.py`, `routing/*`, `context/*`, adapters) when a configuration
value or preset name is invalid.

**Fix:** check the value or preset name against the documented options; the
message names the offending key.

## ValidationError

**Code:** `CW_VALIDATION`

Raised by the pure-data layer (`envelope.py`, `extras/llm_summarizer.py`) when a
core data type fails construction-time validation (issue #463). It also derives
from the builtin `ValueError`, so existing `except ValueError` call sites keep
working.

**Fix:** correct the field that failed validation; the message names the
constraint that was violated.

## DeterminismError

**Code:** `CW_DETERMINISM`

Raised by the context firewall (`context/firewall.py`, `context/ingest.py`) when
a `deterministic=True` path would have to invoke an LLM. Deterministic mode
*fails closed* (issue #404) so regulated callers can prove no data passed
through a summarisation model.

**Fix:** disable deterministic mode if model calls are acceptable, or supply a
deterministic (rule-based) summarizer/extractor so no LLM is needed.

## PathInvalidError

**Code:** `CW_PATH_INVALID`

A `CatalogError` subclass raised by `routing/path.py` when a `tool_browse` path
violates the §3.2 grammar.

**Fix:** correct the path syntax against the grammar in the
[gateway spec](gateway_spec.md).

## PathNotFoundError

**Code:** `CW_PATH_NOT_FOUND`

A `CatalogError` subclass raised by `routing/path.py` when a well-formed
`tool_browse` path resolves to no node.

**Fix:** browse from the root to discover valid paths; the catalog may have
changed since the path was constructed.

## UpstreamError

**Code:** `CW_UPSTREAM`

Signals an upstream MCP tool-call failure for transport/protocol reasons. Note
that the MCP gateway/proxy meta-tools never raise across the MCP boundary —
they return a structured [`GatewayError`](gateway_spec.md) payload with its own
wire codes (`UPSTREAM_TIMEOUT`, `AUTH_FAILED`, …) instead. Catch `UpstreamError`
when calling upstream helpers directly outside the meta-tool boundary.

**Fix:** check upstream connectivity/credentials; retry transient failures
(timeouts, unavailability) per the `retryable` hint on `GatewayError`.

## StoreClosedError

**Code:** `CW_STORE_CLOSED`

Raised by the SQLite-backed stores (`store/sqlite_facts.py`,
`store/sqlite_event_log.py`, `store/sqlite_episodic.py`) when an operation runs
after the backing connection was released via `close()`.

**Fix:** do not use a store after closing it; open a new instance, or use the
store as a context manager so its lifetime is scoped correctly.

## StoreTimeoutError

**Code:** `CW_STORE_TIMEOUT`

Raised when an async store backend, driven through the synchronous bridge
(`store/_async_to_sync.py`), does not complete an operation within its timeout
(default 30s). Without the bound, a single hung backend call would wedge the
private loop thread and, via the `ContextManager` build lock, every subsequent
`build()` — turning one stuck I/O into a permanent manager hang.

**Fix:** check the backend's health/connectivity; for a backend that is
legitimately slow, raise the bound via `_LoopThread(timeout=...)` (or pass
`timeout=None` to wait indefinitely, restoring the pre-fix behaviour).

## UpstreamStartupError

**Code:** `CW_UPSTREAM_STARTUP`

Raised by `adapters/upstream_launch.py` (`launch_upstreams`) when live
multi-upstream startup fails under the configured
`adapters.startup_policy.StartupPolicy` (issue #374): a `required` upstream
failed while `startup.mode: strict`, fewer than `min_healthy_upstreams`
upstreams started, or the effective catalog is empty and
`fail_on_empty_catalog` is set. The exception carries a `report` attribute
(a `StartupReport`) describing every upstream's individual startup outcome.

**Fix:** inspect `exc.report.statuses` for the per-upstream failure reason
(connection refused, auth failure, timeout, …); either fix the failing
upstream or relax `startup.mode` to `degraded` / lower
`min_healthy_upstreams` if partial startup is acceptable.

---

See also the [Troubleshooting guide](troubleshooting.md) for symptom-first
debugging and the [Stability page](stability.md) for the compatibility policy
that codes participate in.

---

<!-- FILE: docs/agent-context/architecture.md -->

# Architecture Guidance

> Deeper architectural detail lives in [docs/architecture.md](../architecture.md).
> This file covers non-obvious design decisions relevant to change-scoping.

## Architectural Intent

contextweaver separates **what to show the LLM** (Context Engine) from **which tools to offer** (Routing Engine). These engines share types and stores but have intentionally different execution models:

- **Context Engine** — async-first. Deals with I/O-bound operations (event log queries, artifact storage).
- **Routing Engine** — sync-only. Pure computation (DAG traversal, beam search). No I/O.

This boundary is intentional. Do not propose making routing async "for consistency" — it adds complexity for zero benefit.

## Non-Goals

contextweaver is **not** an LLM inference layer and **not** a tool execution runtime. It prepares context and routes tools but never calls models or executes tools. Feature proposals that cross these boundaries are out of scope.

## Major Boundaries

### Context Pipeline (8 stages)

The pipeline is a fixed sequence. Each stage has a single responsibility:

1. **generate_candidates** — pulls events from stores into a candidate pool
2. **dependency_closure** — ensures parent items (via `parent_id`) are included alongside their children
3. **sensitivity_filter** — drops or redacts items above the sensitivity floor
4. **apply_firewall** — summarises large outputs, stores raw data as artifacts
5. **score_candidates** — ranks candidates by recency, tag match, kind priority, token cost
6. **deduplicate_candidates** — removes near-duplicates via Jaccard similarity
7. **select_and_pack** — greedily packs highest-scoring candidates into the phase token budget
8. **render_context** — assembles the final prompt with BuildStats metadata

**Why this order matters:**
- Dependency closure must happen before scoring, otherwise parents could be discarded before their children pull them in.
- Sensitivity filtering before the firewall prevents sensitive data from reaching the summarizer.
- Scoring after the firewall ensures scores reflect the summarised (not raw) content.

### Routing Pipeline (4 stages)

1. **Catalog** — register and manage `SelectableItem` objects
2. **TreeBuilder** — convert flat items into a bounded `ChoiceGraph` DAG
3. **Router** — beam-search over the graph for top-k relevant items
4. **ChoiceCards** — render compact LLM-friendly cards (never full schemas)

### Store Layer

All stores use `typing.Protocol` interfaces with in-memory defaults. This enables custom backends (database, Redis, etc.) without changing pipeline code.

- **EventLog** — append-only. The audit trail.
- **ArtifactStore** — raw tool outputs stored by the firewall. Supports drilldown via `ViewRegistry`.
- **EpisodicStore** — short episodic memory entries.
- **FactStore** — key-value facts persisted across turns.

Any backend can prove it honours these protocols with the shipped conformance
kit (`contextweaver.store.testing`, issue #520): each `check_*_conformance`
function takes a factory for an empty backend and asserts the round-trip,
ordering, and not-found semantics the Context Engine relies on. For the
`ArtifactStore` it also asserts that `put()` stamps a sha256 `content_hash`
on the returned ref — the firewall's content-addressed idempotency
short-circuit (#190) depends on it, so it is a protocol contract, not a
backend detail. The bundled in-memory, JSON-file, and SQLite backends are all
run through it in `tests/test_store_conformance.py`.

#### Thread-safety contract (issue #458)

The store protocols make **no concurrency guarantee** in their interface; each
backend documents its own. The bundled backends:

- **`InMemory*` stores** are *not* thread-safe. They are for single-threaded
  use and tests; guard them with your own lock for concurrent access.
- **`JsonFileArtifactStore`** is single-process. Within one process it is
  thread-safe: `put` / `delete` / `list_refs` on a shared instance are
  serialised by an internal lock, and each individual file write is **atomic**
  (temp file + `os.replace`), so a reader never observes a torn or truncated
  artifact and a crash mid-write leaves the previous version intact. There is
  no cross-process advisory locking, so two processes writing the same
  `base_dir` are still unsupported.
- **`SqliteEventLog`** opens its connection in WAL mode for single-process use;
  it is not shared across threads.

The gateway runtime (`ProxyRuntime`) inherits these guarantees through the
store it is given: its read-only `tool_view` (drilldown) is safe to call
concurrently against a `JsonFileArtifactStore`. A gateway that fans out to
real concurrent clients should pick (or wrap) a backend whose contract matches
its load — this is exactly what the protocol seam and conformance kit are for.

### Sensitivity Enforcement

`context/sensitivity.py` is security-grade code. It enforces data classification (`public` → `restricted`) with two actions: drop or redact. The `MaskRedactionHook` is the built-in redactor. Changes to this module require extra review scrutiny — never weaken defaults.

### Progressive Disclosure (ViewRegistry)

`context/views.py` provides a `ViewRegistry` that maps content-type patterns to view generators. When the firewall stores a large tool output as an artifact, the view system generates alternative representations (JSON subset, CSV summary, etc.) the agent can drilldown into without retrieving the full blob. `drilldown_tool_spec()` exposes drilldown as an agent-callable tool.

## Key Tradeoffs

| Decision | Tradeoff | Consequence of reversing |
|---|---|---|
| Protocol-based stores | More files and indirection | Allows backend swaps without pipeline changes |
| `to_dict()`/`from_dict()` + `serde.py` | Two serialization paths | Per-class methods handle class-specific logic; `serde.py` handles shared primitives. Consolidating loses encapsulation. |
| Sync routing / async context | Two calling conventions | Routing has no I/O — async would add overhead for zero benefit |
| 8-stage pipeline | Pipeline is long | Each stage has a single well-defined responsibility. Merging stages creates coupling. |
| ChoiceCards never include schemas | Limits LLM tool-call generation | Keeps routing focused on *which* tool, not *how* — schema is provided at call-time via hydration |

## Structural Mental Model

Think of contextweaver as three layers:

1. **Data layer** (`types.py`, `envelope.py`, `config.py`, `serde.py`, `exceptions.py`) — pure data, no I/O, no side effects.
2. **Store layer** (`store/`, `protocols.py`) — stateful but simple append-only/read interfaces.
3. **Pipeline layer** (`context/`, `routing/`, `summarize/`) — orchestration logic that reads from stores and produces output types.

Adapters (`adapters/`) convert external formats (MCP, FastMCP, A2A) into contextweaver types at the boundary.

Changes should flow within a layer. Cross-layer changes (e.g., adding I/O to the data layer) are red flags.

---

<!-- FILE: docs/agent-context/invariants.md -->

# Invariants

This file separates **mechanically enforced invariants** from **review policy**.
A hard invariant must name the test/gate that enforces it. If a constraint has
no mechanical check, it is review policy and must not be described as if CI
proves it automatically.

## Mechanically enforced invariants

### Canonical MCP `tool_id` round-trip

MCP-facing adapter surfaces that opt into the canonical gateway ID contract
must emit IDs that round-trip through `parse_tool_id` / `format_tool_id`.
The canonical contract carries namespace, name, optional version, and schema
hash identity; the legacy hand-formatted `mcp:{name}` form is not valid for
those surfaces.

Framework adapters are a separate, deliberately loose identifier class. IDs
such as `crewai:{name}`, `langchain:{name}`, or `openapi:{name}` identify an
adapter-local capability and are **not** required to parse as canonical MCP
`tool_id` values unless that adapter explicitly adopts the canonical contract.
Do not silently broaden one ID class into the other.

**Enforcement:** `tests/test_tool_id.py` plus the MCP adapter/gateway tests.

### Module-size ratchet

Ordinary source modules must remain at or below **500 lines**. Named exemptions
remain explicit; pre-existing modules above 500 are frozen at their recorded
ceiling and may shrink but not grow past it.

The limit is a maintainability signal, not a decomposition target. A contributor
must not create a grab-bag helper or duplicate logic merely to move a cohesive
module below the limit. When a split is needed, split along a real responsibility
or dependency boundary.

**Enforcement:** `make module-size-check` / `scripts/check_module_size.py` and
`scripts/module_size_baseline.json`.

### Public API manifest

Changes to the supported public package surface must be intentional and keep the
checked-in API manifest in sync.

**Enforcement:** `make api-check` through `make drift-check` / `make ci`.

### Generated-artifact consistency

Generated schemas, API manifests, scorecards, and other registered generated
artifacts must match their canonical inputs.

**Enforcement:** `make drift-check` through `make ci`, with the individual
registered `*-check` commands documented in `AGENTS.md`.

### D1 evidence-record consistency

Persisted D1 evaluation records must remain internally consistent. Funnel flags
are monotonic; first-success timing and real-source fields must agree with the
setup/success stages; outcome semantics must agree with the funnel; and negative
outcomes (`declined`, `setup_failed`, `removed`) require a structured
`dropoff_reason` other than the sentinel `"none"`, while non-negative outcomes
must use `"none"`.

The reporter rejects inconsistent records rather than repairing or normalizing
them, because silently changing evidence would corrupt the survival experiment.

**Enforcement:** `tests/test_onboarding_report.py` exercises
`scripts/onboarding_report.py` and runs in the normal pytest / `make ci` test
gate.

## Must-preserve contracts

The constraints below are product/architecture contracts. Each names the tests
or gate that exercises the behavior where a focused mechanical gate exists;
otherwise it is explicitly marked **review policy**.

### Minimal core dependencies

Core dependencies (`pyproject.toml` `dependencies`) must stay small, audited,
and broadly useful to the primary surfaces. Heavy or runtime-specific packages
belong under optional dependencies unless a maintainer explicitly accepts the
core cost.

**Enforcement:** dependency-resolution/floor CI jobs exercise compatibility;
whether a new dependency is justified is **review policy**.

### Sync core, bridged public execution

The Context Engine's core build/selection pipeline is synchronous computation.
Public/runtime integration may expose synchronous and asynchronous entry points
and store bridges where I/O requires them. Routing remains synchronous pure
computation. Do not force core context logic into `async` merely to satisfy a
naming convention, and do not remove the supported async-store bridge contracts.

**Enforcement:** context manager/build tests plus async-store bridge tests;
architectural placement is **review policy**.

### Context pipeline ordering

The context build stages remain:

1. `generate_candidates`
2. `dependency_closure`
3. `sensitivity_filter`
4. `apply_firewall`
5. `score_candidates`
6. `deduplicate_candidates`
7. `select_and_pack`
8. `render_context`

Reordering can change correctness and security semantics.

**Enforcement:** context build/manager test suites exercise the pipeline output;
the exact architectural stage boundary is **review policy** unless a dedicated
ordering test is added.

### Dependency closure

If a selected `ContextItem` depends on a parent through `parent_id`, the build
must preserve the required parent relationship so a tool result is not surfaced
without the context needed to understand it.

**Enforcement:** context selection/build tests.

### Append-only event log

The event log is append-only through the store protocol. Callers must not mutate
stored history behind the protocol.

**Enforcement:** store protocol/conformance tests; direct source-level mutation
avoidance is **review policy**.

### Determinism

Core pipelines must be deterministic for identical inputs/configuration: stable
ordering, stable tie-breaking, and no implicit randomness.

**Enforcement:** deterministic/golden/benchmark tests across the context and
routing suites; introducing a new nondeterministic dependency is **review
policy** unless separately gated.

### `ContextManager` public surface, not its mixin layout

`ContextManager`'s supported public method surface and behavior are the contract.
The current `_IngestMixin` / `_BuildMixin` / `_RoutingMixin` decomposition is a
private implementation detail and may be simplified, recomposed, or replaced as
long as the public contract remains compatible.

Do **not** preserve a mixin merely because issue #101 once used it to satisfy an
old line-count target.

**Enforcement:** public API manifest/gates plus `ContextManager` behavioral tests.
Private class composition is **not** an invariant.

### Layer direction: core must not depend on adapters

`adapters/` owns external/protocol integration surfaces. Core modules under
`context/`, `routing/`, stores, data/config, and other provider-neutral layers
must not import `contextweaver.adapters` to obtain implementation logic.

When core needs a protocol-specific pure transform (for example MCP result →
`ResultEnvelope` shaping), place that transform in an appropriate neutral/core
module and let the adapter re-export or call it for compatibility. Lazy imports
are not an acceptable way to hide a `core ↔ adapters` dependency cycle.

**Enforcement:** currently **review policy** pending the import-linter work in
#648. Changes fixing #752 should add a focused regression check where practical.

### Sensitivity defaults

The default sensitivity floor (`confidential`) and action (`drop`) are
conservative. Do not weaken those defaults without explicit security review.

**Enforcement:** configuration/sensitivity tests exercise the defaults; changing
the security posture still requires **review policy**.

### Data-layer purity

`types.py`, `envelope.py`, `config.py`, `serde.py`, and `exceptions.py` are data
and serialization layers: no network/filesystem I/O and no runtime orchestration.

**Enforcement:** currently **review policy**.

### ChoiceCard schema hiding

`ChoiceCard` carries whether a schema exists but not the full input/output schema.
Full schemas are hydrated only at the appropriate boundary. This preserves the
bounded-context property of browse/routing surfaces.

**Enforcement:** gateway/routing ChoiceCard and hydration tests.

### Serialization design

<a name="serialization-design"></a>

`serde.py` contains shared primitives; class-specific `to_dict()` / `from_dict()`
methods retain class-specific serialization semantics. Do not replace the latter
with indiscriminate `dataclasses.asdict()`.

**Enforcement:** serialization round-trip/golden tests; exact placement remains
**review policy**.

### Store protocols remain structural seams

Do not collapse store protocols into the bundled concrete backends. Structural
protocols are what allow user-provided and remote/persistent implementations.

**Enforcement:** store conformance tests exercise bundled implementations;
preserving the abstraction boundary is **review policy**.

## Safe vs unsafe simplifications

| Change | Safe? | Why |
|---|---|---|
| Replace ContextManager mixins while preserving public methods/behavior | Usually safe | The public contract is the invariant; the mixin layout is private |
| Fold a helper back into its natural parent when cohesion improves and the module stays ≤500 | Usually safe | Avoids artificial fragmentation and duplicated hardening paths |
| Add a field to an existing dataclass | Usually safe | Follow `to_dict`/`from_dict`, compatibility, schema/API gates |
| Merge two semantically distinct pipeline stages | **Unsafe by default** | Can change ordering, auditability, and security behavior |
| Replace structural store protocols with concrete classes | **Unsafe** | Removes backend extensibility |
| Duplicate `_utils.py` similarity logic in a caller | **Unsafe** | Creates diverging implementations |
| Put full schemas on `ChoiceCard` | **Unsafe** | Regresses bounded context cost |
| Add `context -> adapters` lazy imports | **Unsafe** | Hides a layer cycle rather than fixing it |

## Cross-cutting review policy

- `from __future__ import annotations` in source files unless a supported Python
  constraint gives a reviewed reason otherwise.
- Google-style docstrings on public classes/functions.
- Type hints on public functions/methods.
- Project exceptions from `contextweaver.exceptions` rather than ad-hoc public
  exception types.
- Reserved `metadata['_contextweaver']` namespace remains owned by the
  weaver-spec adapter contract; caller input must not be silently clobbered.

These are checked by combinations of lint/type/tests/review rather than one
single invariant gate; do not describe them as individually mechanically proved
unless a dedicated check is added.

## Update triggers

Update this file when:

- a new hard invariant gets a concrete enforcing test/gate;
- a previously mechanical invariant loses its enforcement;
- a review-policy constraint becomes mechanically enforced;
- a forbidden shortcut is discovered through a real failure;
- a safe/unsafe determination changes due to architectural evolution;
- the module-size policy or another cross-cutting architecture rule changes.

Every update must keep the stated enforcement status truthful. Documentation is
not evidence that an invariant is enforced.

---

<!-- FILE: docs/agent-context/workflows.md -->

# Workflows

## Authoritative Commands

```bash
make fmt      # ruff format src/ tests/ examples/ scripts/
make lint     # ruff check src/ tests/ examples/ scripts/
make type     # mypy src/ examples/ scripts/  (examples + scripts gated too, #539)
make test     # pytest --cov=contextweaver --cov-report=term-missing -q
make example  # run all example scripts (includes architectures)
make architectures  # run reference architecture scripts under examples/architectures/
make demo     # python -m contextweaver demo
make ci       # fmt + lint + type + test + drift-check + module-size-check + doc-snippets-check + readme-version-check + security-policy-check + example + demo
make ci-full  # make ci + floor-deps + tool-smoke (the two isolated-env CI jobs; #710)
make floor-deps # prove declared dependency floors resolve + pass tests (mirrors the floor-deps CI job; needs uv; #710)
make tool-smoke # build the wheel + run the entry point via uvx/pipx (mirrors the Linux tool-run-smoke CI job; needs uv/pipx; #710)
make drift-check  # one gate over every generated-artifact drift check (#522; in `make ci`)
make module-size-check  # enforce the ≤500-line convention, frozen baseline (#456/#853; in `make ci`)
make doc-snippets-check # execute README + curated docs Python snippets (#526; in `make ci`)
make docs     # mkdocs build --clean (docs site — not part of CI)
make docs-serve  # mkdocs serve (live preview)
make benchmark        # run benchmark harness (non-gating; writes benchmarks/results/latest.json)
make benchmark-matrix # benchmark + per-backend × per-size matrix (#208) and per-namespace breakdown (#209)
make scorecard        # render benchmarks/scorecard.md from benchmarks/results/latest.json
make scorecard-check  # verify scorecard.md is up to date (gating CI step; exits non-zero on drift)
make sweep-scoring    # weight sweep for ScoringConfig (#214); writes benchmarks/sweep_scoring.md
make context-rot       # render context-rot demo JSON + docs/assets/context_rot.svg (#349)
make context-rot-check # verify context_rot.svg matches its committed JSON (gating CI step; exits non-zero on drift)
make readme-version-check  # verify README package/comparison/roadmap refs and Python classifiers match sources (gating CI step; #347/#473/#531)
make security-policy-check # verify SECURITY.md supported series + relative links match pyproject.toml (gating CI step; #691)
make llms        # regenerate llms.txt and llms-full.txt from canonical docs
make llms-check  # verify llms.txt and llms-full.txt are up to date (gating CI step; exits non-zero on drift)
make gateway-scorecard-check  # verify gateway scorecard Markdown matches its committed JSON (gating CI step)
make record-demos-check  # verify committed asciinema casts match demo output (gating CI step)
make smoke-eval  # deterministic, credential-free smoke evaluation (gating CI step on Python 3.12)
make weaver-conformance  # round-trip + JSON-Schema validate the weaver-spec adapter
                         # (fetches schemas from raw.githubusercontent.com; CI runs it as a gate)
```

> As of #474, `make ci` now mirrors the gating CI checks a contributor can run
> offline: the consolidated generated-artifact drift gate `make drift-check`
> (#522 — schemas, scorecards, recorded demos, llms.txt, context-rot SVG, and
> the public-API manifest #518), plus `make module-size-check` (#456),
> `make doc-snippets-check` (#526), `make readme-version-check` (#347/#473/#531),
> and `make security-policy-check` (#691).
> The individual `*-check` targets still exist for granular use, but you no
> longer need to remember to run them separately before a PR.
>
> CI-only gates (not in `make ci` because they need the network or are heavy):
> `make weaver-conformance` (fetches schemas) and the docs build job.
> `make smoke-eval` (#392/#491) is gating on the Python 3.12 CI cell;
> the raw benchmark job remains informational.
> The two gating CI *jobs* that build isolated environments now have local
> equivalents — `make floor-deps` and `make tool-smoke` (#710), bundled into
> `make ci-full`; only the macOS cell of `tool-run-smoke` stays CI-only.
>
> All targets invoke `$(PYTHON)` (defaults to `python3`); override with
> `make <target> PYTHON=python3.11` where no bare `python` exists (#712).
> That now includes the tools themselves — `ruff`, `mypy` and `mkdocs` run as
> `$(PYTHON) -m <tool>`, never bare. Until this was fixed, `make type` ran
> whatever `mypy` PATH resolved first: in one container that was an unrelated
> `uv`-tool install reporting *18 errors in 11 files*, while the project's own
> mypy reported *Success: no issues found in 314 source files*. A gate that
> reports a fact about its environment is worse than no gate, so
> `tests/test_makefile_interpreter.py` pins the convention.

`make ci` runs all declared targets in sequence. The CI-only gates above run on
every PR regardless; run them locally when the affected integrations change.

## Command-Selection Rules

| Goal | Command |
|---|---|
| Quick format check | `make fmt` |
| Quick lint check | `make lint` |
| Full validation | `make ci` (always — do not skip targets) |
| Full validation + isolated-env jobs | `make ci-full` (#710; adds floor-deps + tool-smoke) |
| Verify dependency floors resolve | `make floor-deps` (#710; needs uv) |
| Wheel / entry-point smoke | `make tool-smoke` (#710; needs uv/pipx) |
| Run a single test | `pytest tests/test_<module>.py` or `pytest -k "test_name"` |
| Run all tests | `make test` |
| Verify examples work | `make example` |
| Interactive demo | `make demo` |
| Verify recorded demo casts | `make record-demos-check` |
| Build docs site | `make docs` |
| Live docs preview | `make docs-serve` |
| Run benchmark harness | `make benchmark` (non-gating; writes `benchmarks/results/latest.json`) |
| Run full per-backend × per-size matrix | `make benchmark-matrix` (#208 + #209) |
| Verify gateway benchmark scorecard | `make gateway-scorecard-check` |
| Run deterministic smoke evaluation | `make smoke-eval` (gating on the Python 3.12 CI cell) |
| Run scoring-weight sweep | `make sweep-scoring` (#214; writes `benchmarks/sweep_scoring.md`) |
| Add an eval for a feature | follow [`.github/prompts/add-eval.prompt.md`](../../.github/prompts/add-eval.prompt.md) (#216) |
| Regenerate llms.txt / llms-full.txt | `make llms` (after editing canonical docs) |
| Check llms.txt / llms-full.txt for drift | `make llms-check` (exits non-zero if regeneration needed) |

**Do not** use `make test` alone as a validation gate. Always run `make ci` before declaring a change complete — it includes example and demo verification that catch integration issues `make test` misses.

## Setup (one-time)

```bash
pip install -e ".[dev]"
pre-commit install
```

Pre-commit hooks run `ruff format`, `ruff check --fix`, and file hygiene checks on every commit. Hooks may modify files — re-stage with `git add` if needed.

## Preparing a Release

Before creating the release tag, refresh and commit the deterministic benchmark
history for the package version:

```bash
make benchmark
python scripts/render_trend.py --snapshot <version> --from benchmarks/results/latest.json
make trend
make trend-check
```

The publish workflow verifies that `benchmarks/results/history/<version>.json`
exists and that `benchmarks/trend.md` matches the committed history. Latency is
excluded from release snapshots because it is host-dependent.
## Adding a Feature

1. Identify the relevant module (see module map in [AGENTS.md](../../AGENTS.md)).
2. Modify only the targeted module.
3. Update `protocols.py` if adding a new protocol.
4. Add tests in `tests/test_<module>.py`.
5. Run `make ci` — all declared targets must pass.
6. Update `CHANGELOG.md` under `## [Unreleased]`.
7. Add Google-style docstrings to any new public APIs.
8. Update examples/demos if the feature is user-facing.
9. Update agent-facing docs if the pipeline or public API changed.
10. **If the feature can move `recall@k` / `dropped` / `dedup_removed` /
    `prompt_tokens`**, follow
    [`.github/prompts/add-eval.prompt.md`](../../.github/prompts/add-eval.prompt.md)
    (#216) to extend the gold set or scenarios and regenerate
    `benchmarks/scorecard.md`. The sticky CI benchmark-delta comment
    (#211) surfaces any matrix-cell ⚠️ markers on the PR.

## Definition of Done

A change is complete when **all** of the following are true:

- [ ] `make ci` passes (all declared targets)
- [ ] `CHANGELOG.md` updated
- [ ] Google-style docstrings on all new public APIs
- [ ] Type hints on all new public functions and methods
- [ ] Tests added for new functionality
- [ ] Examples/demos updated if the feature is user-facing
- [ ] Agent-facing docs updated if pipeline, public API, or conventions changed

## Fixing a Bug

1. Write a failing test that reproduces the bug.
2. Fix the bug.
3. Run `make ci`.
4. Update `CHANGELOG.md`.
5. If the bug revealed a reusable lesson, record it per the process in [lessons-learned.md](lessons-learned.md).

## Adding a Store Backend

1. Implement the store class in `src/contextweaver/store/<name>.py`.
2. The class must implement the relevant protocol from `protocols.py`.
3. Export from `src/contextweaver/store/__init__.py`.
4. Add tests in `tests/test_store_<name>.py`.
5. Update `StoreBundle` if appropriate.

## Adding an Adapter

1. Create the adapter in `src/contextweaver/adapters/<protocol>.py`.
2. Pure stateless converter — no state, no core-type leakage.
3. External format dependencies stay at the adapter boundary.
4. Add tests in `tests/test_adapters.py`.
5. Add an example in `examples/`.

## Deprecating an API

Use the runtime machinery in `src/contextweaver/_deprecation.py` (issue #517);
do not call `warnings.warn` ad hoc.

1. Add a `Deprecation(name, since, removal, instead)` entry to the `_SHIMS`
   table in `_deprecation.py` (the single source of truth). Use the next minor
   for `since` and `"1.0.0"` for `removal` unless decided otherwise.
2. At the call site, emit the warning: `warn_deprecated("<name>")` inside a
   property/method/branch, or decorate a callable with
   `@deprecated("<name>", since=..., removal=..., instead=...)`. Keep behavior
   identical; route in-library callers through a private helper so the canonical
   path does not trip the warning.
3. Migrate every in-repo caller (src, examples, docs snippets, tests) off the
   deprecated surface; intentional shim tests assert the warning with
   `pytest.warns(DeprecationWarning)`. The `filterwarnings` gate in
   `pyproject.toml` escalates first-party deprecations to errors, so a leftover
   caller fails CI.
4. Add the surface to the inventory table in `docs/upgrading.md` and a
   `CHANGELOG.md` "Deprecated" entry naming the replacement.

**Documentation-only exception.** Do **not** add a runtime warning when the only
call site would live in a module barred from side effects — a re-export-only
`__init__.py` (hard rule: only re-exports) or a pure-data module such as
`types.py` (invariant: no side effects in the data layer), or an internal
serialization key on a hot path. Keep the surface a plain alias/accessor, skip
steps 1–2 (no `_SHIMS` entry, no `warn_deprecated`), and record it as a
documentation-only row in `docs/upgrading.md`. The `ToolCard` alias is the
reference example.

## Documentation Governance

### When docs must be updated

- Any PR that changes the context pipeline stages, routing pipeline, or public API.
- Any PR that adds, removes, or renames a module.
- Any PR that changes project conventions, commands, or the definition of done.

### Who triggers updates

- The author of the PR is responsible for updating docs in the same PR.
- Reviewers should check the [review checklist](review-checklist.md) for doc-update requirements.

### Resolving contradictions

If two docs disagree:
1. `AGENTS.md` is authoritative for agent guidance and shared rules.
2. `docs/architecture.md` is authoritative for architecture detail.
3. `Makefile` is ground truth for command definitions.
4. Source code is ground truth for implementation details.

Fix the less-authoritative source to match.

### Promoting lessons into canonical docs

When a lesson from [lessons-learned.md](lessons-learned.md) represents a durable pattern:
1. Add it to the appropriate canonical doc (`AGENTS.md`, `invariants.md`, or `workflows.md`).
2. Keep the lesson entry but mark it as promoted with a cross-reference.

### Avoiding duplicate authority

Each piece of guidance should have exactly one canonical home. Use cross-references instead of copies. Exception: hard rules (the 2 auto-reject items) may be briefly restated in tool-specific override files for visibility, since those files may be the only context an agent loads.

### Updating navigation tables

When adding a new canonical doc under `docs/agent-context/`, update the navigation tables in all three routing files:
- `AGENTS.md` (Documentation Map)
- `.github/copilot-instructions.md` (Canonical References)
- `.claude/CLAUDE.md` (Canonical References)

---

<!-- FILE: docs/agent-context/lessons-learned.md -->

# Lessons Learned

This is not an incident archive. It captures reusable patterns from past mistakes
and defines the process for converting incidents into durable guidance.

## Failure-Capture Workflow

When a bad change is caught in review or causes a regression:

1. **Identify the root cause** — was it a missing rule, a misunderstood boundary, or a documentation gap?
2. **Determine if it's reusable** — would a different agent make the same mistake on a different change? If yes, it's a lesson. If no, it's a one-off incident — don't record it here.
3. **Generalize the lesson** — write it as a pattern, not a narrative. "Don't do X because Y" is better than "on date Z, agent A did X to file B."
4. **Choose the right home:**
   - If it's a hard constraint → add to [invariants.md](invariants.md)
   - If it's a workflow fix → add to [workflows.md](workflows.md)
   - If it's an architectural insight → add to [architecture.md](architecture.md)
   - If it's a recurring trap that doesn't fit elsewhere → add to this file
5. **If promoted**, keep the entry here but mark it: "**Promoted →** [target file](link)."

## What Belongs Here

- Recurring mistakes that agents make across different changes
- Generalized lessons with clear "do this instead" guidance
- Patterns where the obvious approach is wrong

## What Does Not Belong Here

- One-off incidents tied to specific dates, PRs, or files
- Narrative history of past bugs
- Lessons that have been fully captured in invariants, workflows, or architecture docs (mark as promoted instead)

## Durable Lessons

### 1. Pipeline stage count drift

**Mistake:** Docs described a 7-stage context pipeline; the actual implementation has 8 stages (missing `dependency_closure`).

**Lesson:** When modifying pipeline documentation, always verify stage count and order against the source code (`context/manager.py`). Do not copy pipeline descriptions from other docs without verification.

**Generalized rule:** Treat pipeline stage documentation like API documentation — verify against implementation, not against other docs.

### 2. "Simplification" proposals that break design intent

**Mistake:** Proposing to merge `serde.py` with per-class `to_dict()`/`from_dict()`, or to collapse store protocols into concrete classes, or to make routing async for "consistency."

**Lesson:** Before proposing a simplification, check [invariants.md](invariants.md) for the "Things That Must Not Be Simplified" section. If the thing you want to simplify is listed, it exists for a reason. Read the rationale before proposing changes.

**Generalized rule:** Things that look redundant in this codebase often exist for extensibility or correctness. Check invariants before proposing consolidation.

### 3. Overstatement in documentation

**Mistake:** "Zero-dependency is a hard constraint" (overstated — extras are acceptable). "Always use X" for things that are strong patterns, not hard rules.

**Lesson:** Distinguish hard rules (auto-reject, 2 items) from strong patterns (recommended, judgment applies). Overstated rules cause agents to either (a) reject valid changes or (b) ignore all rules after discovering false mandates.

**Generalized rule:** Use precise language in constraints. "Must" and "always" should be reserved for actual invariants. Use "prefer" or "strongly recommended" for patterns.

### 4. Module map staleness

**Mistake:** `envelope.py` added in an early version but never added to the module map in agent-facing docs. Agents couldn't find `ResultEnvelope`, `BuildStats`, etc.

**Lesson:** When adding a new module, update the module map in `AGENTS.md` in the same PR.

**Generalized rule:** Treat the module map as part of the public API surface. New modules require map updates just like new functions require docstrings.

### 5. `make ci` composition drift

**Mistake:** `AGENTS.md` described `make ci` with a stale fixed target count and
omitted a target that the Makefile runs.

**Lesson:** Do not describe command composition from memory. Check the `Makefile` for ground truth.

**Generalized rule:** For command documentation, the build system file (`Makefile`, `pyproject.toml`) is always ground truth.

## Update Triggers

Record a new lesson when:
- A review catches a mistake that a well-documented rule would have prevented.
- The same category of mistake recurs across multiple changes or agents.
- A documentation gap directly causes a bad change.

Do not record lessons for:
- Typos, formatting issues, or trivial errors.
- One-off issues that are unlikely to recur.

---

<!-- FILE: docs/agent-context/review-checklist.md -->

# Review Checklist

Use this checklist for both agent self-check (before proposing changes) and
maintainer review (when reviewing PRs). Items are grouped by category.

## Validation

- [ ] `make ci` passes (fmt, lint, type, test, schemas-check, example, demo)
- [ ] No new warnings introduced

## Hard Rules

- [ ] No `print()` in library code (exempt: `__main__.py`)
- [ ] No business logic in `__init__.py` (only re-exports)

## Code Quality

- [ ] Type hints on all new public functions and methods
- [ ] Google-style docstrings on all new public classes and functions
- [ ] `from __future__ import annotations` in any new or modified file
- [ ] Exceptions use custom types from `contextweaver.exceptions`
- [ ] New modules ≤ 300 lines (exempt: `types.py`, `envelope.py`, `__main__.py`)
- [ ] 100-character line length respected

## Testing

- [ ] Tests added for new functionality in `tests/test_<module>.py`
- [ ] Async tests use `pytest.mark.asyncio`
- [ ] No mocking of internal modules — uses real in-memory implementations

## Architectural Consistency

- [ ] No runtime dependencies added to core (`install_requires` stays empty)
- [ ] `context/` code is async-first with `_sync` wrappers
- [ ] `routing/` code is sync-only
- [ ] Store changes implement the relevant protocol from `protocols.py`
- [ ] Adapter changes are pure stateless converters
- [ ] No text similarity logic duplicated outside `_utils.py`
- [ ] `to_dict()` / `from_dict()` added to any new dataclass
- [ ] Event log mutations only via `append()`

## Pipeline Integrity

- [ ] Context pipeline stage order preserved (8 stages — see [invariants](invariants.md))
- [ ] Dependency closure not bypassed or weakened
- [ ] Sensitivity defaults not weakened
- [ ] Changes to `context/sensitivity.py` received extra security scrutiny

## Documentation

- [ ] `CHANGELOG.md` updated under `## [Unreleased]`
- [ ] Module map in `AGENTS.md` updated if modules were added/removed/renamed
- [ ] Agent-facing docs updated if pipeline, API, or conventions changed
- [ ] Examples/demos updated if feature is user-facing
- [ ] No contradictions introduced between `AGENTS.md` and supporting docs

## Cross-File Consistency

- [ ] Pipeline stage count/order matches across all docs and code
- [ ] Command descriptions match `Makefile`
- [ ] Module map matches filesystem
- [ ] Convention changes reflected in both `AGENTS.md` and `CONTRIBUTING.md`

## Invariant Spot-Checks

If the change touches any of these areas, verify the corresponding invariant:

| Area touched | Verify |
|---|---|
| Pipeline stages | 8-stage order preserved, dependency closure intact |
| `sensitivity.py` | Defaults not weakened, security review done |
| Store protocols | Protocol interface unchanged or backward-compatible |
| `serde.py` or `to_dict`/`from_dict` | Both mechanisms still in use, not consolidated |
| `_utils.py` | No similarity logic duplicated elsewhere |
| `__init__.py` files | Only re-exports, no logic |
| `envelope.py` or `types.py` | No I/O added to data layer |

## Update Triggers

Update this checklist when:
- New hard rules or invariants are established.
- New review gates are identified from recurring review feedback.
- Definition of done changes (sync with [workflows.md](workflows.md)).

---

<!-- FILE: docs/guide_agent_loop.md -->

# Agent Runtime Loop Guide

This guide explains the reference runtime loop in
`examples/full_agent_loop.py`.

The loop demonstrates all four context phases in one deterministic flow:

1. `route` - shortlist candidate tools.
2. `call` - inject only the selected tool schema.
3. `interpret` - summarize tool output via the firewall.
4. `answer` - compose the final response context.

## Flow Diagram

```mermaid
flowchart TD
    U[User Query] --> R[Phase.route\nContextManager.build_route_prompt_sync]
    R --> C[ChoiceCards + routed candidate IDs]
    C --> M[Model selects tool_id]
    M --> H[Catalog.hydrate(tool_id)]
    H --> P[Phase.call\nContextManager.build_call_prompt_sync]
    P --> X[Simulated tool execution]
    X --> F[ingest_tool_result_sync\nFirewall stores raw artifact + summary]
    F --> I[Phase.interpret\nContextManager.build_sync]
    I --> A[Phase.answer\nContextManager.build_sync]
```

## Pseudo-code

```python
catalog = build_catalog_with_schemas()
router = Router(TreeBuilder().build(catalog.all()), items=catalog.all())
manager = ContextManager(budget=ContextBudget(route=500, call=800, interpret=600, answer=1000))

manager.ingest(user_turn)

route_pack, cards, route_result = manager.build_route_prompt_sync(goal, query, router)
tool_id = model_select(route_result.candidate_ids)

manager.ingest(tool_call)
call_pack = manager.build_call_prompt_sync(tool_id, query, catalog)

raw_result = simulate_large_json()
manager.ingest_tool_result_sync(tool_call_id, raw_result)
interpret_pack = manager.build_sync(phase=Phase.interpret, query=query)

answer_pack = manager.build_sync(phase=Phase.answer, query=final_query)
```

## Module Pointers

- `src/contextweaver/context/manager.py`
  - `build_route_prompt_sync()` for route-phase prompt + ChoiceCards.
  - `build_call_prompt_sync()` for selected-schema call prompts.
  - `ingest_tool_result_sync()` for firewall interception and envelope creation.
  - `build_sync()` for `interpret` and `answer` context compilation.
- `src/contextweaver/routing/router.py`
  - `Router.route()` to rank candidate tools.
- `src/contextweaver/routing/catalog.py`
  - `Catalog` and `Catalog.hydrate()` for schema hydration.
- `src/contextweaver/routing/cards.py`
  - Choice-card rendering used in route prompts.
- `src/contextweaver/config.py`
  - `ContextBudget` with per-phase token limits.

## When To Use Each Phase

| Phase | Primary goal | Typical contents | Budget posture |
| --- | --- | --- | --- |
| `route` | Choose tools | user intent, policy, compact cards | small |
| `call` | Generate arguments | selected tool schema, examples, constraints | medium |
| `interpret` | Understand result | tool call + summarized result + extracted facts | medium |
| `answer` | Compose final response | relevant history + interpreted findings + policy | largest |

## Running The Example

```bash
python examples/full_agent_loop.py
```

What you should see:

1. Four phase sections (`route`, `call`, `interpret`, `answer`).
2. Compiled prompt text for each phase.
3. BuildStats output for each phase, including token counts.
4. Firewall behavior in `interpret` (raw payload size > summarized text size).

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.