agentleFS
Sign inSign up

mcp-data-platform

txn2/mcp-data-platform/docs/llms-full.txt

MCP server for AI-assisted data exploration. DataHub semantic layer with optional Trino and S3, all three omittable. Cross-enrichment automatically enriches query results with business context: owners, tags, quality scores, deprecation warnings. Implements fail-closed security with OIDC/API key authentication, TLS for HTTP transport, and prompt injection protection. For the security architecture rationale, see: https://imti.co/mcp-defense/ Document version: v1.120.0, revised 2026-08-29. This project releases often. If your copy of this page is older than the current release on https://github.com/txn2/mcp-data-platform/releases/latest, re-fetch it before grounding…

llms.txt9 starsChanged 55 days ago
  • Reads credentials
  • Deletes or force-pushes
  • Installs packages
# mcp-data-platform

> MCP server for AI-assisted data exploration. DataHub semantic layer with optional Trino and S3, all three omittable. Cross-enrichment automatically enriches query results with business context: owners, tags, quality scores, deprecation warnings. Implements fail-closed security with OIDC/API key authentication, TLS for HTTP transport, and prompt injection protection.

For the security architecture rationale, see: https://imti.co/mcp-defense/

**Document version: v1.120.0, revised 2026-08-29.** This project releases often. If your copy of this page is older than the current release on https://github.com/txn2/mcp-data-platform/releases/latest, re-fetch it before grounding an assessment on it; earlier revisions described a narrower system.

**How the documentation site is organized.** Three sections: `home`, `docs` (everything written about the platform -- install, the End User and Administrator references, the API reference, the Go library, evaluation, examples and support), and `portal` (a guided tour of every screen the built-in web portal serves, split into Your Work and Administration). The portal tour lives under /portal/; it replaced the two single-page guides that used to sit at /server/portal-user/ and /server/admin-portal/.

---

# Overview

Your AI assistant can run SQL. But it doesn't know that `cust_id` contains PII, that the table was deprecated last month, or who to ask when something breaks.

mcp-data-platform fixes that. It connects AI assistants to your data infrastructure and adds business context from your semantic layer. Query a table and get its meaning, owners, quality scores, and deprecation warnings in the same response.

Cross-enrichment works from DataHub (https://datahubproject.io/) as the semantic layer. Add Trino (https://trino.io/) for SQL queries and S3 for object storage when you're ready. All three are optional: with none configured the semantic, query, and storage providers resolve to noops, the server starts normally, and the database-backed surfaces (API and MCP gateways, knowledge, memory, portal, search/fetch) run on PostgreSQL alone. DataHub is an adapter behind a provider interface, not a substrate the platform is built on: omitting it costs cross-enrichment and the datahub_* tools, not the platform. semantic:, query:, storage:, and toolkits: are independent config blocks, so Trino and S3 stay available on their own terms and simply stop being enriched. See "Deployment Shapes" below.

## Key Features

- **Semantic-First**: DataHub is the foundation of cross-enrichment. Query a table, get its business context automatically: owners, tags, quality scores, deprecation warnings. No separate lookups.

- **Cross-Enrichment**: Trino results include DataHub metadata. DataHub searches show which datasets are queryable. Context flows between services automatically.

- **Enterprise Security**: Fail-closed authentication model, TLS enforcement for HTTP transport, prompt injection protection, and read-only mode enforcement.

- **Built for Customization**: Add custom toolkits, providers, and middleware. The Go library exposes everything. Build the data platform your organization needs.

- **Personas**: Define who can use which tools and, deny-by-default, which connections. The connection is the authorization boundary rather than the end user: one connection per credential and permission level, several of which may front the same downstream system. Map from your identity provider's roles.

- **Resource Templates**: Browse platform data as parameterized MCP resources using RFC 6570 URI templates. Three built-in templates: table schemas (`schema://catalog.schema/table`), glossary terms (`glossary://term`), and data availability (`availability://catalog.schema/table`).

- **Managed Resources**: Human-uploaded reference material (data, samples, playbooks, templates, references) surfaced directly to AI assistants via MCP `resources/list` and `resources/read`. Three visibility scopes: global, persona, and user. PostgreSQL metadata with S3 blob storage. REST API for CRUD, Admin Portal page for management. Auto-enabled when a database is available.

- **Progress Notifications**: Long-running Trino queries send granular progress updates (executing, formatting, complete) to clients that provide `_meta.progressToken`. Zero overhead when the client doesn't send a token. Enabled by default; set `progress.enabled: false` to opt out.

- **Client Logging**: Server-to-client log messages for platform decisions (enrichment, timing) via MCP `logging/setLevel` protocol. Zero overhead if the client hasn't opted in. Enabled by default; set `client_logging.enabled: false` to opt out.

- **Elicitation**: User confirmation prompts before expensive queries (EXPLAIN IO cost estimation) or PII access (sensitive column detection). Requires client-side elicitation support; gracefully degrades when unavailable. Enabled by default, including `cost_estimation` and `pii_consent`; set `enabled: false` at any level to opt out.

- **Icons**: Visual metadata for tools, resources, and prompts. Upstream toolkits provide default icons; deployers can override via configuration. Enabled by default; set `icons.enabled: false` to opt out.

- **Dynamic Prompts**: Three-tier prompt system: auto-registered `platform-overview` built from description and enabled toolkits, operator-configured prompts with `{placeholder}` argument substitution, and conditional workflow prompts (explore-available-data, create-interactive-dashboard, create-a-report, trace-data-lineage) registered when required toolkits are present. `server.builtin_prompts` is a per-prompt switch (map of prompt name to bool) suppressing a built-in workflow prompt; a name absent from the map registers as usual. Toolkits implement `PromptDescriber` to advertise their own prompts. Operator prompts override auto-registered ones by name. `platform_info` does not list the prompt library. It once did, carrying every prompt with its full body, which made the mandatory first call of a session grow with the library (#1586); the listing predates `show_prompts` and `manage_prompt use`, and a prompt is now reached by handle: `manage_prompt` command `use` resolves a stored name, display name, `mcp:prompt:<id>` reference or free text, command `list` browses, `show_prompts` opens the library for the human, and `fetch` on `mcp:prompt:<id>` returns one in full.

---

# Ecosystem

mcp-data-platform is the orchestration layer for a broader suite of open-source MCP servers designed to work together as a composable data platform. Each component can run standalone or be combined through mcp-data-platform for unified access with cross-enrichment, authentication, and personas.

## mcp-datahub (https://github.com/txn2/mcp-datahub/)

An MCP server for DataHub, the metadata catalog. Provides AI assistants with dataset search, schema exploration, lineage graphs, glossary terms, domains, tags, and ownership information. In the platform, DataHub serves as the semantic layer: every query result is enriched with business context from DataHub before being returned to the AI assistant.

## mcp-s3 (https://github.com/txn2/mcp-s3/)

An MCP server for Amazon S3, providing AI assistants with direct access to object storage. List buckets, browse prefixes, read objects, and generate presigned URLs. Supports multi-server configurations for accessing storage across accounts and regions.

## mcp-trino (https://github.com/txn2/mcp-trino/)

An MCP server for Trino, the distributed SQL query engine. Run read-only SQL queries across any data source Trino connects to, including data lakes, warehouses, and relational databases. List catalogs and schemas, describe tables, explain query plans, and execute analytical queries with configurable timeouts and row limits.

---

# The Data Stack: DataHub + Trino + S3

Modern data platforms need three things: meaning (what does the data represent?), access (how do I query it?), and storage (where does it live?). mcp-data-platform uses DataHub for meaning, Trino for access, and S3 for storage.

These are components the platform composes, not prerequisites an adopting organization must already run. The platform needed a metadata layer and a federating query engine; reimplementing either mature open-source system would have been the wrong call, so it adapts them behind provider interfaces instead. Most organizations have no metadata layer at all today, and DataHub can run as a silent backend component the operator never surfaces. A deployment can attach any subset of the three, including none (see Deployment Shapes).

## DataHub: The Semantic Layer

DataHub (https://datahubproject.io/) is an open source data catalog from LinkedIn. It stores business context: descriptions, owners, tags, glossary terms, lineage, and quality scores.

**Problem it solves**: Data exists everywhere, but understanding what it means requires tribal knowledge. Column `cid` in one system is `customer_id` in another. That deprecated table still gets queried because nobody knows it's deprecated.

**Why DataHub**: Active community (10k+ GitHub stars), rich GraphQL/REST APIs, 50+ integrations (Trino, Snowflake, dbt, Airflow), real-time ingestion, full-text search, lineage tracking.

**For AI**: Without DataHub, an AI sees columns and types. With DataHub, it sees what the data means, who owns it, and whether it's reliable.

## Trino: Universal SQL Access

Trino (https://trino.io/, formerly PrestoSQL from Facebook) is a distributed SQL query engine. It runs SQL against data where it lives, without moving it first.

**Trino connects to everything**: PostgreSQL, MySQL, Oracle, Snowflake, BigQuery, Elasticsearch, MongoDB, S3, HDFS, Delta Lake, Iceberg, Kafka, and more.

**One SQL dialect**: Query across PostgreSQL, Elasticsearch, and S3 in one statement. The AI doesn't need to know where data lives. It's all SQL.

**Why Trino**: Battle-tested at Meta, Netflix, Uber, LinkedIn. Cost-based optimizer, standard ANSI SQL, federated queries without data movement.

## S3: The Universal Data Lake

S3 (Amazon S3, MinIO, or any S3-compatible service) is object storage for any data: files, Parquet, JSON, logs, ML models.

**Problem it solves**: Not all data is structured or in databases. Data platforms must handle structured (tables), semi-structured (JSON, Parquet), and unstructured (PDFs, images) data.

**Why S3**: Infinite scale at low cost, any file format, direct querying via Trino, object versioning, S3-compatible (AWS, MinIO, Ceph).

## Cross-enrichment fills the gaps

| Component | Answers | Limitation Alone |
|-----------|---------|------------------|
| DataHub | "What does this mean?" | Can't query the data |
| Trino | "What's in this table?" | Doesn't know business context |
| S3 | "What files exist?" | Just storage, no meaning |

Cross-enrichment wires them together:
- Trino + DataHub: Query a table → Get schema + owners + tags + deprecation + quality
- DataHub + Trino: Search DataHub → See which datasets are queryable with sample SQL
- S3 + DataHub: List objects → Get matching metadata and ownership

This stack is built for OLAP: S3 stores petabytes at low cost, Trino runs analytical queries across that data, DataHub adds business context.

---

## What's Included

| Toolkit | Tools | Required |
|---------|-------|----------|
| DataHub | 11 tools | No (required for cross-enrichment) |
| Trino | 7 tools | No |
| S3 | 6-9 tools | No |
| Knowledge | 1-2 tools | No |
| Memory | 2 tools | No |

---

# Content Model: Resources, Assets, Knowledge, Memory

Four places content lands, distinguished by who authored it and what the agent should do with it, not by file format.

The canonical statement, held byte-identical in instructions.ResourcePositioning (Go), ui/src/lib/positioning.ts (portal), and the concepts page, and enforced by TestResourcePositioningIsVerbatim:

"Resources are human-uploaded inputs an agent uses as-is: report templates, brand files, data dictionaries, sample payloads, and reference documents. Assets are AI-generated outputs. Knowledge pages are curated facts to search and synthesize. Memory is per-user recall. If it existed before the conversation and the agent should use it verbatim, it is a resource."

| Layer | Holds | Authored by | Agent reaches it via | Portal |
|-------|-------|-------------|----------------------|--------|
| Resources | Files used as-is: templates, brand assets, data dictionaries, sample payloads, runbooks | A person, before the conversation | search then fetch, the resources/read protocol method, or a prompt attachment | Resources |
| Assets | Dashboards, reports, visualizations, documents produced during a session | The agent, during the conversation | save_asset / manage_asset, and search | Assets |
| Knowledge pages | Curated business and domain facts, cross-linked and citable | A person or a promoted insight, reviewed | search then fetch, apply_knowledge | Knowledge |
| Memory | Per-user recall: preferences, corrections, working context | The agent on the user's behalf | memory_capture / memory_manage, and search | Knowledge |

Choosing: existed before the conversation and should be reproduced rather than rewritten -> resource. Made by the agent in a session and worth keeping -> asset. A fact to be found, cited, and synthesized -> knowledge page. About how one person works -> memory.

A resource is filed under a folder path inside its library, and the suggested first-level folders carry the distinction one level down: data are records to read as fact rather than as an example (rosters, mappings, rate tables), templates are layouts a deliverable must be produced in, playbooks are procedures to follow rather than summarize, samples are examples to pattern-match against, references are documents to consult. data and samples are the pair most easily confused: the same CSV is a sample when the agent should copy its shape and data when it should read its rows. The upload dialog states each meaning as the folder is chosen. They are suggestions and not a closed set: a path nests as deep as a library needs.

Some knowledge pages ship with the platform itself: built-in pages documenting the features whose purpose is not readable from tool schemas (managed-script authoring and the Starlark dialect's gotchas, script output identity, the semi-dynamic dashboard pattern, asset references and the refresh loop, provenance and the capture loop, the media type every stored file carries, writing a document that fits the reader's screen). Each page states its mechanism as a mermaid diagram as well as in prose, which the portal renders and an agent reading the page through fetch receives as a labeled graph. They are embedded in the binary and reconciled into the store at startup, so a release that changes them updates every deployment on its next start; unchanged content touches nothing, and a changed page is re-embedded through the ordinary write trigger. They are badged Built-in and read-only where people edit (the portal and apply_knowledge refuse to modify them); a deployment hides one by deleting it — the reconcile respects the tombstone rather than resurrecting it, and a restore action brings hidden pages back refreshed to the running release — and can write its own page on the topic, including under the same slug, which the reconcile also leaves alone. A page a release stops shipping is retired (its tombstone releases the slug), so re-shipping it later resurrects it, unlike an operator hide. The managed-script authoring page derives its dialect section from the same constant `manage_script help` returns, so the two cannot drift. Search ranking is not their only discovery path: the instruction baseline's scripts bullet and `manage_script help`'s see_also list name them by slug, and fetch takes a page's slug in place of its id, because a built-in page's row id is generated per deployment at reconcile time and only the slug is stable across deployments.

## Agent steering

When a deployment has managed resources, platform_info appends a resources section to the agent instructions carrying the positioning statement plus the operating rule:

- Before formatting a deliverable, search for an applicable template or reference resource and follow it rather than inventing a layout.
- When the user names a company file ("our template", "the checklist", "the brand header"), resolve it with search and read it in full with fetch rather than asking them to paste it.
- Material attached to a prompt is authoritative: use it as given.

The section names a tool only when the caller's persona can reach it: a persona denied search is pointed at the resources/list protocol method, which is persona-filtered but is not a tool and is therefore always available. A deployment with no managed resources gets no section. This is steering, not enforcement.

## Making an approved template mandatory

The operator surface for a stricter rule is server.agent_instructions, composed beneath the platform baseline in platform_info. It is set in YAML or edited at runtime from the admin portal's Agent Instructions page, which writes the server.agent_instructions config key to the database.

```yaml
server:
  name: mcp-data-platform
  agent_instructions: |
    Reporting standard: every report, summary, or briefing you produce must use the
    approved template for that deliverable. Before writing one, search for a resource
    in the `templates` category matching the deliverable's name. If one exists, fetch
    it and produce your output in its structure, section order, and headings. Do not
    restructure it, and do not substitute your own layout. If no approved template
    exists, say so in one line before the output so the gap is visible.
```

Scoping the rule to one audience is a persona suffix appended to the same layer:

```yaml
personas:
  analyst:
    display_name: "Data Analyst"
    roles: ["analyst"]
    context:
      agent_instructions_suffix: |
        Client-facing deliverables additionally require the brand header resource. Fetch
        it and place it before the first section.
```

GET /api/v1/admin/config/agent-instructions-baseline renders the platform baseline for this deployment's registered tools, so an operator can see what is already covered before writing their own.

---

# Authorization Model: the connection is the boundary

The unit of access in mcp-data-platform is the connection, not the end user. This is the authorization design, not a filter layered on top of one and not a missing per-user passthrough. An operator configures one connection per credential and permission level, and grants each persona the subset of connections it may reach. A caller's identity selects a persona; the persona selects connections; the connection carries the credential that talks to the downstream system.

## The shape

A connection is a named, operator-authored binding to one downstream system under one credential: a Trino cluster as one service account, a DataHub instance with one token, an S3 account, an upstream MCP server, an HTTP API. Several connections may point at the same system under different credentials, and that is the intended shape rather than a workaround. A read-only Trino account and a write-capable one on the same cluster become two connections; a persona granted the first cannot do what the second can, because the permission level a caller gets is the permission level of the credential bound to the connection they were granted. The HTTP API gateway uses the same split heavily, with api_routes narrowing further by method and path.

## Why the connection and not the end user

Three reasons. First, the downstream systems are heterogeneous: the platform federates Trino, DataHub, S3, arbitrary upstream MCP servers, and arbitrary REST APIs. Trino has session users and can sit behind a system that enforces row policies; DataHub has its own actor model; a third-party MCP server and a vendor REST API typically have neither and expose no token-exchange endpoint to impersonate a caller through. An identity model that works for one of five backends is a special case, not an authorization model. The one construct all of them share is a credential and an endpoint. Second, a credential the operator wrote is auditable ahead of time: an operator can read a connection's service account in the downstream system, see exactly what it reaches, and reason about blast radius before any tool call happens, rather than deferring that reasoning to whatever the identity provider asserted at request time for each user. Third, the distinctions organizations actually draw for AI agents (read the warehouse, read and write the sandbox, read the customer API, write the customer API) are coarse and fit a handful of connections, which keeps each grant explicit and greppable.

## What the boundary enforces

Deny-by-default on both axes, checked on the same call. Connections are deny-by-default: a persona reaches a connection only when a connections.allow glob matches its name, an omitted connections block or an empty allow grants no connections, and deny patterns win over allow (pkg/persona/filter.go, IsConnectionAllowed); a caller whose roles match no persona is denied everything by the built-in deny-all persona the role mapper returns (pkg/persona/mapper.go). Both checks run on every tool call: Authorizer.IsAuthorized refuses unless the tool pattern and the connection both pass. Discovery is bound by the same predicate: search, fetch, list_connections, and the portal search consult one shared scope that delegates to IsConnectionAllowed rather than reimplementing the glob rules, so a caller cannot find an entity behind a connection it was not granted (internal/platform/connscope), and argument completion applies the same predicate directly. search and list_connections report a withheld count and a notice naming the persona rather than silently shortening results. The scope is deliberately permissive in one direction: a catalog dataset whose URN maps to no configured connection is unattributable and stays visible (pkg/knowledge/connscope.go), and a deployment with no persona registry has no scope to apply so discovery is unfiltered there. API routes narrow within a connection: for kind=api connections, api_routes constrains (connection, method, path), and when no rule matches the connection the route check is a no-op so the connection-level grant is the sole gate. A rule's path globs are matched against both the path a call reaches ("/v1/orders/42") and the catalog path its operation declares ("/v1/orders/{id}"), so naming the declared path governs that one operation and every call it serves; wildcards do not cross a "/" and there is no recursive form. Rules are written in the config file under a persona's api_routes, or in the portal at Settings > Personas > Permissions > API endpoints, which lists each api-kind connection and the operations its catalog declares with the persona's decision on each one, and compiles a selection to that operation's own method and declared path so a portal-written rule and a file-written one are the same rule; a hand-written glob is displayed as the glob it was typed as and is not rewritten on save. POST /api/v1/admin/personas/{name}/test-access answers a (connection, method, path) question with the decision and the matched rule. The credential is the enforcement and the toolkit adds what it can: the S3 toolkit's read_only flag is per connection and withholds the mutating tools outright (pkg/toolkits/s3/toolkit.go), and Trino's read_only is per connection too — several Trino instances fold into one multi-connection toolkit whose query interceptor rejects write SQL on the connection each call names, or on the default connection when the call names none, leaving the other connections of the same toolkit untouched (pkg/toolkits/trino/readonly.go); a read/write split across two connections on one cluster is enforced by the toolkit as well as by the downstream service account. Tools belonging to no connection (platform_info, search, and the other platform-level tools) carry an empty connection name, so the connection check admits them and the tool patterns are the gate; the middleware does take the connection from a "connection" argument when the toolkit resolved none (pkg/middleware/mcp.go), so a caller that sends one to a platform-level tool has it recorded and checked like any other.

## What the boundary does not enforce

Every caller granted a connection acts as that connection's credential downstream. The platform performs no per-user token exchange, no impersonation, and no session-user propagation; nothing in the tree swaps a caller's identity for a downstream one. (It does run outbound OAuth per connection, obtaining and refreshing that connection's own credential against upstream MCP servers and APIs via pkg/connoauth/exchange.go; the identity that flow yields belongs to the connection, not to the caller.) Two analysts granted the same connection are indistinguishable to the downstream system. Row-level policies and column masking that key off the end user therefore do not follow a caller through the platform: if a warehouse masks a column for one person and not another and both reach it through one connection, both see whatever that connection's service account sees. Per-person policy means one connection per distinct policy outcome, which is workable when those outcomes are few and unpleasant when they are many. Per-user attribution comes from the audit trail, not from distinct downstream identities: with audit enabled, each tool call writes a row carrying user_id, user_email, persona, tool_name, timing, the connection when the call targets one, and the call arguments subject to redact_keys (pkg/audit/logger.go for the schema, pkg/middleware/mcp_audit.go for the redaction and the write), so the platform can answer who ran what, as which persona, through which connection, even though the downstream system sees only the service account. That makes the audit trail load-bearing and it is not unconditional: audit requires a database, so a deployment with no database.dsn or with audit.enabled: false gets a no-op logger and no rows (pkg/platform/platform.go), log_tool_calls: false keeps audit on but drops per-call rows, log_parameters: false keeps the row without arguments, and async delivery is best-effort under a sustained store outage. A deployment that leans on connection-scoping for authorization should not also run without audit.

## More connections, not more roles

The lever for tightening access is usually a new connection, not a new role. A role divides people; a connection divides reach. Adding a role to split two groups that both land on the same connection changes nothing about what either group can do downstream. Reach for another connection when a group needs a different permission level in the same system (bind the narrower downstream account to a second connection and grant it separately), when a group needs a different blast radius (a connection is the unit a compromised credential is bounded by), or when an API needs a subset of its endpoints exposed (two connections, or one plus api_routes). Reach for another role when the same reach should carry different tools, different agent instructions, or different portal visibility.

## Compared with a warehouse-native MCP server

A server that lives inside one warehouse inherits that warehouse's per-user authorization: row policies, column masks, and grants apply to the calling person because the calling person is the session user. That is a real advantage inside that warehouse and the right choice for a single-warehouse deployment that needs per-person data policy. It does not extend past the warehouse. This platform's premise is a caller reaching a warehouse, a catalog, object storage, third-party MCP servers, and REST APIs through one authenticated, audited, persona-governed endpoint, and the connection is the boundary that spans all of them uniformly. The trade is explicit: uniform enforcement and pre-auditable credentials across heterogeneous backends, in exchange for per-person downstream policy that has to be expressed as connections rather than inherited.

---

# Deployment Shapes

A deployment's shape is the set of backends it attaches. Shape is independent of operating mode, which is determined solely by whether database.dsn is set.

Three shapes:

| Shape | Backends | Core capability |
|-------|----------|-----------------|
| Semantic stack | DataHub, optionally Trino and S3 | Cross-enrichment: every data response carries business context |
| API and knowledge | PostgreSQL | Gateways, knowledge, memory, portal, universal search |
| Combined | Both | Cross-enrichment plus the database-backed surfaces |

## Semantic stack

The shape the platform was built for. DataHub supplies meaning, Trino supplies SQL access, S3 supplies object storage, and cross-enrichment wires them together. DataHub is what cross-enrichment requires; Trino and S3 are optional additions to it.

## API and knowledge

A deployment with no data warehouse and no catalog. PostgreSQL is the one infrastructure requirement, because connections, catalogs, knowledge, memory, and portal assets are all database-backed. It provides the API and MCP gateways (portal-authored connections with encrypted credentials and OAuth grants), API catalogs, the knowledge layer (capture, review, promotion to canonical knowledge pages), the memory layer, the portal (assets, collections, prompts, feedback threads, sharing), search/fetch federating over those sources, and the full auth/persona/audit/observability envelope.

Minimal configuration:

```yaml
server:
  name: mcp-data-platform
  transport: http
  address: ":8080"

database:
  dsn: ${DATABASE_URL}

auth:
  api_keys:
    enabled: true
    keys:
      - key: ${API_KEY_ADMIN}
        name: admin
        roles: ["admin"]

admin:
  enabled: true
  persona: admin

# API connections are authored in the admin portal and stored in the
# database; enabling the toolkit is all the YAML needs to say.
toolkits:
  api:
    enabled: true

personas:
  admin:
    display_name: "Administrator"
    roles: ["admin"]
    tools:
      allow: ["*"]
```

There is no semantic: or query: block. Omitting them selects the noop providers, a supported configuration rather than an error, and enrichment stays enabled at no cost because it no-ops without a semantic provider.

From that configuration and an empty database, tools/list returns twenty tools: api_discover, api_invoke_endpoint, apply_knowledge, fetch, list_connections, manage_asset, manage_feedback, manage_prompt, manage_script, memory_capture, memory_manage, platform_find_tools, platform_info, run_script, save_asset, search, show_prompts, show_scripts.

What this shape does not have:

- Cross-enrichment. With no semantic provider there is no business context to add.
- Trino, S3, and DataHub tools. No trino_*, s3_*, or datahub_* tools are registered.
- Catalog search results. The technical catalog provider registers only when the semantic provider is DataHub.
- Writing knowledge back to the catalog. apply_knowledge with the default sink: datahub refuses on a deployment with no DataHub connection rather than reporting a write it cannot perform; use sink: knowledge_page, the catalog-free destination.

Object storage is optional here too: without an s3_connection, portal assets are stored in the database and managed-resource blob storage is disabled; the platform logs the fallback and starts normally.

## Combined

Attaching a warehouse and catalog to an API-and-knowledge deployment, or API connections to a semantic-stack deployment, yields both sets of surfaces. Neither addition is a migration: providers are selected by the semantic: and query: blocks, and API connections are database rows.

## Replicas

A database-backed deployment of any shape can run several replicas over one database behind a load balancer. State in the database is shared and state in a replica's memory is not, so set sessions.store: database for an MCP session to continue on whichever replica serves its next request. Every HTTP response carries X-Platform-Instance, the hostname and listen port of the process that served it (the hostname separates pods, the port separates two processes on one machine), so two replicas answering one read differently can be told apart.

A connection exists because the connection store holds a row for it, committed before the admin API's save returns; a replica also keeps a live object for each connection it serves (an HTTP client, a parsed spec set, a compiled GraphQL schema, a Trino pool), which it builds for itself. Both questions are answered from the rows: what connections exist (list_connections, the portal's connection pickers, GET /api/v1/apis) is the store unioned with what this replica serves, so a connection saved a moment ago on another replica is named at once with its description, catalog and operation count -- facts about the connection, not about the replica -- and with no health until this replica has called it; a call naming a connection this replica does not serve takes it on from the store and then answers as the saving replica answers. A connection is therefore never listed on one replica and missing on another, never refused as non-existent on one while working on another, and a deleted connection leaves every listing as soon as its row is gone. The startup merge of stored connections into toolkit configuration and the save announcement over the database's notification channel make this cheaper, not correct: a deployment is correct without either.

The local dev stack runs this shape. make dev starts two platform processes over one PostgreSQL (replica A on DEV_API_PORT, replica B on DEV_API_PORT + 1) and an nginx proxy (dev/lb/platform-lb.conf.template, DEV_API_PORT + 2) that round-robins them with no session affinity, so a request and the next are served by different processes. The proxy also replaces the body of any origin 502 or 504 with "error code: <status>" as text/plain, as the CDN in front of a deployment does. make acceptance connects to the proxy by default (MCP_BASE_URL overrides); forEachReplica and connectReplicaPair in test/acceptance discover the replicas from the header and open a session on each directly. The portal's Vite server stays on replica A. DEV_REPLICAS=1 make dev runs one process and no proxy; criteria about two replicas then fail.

## Shape versus operating mode

| | Standalone (no database) | Database-backed |
|---|---|---|
| Semantic stack | DataHub, Trino, S3 tools with cross-enrichment. No knowledge, memory, portal, or gateways. | Everything. |
| API and knowledge | Not available: every surface in this shape is database-backed. | The shape described above. |

---

# Installation

## Prerequisites

- Go 1.26+ (for building from source)
- An MCP-compatible client (Claude Desktop, Claude Code, or custom)
- Access to Trino, DataHub, and/or S3 services you want to connect

## Installation Methods

### Go Install

```bash
go install github.com/txn2/mcp-data-platform/cmd/mcp-data-platform@latest
```

Note: `go install` and plain `go build` do not embed the web portal UI. The
portal SPA is compiled in via `//go:embed` from `internal/ui/dist/`, which is
populated only by the release build or `make frontend-build` and is empty in a
plain source checkout. Such a binary has a working MCP server and admin REST
API but no portal. For the full product use the released binaries, Homebrew, the
Docker images, or run `make frontend-build` (requires Node.js 22) before
`go build`.

### Homebrew (macOS)

```bash
brew install txn2/tap/mcp-data-platform
```

### Docker

```bash
docker pull ghcr.io/txn2/mcp-data-platform:latest
docker run -v /path/to/platform.yaml:/etc/mcp/platform.yaml ghcr.io/txn2/mcp-data-platform:latest --config /etc/mcp/platform.yaml
```

## Client Setup

### Claude Code

```bash
claude mcp add mcp-data-platform -- mcp-data-platform --config /path/to/platform.yaml
```

### Claude Desktop

Add to `~/Library/Application Support/Claude/claude_desktop_config.json` (macOS):

```json
{
  "mcpServers": {
    "mcp-data-platform": {
      "command": "mcp-data-platform",
      "args": ["--config", "/path/to/platform.yaml"],
      "env": {
        "DATAHUB_TOKEN": "your-token"
      }
    }
  }
}
```

## Command Line Options

| Option | Description | Default |
|--------|-------------|---------|
| `--config` | Path to YAML configuration file | None |
| `--transport` | Transport protocol: `stdio` or `http` | `stdio` |
| `--address` | Listen address for HTTP transports | `:8080` |

---

# Configuration

Configuration uses YAML with environment variable expansion (`${VAR_NAME}`).

By default, unrecognized config keys are logged as a `WARN` and ignored so older configs keep loading. Detection applies at every level that maps to a defined field (a stray key under `server:`, `auth.oidc:`, or inside a persona definition is flagged, not just stray top-level keys). Set `config.strict: true` to reject unknown keys with a hard startup error instead (recommended: turns typos and stale keys into an immediate failure rather than a silent no-op). Free-form maps are exempt because they accept arbitrary keys by design: the `toolkits` tree, each toolkit's `config:` map, and the persona *names* under `personas:` (the fields within a persona definition are still validated). A future release will make strict rejection the default, opt-out via `config.strict: false`. Separately, after parsing the server validates the config and refuses to start on a recognized key whose value cannot work: `personas.default_persona` (removed), a missing `auth.oidc.issuer` under enabled OIDC, the `auth.browser_session` requirements (OIDC enabled, `signing_key`, a valid `same_site`, and `secure: true` under `same_site: none`), the `oauth` issuer/upstream fields and `oauth.signing_key` on an HTTP transport without `allow_ephemeral_signing_key`, `database.dsn` and a `LISTEN`-legal `sessions.broadcast_channel` under `sessions.store: database`, and `audit.delivery`. All problems are reported in one error. These checks are newly enforced: they were written alongside the settings they guard, but nothing called them, so a deployment may be running on a config that is refused after upgrading - notably `browser_session.enabled: true` under disabled OIDC (previously started with portal login silently off) and an HTTP-transport OAuth server with no `oauth.signing_key` (previously minted a per-process key that peers reject).

## Config Versioning

Every configuration file should include an `apiVersion` field as the first key. This enables safe schema evolution with deprecation warnings and migration tooling.

```yaml
apiVersion: v1

server:
  name: mcp-data-platform
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `apiVersion` | string | `v1` | Config schema version. Omitting defaults to `v1` for backward compatibility. |

**Supported versions**: `v1` (current)

### Migration Tool

Migrate config files to the latest version:

```bash
# From file to stdout
mcp-data-platform migrate-config --config platform.yaml

# From stdin to file
cat platform.yaml | mcp-data-platform migrate-config --output migrated.yaml

# Specify target version
mcp-data-platform migrate-config --config platform.yaml --target-version v1
```

The migration tool preserves comments and `${VAR}` environment variable references.

### Version Lifecycle

- **current**: Actively supported, no warnings
- **deprecated**: Still works, emits a warning at startup with migration guidance
- **removed**: Rejected at startup with an error pointing to the migration tool

## Minimal Configuration

```yaml
server:
  name: mcp-data-platform
  transport: stdio

toolkits:
  datahub:
    enabled: true
    instances:
      primary:
        url: https://datahub.example.com
        token: ${DATAHUB_TOKEN}
    default: primary

  trino:
    enabled: true
    instances:
      primary:
        host: trino.example.com
        port: 443
        user: ${TRINO_USER}
        password: ${TRINO_PASSWORD}
        ssl: true
        catalog: hive
    default: primary

enrichment:
  trino_semantic_enrichment: true
  datahub_query_enrichment: true
  unwrap_json: true                # Auto-unwrap single-row VARCHAR-of-JSON (default: true)
  column_context_filtering: true   # Only enrich columns referenced in SQL (default: true)
  estimate_row_counts: false       # Run COUNT(*) for availability enrichment (default: false)
  semantic_fallback: false         # Fallback to similarity search on URN miss (default: false, issue #444)
  semantic_fallback_top_k: 1       # Suggested matches per miss (default: 1, max: 10)
  memory_limit: 5                     # Max memory records recalled per tool call (default: 5, issue #761)
  memory_context_budget_bytes: 1500   # Byte budget for rendered summaries; over-budget records become fetchable stubs; 0 disables (default: 1500)
  memory_summary_bytes: 280           # Per-record summary excerpt cap; 0 = full content (default: 280)
```

## Server Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `server.name` | string | `mcp-data-platform` | Platform identity (e.g., "ACME Corp Data Platform") - helps agents identify which business this MCP serves |
| `server.description` | string | - | Explains when to use this MCP - what business, products, or domains it covers. Agents use this to route questions to the right MCP server. |
| `server.tags` | array | `[]` | Keywords for discovery: company names, product names, business domains. Agents match these against user questions. |
| `server.agent_instructions` | string | - | Business/deployment context (data conventions, which backends hold what, domain rules), layered BENEATH the platform-owned instruction baseline (#646). The baseline is an index rather than a manual (#1586): one line per capability naming its entry-point tool and the judgment made before reaching for it (search-first / topology discovery, capture proactively), plus an index of the built-in knowledge pages that carry the depth, fetched when needed rather than carried in every session's first response. It is versioned with the release, names only tools the caller's persona can reach and only pages that caller can read, and is always present (a persona `agent_instructions_override` replaces this admin layer only). The agent receives the baseline as part of the composed `agent_instructions` in the `platform_info` response; the baseline on its own is at `GET /api/v1/admin/config/agent-instructions-baseline` (used by the portal's read-only Agent Instructions panel, which also draws its size meter from the `limit_bytes` and `advisory_bytes` that endpoint reports). This customized layer is BYTE-BOUNDED (#1607) because it is composed into every session's first response: past 12,288 bytes a write succeeds and carries a notice, and past 32,768 bytes it is refused, naming knowledge pages as the home for the overflow. Both writers enforce it -- `PUT /api/v1/admin/config/entries/server.agent_instructions` (400) and the `apply_knowledge` `agent_instructions` sink. |
| `server.prompts` | array | `[]` | Platform-level MCP prompts registered via `prompts/list`. Support `{arg_name}` placeholder substitution. Operator-defined prompts override auto-registered workflow prompts with the same name. |
| `server.prompts[].arguments` | array | `[]` | Typed arguments: name, description, required. Substituted into content as `{name}` placeholders. |
| `server.prompts[].display_name` | string | - | Human-readable title served as the MCP prompt `title` (falls back to the name) |
| `server.transport` | string | `stdio` | Transport: `stdio` or `http` (`sse` accepted for backward compatibility) |
| `server.address` | string | `:8080` | Listen address for HTTP transports |
| `server.streamable.session_timeout` | duration | `30m` | How long an idle Streamable HTTP session persists before cleanup |
| `server.streamable.stateless` | bool | `false` | Disable session tracking (no `Mcp-Session-Id` validation) |

The MCP SDK caps each Streamable HTTP request body at 4 MiB and rejects a larger one with `413 Request Entity Too Large` (`request body exceeds 4194304 bytes`). It bounds inbound JSON-RPC bodies, so it bounds tool-call arguments; tool results, managed-resource uploads, and asset exports travel other paths with their own limits.

## Database Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `database.dsn` | string | - | PostgreSQL connection string |
| `database.max_open_conns` | int | 25 | Maximum open database connections |

Setting `dsn` enables audit logging, knowledge capture, session externalization, and OAuth persistence. Without it, these features degrade to in-memory or noop implementations.

PostgreSQL 16 and 17 are the supported majors, with the `pgvector` extension available. Both are covered by the migration gate on every change (`make migrate-check` applies the full set to each, and CI runs a service per major), because the majors disagree about what SQL a migration may legally contain: since 17, maintenance operations including `CREATE INDEX` run with `search_path` restricted to `pg_catalog` and `pg_temp`, so an index expression whose function body calls a platform-defined function builds on 16 and fails on 17. Upgrading between the two majors needs nothing from this platform beyond the usual PostgreSQL upgrade procedure.

## Config Store

When a database is available (`database.dsn` is set), individual config entries in the `config_entries` table override file defaults for whitelisted keys. The store is the authority: every read resolves the key from it and falls back to the file value when no row exists, so a change via the admin API is in force on every replica as soon as it commits, without restart. Deleting a database entry restores the file default everywhere on the next read. If the store cannot be read, the file-config value is used.

**Whitelisted keys (phase 1):** `server.description`, `server.agent_instructions`

`server.agent_instructions` is byte-bounded on both of its writers (#1607): 12,288 bytes is a soft advisory carried on the response, 32,768 bytes is a hard refusal. The value is composed into the first response of every session on the deployment, so its size is paid for by every caller; the refusal names the size, the limit, the overage, and the knowledge-page alternative. `apply_knowledge action=apply sink=agent_instructions` is the other writer and enforces the same bound.

The config store is selected automatically by database presence: no `database.dsn` means read-only file config; setting it makes config database-backed and mutable.

## Tool Visibility Configuration

Reduce LLM token usage by hiding tools from `tools/list` responses. This is a visibility optimization, not a security boundary — persona-level tool filtering continues to gate `tools/call`.

```yaml
tools:
  allow:
    - "trino_*"
    - "datahub_*"
  deny:
    - "*_delete_*"
  description_overrides:
    trino_query: "Custom description for trino_query..."
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `tools.allow` | array | `[]` | Tool name patterns to include in `tools/list` |
| `tools.deny` | array | `[]` | Tool name patterns to exclude from `tools/list` |
| `tools.description_overrides` | map | `{}` | Override tool descriptions in `tools/list` responses. Config overrides take precedence over built-in defaults. |
| `tools.result_budget.max_bytes` | int | `32768` | The context budget: the most of a rendered tool result a model client over MCP is handed. |
| `tools.result_budget.per_tool` | map | `{}` | Per-tool overrides of `max_bytes` (tool name to bytes). |

No patterns configured means all tools are visible. When both are set, allow is evaluated first, then deny removes from the result. Patterns use `filepath.Match` syntax (`*` matches any sequence of characters).

### Built-in Description Overrides

The platform ships with built-in description overrides for `trino_query` and `trino_execute` that instruct agents to call `datahub_search` first for business context. These are always active and require no configuration. Use `tools.description_overrides` to customize or add more overrides; config entries take precedence over built-in defaults.

#### Tool result context budget

A model reads a tool result into its context, and a client refuses or spills a result past what it accepts (a 64,213-character result was measured refused). `tools.result_budget` is that budget, set once for the platform:

```yaml
tools:
  result_budget:
    max_bytes: 32768        # default
    per_tool:
      trino_query: 65536
```

It is enforced in one place, on the MCP response to a model, and only on a result the model can recover the rest of. Three tools qualify, because each has an export equivalent that returns the whole result without a context cost; each cuts in its own shape and says so. What is measured is the text the client receives, the platform's own additions (the call reference, enrichment) included. A result within the budget is untouched.

| Tool | Past the budget |
|------|-----------------|
| `api_invoke_endpoint` | Re-encoded compactly first; if that is not enough a list body is cut on whole items, `body_truncated` is set, `body_items` reports how many are shown of how many, `next_arguments` is the call that reads on when the operation declares paging parameters, and `export_arguments` carries the `api_export` call that streams the whole response into an asset. A body with no list to cut is cut to a prefix, flagged, and steered to `api_export` |
| `graphql_query` | Re-encoded compactly first; if that is not enough the list in `data` is cut on whole items, `data_truncated` is set, `data_items` reports how many are shown of how many, and `export_arguments` carries the `graphql_export` call. An answer with no list to cut has its `data` withheld whole (a JSON document cut in half cannot be parsed) |
| `trino_query` | The most whole rows that fit are kept, in the format asked for; `result_truncated` and `rows_shown` say how many of `row_count`, and `export_arguments` carries the `trino_export` call |

Every other result reaches the model whole, whatever its size: a knowledge page, `manage_script` help, a prompt, `platform_info`, an uploaded document read through `fetch`, an `s3_object` read, a `datahub_*` result, a proxied MCP tool's output. None of these has a way back to the rest, and a model reasoning from half a manual does worse than one whose client spills an oversized result to a file it can still read. The same holds when one of the three tools cannot cut a result recoverably (an answer whose `errors` alone are past the budget): it is returned whole.

A fitted result's structured content is the fitted value too, so the message a client receives carries the cut result twice rather than the whole one.

Only a model's call over MCP is held to the budget. A REST gateway call (`POST /api/v1/gateway/{connection}/invoke`), a managed script's run and an admin call are never fitted: a program that parses a response cannot read a cut one. Those callers meet only the real resource limits, a connection's `max_response_bytes` and the gateway's in-flight memory budget, and a response past `max_response_bytes` fails with an explicit error (`413` on the REST route) instead of arriving cut.

The budget replaced `max_inline_bytes`, which was set per `api` and `graphql` connection. A connection that still carries it loads normally, the value has no effect, and a warning naming `tools.result_budget` is logged when the connection is read.

## Search-First Gate Configuration

A hard gate that refuses query tools until the agent calls `search`. When a query tool (`trino_query`, `trino_execute`) is called before any discovery tool, the tool handler does not run; a `SEARCH_REQUIRED` error result is returned instructing the agent to call `search` first. Once `search` has been called at least once, every subsequent query tool call by the same authenticated user proceeds normally. Discovery is tracked per authenticated user (not per raw MCP session), so the gate is not falsely re-triggered by clients that open a new session for every tool call (for example claude.ai's web connector); it falls back to the session ID for unauthenticated callers. Enabled by default; set `require_search: false` to disable gating (and hinting) entirely.

```yaml
workflow:
  require_search: false           # Default: true (gate on). false disables it entirely.
  # discovery_tools: []           # Tools that satisfy the gate (defaults to search + the datahub_* tools)
  # query_tools: []               # Tools that are gated (defaults to trino_query, trino_execute)
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `workflow.require_search` | bool | `true` | Enable the search-first hard gate. `false` disables gating with no block and no hint. |
| `workflow.discovery_tools` | array | `search` + the `datahub_*` tools | Tool names that satisfy the gate. `search` is the front door agents are steered toward, and the `datahub_*` discovery tools also count so a persona granted `datahub_*` (but not `search`) is not locked out. |
| `workflow.query_tools` | array | `trino_query`, `trino_execute` | Tool names gated by discovery |

`require_search` replaces the former `workflow.require_discovery_before_query` and its warn-after-execution behavior (breaking rename, not aliased). It is a hard gate: a deployment that never touched workflow gating will begin refusing `trino_query`/`trino_execute` until `search` is called once per session. The older `tuning.rules.require_datahub_check` static hint has been removed.

## Session Gate Configuration

A hard gate that refuses every non-exempt tool until the agent calls the session-initialization tool (`platform_info` by default) once in the session. Before the init tool runs, any other tool call is short-circuited before its handler executes and a `SETUP_REQUIRED` error result is returned (error category `setup_required`) telling the agent to call the init tool first, then retry. Once the init tool has been called, subsequent tool calls proceed until the session TTL expires. Unlike the default-on `*bool` sections, this gate is off by default: `enabled` is a plain bool, so an absent `session_gate` block means disabled.

```yaml
session_gate:
  enabled: true                     # Default: false. true activates the gate.
  init_tool: platform_info          # Tool that initializes the session (default: platform_info)
  exempt_tools:                     # Tools that bypass the gate entirely
    - list_connections
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `session_gate.enabled` | bool | `false` | Activate the session-initialization gate. |
| `session_gate.init_tool` | string | `platform_info` | Tool that initializes a session; always exempt from the gate. |
| `session_gate.exempt_tools` | array | (empty) | Tool names that bypass the gate and may be called before the init tool. |

The gate's session-initialized memory expires after the session TTL, derived from `sessions.ttl` (falling back to the Streamable HTTP session timeout), not from a field on this block. The session gate is distinct from the search-first gate: the session gate requires an init tool (`platform_info`) before any tool; the search-first gate requires a discovery tool (`search`) before query tools. When explicit session handles (`sessions.handles`) are enabled, the session gate is skipped and handle resolution enforces initialization instead, carrying the gate's `exempt_tools` into the handle resolver.

## Purpose Configuration

The one sentence an agent states about WHY a data-access call was made (top-level `purpose:` block, issue #1317). Audit records what a call did and, without this, never why. The platform advertises a `purpose` string property on the input schema of each gated tool via a tools/list decorator, takes the argument off the request in `MCPToolCallMiddleware` before the handler or a gateway-proxied upstream server can see it, and records it on the audit event's own `purpose` column (never in `parameters`, so the parameter redaction policy does not apply to it; the admin events search matches it). Named `purpose`, not `intent`, because `search.intent` is the query text and keeps that meaning. Enabled and required by default.

```yaml
purpose:
  enabled: true      # Default: true. false removes the argument entirely.
  require: true      # Default: true. false records when present, never refuses.
  tools:             # Override the gated set; empty means the default below.
    - search
    - trino_query
    - "datahub_get_*"
    - "kind:mcp"
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `purpose.enabled` | bool | `true` | Advertise, strip, and record the `purpose` argument. |
| `purpose.require` | bool | `true` | Refuse a gated call that states none with `PURPOSE_REQUIRED` (error category `purpose_required`). |
| `purpose.tools` | array | see below | Gated set: tool-name globs (`filepath.Match`) plus `kind:<toolkit-kind>` entries. |

Default gated set: `search`, `fetch`, `trino_query`, `trino_execute`, `trino_export`, `trino_describe_table`, `api_invoke_endpoint`, `api_export`, `datahub_get_*`, `s3_object`, `s3_list`, `save_asset`, `manage_asset`, and `kind:mcp` (every tool an MCP gateway connection proxies, whose names are chosen upstream and cannot be written as a glob). The two asset tools are in the set because an asset is the one output a person opens weeks later, often someone who was not there when it was made, and the first question they ask is what it was for (#1695); `manage_asset` is named whole, so its reads state a purpose alongside its writes, since the gate decides from a tool name and one name covers both halves. `apply_knowledge` stays outside on its own reasoning -- what an apply_knowledge call applies is itself the explanation, while an asset is the output and not the reason for it. The rest of the orientation and platform-management surface (`platform_info`, `list_connections`, `platform_find_tools`, `memory_*`, the other `manage_*` tools) is deliberately outside it: their purpose is their name.

One predicate (`middleware.PurposeResolver.Gates`, deciding from tool name plus toolkit kind) drives both the advertisement and the enforcement, so what the platform asks for and what it enforces cannot drift, and both are replica-stable without either request remembering the other. `require: true` refuses ONLY a caller that threaded an explicit `session_id` handle on the same call, which is the platform's only proof that a caller can thread a platform-injected argument. That single condition subsumes every exemption: an MCP App's sandboxed call is session-adopted from its authenticated identity and threads nothing (#1040); the gateway REST shim, the admin tool runner, and a managed script drive a fresh in-memory session per request and thread nothing (#811); an isolated `dpp_`/`dpx_` run has its session minted server-side (#859). Consequently `sessions.handles.enabled: false` also stops purpose from ever being required, though it is still advertised and recorded. A stated purpose is trimmed and bounded to 1000 runes before it reaches the audit row. The agent instructions name the gated set rather than a category (#1640): the tools the configured set NAMES and the caller can reach, listed, plus a clause per `kind:` entry it gates wholesale ("every tool served by a connection of kind `mcp`"). The kind is named rather than expanded because one MCP gateway connection can proxy more tools than the platform's own data-access surface has, under names the platform did not choose and that change when the upstream does. `middleware.PurposeResolver.GatesByName` is the split, so the description and the enforcement still come off one resolver.

What the gate decides is where the argument is advertised and where a MISSING one refuses the call, not whether a stated one is understood. A `purpose` stated on an UNGATED tool is taken off the request and recorded on the audit row exactly as a gated one is (#1640): tool input schemas are closed to unknown properties (#1057), so leaving it in place refused the call with `invalid_arguments` for stating what the platform's own instructions asked the model to state, and the corrective hint then told the agent to drop a property those instructions required. The argument is taken on every tool the platform DEFINES, gated or not; a tool proxied from an upstream MCP server (`registry.ToolkitMatch.Kind == "mcp"`) is the one place the platform did not choose the parameter names, so when such a tool is outside the gated set its `purpose` is recorded and still delivered to the upstream — which is what makes dropping `kind:mcp` from `purpose.tools` the remedy for an upstream server that defines its own `purpose` parameter.

## Tool Annotations

Every tool the platform registers advertises MCP tool annotations on `tools/list` (issue #1692). `readOnlyHint` is `true` when no call to the tool modifies state; `destructiveHint` is stated on every write (`true` when some action removes or overwrites, `false` when every action only adds, because the specification's default for an absent one is `true`); `idempotentHint` is `true` on a read. The specification tells a client to assume a tool is NOT read-only when `readOnlyHint` is absent, and clients that act on that ask the user to confirm every unannotated call, so the mandatory opening sequence `platform_info` then `search` then `fetch` arrived as three write confirmations before any data.

The hint describes the tool rather than the call: `manage_asset`, `manage_table`, `manage_resource`, `manage_prompt`, `manage_script`, `memory_manage` and `s3_object` expose reads alongside writes and each advertises the most it can do. Read-only: `platform_info`, `search`, `fetch`, `list_connections`, `platform_find_tools`, `show_prompts`, `show_scripts`, `api_discover`, `graphql_discover`, `s3_list`, plus the upstream reads `trino_query`, `trino_explain`, `trino_browse`, `trino_describe_table`, `datahub_browse` and `datahub_get_*`.

`toolkit.ReadOnlyAnnotations()` and `toolkit.WriteAnnotations(destructive)` are the two shared constructors a registration states its classification with. `internal/toolwrite.ReadOnly` is the platform's other per-tool statement about whether a tool writes — it is what a managed script's draft run consults — and a structural gate (`TestAnnotationsAgreeWithToolwrite`) fails the build when the two disagree, so a tool cannot be advertised read-only and refused inside a draft. A second gate (`TestEveryToolLiteralIsAnnotated`) fails when a registration sets no `Annotations` or no `Title`. A tool proxied from an upstream MCP server carries the upstream's annotations unchanged, including none; the trino, datahub and s3 toolkits take theirs from the upstream libraries, overridable per deployment under a toolkit's `annotations:` key.

The admin API returns the same object (#1706): `GET /api/v1/admin/tools/schemas` carries `annotations` on each tool and `GET /api/v1/admin/tools/{name}` carries it on the detail, both read from the server's own `tools/list` after any override, with the MCP keys unchanged (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`) and the key absent for a tool that states none, `destructiveHint` and `openWorldHint` kept absent rather than false when unstated. The portal's Tools page shows them as badges on the tool header: read-only, additive (a write stating `destructiveHint: false`), destructive (any other write, since an unstated `destructiveHint` defaults to true), idempotent, and "no hints" for a tool advertising none, which is where an operator checks that an `annotations:` override took effect.

## Tool-Call Rate Limiting

A per-identity safety net on authenticated `tools/call` requests (top-level `rate_limit:` block). It bounds a runaway agent loop or a compromised account before it can saturate the audit pipeline, the shared database pool, or an upstream (Trino, DataHub, S3, a proxied MCP server); it is not a throughput throttle: the default limit is generous enough that ordinary interactive and agent use never touches it. Over-limit calls are short-circuited before the handler runs and outer to the audit/enrichment middleware, so a refused call consumes no handler, audit, or upstream work; the caller receives a `RATE_LIMITED` in-band error result (error code and category `rate_limited`) with a retry hint so an agent backs off rather than seeing a transport failure, and the structured envelope carries the interval as `retry_after_seconds` beside `code`, `category`, `message` and `hint`. A platform run of a managed script is queued, not refused: its calls cross the limiter as the script principal (`AuthType` `script`), and when the bucket is empty the limiter holds the call until the sustained rate refills a token and admits it, so a run's throughput is governed by the sustained rate and its calls never see `RATE_LIMITED`; a queued call does not count in `mcp_rate_limited_total` or log a warning but counts in `mcp_rate_limit_queued_total` and logs at Info under the run id, and a run canceled or reaching its deadline while a call is held ends then. A draft run carries its author's own identity and is refused as any interactive call is; there the script host waits `retry_after_seconds`, bounded by the run's deadline, and issues the call again, so the script sees only the admitted call's result and the wait is recorded in the run's log (a run whose deadline arrives while waiting fails as a timeout and is not re-queued). `platform_info` is always exempt so a throttled agent can re-read platform guidance. The limit is keyed on the authenticated user (not the client IP: a multi-user connector delivers all traffic from one egress, so per-IP limiting would be useless and would let one user starve the rest); shared/anonymous identities (auth disabled) fall back to a per-session key, and a call with no attributable identity is not limited (fail-open, since it has already passed auth). The token bucket is in-memory per replica, so behind a load balancer the effective ceiling is N times the configured limit (intentional for a backstop, since distributed coordination on the hot tool path is not warranted). Each refusal increments `mcp_rate_limited_total`; each call a script principal was held for increments `mcp_rate_limit_queued_total`. Enabled by default; set `rate_limit.enabled: false` to remove the limiter from the chain, for scripts as for everyone. This is unrelated to `oauth.rate_limit`, the per-IP limiter on the unauthenticated `/token` and `/register` endpoints.

```yaml
rate_limit:
  enabled: true                     # Default: true. false removes the limiter.
  requests_per_minute: 240          # Default: 240. Sustained per-user tools/call rate (4/second).
  burst: 60                         # Default: 60. Largest instantaneous per-user burst.
  exempt_tools:                     # Never limited (platform_info is always exempt)
    - search
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `rate_limit.enabled` | bool | `true` | Enable the per-user tool-call limiter. `false` removes the middleware from the chain. |
| `rate_limit.requests_per_minute` | int | `240` | Sustained per-user `tools/call` rate (token refill). |
| `rate_limit.burst` | int | `60` | Token-bucket depth: largest burst a single user may issue before the sustained rate governs. |
| `rate_limit.exempt_tools` | array | (empty) | Tool names never rate limited, in addition to the always-exempt `platform_info`. |

## Portal Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `portal.enabled` | bool | `false` | Enable portal toolkit and save_asset/manage_asset/manage_feedback tools |
| `portal.s3_connection` | string | - | Name of the S3 toolkit instance for asset storage |
| `portal.s3_bucket` | string | `portal-assets` | S3 bucket for asset content |
| `portal.s3_prefix` | string | `artifacts/` | Key prefix every asset object is written under: content and versions from the portal, the admin console, the asset tools, the exports and script outputs, and collection mosaics (`<prefix>collections/<id>/thumbnail.png`), all through `portaldomain.AssetContentKey` and `portaldomain.CollectionThumbnailKey` (#1903). Releases before that wrote portal and admin edits and mosaics under a fixed `portal/`; those objects keep the key their row recorded and stay where they are, and a mosaic moves under the prefix the next time it is drawn |
| `portal.deleted_retention_days` | int | `30` | How long a soft-deleted asset, collection, feedback thread or knowledge page is kept before it is hard-purged with every version object, tile and mosaic it names, and its shares, collection places and threads (#1904). `0` = default, negative keeps them. A hidden built-in knowledge page is never purged. An asset whose object will not delete keeps its row for the next day's sweep |
| `portal.orphaned_producer_retention_days` | int | `90` | How long a `content_producers` row outlives its asset or resource (#1904). `0` = default, negative keeps them |
| `portal.public_base_url` | string | `""` | Base URL for portal links in save_asset responses |
| `portal.max_content_size` | int | `10485760` | Maximum asset size in bytes (10 MB) |
| `portal.max_versions` | int | `100` | Versions an asset keeps when it carries no override of its own. A version pushed past the cap is deleted along with its stored content and thumbnails; the current version is never pruned. `0` keeps every version, and a negative value is refused at startup. Applied at the write, so an asset already over the cap is trimmed the next time it is written. An asset's owner (or an admin) can override it, from the portal metadata form, `manage_asset` action `update` with `max_versions`, or the admin asset route |
| `portal.implementor.name` | string | `""` | Implementor display name shown in the left zone of the public viewer, the public collection viewer, the guest share landing page, and the access-denied page. Independent of `portal.implementor.logo`: either one alone renders the implementor block |
| `portal.implementor.logo` | string | `""` | URL to the implementor logo in any image format. The viewer and share pages link it with an `<img>` element; its origin is added to the `img-src` of the pages whose policy would otherwise block it. Renders with or without `portal.implementor.name` |
| `portal.implementor.url` | string | `""` | Clickable link wrapping the implementor name and logo |
| `portal.terms_url` | string | `""` | Terms-of-service link rendered in notification email footers |
| `portal.privacy_url` | string | `""` | Privacy-policy link rendered in notification email footers |
| `portal.export.enabled` | bool | auto | Enable trino_export tool (auto-enabled when portal + trino configured) |
| `portal.export.max_rows` | int | `100000` | Hard row cap for exports |
| `portal.export.max_bytes` | int | `104857600` | Hard byte cap for formatted output (100 MB) |
| `portal.export.default_timeout` | string | `"5m"` | Default query timeout for exports |
| `portal.export.max_timeout` | string | `"10m"` | Maximum allowed query timeout for exports |

The asset reference route (`/portal/refs/`) has a limiter of its own sized at `assetrefs.MaxRefs` (20) times the `portal.rate_limit` values after their defaults are applied, so a deployment with no `portal.rate_limit` block gets 1,200/min and a burst of 200 per client rather than the unscaled 60/10 (#1791); the thumbnail renderer's in-process requests to it are not counted.

Portal requires `database.dsn` for metadata storage and at least one S3 toolkit instance for asset content.

## Admin API Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `admin.enabled` | bool | `true` | Enable admin REST API. Set `false` to disable. Routes still require admin-role auth to use. |
| `admin.persona` | string | `admin` | Persona required for admin access |
| `admin.path_prefix` | string | `/api/v1/admin` | URL prefix for admin endpoints |

Branding fields have moved to the `portal:` section:

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `portal.enabled` | bool | `false` | Enable portal SPA frontend and asset API |
| `portal.brand_name` | string | mcpapps `brand_name` | Deployment brand; composes the portal title as `<brand_name> Portal` and names the public viewer header, branded denial pages, and built-in MCP Apps. Falls back to `brand_name` in `mcpapps.apps.platform-info.config` when MCP Apps are enabled |
| `portal.brand_url` | string | mcpapps `brand_url` | Brand home page; the sidebar brand mark (logo and name) links to it in a new tab. Unset leaves the mark inert. Falls back to `brand_url` in `mcpapps.apps.platform-info.config` when MCP Apps are enabled |
| `portal.version_url` | string | `""` | Link target for the version number in the portal header (release notes or changelog). Unset leaves the version as plain text. Served on the unauthenticated branding endpoint |
| `portal.title` | string | `<brand_name> Portal`, else `MCP Data Platform` | Sidebar/branding title text; composed from `brand_name` when unset. A brand already ending in "Portal" is not doubled. Only a `brand_name` set in the `portal` block composes the title, so an mcpapps-inherited brand leaves an existing deployment's title unchanged |
| `portal.tagline` | string | `Sign in to access the platform.` | Login-screen subtitle text |
| `portal.oidc_button_label` | string | `Sign in with OIDC` | Login-screen SSO button text (e.g. "Sign in with ACME Keycloak") |
| `portal.logo` | string | `""` | Logo URL (fallback for both themes) |
| `portal.logo_light` | string | `""` | Logo URL for light theme |
| `portal.logo_dark` | string | `""` | Logo URL for dark theme |
| `portal.logo_email` | string | `""` | Raster PNG logo URL for notification emails (fetched at startup, max 1 MB) |
| `portal.implementor.name` | string | `""` | Implementor display name (left zone of the public viewer, the collection viewer, and the share pages); renders with or without a logo |
| `portal.implementor.logo` | string | `""` | URL to the implementor logo in any image format, linked by the viewer and the share pages; renders with or without a name |
| `portal.implementor.url` | string | `""` | Clickable link wrapping implementor name and logo |
| `portal.terms_url` | string | `""` | Terms-of-service link rendered in notification email footers |
| `portal.privacy_url` | string | `""` | Privacy-policy link rendered in notification email footers |

When `portal.enabled: true`, an interactive web dashboard is served at `/portal/`. It provides audit log exploration, tool execution testing, and system monitoring. The sidebar displays a configurable logo and title. Theme-specific logos are resolved as: light theme uses `logo_light` → `logo` → built-in default; dark theme uses `logo_dark` → `logo` → built-in default. The resolved logo is also used as the browser favicon. Shared asset links (`/portal/view/{token}`) display a two-zone header: the right zone shows the platform brand, and the optional left zone shows the implementor brand configured via `portal.implementor`. Both logos may be any image format, and the pages link them with an `<img>` element at the configured URL rather than embedding them; the guest share pages and the branded denial page add that URL's origin to their `img-src`, and a deployment with no logo configured keeps a policy that loads no remote image. The exception is a built-in MCP App, which runs in a sandboxed iframe that blocks external loads: its config carries the logo inline, fetched once at startup — an SVG as its own markup, any other format as a `data:` URI — and a logo that is unreachable, larger than 1 MB, or not an image logs one warning naming the URL and the reason. Notification emails cannot reuse these logos because mail clients strip inline SVG; set `portal.logo_email` to a raster PNG URL, which is fetched once at startup and attached to each message as an inline part so it renders without remote-image loading. It is additive to the brand wordmark, which also serves as the image's alt text, and leaving it unset renders the wordmark alone.

The public viewer supports light/dark mode (system default with toggle, persisted to localStorage), an expiration countdown notice showing relative time until share expiry, and a configurable per-share notice text. The `notice_text` field defaults to "Proprietary & Confidential. Only share with authorized viewers." — set it to a custom string for different text, or to `""` (empty string) to hide the notice entirely. Only a public link is created with an expiry, so the countdown is driven by the share's own `expires_at` rather than by its access mode — shares created before that rule keep the expiry they were given. Set `hide_expiration: true` to hide the countdown from the viewer.

Every share carries an `access_mode` deciding who its token opens for: `restricted` (only the named recipient and the share's creator), `authenticated` (any signed-in platform user), or `public` (anyone with the link, without signing in). A create-share request that names a recipient and no mode is `restricted`; one that names neither is `authenticated`. `public` is never implied and must be requested; `restricted` without a recipient is rejected with 400. Enforcement is a single gate (`publicShareGate`, `pkg/portal/share_access.go`, deciding via `pkg/portal/shareaccess`) wrapping every route under `/portal/view/` (the page, `/content`, `/thumbnail`, `/collection-thumbnail`, and the three `/items/{assetId}/...` routes), so no route can serve bytes the mode refuses. A caller who is not admitted receives 403 with the reason; revoked and expired tokens still return 410. The gate also writes the response's cache policy, so a per-caller verdict cannot be cached under a per-URL key: `Vary: Cookie` everywhere, `Cache-Control: private` for every mode but `public`, `no-store` on refusals, and `public` only on a fully public share's thumbnail, with `max-age` clamped to the smaller of an hour and the share's remaining life so no stored copy outlives the token (revocation is the residual: a copy a shared cache already holds is served until it goes stale). Migration 000083 backfills existing rows: recipient present becomes `restricted`, recipient absent becomes `authenticated`, so previously distributed links stop resolving anonymously.

Who may act on a portal item is resolved by one seam per authority rather than by an ownership comparison repeated at each route (`internal/portal/access`). `CanManage(ownerID, user)` is owner-or-admin and gates deleting an asset or collection, creating a share, reading a share list, and revoking one; `CanManageEmail` is its form for prompts, whose ownership is an address. `CanEditCollection` is owner-or-admin-or-Editor and gates the four operations that shape a collection as a document — name and description, display config, sections, thumbnail — so an Editor share on a collection means something about the collection and not only about the assets inside it. The split is deliberate: an Editor edits, but destruction and re-granting access stay owner authority, which is why delete, share, and share-list remain on `CanManage`. The admin arm never grants more than an admin already holds, since the admin API reads, edits, and deletes any asset outright; withholding the weaker right to share one was an artifact of the gate, not a policy, and it stranded content owned by a non-human principal (an API-key session owns its assets as `<key name>@apikey.local`, an identity nobody can sign in as). `GET /portal/collections/{id}` reports the resolved `can_edit` and `can_manage` beside `is_owner`, so the page offers exactly the actions that will succeed instead of rendering a form the server refuses on submit. View gates are unchanged: an admin still reaches another principal's collection through the admin surface, not through the portal's own read routes.

## Resource Templates Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `resources.enabled` | bool | `true` | Serve the MCP resource templates and the DataHub-to-Trino resource links. Read-only, so default on; set `false` to disable |

When enabled, three RFC 6570 URI templates are registered:
- `schema://{catalog}.{schema_name}/{table}` — Table schema with semantic context
- `glossary://{term}` — Glossary term definition and related assets
- `availability://{catalog}.{schema_name}/{table}` — Data availability status and row count

## Argument Autocompletion (completion/complete)

The platform serves the MCP `completion/complete` request so completion-capable clients suggest valid values as a user types. The capability is advertised automatically whenever prompts or resource templates are available (no configuration). Prompt arguments are routed by name across built-in and database prompts: `dataset` gives dataset names from the semantic index; `topic` gives domains, data products, and glossary terms; `connection` gives configured connection names. Resource-template variables complete from the query engine: `schema://`/`availability://` `catalog`/`schema_name`/`table` (each variable completes once its predecessors are chosen), and `glossary://` `term` from the glossary index. Completions are persona-filtered like `tools/list` and `search` (dataset/topic/glossary require `search`; catalog/schema/table require `trino_browse`; connections require `list_connections` and are further filtered by persona connection rules); unauthenticated sessions get nothing, and each lookup runs under a short latency budget that degrades to an empty list on upstream failure. A response carries at most 100 values; `hasMore` is set from the catalog's own match count (never inferred from the number of rows a page returned, which a catalog may clamp below what was requested) and `total` only when the returned set is provably complete — a catalog that cannot count leaves both omitted.

## Custom Resources Configuration

Custom resources expose arbitrary static content as named MCP resources. Registered whenever `resources.custom` is non-empty, independent of `resources.enabled`. Content can be inline (`content`) or read from a file on every request (`content_file`, supports hot-reload).

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `resources.custom[].uri` | string | — | Unique resource URI, e.g. `brand://theme` (required) |
| `resources.custom[].name` | string | — | Display name shown in `resources/list` (required) |
| `resources.custom[].description` | string | `""` | Optional description for MCP clients |
| `resources.custom[].mime_type` | string | — | MIME type, e.g. `application/json`, `image/svg+xml` (required) |
| `resources.custom[].content` | string | — | Inline text/JSON/SVG content; mutually exclusive with `content_file` |
| `resources.custom[].content_file` | string | — | Absolute file path; read on every `resources/read` request |

Example — inline JSON brand theme and file-backed SVG logo:

```yaml
resources:
  custom:
    - uri: "brand://theme"
      name: "Brand Theme"
      mime_type: "application/json"
      content: '{"colors":{"primary":"#FF6B35"},"url":"https://example.com"}'
    - uri: "brand://logo"
      name: "Brand Logo"
      mime_type: "image/svg+xml"
      content_file: "/etc/platform/logo.svg"
```

Invalid entries (missing URI, name, or mime_type; both or neither content fields) are skipped with a warning; valid entries in the same list are still registered.

## Managed Resources Configuration

Managed resources are human-uploaded reference files (data, samples, playbooks, templates, references) stored in S3 with metadata in PostgreSQL. They are surfaced to AI assistants via the standard MCP `resources/list` and `resources/read` protocol, and managed through a REST API and the Admin Portal.

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `resources.managed.enabled` | bool* | auto | Enable managed resources. nil = auto (enabled when database is available) |
| `resources.managed.uri_scheme` | string | `"mcp"` | URI scheme prefix for resource URIs |
| `resources.managed.s3_connection` | string | `""` | Name of the S3 toolkit instance used for blob storage |
| `resources.managed.s3_bucket` | string | `managed-resources` | S3 bucket for resource file blobs |
| `resources.managed.max_versions` | int | `10` | Content revisions kept per resource, counting the current one. A revision past the cap deletes the oldest version's blob; the live content is never pruned. Non-positive selects the default; below 2 is raised to 2 |
| `resources.managed.extract.max_member_bytes` / `max_total_bytes` / `max_members` / `max_ratio` | int64 / int64 / int / int64 | 2 GiB / 4 GiB / 10000 / 500 | Limits on `manage_resource extract` (#1879); see "Archive extraction limits" below. Independent of `max_upload_bytes` |
| `resources.managed.max_upload_bytes` | int64 | `104857600` (100 MB) | Largest file `POST /api/v1/resources` and `POST /api/v1/resources/{id}/content` accept. Non-positive selects the default, so a deployment that sets nothing keeps 100 MB. The refusal names this deployment's number, and the portal's file chooser reads it from `GET /api/v1/portal/me` rather than holding a second copy. It bounds bytes streamed rather than bytes held: the upload path streams into the multipart uploader (see below). It is also the size a table registration reads an object by (#1634), so a file this deployment accepts is one it can register a table over |

Example:

```yaml
resources:
  managed:
    enabled: true
    uri_scheme: "mcp"
    s3_connection: "primary"
    s3_bucket: "platform-resources"
    max_versions: 10
    max_upload_bytes: 104857600
    extract:
      max_member_bytes: 2147483648
      max_total_bytes: 4294967296
      max_members: 10000
      max_ratio: 500
```

#### Archive extraction limits (`resources.managed.extract`)

`manage_resource action=extract` (#1879) writes the members of a stored zip, gzip or gzipped tar out as managed resources, and `resources.managed.extract` bounds it. `max_member_bytes` (default 2 GiB) is the largest one member may be uncompressed; `max_total_bytes` (default 4 GiB) is the most the selected members may add up to; `max_members` (default 10000) is how many entries an archive may hold, directories included, and since a zip's directory is read whole to open it (`archive/zip.NewReader`), it is checked against the count the end-of-directory record declares, zip64 locator followed, before the directory is parsed; `max_ratio` (default 500) is the furthest a member may expand past its compressed size, checked once a member passes 1 MiB. Each field takes its default when absent, zero or negative. The struct is `unarchive.Limits` (`internal/unarchive`), used directly as the config section, so the key a `LimitError` names is the key an operator sets. The limits are apart from `max_upload_bytes` on purpose: an archive is delivered by an upstream rather than picked at an upload form, and the file inside a monthly delivery is routinely several times the upload ceiling, so the extractor lands each member with an unbounded lander ceiling (`Lander.landPlanned` with ceiling 0) and enforces its own limits in the archive reader.

Extraction streams both ways: the archive is read by range (`GetObjectRange`) through a two-block cache of 8 MiB blocks, and each member is streamed into `PutObjectStream`. A zip declares every member's name, size, flags and method up front, and a gzipped tar is read through once before its members are written, so for both every refusal (unsafe name, encryption, method, limit, empty selection) comes before the first write; a plain gzip's single member has no declared size, so its limits are enforced as it streams, and a gzip past one leaves nothing stored because its one multipart upload is abandoned.

#### What raising `max_upload_bytes` costs

The upload path streams (#1631). A request walks the multipart form part by part (`r.MultipartReader`, `pkg/resource/handler.go`) and hands the `file` part to the multipart uploader as a reader (`S3Client.PutObjectStream`, `pkg/resource/store.go`), so the object is never assembled: not in one `[]byte`, and not in a temporary file. What a request holds is the uploader's part buffers plus `contenttype.StructuredSniffLen` (8 KiB) for content detection, whatever the ceiling is. Two consequences, both of which were true the other way before: the upload path needs no writable temporary directory, which is what the published `FROM scratch` image has; and blob storage is written through multipart, so a backend that bounds a single `PutObject` — MinIO refuses one above 16 MiB under the `aws-chunked` encoding the AWS SDK uses over HTTPS — takes a file of any size the ceiling allows. The READ path has not changed shape: `GetObject` returns a `[]byte`, so serving a large managed resource still holds it. Every other object write goes the same way since #1863: the portal's blob adapter sends a buffered `PutObject` through the multipart uploader (8 MiB parts, a body under one part still one request), which covers `platform.export` to the portal, `trino_export` and portal content edits, and `s3_object action=put`, which a script's bucket destination delivers through, streams when the connection's client has the uploader. `manage_script help` states the per-output ceiling as `limits.output_max_bytes` (100 MiB).

The ceiling is enforced on the bytes streamed, since a streamed part declares no length: the read that passes it fails, the uploader aborts the incomplete multipart upload, and the route answers 400 naming this deployment's number. The request body is separately bounded at the ceiling plus 64 KiB of multipart framing headroom as a backstop against a body with no end, so a file of exactly the ceiling is still accepted.

A table registration reads by this same ceiling (`registrationMaxBytes`, `internal/httpserver/tablemounts.go`, feeding `tableregister.Deps.MaxBytes`). It was a compiled-in `DefaultMaxBytes` of 100 MB, written to track the upload cap back when that cap was also compiled in; #1628 made the cap configurable and the constant did not follow, so a deployment that raised its ceiling stored CSVs no registration would read (#1634). The registration read holds the whole object, unlike the upload, because the header row and the line-shape check need the bytes (#1441).

Both write routes read the form IN ORDER and stop at the `file` part, because that part is handed to the uploader where it is found. Every other field (`scope`, `scope_id`, `path`, `display_name`, `description`, `tags`) has to arrive before it; a form that puts a part behind the file is refused with `the file part must be the last part of the form`, and because the refusal happens while the file is being read, no object and no record survive it. The portal appends its file last (`draftForm`, `ui/src/pages/resources/modals/UploadModal.tsx`).

Two things a raised ceiling does not change: content indexing still stops at `MaxContentReadBytes` (8 MiB), so a file above that is indexed on its metadata alone whatever the ceiling is; and an ingress or proxy in front of the platform enforces its own body limit, which has to be raised too or a request the platform would accept never reaches it.

### Scope Model

Resources are assigned one of three visibility scopes:

| Scope | Visibility | scope_id | Example URI |
|-------|-----------|----------|-------------|
| `global` | All authenticated users | (empty) | `mcp://global/templates/query-patterns.sql` |
| `persona` | Users operating under the named persona | persona name | `mcp://persona/analyst/playbooks/data-quality-checklist.md` |
| `user` | Only the owning user | user subject ID | `mcp://user/abc123/references/my-notes.txt` |

### URI Scheme

Resource URIs follow the pattern: `{scheme}://{scope}/{scope_id?}/{path}/{filename}`, where `path` is the slash-separated folder path the resource is filed under

- Global: `mcp://global/{category}/{filename}`
- Persona: `mcp://persona/{persona_name}/{category}/{filename}`
- User: `mcp://user/{user_sub}/{category}/{filename}`

### Categories

Resources must be assigned a folder path (#1529): slash-separated, each segment lowercase alphanumeric with hyphens starting with a letter or digit (#1886; the webhook compactor files windows under dated folders) and at most 31 chars, at most 8 segments and 200 characters overall, no empty segment, no leading or trailing slash, no `.` or `..`. `ValidatePath` (`pkg/resource/path.go`) names the rule a refusal broke rather than restating the grammar, and it refuses depth before length because removing folders is what fixes a path that breaks both. Every segment keeps the rule the flat `category` column carried, widened to a leading digit, so every pre-tree row is a legal one-segment path, migration 000130 rewrites no rows and no existing URI changes. Folders are STORED (#1872, migration 000160 `resource_folders`, one row per folder, NULL `scope_id` for global like `resources`, unique on `(scope, COALESCE(scope_id,''), path)`): a folder exists once it is created (`POST /api/v1/resources/folders`, the authority an upload into that library takes) or a resource is filed under it, and it stays when its last file leaves. Every insert and move records the folder chain in the same transaction, the migration recorded every folder in use, and `Store.Folders` is the union of the stored rows and the paths in use, so a writer that skips the chain hides nothing. The file manager's Delete dialog opens at once for a selection holding a folder and counts the files inside it itself, showing "Counting files..." with its confirm disabled until the count is in, and showing a failed count inside the dialog with Retry; a second Delete while a dialog is open starts nothing (#1887; before, the dialog rendered nothing until the folder's listing returned, and a read that took 12-30 s on the reference install looked like a dead button). `DELETE /api/v1/resources/folders` removes a folder and the stored folders beneath it and answers 409 while any resource is filed there; a folder move rewrites the stored folders beneath it in the same transaction as the resources (`FolderStore.MoveFolderTree`) and moves a folder that holds only empty folders. The portal suggests six first-level names (`data`, `visual`, `samples`, `playbooks`, `templates`, `references`, in `SEED_FOLDERS`, `ui/src/pages/resources/parts/PathField.tsx`) as completions rather than as a closed set; the server validates the shape of a path and not its membership.

The library is browsed as a TREE of folders (#1530). `folderView` (`ui/src/pages/resources/parts/tree.ts`) divides whatever is loaded into the folders and files at one location; a folder's count is everything beneath it at every depth, and while further pages remain it is written `12+`, because it is how many have arrived and not how many the folder holds. The location is in the ROUTE, not the query string: `/resources/lib/<tab>/<path>` and `/admin/resources/lib/<tab>/<path>`, always at least two segments after the section so it can never be confused with `/resources/{id}`, which is exactly one -- `readPath()` drops the query string, so a location kept there is one the page cannot see on a nav-link push (#1470). Each level is a distinct address, so a reload returns to it, a link opens it, and Back steps out one folder. Filters (q, tag, sort) stay in the query and are written with replace. Search spans the WHOLE library rather than the folder in view -- the path filter is dropped while a query runs -- and each hit carries the path it was found at with a Reveal control that walks the tree to it. A folder whose resources are all images renders as a grid of tiles instead of rows -- decided by the content, not by the folder name, so a photograph filed under `references` is shown as a photograph and a written note filed under `visual` is shown as a row; on the administrator's page a tile also carries the scope badge and never-read flag its table columns carry. A resource has no stored thumbnail, so a tile is the original object: tiles load lazily, and one past the 2 MB tile cutoff renders a placeholder carrying its name and size and requests no content until the resource itself is opened. Because a tile is a real read of the bytes, it is audited as one, under the `portal_preview` surface the content route accepts as `?preview=1` -- the only surface that does not stamp `last_read_at`, so browsing a library of photographs cannot clear the never-read flag on every image in it or reorder the recently-read sort. A tag filter sits beside the search box, populated from the tags the resources in view carry plus whichever tag is selected, and sets the `tag` query parameter the list endpoint has always supported; tags are orthogonal to the tree. Rows are multi-selectable, and one Move, Tag or Delete covers the selection: each file is its own PATCH or DELETE, so the report names what happened to each and a refused file stays where it was with the server's reason beside it. A file can be dragged onto a folder and a folder onto another folder; both open the same confirmation the menu action does, because dragging is easy to do by accident and either rewrites an address.

### Permission Model

| Action | Global | Persona | User |
|--------|--------|---------|------|
| **Read** | All authenticated users | Users with matching persona | Owner only |
| **Write (create)** | Platform admins | Platform admins or persona admins (`persona-admin:{name}`) | Any authenticated user (own scope) |
| **Modify/Delete** | Original uploader or platform admin | Original uploader, platform admin, or persona admin | Original uploader or platform admin |

Platform admin roles: `admin`, `platform-admin`.

### REST API Endpoints

All endpoints are mounted at `/api/v1/resources` and require authentication (Bearer token or browser session).

| Method | Path | Description |
|--------|------|-------------|
| `POST` | `/api/v1/resources` | Upload a new resource (multipart form) |
| `GET` | `/api/v1/resources` | List resources visible to the caller |
| `GET` | `/api/v1/resources/{id}` | Get resource metadata by ID |
| `GET` | `/api/v1/resources/{id}/content` | Download resource file content |
| `PATCH` | `/api/v1/resources/{id}` | Update resource metadata (display_name, description, tags), refile it in another folder (`path`), and/or move it to another library (`scope`, `scope_id`) |
| `POST` | `/api/v1/resources/folders/move` | Rename a folder, or nest it under another, by rewriting the path prefix of every resource beneath it in one transaction |
| `DELETE` | `/api/v1/resources/{id}` | Delete resource (metadata and S3 blob) |

List supports query parameters: `scope`, `scope_id`, `path` (a folder and everything beneath it), `tag`, `q` (text search), `offset`.

Upload (POST) accepts multipart form fields: `file` (required), `scope`, `scope_id`, `path`, `display_name`, `description`, `tags`. Maximum upload size: 100 MB. Executable file extensions and MIME types are blocked.

### MCP Protocol Integration

When managed resources are enabled, the `MCPManagedResourceMiddleware` intercepts MCP `resources/list` and `resources/read` requests:

- **resources/list**: Appends managed resources (filtered by caller's visible scopes) to the SDK's static resource list.
- **resources/read**: For URIs matching the configured scheme (default: `mcp://`), looks up the resource in the database, checks read permission, and returns the file content from S3. Text types under 1 MB are returned inline; larger or binary content is returned as a blob. Falls through to the SDK handler for non-matching URIs.

MCP resource capabilities are advertised whenever resource templates, custom resources, or managed resources are enabled.

### Portal UI

A batch of files is loaded through the same create route (#1862). `POST /api/v1/resources` takes an optional `if_exists` form field ahead of the file part: `fail` (the default) answers 409 for an address that already holds a resource, and `skip_unchanged` compares the upload with that resource by SHA-256: identical bytes answer 200 with `outcome: unchanged` and write nothing (the blob written for the comparison is removed), and different bytes are recorded as its next version (200, `outcome: revised`) with its id, URI, name, description and tags kept. Every create answers `outcome` (`created` on 201). The hash is taken as the upload streams to storage and recorded on each version as `content_sha256` (migration 000158); a version written before that has none, so its first re-upload is recorded as a version. The portal's Upload many dialog sends files, a folder, or a `.zip` unpacked in the browser, four at a time, at most 5,000 per batch, each within `max_upload_bytes`, with shared library, folder, tags and a `{name}` description template, and marks the files a document can reference directly as an image.

The Admin Portal includes a Resources page for uploading, browsing, editing, and deleting managed resources. Users browse the folder tree, filter by scope and tags, search the whole library by name or description, upload files, and edit metadata inline.

One resource opens at a route of its own: `/resources/{id}` in the reader's portal, `/admin/resources/{id}` in the administrator's, both registered in the shell's route table so an id that names nothing gets the section's not-found rather than a blank page. Clicking a row navigates there. The page takes the same chrome a portal asset takes, through one shared `ViewerLayout`: the content at the full width of the page with the page area as its only scroll region, and the metadata, tags, canonical URI, read-activity rollup, version history, table registration panel and attaching prompts in a sidebar beside it. Download, Edit and Delete sit in the page header under the same authority the API applies — the uploader or an administrator. Editing, deleting and uploading stay dialogs, being bounded forms. The library's LOCATION -- which library, which folder -- is carried in the route (`/resources/lib/global/data/media-manager`) and its filters in the query string, so Back from a resource returns to the folder it was left in and any level of the tree can be linked to. The viewer's subtitle renders the same `FolderBreadcrumbs` component the library's header does, so the folder trail on a resource's own page is clickable back into the library it lives in.

## Thumbnails Configuration

The platform draws a thumbnail of every portal asset, managed resource and collection itself, in a headless Chrome (`chromedp/headless-shell`, pinned by digest) that runs as a second container in the platform's pod with no Service or port (#1787). Enabled by default; nothing needs to be set when the renderer runs beside the platform.

```yaml
thumbnails:
  enabled: false                        # only needed to opt out; defaults to true
  renderer_url: "http://127.0.0.1:9222" # the renderer's DevTools address; this is the default
  concurrency: 1                        # documents one replica draws at once
  render_timeout: 45s                   # one variant of one document
  batch: 1                              # rows of each kind one pass claims; defaults to concurrency
  lease: 4m10s                          # default worked out from batch, concurrency and render_timeout
  poll: 5s                              # idle wait between passes
  max_attempts: 5                       # tries before a document that never finishes is recorded
  retry_backoff: 1m                     # hold after the first unfinished try; x4 each try, at most 1h
```

The platform dials the renderer and answers every request a page makes itself, so the renderer is never given an address to call and a document reaches no network through it. With no renderer answering, the platform serves normally, files keep their content-type icons, and it logs once that no renderer answers and once when one does. A document that cannot be drawn within `render_timeout`, an image no browser decodes, or an artifact whose linked files did not load is recorded with the reason, shown on the file's Thumbnail panel, and not tried again until the file changes or its owner asks. Every claim charges the row an attempt (`thumbnail_attempts`, migration 000159, on assets, resources and collections); an attempt that does not finish -- the renderer stops answering or closes the page while the document is loaded, the stored file cannot be read, the tile cannot be written -- holds the row back for `retry_backoff`, four times longer each time, and at `max_attempts` records it as not drawable with the last reason, so one document cannot hold the renderer or the queue (#1868). The worker asks the renderer before each document and hands the untried rows back with their attempt returned when it does not answer. A collection mosaic that cannot be composed is recorded against its source (`thumbnail_failure`, `thumbnail_failed_source`) and owed again when the source changes. Pages are loaded with `prefers-reduced-motion: reduce` and their animation timeline stopped in every frame, and just before the capture an animation that ends is shown ended and an infinite one removed. A lease shorter than drawing the batch, or a negative value, is refused at startup.

## Progress Notifications Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `progress.enabled` | `*bool` | `true` (nil = enabled) | Enable progress notifications for Trino queries |

Enabled by default; set `enabled: false` to opt out. When enabled, Trino query tools send three progress notifications per query: before execution, after query returns, and after formatting. Clients must include `_meta.progressToken` in their tool call to receive notifications.

## Client Logging Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `client_logging.enabled` | `*bool` | `true` (nil = enabled) | Enable server-to-client log messages |

Enabled by default; set `enabled: false` to opt out. When enabled, the platform sends log notifications to clients after enrichment is applied (tool name, duration). Uses MCP `logging/setLevel` protocol, with zero overhead if the client hasn't called `setLevel`.

## Elicitation Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `elicitation.enabled` | `*bool` | `true` (nil = enabled) | Enable user confirmation prompts |
| `elicitation.cost_estimation.enabled` | `*bool` | `true` (nil = enabled) | Prompt when EXPLAIN IO estimates exceed the row threshold |
| `elicitation.cost_estimation.row_threshold` | int | `1000000` | Row estimate threshold for cost prompts |
| `elicitation.pii_consent.enabled` | `*bool` | `true` (nil = enabled) | Prompt when query accesses PII-tagged columns |

Enabled by default (including `cost_estimation` and `pii_consent`); set `enabled: false` at any level to opt out. Elicitation requires client-side support (MCP `elicitation/create` capability). When the client doesn't support it, elicitation gracefully degrades to a no-op. If a user declines the confirmation, the tool returns an informational message instead of executing the query.

Behavior change: with no `elicitation` block at all, cost-estimation and PII-consent prompts now fire out of the box (`cost_estimation` still respects `row_threshold`, default 1,000,000 rows, so it only prompts on large queries). Deployments that relied on the previous silent-off default should add `elicitation.enabled: false` to keep the prior behavior.

## Icons Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `icons.enabled` | `*bool` | `true` (nil = enabled) | Enable icon enrichment middleware |
| `icons.tools.<name>.src` | string | - | Icon URI for a tool |
| `icons.tools.<name>.mime_type` | string | - | Icon MIME type |
| `icons.resources.<uri>.src` | string | - | Icon URI for a resource template |
| `icons.prompts.<name>.src` | string | - | Icon URI for a prompt |

Enabled by default; set `enabled: false` to opt out. Upstream toolkits (Trino, DataHub, S3) provide default icons on all tools. This configuration overrides or adds icons for tools, resource templates, and prompts in list responses.

## Audit Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `audit.enabled` | bool | `true` (when a database is available) | Enable audit logging. Set `false` to disable. |
| `audit.log_tool_calls` | bool | `true` | Log MCP tool call events. Set `false` to keep audit on but skip per-tool-call rows. |
| `audit.log_parameters` | bool | `true` | Capture tool-call arguments on each event. Set `false` to drop them; `parameters` then holds only a `result` a tool reports about the call's outcome, or null. |
| `audit.redact_keys` | list of strings | `[]` | Top-level argument keys whose values are replaced with `[REDACTED]` before the event leaves the request path. Case-insensitive; top-level only. |
| `audit.delivery` | string | `async` | Store-write path: `async` (best-effort, never blocks the tool call) or `sync` (writes on the request goroutine for backpressure and zero queue drops). |
| `audit.retention_days` | int | 90 | Days to retain audit events |

Requires `database.dsn`. With a database available and no `audit:` block, both audit and per-tool-call logging are on by default. `enabled: false` disables audit entirely; `log_tool_calls: false` keeps audit on but stops recording per-tool-call events.

Delivery semantics and data captured. Each event records identifiers (`id`, `request_id`, `session_id`, `user_id`, `user_email`, `persona`), the call target (`tool_name`, `toolkit_kind`, `toolkit_name`, `connection`, `event_kind`), the agent's stated reason for the call (`purpose`, see Purpose Configuration), the raw tool-call arguments (`parameters`), the outcome (`success`, `error_message`, `authorized`), timing and size (`timestamp`, `duration_ms`, `request_chars`, `response_chars`, `content_blocks`), transport (`transport`, `source`), and enrichment accounting (`enrichment_applied`, `enrichment_tokens_full`, `enrichment_tokens_dedup`, `enrichment_mode`, `enrichment_match_kind`). The `parameters` field stores arguments verbatim (including full SQL); use `redact_keys` or `log_parameters: false` for sensitive values. Async delivery is best-effort: under sustained backpressure or a crash, queued events are dropped and counted by the `audit_events_dropped_total` metric (which also covers writes that fail or exceed the per-write timeout). Sync delivery writes on the request goroutine with a per-write timeout (5s), trading latency for zero queue-overflow drops; a failed write is still logged, counted, and never fails the tool call. Under a stalled store, sync mode blocks each tool call up to the timeout and draws from the same connection pool as OAuth/sessions/portal (the async writer's single drain goroutine avoids both); graceful shutdown cancels in-flight sync writes.

## Session Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `sessions.store` | string | `memory` | Backend: `memory` or `database` |
| `sessions.ttl` | duration | streamable session_timeout | Session lifetime |
| `sessions.cleanup_interval` | duration | `1m` | Cleanup routine interval |
| `sessions.handles.enabled` | bool | `true` | Explicit session handles (#792): platform_info mints a session_id the model threads on every subsequent tool call. Set false for legacy transport-session behavior. |
| `sessions.handles.ttl` | duration | `8h` | Handle lifetime, refreshed on use |
| `sessions.handles.require` | bool | `true` | Refuse any gated call without a valid platform_info-minted handle (SESSION_REQUIRED); a transport Mcp-Session-Id/stdio sentinel is NOT a fallback (issue #800). `false` falls back to the transport session |

The `database` store requires `database.dsn`. Database-backed sessions survive restarts and support multi-replica deployments.

Explicit session handles (`sessions.handles`, issue #792) adopt the pattern the MCP 2026-07-28 release candidate recommends after removing the `Mcp-Session-Id` header (SEP-2567): `platform_info` mints a `session_id`, every tool advertises it as an input argument (except `platform_info`), and `MCPToolCallMiddleware` validates it against the session store (must exist, be unexpired, and belong to the same authenticated identity), adopts it onto `PlatformContext.SessionID`, and strips it before the handler runs. This makes `platform_info` structurally unskippable and gives audit and provenance a deliberate session key. Unknown/expired/cross-identity handles are refused with `SESSION_EXPIRED`; a missing handle on a gated tool with `require: true` is refused with `SESSION_REQUIRED`. A transport-level session (`Mcp-Session-Id` or the stdio sentinel) is not accepted as a fallback for the requirement (issue #800): it is the churning per-call value the handle exists to replace, so `platform_info` mints and threads a handle on every transport with no stdio carve-out. When `require: false`, a handle-less call falls back to the transport session. The `mcp_session_resolution_total{source}` metric tracks how much traffic still relies on a transport session.

## Toolkit Configuration

### Trino

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `host` | string | required | Trino coordinator hostname |
| `port` | int | 8080/443 | Trino coordinator port |
| `user` | string | required | Trino username |
| `password` | string | - | Trino password |
| `catalog` | string | - | Default catalog |
| `schema` | string | - | Default schema |
| `ssl` | bool | false on the default connection, auto-detected on the others | Enable SSL/TLS. Omitting it on a non-default connection leaves the choice to the host: anything that is not `localhost` or `127.0.0.1` is assumed HTTPS on 443. Write `ssl: false` for plain HTTP (#1436) |
| `ssl_verify` | bool | true on the default connection, inherited on the others | Verify SSL certificates; a non-default connection that omits it takes the default connection's setting |
| `timeout` | duration | 120s | Query timeout |
| `default_limit` | int | 1000 | Default row limit |
| `max_limit` | int | 10000 | Maximum row limit |
| `read_only` | bool | false | Reject write SQL on this connection; set per instance, so the other instances of the same toolkit are unaffected and a call that omits `connection` is judged by the default instance's setting |

`read_only` became per connection in #1269. Before that it was read from the default instance alone and applied to the whole Trino toolkit, which cut both ways: `read_only: true` on the default instance refused write SQL on every Trino connection, and `read_only: true` on any other instance did nothing. A deployment relying on the default instance's `read_only` to cover its other Trino connections must now set `read_only: true` on each connection it wants refused. Connections stored in the database (Admin > Connections) carry the same key, with a Read Only toggle in the Trino connection form; a connection the toolkit holds no setting for - one just added, or a name that is not configured - refuses write SQL until its setting is recorded.
| `scratch.catalog` | string | - | Catalog a registered table is created in on this connection; unset (or set without `schema`) means table registration is unavailable here |
| `scratch.schema` | string | - | Schema a registered table is created in; required alongside `catalog`, and a block naming only one is ignored with a warning |
| `descriptions` | map | {} | Override tool descriptions (key: tool name, value: description text) |

`scratch:` names a TARGET, not a boundary. Nothing in the toolkit restricts a catalog or a schema, and `catalog`/`schema` on a connection are session defaults; what keeps a registration off the warehouse is the Trino identity the connection authenticates as. A read-only warehouse connection beside a write-capable scratch one on the SAME cluster is the worked example, and the difference that matters between them is the Trino account, not the platform flag: `read_only` is a statement-prefix denylist evaluated per connection name, so a scratch connection authenticating as the warehouse's user can write `INSERT INTO warehouse...`. Adding a catalog allow-list to the toolkit is deliberately not done, because parsing SQL is the wrong layer for a boundary the engine already enforces on its own identities. Separately, a SHORT external list of keys needs no table at all: a few hundred join inline with `JOIN (VALUES ('a'),('b')) AS t(id)` or `WHERE id IN (...)` through `trino_query` on a read-only connection, and the platform's instruction baseline says so wherever `trino_query` is accessible, so a handful of pasted keys is joined rather than refused or turned into a request for a table.

### DataHub

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `url` | string | required | DataHub GMS URL |
| `token` | string | - | DataHub access token |
| `timeout` | duration | 30s | API request timeout |
| `default_limit` | int | 10 | Default search limit |
| `max_limit` | int | 100 | Maximum search limit |
| `descriptions` | map | {} | Override tool descriptions (key: tool name, value: description text) |

### S3

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `region` | string | us-east-1 | AWS region |
| `endpoint` | string | - | Custom S3 endpoint for data operations (for MinIO, SeaweedFS) |
| `public_endpoint` | string | - | Public endpoint used only to sign presigned URLs (`s3_object` `presign`); set it to an externally resolvable address. Empty falls back to `endpoint`. |
| `access_key_id` | string | - | AWS access key ID |
| `secret_access_key` | string | - | AWS secret access key |
| `read_only` | bool | false | Restrict to read operations |

## Cross-Enrichment Configuration

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `trino_semantic_enrichment` | bool | true | Add DataHub context to Trino results (default on; set false to disable) |
| `datahub_query_enrichment` | bool | true | Add Trino availability to DataHub results (default on; set false to disable) |
| `s3_semantic_enrichment` | bool | true | Add DataHub context to S3 results (default on; set false to disable) |
| `unwrap_json` | bool | true | Auto-unwrap single-row VARCHAR-of-JSON results (e.g. from raw_query) |
| `column_context_filtering` | bool | true | Limit column enrichment to SQL-referenced columns |
| `estimate_row_counts` | bool | false | Run `SELECT COUNT(*)` when reporting table availability so enriched DataHub results carry an estimated row count, and so the insight review path can state a row count beside a pending claim (and flag one that disagrees). Off by default: `COUNT(*)` can trigger a full table scan and make search enrichment very slow |
| `search_schema_preview` | bool | true | Add bounded column-name+type preview to datahub_search query_context |
| `schema_preview_max_columns` | int | 15 | Max columns per entity in schema preview |
| `semantic_fallback` | bool | false | When a Trino-table URN lookup misses on the semantic provider, fall back to a similarity search and surface the top-K hits as SUGGESTED matches annotated with `match_kind: semantic`. Default off; operators opt in. Issue #444. |
| `semantic_fallback_top_k` | int | 1 | Number of similarity-search hits surfaced per URN miss when `semantic_fallback` is on. Clamped to [1, 10]. |
| `memory_limit` | int | 5 | Max memory records recalled and rendered into `memory_context` per tool call. Issue #761. |
| `memory_context_budget_bytes` | int | 1500 | Byte budget for the rendered memory summaries; records beyond it are listed as compact `id`+`reference` stubs in `memory_context_omitted` (still fetchable, at least one always rendered). `0` disables the budget. |
| `memory_summary_bytes` | int | 280 | Per-record summary-first excerpt cap; the full record is fetchable via its `mcp:memory:<id>` reference. `0` renders full content. |

When `semantic_fallback` fires, the appended content carries `semantic_fallback.match_kind: semantic`, a human-readable note, the queried table identifier, and a list of suggested matches (urn, name, platform, description, tags, domain). The audit row's `enrichment_match_kind` column records `semantic` for these calls and `urn` for exact lookups, so operators can compute the false-positive rate of similarity-based suggestions with a SQL aggregate.

The fallback fires only on the single-table enrichment path (e.g. `trino_describe_table`); multi-table SQL queries fall through with no suggestions until a follow-up extends the same hook to `enrichTrinoQueryResult`.

## URN Mapping Configuration

When Trino catalog or platform names differ from DataHub metadata, configure bidirectional URN mapping:

```yaml
semantic:
  provider: datahub
  instance: primary
  urn_mapping:
    platform: postgres           # DataHub platform (e.g., postgres, mysql, trino)
    catalog_mapping:
      rdbms: warehouse           # Trino "rdbms" → DataHub "warehouse"
      iceberg: datalake          # Trino "iceberg" → DataHub "datalake"

query:
  provider: trino
  instance: primary
  urn_mapping:
    catalog_mapping:
      warehouse: rdbms           # DataHub "warehouse" → Trino "rdbms"
      datalake: iceberg          # DataHub "datalake" → Trino "iceberg"
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `semantic.urn_mapping.platform` | string | trino | Platform name in DataHub URNs |
| `semantic.urn_mapping.catalog_mapping` | map | {} | Map Trino catalogs to DataHub catalogs |
| `query.urn_mapping.catalog_mapping` | map | {} | Map DataHub catalogs to Trino catalogs (reverse) |

This translates URNs in both directions:
- **Trino → DataHub**: `rdbms.public.users` becomes `urn:li:dataset:(urn:li:dataPlatform:postgres,warehouse.public.users,PROD)`
- **DataHub → Trino**: URN with `warehouse.public.users` resolves to Trino table `rdbms.public.users`

---

# Authentication & Security

mcp-data-platform implements a **fail-closed** security model. Missing or invalid credentials deny access, never bypass. For the security architecture rationale, see: https://imti.co/mcp-defense/

## A broker, not an identity provider

**mcp-data-platform is an OAuth 2.1 broker, not an identity provider.** No person authenticates to it: there is no login form, no user password to verify, and no MFA. A human's identity comes from an existing IdP (Keycloak, Auth0, Okta, Azure AD) over OIDC; `/authorize` redirects the browser there and refuses the flow outright when no upstream IdP is configured, and the roles and email that person is authorized against are the ones the IdP asserts. Service accounts authenticate with API keys instead, and their roles come from local configuration. It stores no human passwords, and no migration in the tree defines a password column. The secrets it does hold are machine credentials: API keys and the client secrets Dynamic Client Registration issues to MCP client software are bcrypt hashes, the authorization codes and tokens the platform itself issues are SHA-256 digests, and refresh tokens for upstream services are encrypted at rest (AES-256-GCM when `ENCRYPTION_KEY` is set; the server warns loudly at startup when it is not).

The platform presents an authorization server toward MCP clients because the MCP specification requires a discoverable authorization server supporting Dynamic Client Registration, which upstream IdPs generally do not expose. The broker shape is what the spec requires, not a decision to reimplement identity.

| Concern | Implementation |
|---|---|
| `redirect_uri` matching | Exact match for non-loopback; RFC 8252 section 7.3 handling for loopback (`pkg/oauth/storage.go`) |
| DCR abuse | Plain HTTP to non-loopback hosts refused regardless of configuration; private-use schemes excluded from `AllowAllRedirectURIs` (`pkg/oauth/dcr.go`) |
| Brute force and registration flood | Per-IP token-bucket limits on `/token` and `/register`, applied before the bcrypt work they would otherwise burn (`pkg/oauth/ratelimit.go`) |
| Authorization | Deny-before-allow, default deny, fail-closed on unresolved persona (`pkg/persona/filter.go`) |
| Prompt injection carried in catalog metadata | Untrusted descriptions, tags, and owner notes sanitized before reaching the model; detected attempts logged (`pkg/semantic/sanitize.go`, `pkg/semantic/injection_logger.go`) |

## Engineering posture

More than 1.25 lines of test code per line of production Go, with the security-critical packages carrying the highest ratios in the tree: `pkg/oauth` and `pkg/middleware` are both above 2:1. Fuzz suites cover `pkg/oauth`, `pkg/auth`, `pkg/platform`, and `pkg/middleware`. Every PR passes race-detector tests, `golangci-lint`, `gosec`, and Semgrep and CodeQL SAST, under a coverage floor enforced in CI. Release artifacts are Cosign-signed with GitHub build-provenance attestations, and supply-chain posture is tracked by OpenSSF Scorecard. These ratios are kept true mechanically: `make posture-check` recomputes them and fails when the tree crosses a stated line.

## Transport Security

| Transport | Authentication | TLS | Why |
|-----------|---------------|-----|-----|
| stdio | Not needed | N/A | Local execution with your own credentials |
| HTTP | **Required** | Recommended | Shared server needs to identify users |

## Security Model

- **Required JWT Claims**: Tokens must include `sub` (subject) and `exp` (expiration)
- **Self-Healing JWKS Cache**: OIDC signing keys are fetched at startup and cached for one hour. The key-lookup path refreshes on demand (during that request's validation, honoring the request's deadline) when the cache expires or a token bears an unknown key ID (IdP key rotation), so validation survives rotation without a restart. Concurrent requests collapse into a single fetch that runs independently, so a slow issuer cannot pin request goroutines. Refreshes are throttled by the last fetch's outcome: at most once per minute after a success (anti-hammer), and a short recovery window after a failure so a brief issuer outage heals within seconds. A failed refresh on an expired cache rejects the token rather than accepting it unverified.
- **Default-Deny Personas**: Users without explicit persona have no tool access, and connection rules are deny-by-default on the same call, so access is scoped at the connection rather than the end user (see the Authorization Model section)
- **Prompt Injection Protection**: DataHub metadata is sanitized before exposure
- **Read-Only Enforcement**: Trino `read_only: true` blocks write queries at query level, per connection: the interceptor resolves the connection the call names (the default when it names none) and refuses only when that connection sets it
- **Cryptographic Request IDs**: Secure random identifiers for audit trails

## stdio Transport (Local)

No MCP authentication required. The server runs on your machine using credentials you configured.

## HTTP Transport (Remote/Shared)

Authentication identifies who is making requests. Anonymous access is disabled by default.

- **OIDC**: Human users via Keycloak, Auth0, Okta
- **API Keys**: Service accounts, automation
- **OAuth 2.1 (inbound)**: Claude Desktop authentication via upstream IdP — see "OAuth 2.1 for Claude Desktop" below
- **OAuth 2.1 (outbound)**: Gateway connections to upstream MCPs — `client_credentials` (M2M) and `authorization_code` + PKCE (browser sign-in). Encrypted refresh tokens persist across restarts. See "OAuth to Upstream MCPs" later in this document.

### OAuth 2.1 for Claude Desktop

For Claude Desktop connecting to a remote MCP server, the built-in OAuth 2.1 server bridges authentication to any OIDC-compliant upstream identity provider (Keycloak shown):

1. Claude Desktop calls `/oauth/authorize`
2. MCP server redirects to the upstream IdP's discovered authorization endpoint
3. User logs in to the upstream IdP
4. Upstream IdP redirects back to MCP server's `/oauth/callback`
5. MCP server exchanges code at the upstream IdP's discovered token endpoint, extracts user info
6. MCP server issues its own token to Claude Desktop

Configuration:
```yaml
oauth:
  enabled: true
  issuer: "https://mcp.example.com"
  signing_key: "${OAUTH_SIGNING_KEY}"  # Required on http transport: openssl rand -base64 32
  previous_signing_keys:               # Optional: verify-only keys retained across a rotation
    - "${OAUTH_SIGNING_KEY_OLD}"
  clients:
    - id: "claude-desktop"
      secret: "${CLAUDE_CLIENT_SECRET}"
      redirect_uris:
        - "http://localhost"
        - "http://127.0.0.1"
  upstream:
    issuer: "https://keycloak.example.com/realms/your-realm"
    client_id: "mcp-data-platform"
    client_secret: "${KEYCLOAK_CLIENT_SECRET}"
    redirect_uri: "https://mcp.example.com/oauth/callback"
    # Optional overrides; empty = discover from the issuer's well-known document.
    # authorization_endpoint: "https://idp.example.com/authorize"
    # token_endpoint: "https://idp.example.com/token"
```

The broker resolves the upstream authorization and token endpoints via OIDC discovery: on first use it fetches `<upstream.issuer>/.well-known/openid-configuration` and reads `authorization_endpoint` and `token_endpoint`, so the brokered flow works with any OIDC-compliant IdP rather than assuming Keycloak's URL shape. Resolved endpoints are cached for the process lifetime; a discovery failure is not cached, so `/oauth/authorize` returns a retryable `server_error` and retries on the next request rather than falling back to a fixed path. Set `upstream.authorization_endpoint` and `upstream.token_endpoint` to bypass discovery for an IdP with a broken discovery document; discovery is skipped entirely only when both are set.

Access tokens are signed JWTs whose `aud` claim is the issuer URL (the platform is both authorization server and resource server); the requesting client is carried in the `client_id` claim, and the authenticator rejects tokens minted for any other audience.

Access tokens are HS256-signed and carry a `kid` header derived from the signing key, so verifiers select the right key after a rotation. `oauth.signing_key` is **required** when OAuth is enabled on `server.transport: http`/`sse` (startup fails otherwise, naming the key) because an auto-generated per-process key makes each replica reject its peers' tokens; set `oauth.allow_ephemeral_signing_key: true` to override for a single-replica dev setup. On `stdio` the key is auto-generated when omitted. Rotate in two phases so every replica can verify the new key before any replica signs with it (a single-phase swap causes intermittent 401s during the rolling restart): (1) add the new key to `oauth.previous_signing_keys` as verify-only while `signing_key` stays the old key, roll-restart; (2) promote by setting `signing_key` to the new key and moving the old key into `previous_signing_keys`, roll-restart; then drop the old key after the access-token TTL (1h) elapses. Tokens minted before `kid` support carry no header and fall back to current-then-previous key verification, so an in-place upgrade preserves live sessions.

Dynamic Client Registration (`oauth.dcr`) is off by default. Enabling it requires `allowed_redirect_patterns` (regex list): the registration endpoint is unauthenticated, so registration is denied when no patterns are configured (a startup warning flags the misconfiguration). `allow_all_redirect_uris: true` explicitly opts out of pattern matching. Plain HTTP redirect URIs are rejected for non-loopback hosts regardless of configuration; private-use schemes (RFC 8252 native apps) are accepted only through an explicit pattern.

The unauthenticated `/token` and `/register` endpoints are rate limited by default (`oauth.rate_limit`, disable with `enabled: false`). `/token` runs a bcrypt compare per attempt and `/register` a bcrypt hash plus a DB insert per request, so both are CPU/storage amplification levers. Each endpoint has a per-client-IP token bucket (`/token` 60 rpm / burst 10, `/register` 10 rpm / burst 3) plus a global backstop bucket sized at ten times the per-IP rate that bounds total throughput even when per-IP attribution is defeated or collapses behind a proxy. Client attribution is trusted-proxy-aware: `X-Forwarded-For` is consulted only when the direct peer is in `rate_limit.trusted_proxies` (a CIDR list), taking the rightmost untrusted hop; with no trusted proxies configured the header is ignored and the direct peer address is used (the spoof-safe default). On limit the endpoint returns HTTP 429 with `Retry-After` and `{"error":"slow_down"}`. Separately, DCR-registered clients never issued a token are reaped 24h after registration by the OAuth store's cleanup loop (a `dcr` column marks them; pre-registered config-file clients are exempt), bounding `oauth_clients` growth. The per-IP token-bucket implementation lives in the shared `pkg/ratelimit` package, reused by the public portal viewer, which composes the same trusted-proxy-aware resolver and global backstop (per-IP 60 rpm / burst 10 by default) and reads its trusted-proxy CIDRs from `portal.rate_limit.trusted_proxies` (empty trusts none: the direct peer address is used and `X-Forwarded-For` is ignored).

When a database is configured, in-flight authorization state (linking `/oauth/authorize` to the upstream IdP callback) is stored in PostgreSQL so multi-replica deployments work behind a load balancer: the callback can land on a different replica than the one that started the flow.

Refresh tokens and authorization codes are stored only as hex-encoded SHA-256 digests (in both the PostgreSQL and in-memory stores); the server hashes the presented value on lookup, so a database read, backup, or replica never yields a usable bearer credential. Migration 000078 converts pre-existing plaintext rows in place so live sessions survive the upgrade; rolling it back deletes all outstanding refresh tokens and authorization codes, forcing interactive re-authorization.

## Browser Sessions (OIDC Login for Portal UI)

When `auth.oidc` and `auth.browser_session` are both enabled, the portal UI offers SSO login. The flow uses OIDC authorization code with PKCE. Sessions are stored as HMAC-SHA256 signed JWT cookies (stateless, no server-side store).

```yaml
auth:
  oidc:
    enabled: true
    issuer: "https://auth.example.com/realms/platform"
    client_id: "mcp-data-platform"
    client_secret: "${OIDC_CLIENT_SECRET}"
    scopes: [openid, profile, email]
  browser_session:
    enabled: true
    signing_key: "${SESSION_SIGNING_KEY}"  # openssl rand -base64 32
    ttl: 8h
    secure: true
    same_site: lax  # lax (default), strict, or none
```

Endpoints: `/portal/auth/login` (initiate; accepts optional `return_to` site-relative path), `/portal/auth/callback` (process; redirects to `return_to` when set, otherwise `PostLoginRedirect`), `/portal/auth/logout` (clear). API key fallback always available. MCP clients unaffected.

Limitations: no individual session revocation, no key rotation support (rotating invalidates all sessions), no session refresh (re-auth required after TTL).

CSRF protection: cookie-authenticated state-changing requests (POST/PUT/PATCH/DELETE) to the portal, admin, and managed-resources APIs require an `X-CSRF-Token` header; missing/invalid tokens are rejected with 403. The token is stateless (an HMAC over the session subject under the signing key), returned by `GET /api/v1/portal/me` as `csrf_token` and echoed by the SPA on every mutation. Read-only requests and API-key/Bearer-authenticated requests are exempt (those credentials are not attached automatically by the browser). `SameSite=Lax` is kept as defense-in-depth; `same_site: none` disables that browser defense (requires `secure: true`) and logs a startup warning.

## Threat Model

The threat model states the security model as a whole (reviewed at v1.101.x): trust boundaries, the attackers each mechanism is designed to stop, and what is explicitly out of scope. Every mitigation claim carries a package or config citation a reviewer can verify.

Trust boundaries fall into four groups. Inbound surfaces: MCP over stdio (local-process trust, no MCP-layer auth) versus HTTP (streamable + SSE, bearer/API-key gate that returns 401 with `WWW-Authenticate` and fails closed on an invalid token); the OAuth 2.1 endpoints; the admin REST API (admin persona required, CSRF on the cookie path); the portal API and its share-token viewer (rate-limited, access-mode gated, 403 for a caller the mode does not admit, 410 on revoked/expired; refused browser navigations get a branded landing page, and email-share recipients can request single-use guest view links: uniform response, per-share issue cap, hashed 15-minute tokens claimed atomically, view-only guest cookie scoped to one share under a key domain-separated from the browser-session signing key); the no-login notification unsubscribe endpoint (HMAC token bound to the recipient address, can only opt that address out); the gateway REST shim (runs through an in-memory MCP session so persona and audit apply); the observability PromQL proxy (`observability:read` capability, admin-only); unauthenticated health endpoints (liveness/readiness only); and managed resource upload/download (per-request auth plus scope checks). Identity mechanisms: OIDC bearer validation with a self-healing JWKS cache; API keys (constant-time compare for config keys, bcrypt for DB keys); the built-in OAuth 2.1 server (HS256 with a `kid` key ring); browser cookie sessions (HttpOnly, Secure, SameSite, HS256-signed); and persona default-deny authorization. Outbound dependencies: Trino, DataHub, S3 with per-connection operator-authored credentials; upstream MCP servers and HTTP APIs; outbound OAuth (one shared upstream identity per connection, #374); and the embedding provider. At rest: Postgres holding SHA-256-hashed OAuth tokens (single-use via atomic `DELETE ... RETURNING`), bcrypt-hashed client secrets, AES-256-GCM-encrypted connection secrets (single symmetric key from `ENCRYPTION_KEY`, not a DEK/KEK envelope), and audit rows subject to `redact_keys` and the `log_parameters` opt-out.

Identity-provider outage is a recorded decision rather than an emergent property. Validation has three outcomes: valid, definitively invalid, and undetermined (the IdP is unreachable, so signing keys cannot be fetched — the `ErrValidationUnavailable` sentinel, wrapped only on a JWKS fetch failure with an expired cache). On the third, the HTTP edge passes the request through instead of answering 401, because a 401 asserts a bad credential and would drive every valid user into a re-auth flow that cannot complete during the outage; the protocol layer then refuses each tool call with `feature_unavailable` and a retryable hint, so no tool handler runs. Access is never granted on an unvalidated credential; the cost is depth, since an unauthenticated request reaches the protocol layer before refusal. A client sees successful transport, an established MCP session, and in-band retryable errors on every tool call, with nothing signaling re-auth. Both halves are one control, pinned end to end by `TestStreamableHTTP_ValidationUnavailable_EdgePassesThroughProtocolRefuses` (`internal/httpserver/validation_outage_test.go`) with `TestStreamableHTTP_DefinitiveRejection_EdgeFailsClosed` holding the boundary.

Attacker analysis covers six personas: an unauthenticated network attacker (rate-limited OAuth endpoints with a global backstop, trusted-proxy-aware client-IP resolution that ignores spoofed `X-Forwarded-For`, access-mode-gated share viewer, anonymous only for shares explicitly created public); an authenticated low-privilege persona (default-deny on tools, connections, and API routes; admin and observability surfaces gated separately; Trino cost/PII elicitation consent); a malicious or compromised upstream (tool descriptions and responses flow through without content sanitization, mitigated operationally by per-connection encrypted credentials and audit, not by neutralizing adversarial content); malicious data in query results (prompt-injection content reaches the client; the platform reduces blast radius via read-only mode, persona allowlists, the search-first gate, and audit, but does not scrub content); a database reader (hashed tokens and encrypted secrets yield no usable credentials, but audit content and metadata remain readable, and connection secrets are plaintext if `ENCRYPTION_KEY` is unset); and a compromised downstream credential (blast radius bounded by the downstream service account's own privileges, which least-privilege scoping is a deployment responsibility).

Explicit non-goals and residual risks: stdio transport trusts the invoking process with no in-process sandboxing; the platform does not defend against a malicious admin (who can hold `ENCRYPTION_KEY`); audit delivery is best-effort async by default with a documented loss model; downstream identity is per connection rather than per user, which is the authorization design rather than an unmitigated gap (no token exchange, no impersonation, no session-user propagation to Trino, S3, DataHub, upstream MCP servers, or HTTP APIs, because the construct all of those backends share is a credential and an endpoint, not a caller identity to pass through; the cost, stated without softening, is that all callers granted a connection are indistinguishable downstream and warehouse row policies or column masks keyed off the end user do not follow a caller through, so per-person policy is expressed as one connection per distinct outcome and per-user attribution comes from the audit trail); content sanitization is partial (DataHub semantic metadata is sanitized via `pkg/semantic/sanitize.go`, but raw query row values and gateway-upstream tool descriptions/responses are not); and TLS termination, network segmentation, Postgres transport security, and S3 blob-at-rest are deployment responsibilities. Full document with the mitigations table and citations: https://mcp-data-platform.txn2.com/security/threat-model/

---

# Audit Logging

Every tool call flows through the audit logging middleware, which records who called what, when, how long it took, and whether it succeeded. Audit logs are stored in PostgreSQL and automatically cleaned up based on a configurable retention period.

## Prerequisites

Audit logging requires:
1. A PostgreSQL database (version 13+)
2. The `database.dsn` configuration set

With a database available, both `audit.enabled` and `audit.log_tool_calls` default to on, so no `audit:` block is needed to record tool calls. Set either to `false` to opt out.

## Configuration

```yaml
database:
  dsn: "${DATABASE_DSN}"

audit:
  enabled: true
  log_tool_calls: true
  log_parameters: true
  redact_keys: ["password", "token"]
  delivery: async          # async (default) | sync
  retention_days: 90
```

| Option | Type | Default | Description |
|--------|------|---------|-------------|
| `database.dsn` | string | - | PostgreSQL connection string. Required for audit logging. |
| `audit.enabled` | bool | `true` (when a database is available) | Master switch for audit logging. Set `false` to disable. |
| `audit.log_tool_calls` | bool | `true` | Log every `tools/call` request. Set `false` to keep audit on but skip per-tool-call rows. |
| `audit.log_parameters` | bool | `true` | Capture tool-call arguments. Set `false` to drop them; `parameters` then holds only a `result` a tool reports about the call's outcome, or null. |
| `audit.redact_keys` | list of strings | `[]` | Top-level argument keys whose values become `[REDACTED]` before the event leaves the request path. Case-insensitive; top-level only. |
| `audit.delivery` | string | `async` | Store-write path: `async` (best-effort, never blocks the tool call) or `sync` (writes on the request goroutine for backpressure and zero queue drops). |
| `audit.retention_days` | int | `90` | Days to keep audit logs before automatic cleanup. |

## What Gets Logged

Every tool call produces one row in `audit_logs` with these fields:

| Field | Type | Description |
|-------|------|-------------|
| `id` | VARCHAR(32) | Cryptographically random event ID (base64url, 16 bytes). |
| `timestamp` | TIMESTAMPTZ | When the tool call started. |
| `duration_ms` | INTEGER | Wall-clock time in milliseconds. |
| `request_id` | VARCHAR(255) | Request ID from auth middleware. |
| `user_id` | VARCHAR(255) | The principal that made the call: an OIDC subject for a person, `apikey:<name>` for a key, `script:<name>` for a managed-script run. |
| `user_email` | VARCHAR(255) | The address the principal acts for (the script owner's on a run). Not an identity: an owner's every script writes the same address, so `user_id` is what separates two rows, and `GET /api/v1/admin/audit/events/filters` returns both the distinct ids and a `user_labels` map between them. |
| `persona` | VARCHAR(100) | Resolved persona name. |
| `tool_name` | VARCHAR(255) | MCP tool name. |
| `toolkit_kind` | VARCHAR(100) | Toolkit type: `trino`, `datahub`, or `s3`. |
| `toolkit_name` | VARCHAR(100) | Toolkit instance name from config. |
| `connection` | VARCHAR(100) | Connection name from toolkit config. |
| `parameters` | JSONB | Tool call arguments with sensitive values redacted. |
| `success` | BOOLEAN | `true` if tool handler succeeded. |
| `error_message` | TEXT | Error description if `success` is `false`. |
| `event_kind` | VARCHAR(64) | High-level category: `mcp_tool_call` and `apigateway_invoke` for tool calls, plus `prompt_serve`, `resource_read`, `resource_move`, `script_run` and `admin`. See [Event kind](#event-kind). |
| `created_date` | DATE | Partition key derived from `timestamp`. |

## Sessions read back from the log

Every audit row carries the `session_id` of the call that wrote it, so grouped by that id the log is a record of sessions. Sessions are DERIVED, never stored: the platform's own session rows are working state with a TTL that the store deletes on expiry, while audit rows persist for the retention window, so there is no new table, no second write path, and nothing to backfill. A session's kind comes from its id prefix, the only classification available since isolated runs' ids are never persisted: `dps_` agent handle, `dpp_` portal run, `dpx_` managed-script run, and bare hex for a transport session or a call recorded before handles existed. A summary carries the caller and persona of the session's first event, its first and last call, call and failure counts, the distinct tools and connections it touched, and what it produced — with one exception: while the live session record still exists, the persona the handle was MINTED under outranks the per-call persona, because that is what the session was authorized as. Output is read from the two places the platform records it: `portal_assets` rows whose `session_id` matches (deleted ones excluded), and knowledge-dimension `memory_records` whose metadata carries the capturing session (migration 000031 folded `knowledge_insights` into `memory_records`). Opened, a session reads as its summary, those outputs, and the ordered record of its calls, each with the purpose the agent stated, the connection it reached, its outcome, and its duration — the reading order the flat event list cannot give. Served by `GET /api/v1/admin/sessions` (filters: user, kind, time range, has_assets, has_failures) and `GET /api/v1/admin/sessions/{id}` (summary, assets, insights, paged timeline), with `session_id` accepted on `GET /api/v1/admin/audit/events`; in the portal under Admin > Sessions. Migration 000106 adds the two supporting indexes (`idx_portal_assets_session_id` on live assets, `idx_memory_records_session_id` on the metadata expression); `audit_logs` was already indexed on `session_id`.

## Parameter Sanitization

A built-in baseline replaces the values of these well-known keys with `[REDACTED]`: `password`, `secret`, `token`, `api_key`, `authorization`, `credentials` (case-sensitive, exact match). This is a safety net, not a substitute for configuring your own sensitive argument names.

For your own keys, set `audit.redact_keys`, a case-insensitive list of top-level argument keys whose values are masked in the middleware, before the event ever leaves the request path (nested keys are not matched). To drop arguments entirely, for tools whose inputs cannot be made safe to retain even with redaction, set `audit.log_parameters: false`. A `result` key a tool reports about the call's outcome (an API gateway page walk records `pages_fetched`, `items_merged`, and `stopped_by` there) is not an argument value and is recorded either way; with arguments dropped it is all `parameters` holds.

## Middleware Chain

The middleware sits in a layered execution chain:

```
ResultType → Icons → DescOverrides → Visibility → OutputSchema → Apps Metadata → Auth/Authz → Audit → Rules → Client Logging → Enrichment → Tool Handler
```

Blocks the chain appends — enrichment context, the proven queries a describe carries, the `call_reference` stamped on a data call — are added as JSON text content and merged into `structuredContent`, but only into a structured result the tool's own handler set, and only when it is a JSON object. A handler that returned none keeps its response as written: a structured result synthesized from the appended blocks alone is not context added to the tool's output, it is the tool's output replaced by the platform's additions. `trino_export` is the case this protects, its asset metadata (`asset_id`, `portal_url`, `row_count`, `size_bytes`) being a text block and nothing else; the same holds for a gateway-proxied tool whose upstream answered in text.

For `tools/call` requests: Auth creates the `PlatformContext` (user identity, persona, toolkit metadata) and records tool calls for workflow tracking. Audit reads it to build the audit event. Rules checks session workflow state and prepends warnings when needed. Enrichment adds cross-service context and optional discovery notes. Unauthorized requests are rejected before reaching audit, so only authenticated tool calls appear in audit logs. ResultType, the outermost layer, types every `tools/call`, `prompts/get` and `resources/read` result with the `resultType` the negotiated protocol revision (`2026-07-28` and later) requires: the SDK stamps only the result its own handler returns, so a gate's refusal, the error contract's rebuild of a bare error, and a managed resource read, all built by the layers below, are typed here; an older client receives the field unset, as the SDK leaves it.

For `tools/list` responses: Icons, DescOverrides, and Visibility modify tool metadata and filter the tool list, and OutputSchema opens the top level of every advertised output schema (`additionalProperties` allowed, nothing required, the error envelope documented under `error`) so a client validating a result against the schema it was advertised accepts the keys the chain adds to `structuredContent` — the error envelope, `call_reference`, enrichment. Every nested object is advertised open as well (#1945), so a key a later release adds to a nested type (platform_info's `notices`) validates against a schema a client cached before the upgrade; nested `required` lists are kept. A toolkit registering typed handlers without an explicit schema (mcp-trino before v1.4.0) otherwise advertises the closed schema the SDK infers from its struct; the server still validates the handler's own output against that inferred schema before any middleware runs.

Audit delivery is configurable (`audit.delivery`). In the default **async** mode, events are enqueued on a bounded writer drained by a single background goroutine, so store latency never blocks a tool response, and a graceful shutdown drains that queue (bounded by a 10s deadline) before the store and database are closed; under sustained backpressure or a crash, queued events are dropped and counted by the `audit_events_dropped_total` metric. In **sync** mode, the event is written on the request goroutine with a per-write timeout, trading latency for backpressure and zero queue-overflow drops. In either mode, a write that fails is logged via `slog.Error` but the tool call still succeeds.

### Error contract

Every failed `tools/call` returns a uniform, self-describing error so an agent consumer can tell a correctable mistake from a platform problem instead of guessing (and falsely concluding the platform is down). A failure sets `isError: true`, a self-describing text message (`<message> (code: <code>) Hint: <hint>`), and a machine-readable `structuredContent.error` object `{code, category, message, hint}`. `code` is a stable identifier (`missing_required_parameter`, `invalid_arguments`, `not_found`, `unauthenticated`, `unauthorized`, `setup_required`, `feature_unavailable`, `internal_error`, `tool_error`); `category` is the broad class an agent branches on (`client_input`, `not_found`, `authentication_failed`, `authorization_denied`, `user_declined`, `setup_required`, `feature_unavailable`, `internal`, `tool_error`). A central normalization middleware (`MCPErrorContractMiddleware`, registered inner to Audit/Metrics) guarantees the envelope is present even when a tool returns only a bare message, recovers a panicking handler into a categorized `internal_error` result rather than a dropped connection, and leaves source-categorized results (auth/authz, missing-parameter, not-found, feature-unavailable, setup-required) untouched. The `category` continues to feed the audit log (`error_category`) and metrics (`status_category`). It is built by extending the existing `PlatformError` / `CategorizedError` rather than a parallel type, and covers every toolkit (including the proxied trino/datahub/s3 tools) at the shared chain, so no upstream-library change is required. The platform's own tools publish input schemas closed to unknown top-level arguments (`"additionalProperties": false`), so a misnamed argument is refused at the boundary before the handler runs, with `code: invalid_arguments`, `category: client_input`, and a message naming the offending property; nested maps that carry a foreign namespace (`query_params`, `headers`, `body` on the `api_*` tools) stay open, and gateway-proxied tools keep the upstream server's own schema.

## Caller class via `source`

Tools on this platform are reachable through three entry points; the audit row's `source` field records which one was used:

| `source` | Caller |
|----------|--------|
| `mcp` | Real MCP transport (stdio or HTTP/SSE). Agents. |
| `rest` | Gateway REST shim at `POST /api/v1/gateway/{connection}/invoke`. Apache NiFi, cronjobs, integrations, anything HTTP that wraps the platform. |
| `admin` | Admin REST API at `POST /api/v1/admin/tools/call`. Portal-driven tool runs. |

Both REST entry points open an in-memory MCP session against the assembled server and tag the context (`middleware.WithSource`) before the call so the existing audit middleware records the originating class. Filter via `GET /api/v1/admin/audit/events?source=...` or the **All Sources** dropdown in the portal.

Portal-driven runs (`source=admin`) get a distinct portal session ID (`dpp_…` prefix), minted per request by the tool-call middleware, so they are attributable and their search-first gate, provenance, and dedup state are isolated from the operating admin's own agent session. Both stateless shims (`admin`, `rest`) are exempt from the search-first and `SESSION_REQUIRED` gates, so a portal "test this tool" / replay run of a query tool always executes.

## Event kind

The `event_kind` field separates the classes of audited activity that share the row shape. A tool call's kind is derived at write time from the toolkit kind — `apigateway_invoke` when the toolkit kind is `api`, `mcp_tool_call` otherwise — so the split does not rely on tool-name matching, and it lets the MCP Activity view exclude high-volume gateway traffic by default while a dedicated gateway view includes it. The other kinds are stamped by the surface that writes them. These seven are the complete set the column holds and the set the filter accepts.

| `event_kind` | Meaning |
|--------------|---------|
| `mcp_tool_call` | Tool routed through the trino, datahub, s3, or MCP gateway toolkits. |
| `apigateway_invoke` | HTTP API call through the apigateway toolkit (`api_*` tools). |
| `prompt_serve` | A database prompt served to an agent (`prompts/get`, or a resolved `manage_prompt use`). |
| `resource_read` | A managed resource's content served through one of four doors, named by `surface`: `mcp_read`, `fetch`, `rest_download`, or `portal_preview`. |
| `resource_move` | A managed resource refiled in another library; carries the scope, scope id and URI on both sides. |
| `script_run` | One execution of an approved managed script, written by the run worker when the run finishes; carries the run id as its `session_id`. |
| `admin` | An administrative act against somebody else's object; today, an administrator transferring a managed script to another owner. |

Accepted on the event list, stats, and all metrics endpoints (timeseries, breakdown, overview, performance, enrichment, discovery): `GET /api/v1/admin/audit/events?event_kind=apigateway_invoke`. `GET /api/v1/admin/audit/events/filters` returns the kinds actually present in the data, which is narrower on a deployment that has not used every surface; a filter naming a kind with no rows returns `200` with an empty result rather than an error.

## Retention and partition rotation

A maintenance routine runs every 24 hours and performs three ordered steps:

1. **Ensure upcoming partitions**: `CREATE TABLE IF NOT EXISTS audit_logs_YYYY_MM PARTITION OF audit_logs FOR VALUES FROM (...) TO (...)` for the next two months. The current month is intentionally skipped on brownfield deployments so existing rows in `audit_logs_default` do not block the new named partition. Idempotent.
2. **Delete expired rows**: `DELETE FROM audit_logs WHERE timestamp < NOW() - INTERVAL '<retention_days> days'`. PostgreSQL prunes to overlapping partitions.
3. **Drop fully expired partitions**: `DROP TABLE IF EXISTS audit_logs_YYYY_MM` for any monthly partition whose entire date range ends at or before the retention cutoff. Constant-time reclamation versus row-level DELETE.

Step failures are logged and isolated; a partition operation that errors does not skip the retention DELETE on the same tick. The eager pre-tick `EnsureMonthlyPartitions` also runs once at startup. This pattern keeps audit_logs bounded even at high write volume (e.g. Apache NiFi calling the gateway REST shim at order-of-magnitude-per-second).

In multi-replica deployments, the entire maintenance tick (including the eager startup ensure) is guarded by `pg_try_advisory_lock` on a stable key. Every replica still ticks every 24h, but only the pod that wins the lock runs the work; the rest exit silently and retry next tick. This makes the DELETE scan, partition CREATE, and partition DROP run exactly once per tick across the cluster, regardless of replica count.

---

# Observability (Prometheus Metrics)

Audit answers "who did what." Metrics answer "how is the system performing." Phase 1 of the OpenTelemetry instrumentation in mcp-data-platform exposes Prometheus-format metrics on a dedicated `/metrics` HTTP listener, separate from the main MCP/HTTP listener so scrape traffic and tool traffic do not share an auth path.

## Configuration

Metrics are enabled by default. Configuration is environment-only for Phase 1 to keep the surface small. The full surface lands in subsequent phases:

| Variable | Default | Purpose |
|---|---|---|
| `OTEL_METRICS_ENABLED` | `true` | Master switch. Set to `false` to skip MeterProvider construction and not start the listener. |
| `OTEL_METRICS_ADDR` | `:9090` | Bind address for the `/metrics` listener. |

## Exposed metrics

| Name | Type | Labels |
|---|---|---|
| `mcp_tool_calls_total` | counter | `tool`, `toolkit_kind`, `persona`, `status_category` |
| `mcp_tool_call_duration_seconds` | histogram | `tool`, `toolkit_kind`, `persona`, `status_category` |
| `mcp_inflight_tool_calls` | gauge | (none) |
| `mcp_enrichment_bytes_total` | counter | `tool`, `toolkit_kind`, `persona` |
| `apigateway_outbound_total` | counter | `connection`, `http_status_class`, `status_category`, `persona` |
| `apigateway_outbound_duration_seconds` | histogram | `connection`, `http_status_class`, `status_category` |
| `apigateway_inbound_requests_total` | counter | `connection`, `operation_id`, `method`, `status_class`, `identity` |
| `apigateway_inbound_duration_seconds` | histogram | `connection`, `operation_id`, `method`, `status_class` |

Plus the free Go runtime + process metrics (`go_*`, `process_*`).

The `apigateway_inbound_*` pair measures inbound requests to the REST shim (`POST /api/v1/gateway/{connection}/invoke`, the NiFi-class ETL path), as opposed to `apigateway_outbound_*` which measures the platform's calls to the upstream API. `operation_id` is the OpenAPI operationId resolved from the connection's catalog via path-template matching (`unknown` for no-catalog or no-match). `identity` is the API key name or OIDC subject (`unknown` when unauthenticated), recorded on the counter only (never the histogram) to bound bucket cardinality. `apigateway_outbound_total` carries `persona`, the persona the call was authorized under, which separates an automated principal's traffic from an analyst's on a shared connection; it is bounded by the deployment's persona definitions, records `unknown` for a call the platform could not attribute (the same value `mcp_tool_calls_total` uses for that case), and is on the counter and not the duration histogram -- a scope decision, not the cardinality one that keeps `identity` off the inbound histogram, since persona is bounded and `mcp_tool_call_duration_seconds` carries it. Upgrading: a call that never reached persona resolution used to record an empty `persona`, which Prometheus drops; it now records `persona="unknown"` on `mcp_tool_calls_total`, `mcp_tool_call_duration_seconds` and `mcp_enrichment_bytes_total` too, so a rule matching the label's absence must match `persona="unknown"` instead.

## Chokepoints

Phase 1 instruments two points that cover every tool call through the platform:

1. **`MCPToolCallMiddleware`** — every tool call (MCP transport, SSE, Streamable HTTP, REST gateway shim, admin tools/call) passes through here. One histogram + one counter + one gauge here covers every toolkit at once.
2. **apigateway transport** — every outbound HTTP call from `api_invoke_endpoint`, `api_export`, and the REST gateway shim is recorded at the `http.RoundTripper` level, so adding a new apigateway code path picks up metrics automatically. The calling persona reaches it on the request context, stamped by `MCPToolCallMiddleware` once the authorizer resolves it.

Phase 2 adds OTLP tracing on the same chokepoints. Phase 3 extends to toolkit adapters (Trino, DataHub, S3, OAuth refresh, audit, enrichment).

## Cardinality budget

`status_category` is a closed set: `ok`, `auth_err`, `authz_err`, `validation_err`, `upstream_err`, `internal_err`. `http_status_class` is `2xx` / `3xx` / `4xx` / `5xx` / `other`. High-cardinality fields (user id, email, request id, session id, raw upstream URLs, raw error messages, free-text tool arguments) are NOT recorded as Prometheus labels — they belong on trace spans (Phase 2) and on audit log rows.

Worst-case counter cardinality for `mcp_tool_calls_total` with ~40 tools × 8 toolkit_kinds × 5 personas × 6 status_categories = 9,600 series; for `apigateway_outbound_total` with ~10 connections × 5 status_classes × 6 status_categories × ~5 personas = 1,500 series (the persona dimension is on the counter only, so the duration histogram keeps its 300-series bound). Both well within typical Prometheus and managed-observability backend per-metric series budgets.

## Toolkit, provider, and DB-pool metrics

Beyond the MCP and apigateway chokepoints, the toolkit and provider call paths are instrumented so the cross-enrichment fan-out and dependency health are visible:

- **Trino** (query provider): `trino_queries_total{status, query_kind}` and `trino_query_duration_seconds{query_kind}`. `query_kind` is the SQL verb (`select`, `show`, `insert`, ...) for queries, or the metadata operation (`list_catalogs`, `list_schemas`, `list_tables`, `describe_table`) for catalog calls; unknown SQL maps to `other`. (No `trino_bytes_scanned_total`: the mcp-trino v1.3.0 client exposes no bytes-scanned figure.)
- **DataHub** (semantic provider): `datahub_requests_total{operation, status}` and `datahub_request_duration_seconds{operation}`, instrumented via a client decorator so metrics record on the underlying request even when a cache wraps the provider. `operation` ∈ `get_entity`, `get_schema`, `get_schemas`, `get_lineage`, `get_column_lineage`, `get_glossary_term`, `get_queries`, `list_tags`, `list_domains`.
- **S3** (S3 toolkit): `s3_operations_total{operation, status}` and `s3_operation_duration_seconds{operation}`, recorded by an mcp-s3 `ToolMiddleware` installed before tool registration. `operation` is the S3 tool name (`list_buckets`, `get_object`, ...).
- **OAuth** (authorization server): `oauth_token_issuance_total{grant_type, status}`, `oauth_token_refresh_total{status}`, and `oauth_token_refresh_duration_seconds`.
- **Audit** (async writer): `audit_events_dropped_total` (no labels) counts audit events lost by the bounded async writer: queue-full drops plus writes that failed or were abandoned at the per-write timeout. Audit writes run through a single background goroutine with a per-write timeout, so a stalled database sheds audit load instead of blocking tool calls or leaking goroutines; a growing counter means rows are being lost by design (best-effort delivery).
- **Background indexing** (`pkg/indexjobs`, #1837), recorded as each job settles so the figures cover a backlog of any size: `indexjob_jobs_total{kind, trigger, outcome}` (outcome `succeeded`, `retried`, `failed`, `source_gone`, `lease_lost`, `store_error`), `indexjob_duration_seconds{kind, outcome}` (claim to settling write, buckets to 3600s), `indexjob_running{kind}` (per replica; sum), `indexjob_enqueued_total{kind, trigger, result}` (`created` or `folded` into an open job), `indexjob_items_total{kind, result}` (`embedded` or `reused`), `indexjob_embed_calls_total{kind, status}` (`ok`, `timeout`, `error`), `indexjob_embed_texts_total{kind}`, `indexjob_embed_call_duration_seconds{kind}`, `indexjob_leases_released_total`, `indexjob_units_deferred_total{kind}` (parked units a reconciler sweep skipped). Read from the database at scrape time (one statement over the open rows plus each kind's coverage, cached 30s; every replica reports the same value, so read with `max by (kind)`): `indexjob_queue_jobs{kind, state}` (`pending`, `running`, `retrying`), `indexjob_failed_units{kind}`, `indexjob_oldest_runnable_wait_seconds{kind}`, `indexjob_vectors_indexed{kind}`, `indexjob_vectors_expected{kind}` (only for a kind with a known denominator). The admin Indexing dashboard draws Throughput and Embed latency from these.
- **DB pool** (`database/sql`): `db_pool_open_connections`, `db_pool_in_use`, `db_pool_idle`, `db_pool_wait_count_total`, and `db_pool_wait_duration_seconds_total`, all labeled `pool`. These are OTel observable instruments read from `(*sql.DB).Stats()` at scrape time via a single callback registered once at startup; `RegisterDBPool(*sql.DB, name)` adds a handle to the observed set (the platform registers its shared pool as `pool="platform"`).

`status` on these metrics reuses the platform's bounded status taxonomy (`ok` for success, `upstream_err` for a failed external call). Per-operation/query_kind label sets are bounded by the closed tool/operation lists.

## Deployment manifests

Plain Kubernetes manifests (no Helm, no Operator CRDs) ship in `deployments/observability/`: `pod-annotations.yaml` (enable the listener + `prometheus.io/*` scrape annotations), `recording-rules.yaml` and `alert-rules.yaml` (ConfigMaps with starter rules in the `level:metric:operations` convention — `mcp:tool_call_duration:p95_5m`, `apigateway:inbound_error_rate:5m`, a DB-pool saturation alert, an `up{job="mcp-data-platform"} == 0` alert, etc.), and a README covering how to mount the rules and confirm scraping. The full metric reference is `docs/server/observability.md`.

## What metrics do NOT replace

- **Audit logs** remain the source of truth for "who called what entity, with what result" — compliance and user-level analytics.
- **Application logs** (stderr / structured slog) remain the source of truth for free-text diagnostic detail and stack traces.

Metrics answer "how is the system performing" — they complement, they do not replace.

## PromQL query proxy

The portal reads the scraped metrics back through an authenticated proxy rather than talking to Prometheus directly (which would require CORS, a separate auth path, and exposing an internal service). The platform serves `GET /api/v1/observability/query` and `GET /api/v1/observability/query_range`, forwarding to the configured Prometheus `/api/v1/query` and `/api/v1/query_range` and returning the upstream response body unchanged.

Configured under `observability.prometheus` in platform.yaml (`url`, `timeout`, `basic_auth.{username,password}`, `rate_limit_per_second`). An empty `url` leaves the proxy unconfigured: endpoints return 503 with `observability backend not configured` so the portal renders a clean empty state.

Access is gated by the `observability:read` persona capability, checked through the same persona tool-allow filter that gates tools (operators grant it by adding `observability:read` to a persona's tools.allow; default-deny). A per-persona rate limit (default 10 queries/second) returns 429 when exceeded. Proxy queries are not audited: the dashboards poll these endpoints on a refresh interval, so auditing each one flooded the audit trail and tool-usage analytics with dashboard-internal reads that are not MCP tool calls. Responses are not cached on the platform; Prometheus is the cache and the portal applies React Query stale-time.

## Distributed Tracing (Phase 2)

Where metrics answer "how is the system performing" in aggregate, traces answer "why was THIS call slow" by capturing one MCP request as a single span tree. The platform exports OpenTelemetry spans over OTLP/gRPC to a collector (Tempo, Jaeger, or any OTLP backend). Tracing is OFF by default and independent of metrics: unlike the always-available `/metrics` scrape endpoint, traces need a collector to receive them, so enabling without one is pointless. When off, every span call site is a single span-context check; the tool-call tracing middleware benchmarks at ~0.3 ns/op disabled and ~1.8 µs/op when it samples a span, negligible against millisecond-scale tool calls.

## Load testing and measured limits

A Go load harness lives in `test/load` (a separate module, kept out of the root coverage/test/lint gates). It drives named workloads against a running platform over the MCP streamable-HTTP protocol and the REST surfaces, scrapes the platform's own Prometheus metrics before/during/after each run, optionally captures pprof profiles, and writes a self-contained JSON report. It is written in Go against the official MCP SDK because the hot path is the MCP protocol (initialize handshake, session, `tools/call`), which generic HTTP load tools cannot exercise. Like mutation testing, it is deliberately NOT part of `make verify`; run it via the `load-up` / `load-run` / `load-down` Makefile targets.

Scenarios: `mcp-tool-call` (the primary hot path: `search` then `trino_query`), `mcp-session-churn` (session-store create/destroy), `oauth-token` (the bcrypt-bound, rate-limited OAuth DCR path), `portal-read` (portal REST reads plus the public viewer), `audit-burst` (sized past the async audit queue to observe `audit_events_dropped_total`), and `soak` (a long fixed-rate run asserting flat memory and goroutines).

An opt-in pprof listener supports the harness: setting the `PPROF_ADDR` environment variable starts a dedicated `net/http/pprof` server on a private mux (off by default, never mounted on a client-facing listener, since it exposes process internals). Published single-replica throughput and latency ceilings, the saturation narrative (the async audit writer sheds first; CPU saturates on OAuth bcrypt, not the data path), and copy-pasteable reproduction commands are in `docs/reference/tuning-and-scaling.md` under "Measured limits".

## Agent-effectiveness benchmark

A second measurement harness lives in `bench/` (also a separate Go module): the agent-effectiveness benchmark (issue #930). Where the load harness answers "how much" (throughput, latency, memory), the benchmark answers "how well": it holds the model, prompt scaffold, seed data, and task set constant and ablates only the platform configuration (arms as config profiles: `a0` raw toolkit tools vs `a2` the full semantic-first platform), measuring answer accuracy, pass^k reliability, and tool-call efficiency. Ground truth is generated from a fixed-seed dataset (Trino tables, DataHub metadata, knowledge pages) with the disambiguating facts for its knowledge-trap tasks (units-in-cents, gross-vs-net revenue policy) living only in the metadata and knowledge layers. Efficiency metrics are read back from the platform's own admin audit API (`audit.delivery: sync`), and a run fails loudly when a session's audit rows are incomplete. A deterministic scripted adapter validates the whole pipeline against a live stack with no model API key (`make bench-up`, `make bench-smoke`, `make bench-down`); real runs use a pinned Anthropic model. Like load testing, it is deliberately NOT part of `make verify`. `bench/README.md` is the index to the report series (each study's published page, DOI, protocol, toolchain, and run data); the method for this study is `bench/docs/knowledge-layer-protocol.md` and the sibling knowledge-use study's is `bench/docs/knowledge-use-protocol.md`; every archived run family under `bench/results/` carries a README stating what it does and does not establish, indexed by `bench/results/README.md`. The published reports cite these protocol documents by section name, and `TestHarnessCitationsResolve` fails the build when a cited section stops existing. The S5 suite adds multi-episode memory-insight-knowledge lifecycle protocols (`make bench-lifecycle`, on the `a3` arm): each protocol teaches a novel definition in one session and, across fresh sessions, measures personal recall, cross-identity transfer after the reviewer promotes the insight via `apply_knowledge` (to an entity description or a knowledge page), recall-first supersede when the fact is later corrected, and abstention on facts never taught -- every lifecycle state transition verified through the admin insights and changesets APIs, not inferred from transcripts. A cold-start knowledge-growth suite (`make bench-cold-start`, issue #963) closes the loop between the two halves: it boots the `a3` arm against an empty enrichment layer (undocumented DataHub via `bench_mces_empty.json`, no knowledge pages via `BENCH_SEED_PAGES=0`), teaches a six-lesson curriculum (one fact per S3 trap class, promoted to the same DataHub descriptions and knowledge pages the A2 seed pre-loads), and re-runs the fixed S3 trap suite with a fresh, never-taught evaluator identity after each promotion -- producing a learning curve of accuracy, per-trap-class resistance, and enrichment coverage as a function of accumulated promoted knowledge (the ceiling is the fully-documented A2 level, so the capture-and-propagation half visibly populates what the surfacing half delivers). The canonical, citable results live in `docs/reference/benchmark-report.md`, recomputed from committed run JSON under `bench/results/`; `docs/reference/benchmarks.md` is the product-framed introduction that names the measured surface (the knowledge layer, not the whole platform) and points to the report. The benchmark measures one knowledge system from two coupled angles: a surfacing half (cross-enrichment and search deliver curated knowledge from its sinks) and a capture-and-propagation half (memory -> insight -> apply_knowledge promotes facts into those sinks -- DataHub entity/column descriptions or knowledge pages); the lifecycle populates what enrichment and search deliver, so they are two ends of one pipe, not independent features. Surfacing half, validated by S1-S3 (four arms a0/a1/a2/a3, claude-sonnet-5, k=3, 261 graded attempts per arm, zero harness failures): S3 knowledge-trap accuracy climbs 42.7% (raw tools) -> 57.3% (enrichment) -> 98.7% (full platform), a +56-point gain for the platform over bare tools (95% CI +44 to +67), while S1 discovery and S2 analytical accuracy are near ceiling for every arm; median knowledge-trap tool calls drop 16 -> 10. Arm a3 (platform + memory lifecycle) ties a2 (platform) on S1-S3 because these single-session tasks exercise only the surfacing half, NOT because capture lacks value; capture-and-propagation is what produces the sink contents S1-S3 shows are effective once present, and it is measured separately by S5. Capture-and-propagation half (S5 lifecycle protocols, a3, k=3, 15 multi-session protocols): a capture/personal-recall/unprompted-surface/cross-identity-transfer/update-correctness/duplicate/abstention scorecard plus full-lifecycle pass^3, measuring whether the teach-once-answer-forever mechanism works reliably across fresh sessions rather than an accuracy delta over a (trivially-zero) no-memory baseline, with every state transition verified through the admin insights and changesets APIs. The S5 lifecycle is reported as three runs: a shared-store run (45 attempts, one accumulating knowledge store, k-repeats coupled) and two isolated replicates (each three independent k=1 passes with a full clean platform reset between passes, merged to k=3). The two isolated replicates, identical in configuration, disagree materially (duplicate rate 0% vs 42.9%, pass^3 26.7% vs 20.0%, capture/recall a few points each), which shows most metric-by-metric differences at this scale (15 protocols, small supersede denominators) are run-to-run sampling noise rather than the shared-store confound; the apparent isolation "lift" seen in the first replicate is not reproduced by the second. What is stable across all three runs and is read as the firm result: unprompted surface 100%, update correctness 100%, abstention 96-100%, and cross-identity transfer stuck at ~42-47% (the clearest reproducible reliability ceiling). Because S1-S3 already validates surfacing, the transfer gap is a capture/propagation limit upstream (the apply_knowledge promotion landing in an aspect enrichment reads -- mcp-datahub GetEntity fully populates only datasets/dashboards -- or the second identity surfacing the fact but not using it), not a delivery failure. All three runs recorded zero harness failures and zero episode errors, so transient Claude API outages (which surface as excluded harness failures, never mis-graded answers) did not affect the data. Spend is published as token counts (not dollars), computed from committed per-attempt records: ~158.6M tokens total across all runs (25,182 fresh input, 2,664,630 output, 147,183,943 cache-read, 8,676,898 cache-write), of which cache reads are ~93% -- apply current claude-sonnet-5 per-token pricing to these counts for a cost. S1-S3 and the shared-store S5 run used development build v1.102.0-9-gadfb9d90-dirty; the isolated runs used v1.102.0-10-g32d61254-dirty (platform code byte-identical between the two build commits, verified by diff; only benchmark-harness commits differ); each is pinned in its manifest. A `benchrun -baseline <results.json>` regression gate exits nonzero when a fresh run falls below a committed baseline beyond per-suite thresholds; an earlier phase-1 pilot and phase-3 S5 pilot remain under `bench/results/`.

Two derived, peer-reviewable deliverables read from the committed run data. `docs/reference/benchmark-report.md` is a neutral evaluation report: it recomputes every statistic from the raw `results.json` files and presents two complementary sub-studies that are never mixed (they used different client paths, the Anthropic API for the ablation and `claude-cli` for the cold-start, which the harness itself refuses to compare on absolute accuracy). Ablation headline: the knowledge layer lifts S3 knowledge-trap accuracy 42.7% -> 98.7% (+56 points, 95% bootstrap CI +44 to +67) and is statistically neutral on S1 discovery and S2 numeric tasks where no business context is needed. Cold-start headline: from a reproducible empty-layer floor (five baseline replicates: 44.0/44.0/47.8/48.0/52.0 percent, mean 47.2%), teaching and promoting six facts lifts trap accuracy to 90.7% at K=3 and 100% at K=1, and the per-trap-class trajectories show each fact's gain arriving at its own promotion checkpoint (the two knowledge-page traps that start at the 0% floor unlock cleanly at their checkpoints). An S5 lifecycle scorecard, replaced in report v2.0 by a k=5 re-run over 30 protocols (149 runs) on platform v1.118.0, reports capture 91.9%, personal recall 95.3%, unprompted surface 100%, cross-identity transfer 98.9% CI [96.8-100.0], abstention 92.6%, and supersede metrics as point estimates with intervals (update correctness 41/41, duplicate rate 22.0% CI [9.8-34.1]) rather than the v1.1 ranges on 7 runs; Section 5 is measured on a later build than Sections 3 and 4, so its differences from v1.1 are across-code (#1129, #1141), not sample-size effects. The report carries a mandatory threats-to-validity section (single model, non-comparable client paths, build provenance, small seed dataset, bootstrap scope, one capture miss, one excluded transient API 500, the distinctive-needle caveat) and a data-availability table. It is written to be cited: an author line (Craig Johnston), a publication date, a Report v2.0 version tied to the platform build, commit, and release tag, and a How-to-cite section with a ready-to-paste citation and a BibTeX entry pointing at a tagged/permalinked artifact; the repo ships `CITATION.cff` and `.zenodo.json` recording the Zenodo concept DOI 10.5281/zenodo.21438044 (published, resolves to the latest version; v2.0.1 is version DOI 10.5281/zenodo.21751635 and the v1.0 snapshot is 10.5281/zenodo.21438045). It also carries a related-work section and a references list (BIRD for the execution-match grading, tau-bench for pass^k, and LOCOMO/LongMemEval/Mem0/Zep for the long-term-memory framing) crediting the methods it borrows without drawing any cross-study comparison. `bench/reports/knowledge-layer/report.ipynb` is the reproducible supplement: a Python/pandas/matplotlib notebook that loads only the committed JSON (no API key, no running platform, no network), recomputes every number, and regenerates the four figures under `bench/reports/knowledge-layer/figures/`; dependencies in `bench/reports/knowledge-layer/requirements.txt`. The follow-up roadmap the report motivates is tracked as a GitHub issue rather than committed documentation (the report stays neutral and makes no recommendations; all direction lives in the issue, which references the report): Track A benchmark improvements (vary teaching order, adversarial/conflicting curricula, multi-model climb, capture-rate as a first-class metric, forgetting/decay after supersede, the channel-ablation experiment that teaches the same fact three ways across three discovery modes to resolve the sink-delivery confound the report was careful not to over-claim, and larger-protocol lifecycle scale) and Track B product improvements each citing the report finding that motivates it (give DataHub-resident facts topic-discoverability at search time, make promoted insights searchable cross-identity to close the ~47% transfer gap, and harden teach-time capture as the loop's rate-limiter). An Errata subsection records post-publication corrections; the 2026-08-02 entry documents that every run executed with fetch (and list_connections) absent from the arm personas' allow-lists (19 fetch attempts across all 4,173 archived transcripts, 19 denials, #1176): the arms measured search-only, single-hop delivery rather than the shipped search-then-fetch surface, uniformly across arms, so the contrasts stand and no statistic changed; the bench configs now grant both tools, guarded by a config test.

A second report, `docs/reference/benchmark-report-knowledge-use.md` (version 1.0, 2026-07-26; headline tables replicated on a v1.116.0 tag build), asks when agents USE stored knowledge at all. It grew from a pre-registered study (bench/docs/perishable-knowledge-study-design.md, issue #1054) whose primary hypothesis - that agents under-verify perishable stored beliefs and epistemic metadata would fix it - was falsified at first empirical contact, so everything in it is exploratory with per-rate Wilson intervals and the pre-registered confirmatory matrix never ran. Apparatus: a deterministic social-analytics-shaped fixture behind the API gateway (bench/apisvc -surface perishable) with a between-session world-change control plane, a committed world registry whose workspace scoping makes the cost of re-establishing state a controlled variable (1 to 11 calls), frozen minimal-pair beliefs planted through the real memory_capture tool as the identity later queried, cells whose correct behavior (answer/refuse/verify-then-answer/verify-then-refuse/probe-then-refuse) is derived from computed answerability and belief truth rather than hand-labels, and deterministic grading with a pre-analysis definition of verification (direct state observations only; state-presupposing calls are a separate sensitivity measure). Findings: (1) Sonnet 5 re-derives every checkable delivered belief - verification at ceiling across an elevenfold cost sweep, both belief directions, and matched no-knowledge controls showing a zero-call median effort delta - so delivered world-state observations are redundant for the strong tier; (2) the derivability bridge (an authored reporting convention no endpoint states: positive coverage = sentiment_score >= 70) shows the same model relying on non-derivable testimony 8/8 and, without it, fabricating a plausible threshold in most control attempts - so the knowledge layer's value concentrates in conventions/definitions/policies, consistent with report 1's +56 trap lift; (3) Haiku 4.5 inverts the derivable result (trusts delivered state 29/32), making weak-tier deployments staleness-exposed where strong-tier ones are staleness-immune; (4) neither headline is a client artifact - the tier flip happens within one claude-cli version and the sonnet results replicate exactly (48/48 verified, 0 trusted) on a raw Messages API in-process loop with no agent client; (5) a companion lifecycle probe found supersede at ceiling conditional on capture but strict capture at 33%, decomposed into deterministic mis-filing against the semantically nearest entity and silent non-capture - the measured loss stages of the knowledge lifecycle and a candidate mechanism for report 1's 46.7% transfer. Retired platform features (with evidence): volatility/valid-until schema fields, freshness/recheck-cost enrichment, capture steering toward dated observations (capture already writes dated self-refreshing notes - the archived capture corpus shows it). Every table recomputes offline via bench/reports/knowledge-use/pk_tables.py from bench/results/knowledge-use/; run manifests pin commit, model, driver, client version, seed-set hash, and k. A 2026-08-02 threats-to-validity entry records that every run executed with fetch denied by the pk arm persona, so every rate is a rate under search-only, single-hop delivery; the denial was uniform across cells, the contrasts stand, and the configs now grant fetch (#1176).

A third report, `docs/reference/benchmark-report-knowledge-pollution.md` (version 1.0, 2026-08-07; DOI 10.5281/zenodo.21834813), prices the failure mode the first two left open: what a wrong claim costs once the platform's own curation gate has admitted it. Pre-registered (bench/docs/knowledge-pollution-study-design.md, issue #1166) with hypotheses, decision rules, and falsifiers fixed before the confirmatory data; grading is fully deterministic (exact discriminant values computed from the fixture, no judge); the framing is organic error, not adversarial. Design: a wrong claim planted through the real capture-approve-apply path into the shared applied tier beside a co-present correct source, crossed over derivability class (convention: a fiscal boundary nothing in the fixture refutes; checkable: an order count one query settles) x arm (absent/correct/wrong) x tier (Haiku 4.5 / Sonnet 5 / Opus 5), 24 episodes per cell, 432 confirmatory episodes plus pre-registered follow-ups. Headline, inverting the study's own primary hypothesis: the only claim adopted anywhere was the CHECKABLE one - 16/24 on the weak tier, 0/24 on both strong tiers, while the convention was adopted nowhere on the agent client. Mechanism, exact in 120 episodes with no exception: every episode that observed the refuting count answered correctly, every episode that did not adopted, and with nothing planted the weak tier runs the query 24/24 - the claim does not out-argue the world, it removes the impulse to consult it (verification displacement). The effect is unchanged at three directive strengths (a bare statement adopts like an imperative, so it is adoption, not compliance), at least as strong with the claim on a knowledge page as on the catalog entity (search delivery alone carries it; the applied sink is not the load-bearing channel), replicates on a second fixture (24/24 vs 0/24 controls, capability split intact), and replicates on the raw Messages API with no agent client. Two narrowings: the convention's immunity is client-scaffold-bound (raw-API weak tier adopted it 4/8), and the plant's reviewer note disclosed the plant to any episode that fetched the insight - provenance inspection proved capability-graded (convention cells: haiku 9/24, sonnet 18/24, opus 24/24), the strongest tier's convention null is confounded by that disclosure, and the weak tier adopted straight through it. Design consequences: derivability-aware promotion (the claims worth guarding are the ones the platform could verify itself), capability-aware delivery for weak tiers, and provenance-on-fetch as a load-bearing surface. Retraction (RQ3) never ran and the report says so. Every table recomputes offline via bench/reports/knowledge-pollution/pollution_tables.py and every figure via figures.py beside it (both runnable from bench/reports/knowledge-pollution/report.ipynb) from bench/results/knowledge-pollution/, and make bench-report-check (part of make verify and CI's harness job) pins the published headline numbers to the archives; builds are commit-pinned to the v1.118.0-v1.119.0 lineage with every manifest carrying its exact commit.

A fourth report, `docs/reference/benchmark-report-graph-completion.md` (version 1.0, 2026-08-10; DOI 10.5281/zenodo.21881798), asks what the structure BETWEEN knowledge pages is for: when a completion task's constraints are spread across a page graph, do authored cross-references deliver what retrieval structurally cannot - constraints search cannot rank (semantic discontinuity) and a decidable notion of done (a graph closure terminates; a ranked list never certifies coverage)? Pre-registered as a stage-3 separation design (bench/docs/graph-completion-study-design.md, issue #1250; confirmatory matrix under #1251 running only what the design froze), after a premise probe (#1241) validated the instruments. Apparatus: a deterministic generated corpus - a fixed 27-page hand-authored core, byte-identical at every scale, inside generated operations-wiki filler at 50/500/5000 total pages (Seed 1250, EdgeDensity 3), graded by minted signatures that structurally cannot occur off their source pages - planted through the platform's knowledge-page surface in two arms holding page meaning constant: graph (authored references as real edges, verified exact) vs stripped (each reference rendered as its meaning-preserving prose fallback, reference table verified empty). Two discontinuity constraints per cell were certified unreachable twice per scale before any episode: offline embedding rank outside the exclusion horizon for every task phrasing, and a live sweep gate requiring absence from every hit list. 99 episodes, one agent configuration, commit-pinned build, every manifest carrying the generator spec and corpus fingerprint. Headline - the pre-registered instrument kill fired, and it is the finding: stripped-arm episodes grounded the certified-unreachable constraints at 1.00 (500 pages) and 0.93 (5000 pages) by reading the prose fallback on an ordinary page and re-searching in the named institution's own vocabulary, which ranks the "unreachable" page instantly. The certifications proved unreachability for task-derived queries; the defeating queries are read-derived - the agent spontaneously runs the classical pseudo-relevance-feedback loop, traversing in query space. Consequence for benchmark design, recorded as the series' instrument rule: "unreachable by search" must hold against read-derived queries, and a meaning-constant stripped arm can never provide that, because meaning-preserving prose names what it points at. What the archives affirmatively show: discovery is not enumeration (grounded coverage at ceiling in every cell of the matrix; roughly eleven fetches - 0.2 percent of the corpus - ground every constraint against 5000 pages; scale moved cost, not coverage); authored edges buy cost and robustness, not coverage (searches per grounded constraint 0.59/0.67/0.63 graph vs 0.73/1.02/1.44 stripped across two orders of magnitude - flat vs roughly doubling - with the matrix's only failed episode and only sub-ceiling cell in stripped/5000, and the share of graph-arm fetches dereferencing a page-learned edge rising 0.09/0.28/0.34 with scale); with search removed entirely the graph arm walked every closure at full depth (1.00 grounded, zero searches, 9 fetches per episode), scale-invariant with the pilot's no-search floors (graph 0.96/0.42 across two reading budgets, stripped 0.00 - no edges and no search means no route at all); and the elicited completeness claim was uniformly conservative (0 of 98 surviving episodes claimed complete, 91 declared open items), so the overclaim channel is unmeasured, not null, and the report says so. The completeness-delivery framing for authored edges is retired with the discontinuity construct; the evidence supports writing cross-references for economy and resilience, not reach. Every table recomputes offline via bench/reports/graph-completion/graph_tables.py and every figure via figures.py beside it (both runnable from bench/reports/graph-completion/report.ipynb) from bench/results/graph-completion-confirmatory/ and bench/results/graph-completion-probe/, and make bench-report-check (part of make verify and CI's harness job) pins the published headline numbers to the archives, including the instrument kill's presence and the archived analyzer's deliberate non-zero exit.

Environment-only configuration: `OTEL_TRACES_ENABLED` (default false), `OTEL_EXPORTER_OTLP_ENDPOINT` (OTLP/gRPC `host:port`, default `localhost:4317`), `OTEL_EXPORTER_OTLP_INSECURE` (default true — the common in-cluster topology; set false for a TLS remote collector), `OTEL_TRACES_SAMPLER_ARG` (head-based sampling ratio in [0,1] applied to root spans, default 0.1), and `OTEL_SERVICE_NAME` (service.name resource attribute, default `mcp-data-platform`). The OTLP exporter connects lazily: an unreachable or unconfigured collector never blocks or fails startup; spans are batched and dropped if undeliverable.

Each tool call produces one trace. The ROOT span has the fixed, low-cardinality name `tool_call` (the specific tool is on the `mcp.tool` attribute, not the span name, so all tool calls share one queryable name) and is opened by the tracing middleware inner to auth, so it carries the request's identity. It holds the bounded attributes that mirror the metric labels (`mcp.tool`, `mcp.toolkit_kind`, `mcp.persona`, `status_category`) PLUS the high-cardinality fields deliberately kept off Prometheus labels — `mcp.user_id`, `mcp.user_email`, `mcp.session_id`, `mcp.request_id`, `mcp.connection`, `mcp.transport`, `mcp.source`, and the enrichment summary. CHILD spans nest under the root via context propagation: the cross-service `enrichment` fan-out (the original motivation for tracing — the deep, easily-misordered middleware chain), and one span per upstream call to Trino (`trino.<query_kind>`), DataHub (`datahub.<operation>`), and S3 (`s3.<operation>`). Those toolkit spans are emitted by the same instrumenting decorators that record the toolkit metrics, installed when EITHER metrics or tracing is enabled (the decorator's metric record is nil-safe and its span is a no-op outside an active trace, so whichever subsystem is off costs effectively nothing; when both are off the decorators are not installed at all). Span status is `Error` for any non-`ok` status_category, with the error recorded as a span event so error traces stand out. The inbound OAuth 2.1 server and the asynchronous audit write run outside a tool call's request context and so are not part of the tool-call trace.

Head-based sampling is in-app via a ParentBased ratio sampler (a sampled caller's whole trace is always kept). Tail-based sampling — keeping 100% of error and slow traces and down-sampling only the fast successful majority — belongs in the collector, not the application, so it can be tuned without redeploying. An example collector pipeline (OTLP receiver, `tail_sampling` keeping ERROR and >2s traces, OTLP export to Tempo) ships in `deployments/observability/otel-collector.yaml`. Full reference, the span-tree diagram, and example Tempo/Jaeger queries are in [docs/server/observability.md](https://github.com/txn2/mcp-data-platform/blob/main/docs/server/observability.md).

---

# Cross-Enrichment

When you query a Trino table, you get DataHub context in the response. When you search DataHub, you see which datasets are queryable. No extra calls.

## The Problem It Solves

Without cross-enrichment:
1. Query a table
2. Search DataHub for that table
3. Get entity details for owners and tags
4. Check deprecation status
5. Look up quality score

Five calls to understand one table. With cross-enrichment, step 1 gives you everything.

## What Gets Injected

| When you use | You also get |
|--------------|--------------|
| Trino | DataHub metadata (owners, tags, quality, deprecation) |
| DataHub search | Which datasets are queryable in Trino |
| S3 | DataHub metadata for matching datasets |

## Semantic Context (added to Trino/S3 results)

```json
{
  "semantic_context": {
    "description": "Customer orders with line items and payment info",
    "owners": [{"name": "Data Team", "type": "group"}],
    "tags": ["pii", "financial"],
    "domain": {"name": "Sales"},
    "quality_score": 0.92,
    "deprecation": {
      "deprecated": true,
      "note": "Use orders_v2 instead"
    }
  }
}
```

## Query Context (added to DataHub results)

When `search_schema_preview` is enabled (default), available tables also include a bounded column preview so agents can write SQL without an intermediate `fetch` or `trino_describe_table` call:

```json
{
  "query_context": {
    "urn:li:dataset:orders": {
      "available": true,
      "query_table": "hive.sales.orders",
      "connection": "production",
      "estimated_rows": 1500000,
      "schema_preview": [
        {"name": "order_id", "type": "integer"},
        {"name": "customer_id", "type": "integer"},
        {"name": "order_date", "type": "date"},
        {"name": "total_amount", "type": "decimal(10,2)"},
        {"name": "status", "type": "varchar"}
      ],
      "total_columns": 42
    }
  }
}
```

Primary key columns are listed first. `total_columns` indicates when the preview is truncated. The preview is omitted (not empty) when schema lookup fails or the table is unavailable.

## Lineage-Aware Column Inheritance

Downstream datasets (Elasticsearch indexes, Kafka topics) often lack documentation even when their upstream sources (Cassandra, PostgreSQL) are well-documented. The platform automatically inherits column metadata from upstream tables via DataHub lineage.

### How It Works

1. Query a table with undocumented columns
2. Platform checks DataHub lineage for upstream sources
3. Matches columns using column-level lineage, name matching, or configured aliases
4. Inherits descriptions, glossary terms, and tags from upstream
5. Returns enriched response with provenance tracking

### What Gets Inherited

| Metadata | Inherited When |
|----------|----------------|
| Descriptions | Target column has no description |
| Glossary Terms | Target column has no glossary terms |
| Tags | Target column has no tags |

### Match Methods

| Method | Description |
|--------|-------------|
| `column_lineage` | DataHub has explicit column-level lineage edges |
| `name_exact` | Column names match exactly |
| `name_transformed` | Names match after applying transforms (strip prefix/suffix) |
| `alias` | Explicit alias configuration bypasses lineage lookup |

### Configuration

```yaml
semantic:
  provider: datahub
  instance: primary

  lineage:
    enabled: true              # Enable lineage inheritance
    max_hops: 2                # Maximum upstream traversal depth (1-5)
    inherit:                   # Metadata types to inherit
      - glossary_terms
      - descriptions
      - tags
    conflict_resolution: nearest   # "nearest" (closest wins), "all" (merge), "skip"
    prefer_column_lineage: true    # Use column-level lineage when available

    # Column transforms for nested JSON paths. target_pattern is a glob matched
    # against the target dataset name (same syntax as aliases[].targets) and
    # confines the transform to matching datasets; omitting it applies the
    # transform to every dataset. A pattern that matches nothing, including a
    # malformed one, disables that transform rather than widening it.
    column_transforms:
      - target_pattern: "elasticsearch.default.jakes-sale-*"
        strip_prefix: "rxtxmsg.payload."
      - strip_prefix: "rxtxmsg.header."
      - strip_suffix: "_v2"

    # Explicit aliases when lineage isn't in DataHub
    aliases:
      - source: "cassandra.prod_fuse.system_sale"
        targets:
          - "elasticsearch.default.jakes-sale-*"
          - "elasticsearch.default.pos-sale-*"
        column_mapping:
          "rxtxmsg.payload.initial_net": "initial_net"

    cache_ttl: 10m             # Cache lineage graphs
    timeout: 5s                # Timeout for inheritance operation
```

### Response Format

When lineage inheritance is active, responses include column context with provenance:

```json
{
  "columns": [
    {"name": "rxtxmsg.payload.amount", "type": "DOUBLE"}
  ],
  "semantic_context": {
    "description": "Elasticsearch index for sales data",
    "urn": "urn:li:dataset:elasticsearch.default.jakes-sale-2025"
  },
  "column_context": {
    "rxtxmsg.payload.amount": {
      "description": "Net sale amount before adjustments",
      "glossary_terms": [
        {"urn": "urn:li:glossaryTerm:NetSaleAmount", "name": "Net Sale Amount"}
      ],
      "tags": ["financial"],
      "is_pii": false,
      "inherited_from": {
        "source_dataset": "urn:li:dataset:cassandra.prod_fuse.system_sale",
        "source_column": "initial_net",
        "hops": 1,
        "match_method": "name_transformed"
      }
    }
  },
  "inheritance_sources": [
    "urn:li:dataset:cassandra.prod_fuse.system_sale"
  ]
}
```

### Use Cases

**Elasticsearch indexes**: JSON documents from Cassandra have nested paths like `rxtxmsg.payload.amount`. Configure `strip_prefix: "rxtxmsg.payload."` to match upstream column `amount`.

**Kafka topics**: Event streams derived from source tables. Column-level lineage (if available) provides precise mapping.

**Data lakes**: Parquet files derived from operational databases. Aliases provide explicit mapping when lineage isn't tracked.

### Knowledge Pages (entity to knowledge)

When a tool result names entities, the response also carries the canonical knowledge pages that document those entities, so an agent sees the business and domain knowledge about what it just fetched. Entity URNs are read from the result, the reverse lookup finds the pages that reference them, and a bounded `knowledge_pages` block is appended (capped so a widely documented entity does not balloon the response):

```json
{
  "urn": "urn:li:dataset:(urn:li:dataPlatform:trino,iceberg.retail.daily_sales,PROD)",
  "knowledge_pages": [
    {"id": "kp_5cdde7bb", "slug": "retail-sales-vocabulary", "title": "Retail Sales Vocabulary"}
  ]
}
```

The same reverse lookup surfaces in two more agent-facing places: `search(entity_urns=[...])` returns the knowledge pages that reference each entity (datasets and connections) merged with the text-relevance results, and `list_connections` carries per connection a `knowledge_page_count` and a bounded sample of the pages that document it.

## Session Metadata Deduplication

When Trino enrichment is enabled, every tool call targeting a table receives ~2KB of semantic metadata. In a typical session querying the same table multiple times, this wastes LLM context tokens with repeated information.

Session dedup tracks which tables have been enriched per client session. First call gets full metadata; repeat calls within the TTL get reduced content based on the configured mode.

### Configuration

```yaml
enrichment:
  trino_semantic_enrichment: true
  column_context_filtering: true    # Only include SQL-referenced columns (default)
  session_dedup:
    enabled: true          # Default: true
    mode: reference        # reference (default), summary, none
    entry_ttl: 5m          # Defaults to semantic.cache.ttl
    session_timeout: 30m   # Defaults to server.streamable.session_timeout
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `enabled` | bool | `true` | Whether session dedup is active |
| `mode` | string | `reference` | What to send for repeat queries |
| `entry_ttl` | duration | semantic cache TTL | How long a table stays "already sent" |
| `session_timeout` | duration | streamable session timeout | Idle session cleanup |

### Modes

- **`reference`** (default): Repeat calls return `{"metadata_reference": {"tables": [...], "note": "Full semantic metadata was provided earlier..."}}`. Minimal tokens, LLM can refer back.
- **`summary`**: Repeat calls return table-level `semantic_context` without column details. LLM gets a reminder of ownership and tags.
- **`none`**: Repeat calls return raw tool results with no enrichment. Maximum token savings.

### Behavior

- Enabled by default when `trino_semantic_enrichment: true`
- Trino-only (DataHub/S3 enrichment is not deduplicated)
- In-memory state by default; persisted to database session store if `sessions.store: database`
- Each client session has independent dedup state
- SQL parsing tracks multiple tables independently (`JOIN` queries)

## Session Externalization

Externalize MCP session state to PostgreSQL for zero-downtime restarts and horizontal scaling.

### Configuration

```yaml
database:
  dsn: "${DATABASE_URL}"

sessions:
  store: database              # "memory" (default) or "database"
  ttl: 30m                     # session lifetime
  cleanup_interval: 1m         # cleanup routine frequency
```

When `store: database`:
- Platform forces `server.streamable.stateless: true` on the SDK
- Sessions are managed in PostgreSQL via `SessionAwareHandler`
- On shutdown, enrichment dedup state is flushed to the session store
- On startup, dedup state is restored from persisted sessions
- Multiple replicas share the same session store

Session hijack prevention: each session stores a SHA-256 hash of the creation token. Requests with a different token get HTTP 403.

Session termination is served by `SessionAwareHandler` itself: `DELETE /` with an `Mcp-Session-Id` removes the row and answers `204 No Content`; a `DELETE` without the header answers `400`. The request is not forwarded to the SDK, whose stateless handler serves `POST` only (SEP-2575, go-sdk v1.7.0) and would answer `405 Method Not Allowed` for a termination that had already succeeded.

### Server-pushed notifications (`tools`/`prompts`/`resources` `list_changed`)

In stateless streamable HTTP mode (the production shape when `sessions.store: database`), the SDK refuses GET requests for SSE streaming and closes each session at end-of-request. Without intervention this means downstream agents (Claude.ai, Claude Desktop) never receive `notifications/tools/list_changed` — when a gateway upstream re-authenticates, agents still show the old tool list until they disconnect and reconnect.

The platform handles this with a session broadcaster:

- `pkg/session.Broadcaster` is the abstract fan-out interface; the AwareHandler subscribes per-session.
- In-memory implementation for single-replica or no-DB deployments.
- Postgres `LISTEN/NOTIFY` implementation (channel `mcp_notifications`) for multi-replica. Each replica `LISTEN`s once at startup; every replica re-publishes received events to its local SSE subscribers.
- The `AwareHandler.ServeHTTP` GET branch opens a long-lived `text/event-stream` response, subscribes to the broadcaster bound to the request's session ID, and streams every event as a JSON-RPC 2.0 notification (`data: {"jsonrpc":"2.0","method":"notifications/tools/list_changed","params":{}}\n\n`). A 25-second comment-frame heartbeat keeps the stream alive through proxy idle timeouts.
- The gateway toolkit publishes `tools/list_changed` (debounced 50 ms — a longer window than the MCP SDK's own 10 ms internal debounce, chosen to absorb the postgres LISTEN/NOTIFY round-trip cost across replicas) after every aggregate tool-inventory change that registers or removes at least one tool: a connection coming up after re-auth, a connection being removed, the SetTokenStore retry promoting a placeholder. Zero-tool changes (placeholder upstreams, removals on connections with no live tools) are silently skipped.
- The same broadcaster carries `prompts/list_changed` and `resources/list_changed` (also debounced 50 ms). `notifications/prompts/list_changed` fires on every prompt store write (create, update, delete, approve, scope promotion) across all write paths (`manage_prompt`, admin/portal REST, knowledge `add_prompt`), emitted from a notifying wrapper over the shared prompt store so no write path can change the prompt set without notifying; the advertised `prompts.listChanged` capability is honest. `notifications/resources/list_changed` fires when a managed resource is created or deleted through the admin REST API (startup registration of existing resources does not notify). `resources/subscribe` (per-resource `notifications/resources/updated`) is intentionally not implemented: no current client watches individual managed resources, and the coarse list_changed signal already prompts a re-list, so the capability flag is left unset to keep the contract honest.

Broadcaster lifecycle is owned by the platform: `initSessions` builds it (postgres if a DSN is configured, in-memory otherwise) and `Close` shuts it down before the database connection closes. `WireGatewayBroadcaster` plugs it into every gateway toolkit, mirroring `WireGatewayTokenStore`; the prompt and managed-resource notifiers are bound during `finalizeSetup` and stopped via the platform lifecycle.

If the postgres broadcaster fails at startup (e.g. role missing `LISTEN` privilege), the platform falls back to in-memory and continues: `tools/list_changed`, `prompts/list_changed`, and `resources/list_changed` propagation degrades to single-replica scope rather than failing platform startup.

---

# Tools API Reference

## Trino Tools

### trino_query

Execute a read-only SQL query against Trino. Write operations are rejected. Annotated with `ReadOnlyHint: true`.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `query` | string | Yes | - | SQL query to execute (read-only) |
| `limit` | integer | No | 1000 | Maximum rows to return |
| `connection` | string | No | the kind's `default:` | Trino connection name |

### trino_execute

Execute any SQL against Trino including write operations (INSERT, UPDATE, DELETE, CREATE, DROP). Annotated with `DestructiveHint: true`. `read_only` is set per instance and the block applies to the connection the call names, or to the default connection when it names none: a call routed to an instance with `read_only: true` is refused while the other instances of the same toolkit still accept writes.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `query` | string | Yes | - | SQL query to execute |
| `limit` | integer | No | 1000 | Maximum rows to return |
| `connection` | string | No | the kind's `default:` | Trino connection name |

### trino_explain

Get the execution plan for a SQL query.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `query` | string | Yes | - | SQL query to explain |
| `connection` | string | No | the kind's `default:` | Trino connection name |

### trino_browse

Browse the Trino catalog hierarchy. Omit all parameters to list catalogs. Provide `catalog` to list schemas. Provide `catalog` and `schema` to list tables (with optional `pattern` filter). Parameters: `catalog`, `schema`, `pattern`, `connection` (all optional).

### trino_describe_table

Get table schema and metadata.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `table` | string | Yes | - | Table name (can be `catalog.schema.table`) |
| `connection` | string | No | the kind's `default:` | Trino connection name |

### trino_export

Export query results directly to a portal asset (CSV, JSON, Markdown, text), bypassing the LLM token budget. Use after validating the query shape with `trino_query`. Only metadata is returned to the agent. Requires portal + trino configured. SQL runs through the read-only interceptor. CSV cells hold each value exactly as the query returned it, including one starting with `=`, `+`, `-` or `@`. Sensitivity tags inherited from source datasets. The response is one JSON object, returned as the tool's structured result and as a single text block; the platform's `call_reference` is merged into that object, so a client reading only the structured result sees `asset_id`, `portal_url`, `row_count`, `size_bytes` and `call_reference` together, as it does for `api_export`.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `sql` | string | Yes | - | SQL query (read-only enforced) |
| `format` | string | Yes | - | csv, json, markdown, or text |
| `name` | string | Yes | - | Display name for the asset |
| `connection` | string | No | default | Trino connection name |
| `description` | string | No | - | Asset description |
| `tags` | array | No | [] | Lowercase kebab-case tags (max 50 chars, max 20; `_sys-` reserved) |
| `limit` | integer | No | deployment max | Row limit (subject to deployment cap) |
| `idempotency_key` | string | No | - | Dedup key to prevent duplicate assets on retry |
| `timeout_seconds` | integer | No | deployment default | Query execution timeout |
| `create_public_link` | boolean | No | false | Generate a public share link for automation |
| `resource` | object | No | - | Land the result in the managed resource at `{path, filename}` instead of a new asset, create-or-replace by path (#1663). Mutually exclusive with `idempotency_key` and `create_public_link`. `api_export` and `graphql_export` take the same block; see "Landing a response in a managed resource". |

Configuration: `portal.export.enabled` (auto), `portal.export.max_rows` (100000), `portal.export.max_bytes` (100MB), `portal.export.default_timeout` (5m), `portal.export.max_timeout` (10m).

### trino_list_connections

List configured Trino connections. No parameters.

## DataHub Tools

Catalog relevance search lives in the universal `search` tool and reading one catalog entity by URN is `fetch` (see Knowledge Tools). The DataHub toolkit registers the two reads a reference read cannot do, `datahub_browse` (a hierarchy walk) and `datahub_get_lineage` (a graph query with direction and depth), plus the write tools. The five by-URN reads it used to register are gone, not aliased; each is one `fetch` call on the same URN:

| Retired tool | Replacement |
|--------------|-------------|
| `datahub_get_entity` | `fetch urn:li:dataset:<id>` |
| `datahub_get_schema` | `fetch urn:li:dataset:<id>` (the record's `schema`) |
| `datahub_get_queries` | `fetch urn:li:dataset:<id>` (the record's `queries`) |
| `datahub_get_glossary_term` | `fetch urn:li:glossaryTerm:<id>` |
| `datahub_get_data_product` | `fetch urn:li:dataProduct:<id>` |

A fetched dataset carries at least what the three dataset reads returned together: business context (owners, tags, glossary terms, domain, deprecation, custom and structured properties, incidents, data contract), identity (name, type, platform, sub-types, created), the declared `schema` (fields with types, nullability, field-level tags and terms, primary and foreign keys), the catalog's saved `queries`, `related_documents`, and `query_availability` (whether it is queryable, the table to query, the connection, the estimated row count) when a query provider is configured. A part the catalog could not serve is named in `unavailable` rather than dropped. A fetched glossary term carries its definition, `parent_node`, `owners`, `custom_properties`, and the datasets that carry it; a fetched data product carries its domain, owners, properties, and member datasets.

### datahub_get_lineage

Get dataset or column-level lineage. Set `level=column` for column-level lineage showing which upstream columns feed each downstream column. Parameters: `urn` (required), `level` (dataset/column), `direction` (UPSTREAM/DOWNSTREAM, dataset only), `depth` (max 5, dataset only), `connection`.

### datahub_browse

Browse the DataHub catalog by category. Set `what=tags` to list tags, `what=domains` to list data domains, or `what=data_products` to list data products. Parameters: `what` (required: tags/domains/data_products), `filter` (optional, tags only), `connection`.

### datahub_create

Create a new entity in DataHub. Uses `what` discriminator to select entity type: tag, domain, glossary_term, data_product, document (1.4.x+), application, query, incident, structured_property, data_contract. Parameters: `what` (required), `name` (required for most types), additional fields vary by type, `connection`. Only available when `read_only: false`.

### datahub_update

Update metadata on a DataHub entity. Uses `what` discriminator: description, column_description, tag (add/remove), glossary_term (add/remove), link (add/remove), owner (add/remove), domain (set/remove), structured_properties (set/remove), structured_property, incident_status, incident, query, document_contents (1.4.x+), document_status (1.4.x+), document_related_entities (1.4.x+), document_sub_type (1.4.x+), data_contract. Parameters: `what` (required), `urn` (varies), `action` (required for tag/glossary_term/link/owner), `connection`. Only available when `read_only: false`.

### datahub_delete

Delete an entity from DataHub. Uses `what` discriminator: query, tag, domain, glossary_entity, data_product, application, document, structured_property. Parameters: `what` (required), `urn` (required), `connection`. Only available when `read_only: false`.

### datahub_list_connections

List configured DataHub connections. No parameters.

## S3 Tools

Two tools (#1591): `s3_list` lists the buckets of a connection when no bucket is named, or the objects in one bucket when one is; `s3_object` performs one action over a `(bucket, key)`. Both are registered on every S3 connection. `put`, `copy` and `delete` on a connection configured `read_only: true` are refused with an error naming the connection; `get`, `metadata` and `presign` still work there.

### s3_list

List buckets, or the objects in one bucket.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `bucket` | string | No | - | Bucket to list objects in; omit to list the connection's buckets (filtered by its `bucket_prefix`) |
| `prefix` | string | No | - | Key prefix filter |
| `delimiter` | string | No | - | Groups keys at this character (`/` lists one folder level) into `common_prefixes` |
| `max_keys` | integer | No | 1000 | Objects per page (1-1000) |
| `continuation_token` | string | No | - | The previous page's `next_continuation_token` |
| `connection` | string | No | first configured | S3 connection name |

### s3_object

One action over an object. Parameters: `action` (`get`, `metadata`, `put`, `copy`, `delete`, `presign`), `bucket`, `key` (required), `connection`; for `put`: `content` (text, or base64 with `is_base64: true`), `content_type`, `metadata`; for `copy`: `dest_key` (required), `dest_bucket` (default: the same bucket), `metadata`; for `presign`: `method` (`GET` default, or `PUT`), `expires_in` (seconds, default 3600, maximum 604800). `get` is bounded by the connection's `max_get_size` and returns binary content as base64; `put` is bounded by `max_put_size`; `presign` signs against `public_endpoint` when the connection sets one.

## Knowledge Tools

### memory_capture

The one way to record knowledge, in the memory toolkit. The `type` (sink-class) drives routing: `personal_preference` and `episodic_event` write live to memory; `business_knowledge`, `schema_entity`, and `operational_rule` write a knowledge-dimension record carrying a pending insight overlay (`insight_status=pending`) that `apply_knowledge` can promote. Capture is recall-first over the caller's memory (superseded rows excluded so a dead predecessor never absorbs the supersede; stale rows stay matchable, since a restatement corrects them): every near-duplicate at or above the supersede threshold (0.9 cosine) is superseded by the new capture, not appended (all of them, so a capture over an already-duplicated pair consolidates the whole set); the response's `superseded` field carries the best match id and `superseded_ids` the complete list. Matches in the 0.75-0.9 band are returned as `similar_existing` candidates (id + score) so the agent can consolidate (`memory_manage` update/consolidate) instead of leaving a near-duplicate. When the capture carries `entity_urns`, only records sharing an entity URN can match. Recall requires an embedding provider; without one, captures append.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `type` | string | Yes | - | personal_preference, episodic_event (live); business_knowledge, schema_entity, operational_rule (reviewed) |
| `content` | string | Yes | - | Knowledge to record (10-4000 chars) |
| `entity_urns` | array | No | [] | Related DataHub entity URNs (schema_entity); max 10 |
| `suggested_actions` | array | No | [] | Proposed catalog changes for apply_knowledge (schema_entity) |
| `confidence` | string | No | medium | high, medium, low |
| `source` | string | No | user | user, agent_discovery, enrichment_gap |
| `thread_ids` | array | No | [] | Feedback threads this capture resolves |

Captures are owned by the user's email (`memory_records.created_by`), the same key `memory_manage`, search, and the portal use, so a person's knowledge and memory appear together under their **My Knowledge** view. The sink-class is stored on `memory_records.sink_class` (derived from the record's dimension when absent).

### search

The universal, topology-free discovery entry point (renamed from `knowledge_search` in #645; corpus widened to API endpoints and connections). Call it FIRST. One query fans across every searchable source the caller can access and returns results grouped by source with a coverage summary, so the agent sees the shape of the answer space instead of tunneling into the first tool that comes to mind. Structured catalog navigation stays in `datahub_browse`; the scoped API drill-down stays in `api_discover`. Registered whenever at least one source is available.

Corpus (everything the persona can access): the technical catalog (DataHub, when configured), the governance vocabulary (DataHub glossary terms, tags, and domains as first-class entities, #1160), context documents (the non-dataset knowledge home that predates knowledge pages, surfaced both by relevance and by the entities they document, the same way knowledge pages are, #692), canonical knowledge pages (the internal-knowledge home for business/domain ontology, stored as markdown in the portal database and searched over their full content), the caller's personal memory (non-knowledge dimensions; captured/remembered knowledge surfaces via the insights provider, so a record is never double-listed), captured insights, the caller's feedback threads (lexical, since threads carry no embedding), saved assets, managed resources (human-uploaded reference material, indexed over their metadata AND a bounded text prefix extracted from the uploaded file, so a data dictionary is found by a column name that appears only inside it; #1012), prompts, managed scripts (found by name, description, tags, and typed parameter contract, never by their source code; #1302), the caller's own recorded calls (each hit carrying its derived outcome and its reuse count, so a proven statement outranks a guess; #1321), the caller's own sessions (matched against the purposes their calls stated and the names of the assets they saved, ranked lexically because a session is derived from the audit log and so has no row to carry an embedding; the two arms use the indexes that do exist — the audit purpose full-text index added by migration 000108, and portal_asset_fts; #1322), API endpoints (aggregated across every API gateway connection, reusing `api_discover`'s per-connection semantic/hybrid ranking, each gateway applying its own route policy fail-closed), and connections. API endpoints and connections are in the default corpus, not behind an opt-in. Memory, insights, feedback, and assets are per-user, scoped server-side to the caller (memory and insights by email `created_by`/`captured_by`, feedback by author email, assets by `owner_id`) and fail closed when their scope key is absent; the catalog, the governance vocabulary, knowledge pages, prompts, endpoints, and connections are shared. Managed scripts are per-user: a script is its owner's, so an unidentified caller reaches none. Managed resources are visibility-scoped rather than per-user: the provider derives the caller's visible (scope, scope_id) set exactly as the `resources/list` middleware does (global to everyone, persona to its members, user to its owner) and passes it into the SQL, so a resource the caller could not list is never ranked. A managed script is scoped the same way from `script.Script.VisibleTo` — global to everyone, persona to a caller BELONGING to it, personal to its owner — applied as a store predicate rather than a filter over the answer, and the ranking additionally skips dead ends (disabled, deprecated, superseded), which is a ranking rule and not an access rule: a caller holding an `mcp:script:<id>` reference to a retired script still fetches its contract, which states plainly that nothing will run it. Persona visibility keys on MEMBERSHIP — `knowledge.Caller.Personas`, resolved from roles by the same resolver the resources middleware uses and bound via `search.Toolkit.SetPersonasForRoles` — never on `Caller.Persona`, the persona the request resolved to: resolution substitutes the configured default persona when a caller's roles match none, so using it would hand every unmatched caller the default persona's material. With no resolver bound the set is empty and the caller sees only global plus their own user-scoped material (fail closed). A resource hit additionally carries an MCP `resource_link` content block with the canonical `mcp://` URI so a client with native resource support can attach the file itself. Knowledge pages are org-shared and editable by personas with `apply_knowledge` access; everyone can read them, and anyone can add feedback (threads). A search never surfaces another user's private records or a route a persona could not invoke; an anonymous caller still sees shared sources but no per-user data. The three topology sources (catalog, connections, endpoints) are additionally narrowed by the caller's persona `connections.allow` rules through the same predicate that authorizes a tool call (#1108), so discovery and authorization cannot drift: a catalog dataset is attributed to a connection through its DataHub platform name and hidden when the persona reaches none of the candidates, while a dataset that maps to no configured connection stays visible rather than being hidden on a guess. `fetch` applies the identical boundary, returning not-found for a denied dataset URN or connection reference so a citation cannot read around what search omitted, and `list_connections` enumerates only granted connections.

A query may be text (`intent`), entity-keyed (`entity_urns`), or both. The entity path unions every source linked to those DataHub URNs (the catalog entity, URN-linked insights, and the caller's URN-linked memory), expanded along lineage; per-user scope and the catalog's access rules still apply to each source.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `intent` | string | Conditional | - | Natural-language description of what is sought; provide intent, entity_urns, or both |
| `context` | string | No | - | Optional surrounding context, folded into the intent to sharpen relevance |
| `entity_urns` | array | Conditional | - | Exact entity-keyed lookup: everything linked to these DataHub URNs (catalog entity, URN-linked insights, the caller's URN-linked memory), lineage-expanded |
| `status` | string | No | - | Optional filter by insight review status (pending, approved, rejected, applied, superseded) |
| `sources` | array | No | - | Narrow to named sources (catalog, governance, context_documents, knowledge_pages, memory, insights, feedback, assets, resources, prompts, scripts, calls, sessions, endpoints, connections); only narrows, never opts past scope; an unrecognized name is echoed back in `unknown_sources` rather than silently ignored |
| `limit` | integer | No | 10 | Total results to display across all sources (max 50). Also the per-source ranking depth when it exceeds 25, so a search narrowed to one source returns up to `limit` from it |

The display set is balanced rather than a flat relevance list: a total budget with a per-source floor (every matching source stays visible), a per-source ceiling (none runs away), and redistribution of unused budget to sources with more relevant hits. Ranking is hybrid (semantic vector + lexical) when an embedding provider is configured, falling back to lexical-only otherwise; an entity-only query reports ranking `entity`. The response carries `ranking`, `count` (total shown), a `groups` array (`{source, hits[]}` where each hit pairs the matched `text` with its `source`, a `ref` (record id within that source), a relevance `score`, a canonical `reference` (the citation `fetch` dereferences), and where present `status`, `entity_urns`, and `dimension`), and a `coverage` array (`{source, matched, shown, matched_capped, withheld}`) reporting how many matched beyond what is displayed (the anti-tunnel signal) and how many the persona connection boundary removed. Each source ranks up to `limit` candidates (25 when `limit` is smaller) and `matched` counts within that bound; a source with more matching records than the bound carries `matched_capped: true`, making its `matched` a floor rather than a total (#1585). Without it a source that matched exactly the bound and one that matched thousands both read as `matched: 25, shown: 25`, which tells the agent it has seen everything at the moment it has not; raise `limit` to rank further into one source, or browse it for a true `total`. A count the persona connection boundary shortened is exact for what it says and is not flagged. `withheld` is present only when something was hidden, is accompanied by a top-level `withheld_notice` naming the persona and the remedy, and a source filtered down to nothing still reports its withheld count -- so an agent reads "present, but not yours to see" instead of concluding the data does not exist and re-deriving it.

The catalog source ranks two ways (#1131). The platform keeps its own index of every catalog dataset's text (name, description, tags, domain) in `catalog_datasets` and ranks the query against it first (hybrid vector + full-text); DataHub's own keyword search follows as the recall tail, covering fields the mirror does not carry (column names, glossary terms, ownership), and a dataset both return is emitted once at the index's rank. This is what gives a fact written into a dataset description — by `apply_knowledge` with sink `datahub`, or by a steward in DataHub — a search-time route that does not require a tool result to already name its entity, closing the asymmetry with knowledge-page sinks, which always rode the platform's own search. The corpus is refreshed by the `catalog-datasets` indexjobs consumer (internal/platform/datasetindex): its Source enumerates the catalog, materializes the mirror, and the Sink's atomic replace prunes datasets the catalog no longer returns, so a deleted dataset cannot linger as a hit. Scheduling needs no goroutine — `FindGaps` reports the corpus stale once `knowledge.catalog_index.sync_interval` (default 30m) has passed since the last enumeration, and the queue's own claim/lease makes the sweep single-flight across replicas. `knowledge.catalog_index.max_entries` (default 5000) caps the mirror and a truncated sweep is logged. Every hit is still dereferenced against DataHub by URN, the persona connection boundary applies to index hits exactly as to DataHub-ranked ones, and an index failure degrades to the DataHub arm rather than blanking the catalog source. The index needs a DataHub semantic provider, a database, and an embedding provider; without them the catalog is ranked by DataHub's keyword search alone.

The `governance` source (#1160) makes DataHub's glossary terms, tags, and domains searchable and fetchable as entities in their own right rather than as attributes of a dataset: before it, asking what a business term meant returned the datasets tagged with it, and a `urn:li:glossaryTerm:` reference a knowledge page legitimately carries (#1159) had no `fetch` owner. It is a sibling of `catalog`, not an arm of it -- the two answer different questions from different corpora ("what does X mean here" versus "which datasets are about X"), and the allocator balances the display budget per source, so a definition is never crowded out by a broad dataset match; separate ownership also keeps the `fetch` reference partition one-prefix-one-owner. It is text-only: a governance URN handed to `search` is a reference, and `fetch` is the verb that dereferences one. Each vocabulary is ranked by the best read upstream offers for it, an asymmetry that is upstream's rather than a design choice: glossary terms and tags have a name search that receives the intent and leads (it ranks against DataHub's index and is not bounded by an enumeration page, so a large glossary stays fully reachable), falling back to enumerate-and-rank-locally when it returns nothing, because `search` receives a natural-language intent while those upstream searches match a NAME -- "what does net revenue mean here" is a question, not a label. Domains have no upstream search at all, so the whole set (DataHub returns at most 100) is enumerated and ranked by the same lexical token-overlap rule the `connections` source uses for the same reason. A read that fails fails its vocabulary, with one exception: a fallback enumeration failing AFTER a successful but empty search does not, since the vocabulary already answered and the fallback was only trying to do better. The three reads run concurrently, so the source costs one round trip rather than three, a vocabulary that fails costs its own recall, and the source is blanked only when every vocabulary failed. Results are interleaved so no vocabulary crowds out the others, and each hit carries the definition in its text so a term match needs no follow-up `fetch`. Connection boundary: a governance URN carries no `dataPlatform` segment and so is attributable to no connection, which by the documented rule for an unattributable URN leaves it visible (the mapping failed, not the permission check); the datasets a `fetch` lists under an entity are ordinary catalog hits and are filtered by the caller's boundary, with the removed count and reason carried on the fetched entity. A fetched entity fills `content` with `{urn, kind, name, description, datasets[], more_datasets?, datasets_withheld?, notice?}`; the carrier list is capped at 25 and `more_datasets` reports that the list is not known to hold every carrier (the catalog counted more, or could not count at all), so a bounded list never reads as the whole membership. A tag and a domain have no by-URN read upstream, so each is resolved by listing its vocabulary and matching -- an entry past the page DataHub returns is a clean not-found, the same bound the portal's reference labels carry.

Browse mode (#695) is `search`'s exhaustive counterpart: relevance ranking with per-source flooring/capping cannot list a source in full, so browse pages the complete set of one source with a total count and NO relevance threshold, for auditing, deduping, governing, or migrating a corpus an agent must first obtain whole. It is the same `search` tool. A call enters browse mode when it carries exactly one `sources` entry, no `intent`, and no `entity_urns`; `offset`/`limit` page it (limit default 50, max 100). Browsable sources are `knowledge_pages` and `context_documents` (the two tiers with no prior MCP enumeration); browsing zero or several sources, an unknown source, or a non-browsable one is a tool error that names what can be browsed. The response is a flat, unranked page `{source, total, offset, limit, count, items[]}`: `total` is the source's full member count (how many pages remain) and each item carries the same `reference` that `fetch` reads in full. Context-document enumeration includes every document (drafts and hidden ones included), so the page and `total` describe the same complete set. Scope mirrors `search`: the two browsable sources are org-global so any caller may enumerate them, and a per-user source is never browsable for an anonymous caller. The document total comes from a DOCUMENT-scoped `searchAcrossEntities` count and the page from a `*` document listing with offset/limit, so no mcp-datahub change is required.

### fetch

The companion read verb to `search` (#694). `search` returns navigational pointers with truncated snippets; `fetch` dereferences one pointer's `reference` back to its complete content, so an agent can read in full what it found. It is the single consumer of the `reference` every search hit already carries, and it collapses the previously fragmented scoped readers (the five `datahub_get_*` by-URN reads, `manage_asset` get, `manage_prompt` get) into one verb. Registered alongside `search` (same toolkit) whenever search exists.

A reference comes in one of two namespaces: `urn:li:...` is the external DataHub catalog scheme, `mcp:...` is the internal-platform scheme. `fetch` accepts both, routing each well-formed reference by its form to the owning source: knowledge pages (`mcp:knowledge_page:<id>` -> the full markdown body; the page's slug is accepted in place of the id, with the id resolved first so a slug can never shadow a page asked for by id), context documents (`urn:li:document:<id>` -> the full document body, the only MCP path to it), catalog datasets (`urn:li:dataset:<id>` -> the dataset's catalog context), governance entities (`urn:li:glossaryTerm:<id>`, `urn:li:tag:<id>`, `urn:li:domain:<id>` -> the entry's name and definition plus the datasets that carry it, #1160), saved assets (`mcp:asset:<id>` -> the asset's metadata record; the blob bytes stay in S3, reached with `s3_object` `get` or `presign`), prompts (`mcp:prompt:<id>` -> the full prompt), connections (`mcp:connection:(kind,name)` -> the connection descriptor), the caller's captured insights (`mcp:insight:<id>` -> the full insight, scoped to the caller), the caller's personal memory (`mcp:memory:<id>` -> the full memory record, scoped to the caller; fetch-only, not citable on a page) (#699), and managed resources (`mcp:resource:<id>` -> the metadata record plus the file itself, in whichever form the file admits: a text file's contents inline, a PDF's extracted text, the XML parts of an Office document (`.docx`, `.xlsx`, `.pptx`, `.odt`, `.ods`, `.odp`, and any zip) each under a banner naming the part and its size followed by an inventory of what was not rendered, an image as an MCP `image` content block a model looks at directly, and any other family as an MCP `resource` content block carrying the bytes for the client's own tools to open. Only a file above the 1 MB inline threshold the MCP read path uses returns metadata alone, with the canonical `mcp://` URI, MIME type, and size, which `resources/read` reads the whole file by with no limit. Until #1657 every non-textual family returned that metadata row, and an agent handed a PDF's size and media type reported that the platform was refusing it the file. PDF text is read by PDFium compiled to WebAssembly and run under wazero, cgo-free, with its module compiled on the first PDF a deployment reads rather than at startup; the same reader feeds the content index, so a PDF and a presentation are found by words that appear only inside them. Fetch-only, not citable on a page since a scoped resource would be a broken citation outside its scope) (#1012, #1657), and managed scripts (`mcp:script:<id>` -> the script's CONTRACT: name, description, owner, typed parameter contract, whether a run requested now would be admitted (and the run gate's own refusal when it would not), schedule cadence, and the last successful run with what it produced, each output naming its shape — a portal asset the platform still serves, or an object delivered to a configured bucket destination that it wrote and does not hold; never the source code, which is read with `manage_script get` and on the portal script page, by anyone signed in (#1866); fetch-only, not citable on a page for the same reason a resource is not, and resolved by ID so renaming a script never breaks a stored reference) (#1302), recorded calls (`mcp:call:<id>` -> the cataloged call: the statement or request that ran, the purpose stated for it, what came of it, and how many later sessions re-ran it; reading one is what makes the caller's own re-run of it creditable as reuse) (#1321), and sessions (`mcp:session:<id>` -> the caller's own session: its summary, the assets and insights it produced as references to follow, and its call timeline in order, each call carrying the purpose stated for it and, where the catalog recorded one, its kind, its outcome, and the `mcp:call:` reference that reads the record in full; a call with no record stays on the timeline and simply carries nothing to fetch) (#1322). The usual source of a reference is a `search` result's `reference` field, but `fetch` is not limited to references `search` produced: a well-formed reference held from another tool works too (for example a `urn:li:dataset:...` from `datahub_get_lineage` or an `entity_urns` lookup). Feedback threads and API endpoints emit no reference and are not fetch targets. The per-user sources (assets, memory, insights, recorded calls, sessions) are read only for the identity that owns the record, so a non-owner gets a structured not-found, exactly as search would never have surfaced the record. A session id is guessable, being the handle the agent threads, so the scoping is a predicate on the audit rows themselves: another caller's session groups to no rows and is indistinguishable from one that never ran.

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `reference` | string | Yes | A reference in the `urn:li:...` (DataHub) or `mcp:...` (internal) namespace; usually a `search` result's `reference` field passed verbatim, but any well-formed reference held from another tool works |

Scope mirrors `search` exactly: a per-user source (assets, recorded calls, sessions) is read only for the identity that owns the record, and a persona/personal-scoped prompt only for the matching caller, so `fetch` never returns content the same caller could not have found with `search` (a reference outside the caller's scope is reported as not-found, indistinguishable from a missing one, so existence does not leak). The response is `{found, reference, document?, message?}`: a resolved reference returns `found: true` with `document` (`{reference, source, title, body?, content?, entity_urns?}`, where text-bodied sources fill `body` and structured sources fill `content` with the source-native payload); a stale, unknown, or out-of-scope reference returns `found: false` with an explanatory `message` (a structured not-found, not a tool error, so a dangling citation is a normal answer). A malformed call (empty reference) and a real backend failure are tool errors. A catalog or governance reference the catalog has never ingested is one of those not-founds, and the catalog reports it: a dataset and a glossary term by DataHub's exists field, a data product by the absence of the properties aspect a product cannot be created without. Whether the record is documented has no bearing on it, so a dataset with no description, owner or schema, and a glossary term at the root with no definition, resolve like any other. The same read backs search by entity_urns, so a URN that names nothing reports no match rather than a hit whose reference then fails to fetch. This needs mcp-datahub v1.15.1 or later against a DataHub that reports exists; on a DataHub that omits the field, a URN with no entity resolves to a stub built from that URN and is reported as found (#1605, #1610).

### apply_knowledge

The sink router (#633, #1607). Review, synthesize, and promote captured insights to their canonical home. Admin-only. Requires `knowledge.apply.enabled: true`. The `apply` action's `sink` selects the target: `datahub` (default) applies `changes` to a catalog entity; `knowledge_page` promotes the insight(s) to a canonical portal knowledge page, found-or-created by `page.slug`; `agent_instructions` promotes an operating rule into the deployment's own customized agent instructions, found-or-created by `instructions.section`. The destination is chosen at apply, suggested by whether the insight is entity-anchored (carries entity_urns) rather than frozen at capture: the capture-time sink-class is a non-binding hint, and any insight can be applied to either sink. Both record a changeset (page promotions target `kp:<slug>`) reversible via rollback (a page rollback soft-deletes a new page or restores a prior version, refused if the page was edited since). A promotion that wrote a page reports `portal_url`, the address that page is read at, in the payload and in the message (#1696): the response carried the page's id and slug and no address at all, so an agent handing someone a link composed one, and the one it composed opened nothing. The address is `<portal.public_base_url>/portal/knowledge/pages/<page_id>`, composed in one place for both sinks that write a page, and omitted on a deployment that declared no public address. `GET /api/v1/portal/knowledge-pages/{id}` resolves an id then a slug, the order `pkg/knowledge` resolves an `mcp:knowledge_page:` reference in, so a slug-shaped address opens too; the portal's page view keys the panels beside the body on the id the read returned rather than on the key the address carried, since their own routes take an id. operational_rule is stored as a page like business_knowledge; active rules-engine enforcement is a separate follow-up. To cite an entity on a page, pass it in `page.references` or write it in the body as plain text or a markdown link (not in backticks or a code block, which are treated as examples and ignored). Each entry in `page.references` is existence-checked before the page is written: a missing internal (mcp:) entity rejects the apply (a DataHub urn:li: reference is free text, stored as given). References in `page.references` and those carried from source insights attach with the promotion, so a rollback undoes them; a stale insight-carried reference is skipped rather than blocking. Inline body references are also filtered to those that exist, so a stale mcp: token in prose is skipped rather than blocking the page. A dropped insight-carried or inline-body reference is reported on the apply response: references_dropped lists the references that did not land (a deleted target, or an insight-carried reference not citable on a shared page such as mcp:memory:/mcp:insight:) and references_attached always carries the count of references that landed, so an agent can reconcile what it cited against what was attached (#696). Page creation is dedup-gated (#705): on a slug miss, the backend ranks existing pages by pure embedding cosine similarity against the candidate's content (embedded over the same title+body+tags composition the stored page vectors use), and a near-duplicate (cosine at or above a configurable threshold, default 0.85, active only when a real embedding provider is configured) is rejected with the candidate pages returned (id, slug, title, score) rather than created; the agent then re-applies against a candidate's slug to update it, or sets page.force_new to create a distinct page (a deliberate, auditable act separate from confirm). The identical gate runs on the portal REST create path (POST returns 409 with the candidates), so the two write surfaces cannot drift. A slug hit is always an update and is never gated. An oversized page (body bytes or markdown-heading count past a configurable threshold) returns a non-blocking split_suggestion on the apply response, nudging the agent to split it into focused, cross-linked sub-pages linked from a thin index page; the write still succeeds. The guidance steers toward many focused, cross-linked pages (progressive revelation) over one sprawling page. On a deployment with no DataHub connection the datahub sink is refused rather than silently accepted: the apply returns an error naming the blocked change types, knowledge.apply.datahub_connection, and sink: knowledge_page as the catalog-free destination, and records no changeset; rollback of a DataHub changeset is refused the same way. The add_prompt change type creates a platform prompt rather than a catalog write and still applies. The knowledge-page sink is unaffected.

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `action` | string | Yes | bulk_review, review, synthesize, apply, approve, reject, rollback, list_changesets, bulk_untag |
| `sink` | string | No | apply target: datahub (default), knowledge_page, or agent_instructions |
| `entity_urn` | string | Conditional | Required for review, synthesize, list_changesets, and apply with sink=datahub |
| `tag_urn` | string | Conditional | Required for bulk_untag; the tag (name or urn:li:tag:...) to remove from every entity that carries it |
| `page` | object | Conditional | {slug, title, body, summary?, tags?, references?} for apply with sink=knowledge_page. references is a list of serialized reference strings (mcp:<type>:<id> / urn:li:...) attached to the page independent of body text. The per-user forms, personal memory (mcp:memory:<id>) and captured insight (mcp:insight:<id>), are fetchable but NOT citable on a shared page (they are private to their owner, so a citation would resolve for no one else) and are rejected; promote an insight to the catalog and cite the resulting urn:li:... entity instead (#699). A managed script (mcp:script:<uuid>, #1855) IS citable: existence-checked like every internal reference, stored with a foreign key that removes the citation when the script is deleted, and resolved per reader, so the script's owner and administrators see the reference and any other reader reads the page without it (fetch answers them found=false) |
| `insight_ids` | array | Conditional | Source insights; required for approve, reject. On apply, pass the promoted insights so they are marked applied and the changeset is linked to them (closes the review loop; the queue then reflects what is live). Sink-class is a non-binding hint; any insight can be applied to either sink |
| `instructions` | object | Conditional | {section, body, slug?, title?, summary?, tags?, references?} for apply with sink=agent_instructions (#1607). section is the find-or-create key: a second promotion of the same section rewrites that section and leaves every other section byte-identical (the edit runs through pkg/textpatch, the same anchored-edit engine behind manage_asset patch). Recorded as a changeset with target ai:<section>, listed by list_changesets and reverted by rollback, which restores the previous text byte for byte and is refused when the layer was written again after the promotion. Refused when the rule names a tool this deployment does not register, naming the tool: the recognized prefixes are derived from the deployment's own registered inventory plus a floor list, so a retired api_/manage_/memory_/save_ name is caught as readily as a trino_ one. A body over 1,500 bytes is written to a knowledge page instead (slug and title defaulted from the section) and the section keeps one mcp:knowledge_page:<slug> index entry pointing at it, in the same form the platform baseline's page index uses; both halves are one changeset, so one rollback undoes both. Refused before any write when the layer would pass its 32,768-byte limit, and carries a size_notice past the 12,288-byte advisory. Needs a database-backed config store; without one the promotion is refused with sink=knowledge_page named as the alternative. Route domain knowledge to knowledge_page: this layer should carry at most a pointer to it |
| `changes` | array | Conditional | Required for apply with sink=datahub |
| `changeset_id` | string | Conditional | Required for rollback |
| `confirm` | bool | No | Required when `require_confirmation` is true (apply and rollback) |
| `review_notes` | string | No | Notes for approve/reject actions |
| `itemize` | bool | No | With bulk_review, also return the pending insights themselves (full insight_text body, captured_by, sink_class, created_at, suggested_actions_count, ...; full suggested_actions omitted, fetch for it), paginated by offset/limit. The response is bounded so it stays under the output limit: page_size_capped:true flags a short insights page (continue with next_offset) and by_entity_truncated:true flags a capped by_entity |
| `limit` | int | No | Page size for itemized bulk_review (default 20, max 100) |
| `offset` | int | No | Page start for itemized bulk_review; pass the previous next_offset to continue |

The `rollback` action reverts a changeset's changes to their before-image (removing added tags/glossary terms/documentation links, restoring a changed description), returns the source insights to the review queue as `pending` (`insights_returned_to_review`), and marks the changeset rolled back. It is refused if the changeset is already rolled back, if a newer changeset has since modified the same aspect, or if the changeset touched change types whose prior state was not captured or is irreversible (column descriptions, structured properties, custom properties, incidents, curated queries, context documents, prompts, delete_tag, bulk_untag). The `list_changesets` action lists an entity's changesets for discovering rollback targets. The apply response and each `list_changesets` entry carry a `revertible` boolean (computed with the same all-or-nothing gate rollback enforces, so any single unrevertible change makes the whole changeset false; there is no partial rollback) plus `unrevertible_change_types` when false, and the apply message instructs a rollback only when the changeset can actually be reverted, so a caller never plans on reversibility a changeset lacks (a column-level `update_description`, whose before-image is only the entity-level description, is the common false case). The field is structural revertibility; a structurally revertible changeset can still be refused at rollback time if already rolled back (list_changesets folds this in) or a newer changeset touched the same aspect. The `bulk_untag` action removes a tag (tag_urn) from every entity a catalog search finds carrying it, recording one changeset for audit; it requires confirm and is not auto-revertible (re-apply add_tag to restore).

## Memory Tools

Persistent memory for agent and analyst sessions. Requires `memory.enabled: true`. Tools are opt-in per persona (`memory_*` in `tools.allow`).

### memory_manage

Lifecycle operations for existing memory records; create new records with `memory_capture`. Commands: `update`, `forget` (soft-delete), `list` (persona-scoped), `review_stale` (admin review of stale memories; the staleness watcher checks only records carrying at least one entity URN, so `last_verified` is null on a record no catalog can be asked about rather than reading as a check nothing performed, and verification never moves `updated_at`, which stays the last content change), `review_duplicates` (list the caller's own active pairs at or above 0.75 cosine similarity, the backstop for near-duplicates the capture-time recall gate missed; requires the database-backed store with vector search), `consolidate` (supersede a duplicate by the record kept: `id` = keep, `duplicate_id` = supersede; both must belong to the caller, the record kept must be active, and the correction chain is preserved via `metadata.superseded_by`).

`review_duplicates` is summary-first and byte-bounded: each pair returns ids, `score`, `status`, timestamps, owner, and a bounded `content_preview` (first ~200 characters per side), not the two full records, so its listing never overruns the MCP output budget. It is not offset-paginated: the candidate set is score-ordered and shrinks from the top as pairs are consolidated, so offset paging would skip pairs; instead it returns the current highest-similarity pairs (at most `limit`, default 20) and sets `more_pairs: true` when the byte budget or the page limit hid lower-scored pairs. The pagination is the review loop itself: consolidate the surfaced pairs and re-run to surface the rest. Read a record in full with `fetch mcp:memory:<id>` or `memory_manage list` before consolidating.

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `command` | string | No | update, forget, list, review_stale, review_duplicates, consolidate (create with memory_capture) |
| `id` | string | For update/forget/consolidate | Memory record ID (for consolidate: the record to keep) |
| `duplicate_id` | string | For consolidate | The duplicate record the kept record supersedes |
| `dimension` | string | No | LOCOMO: knowledge, event, entity, relationship, preference |
| `category` | string | No | correction, business_context, data_quality, usage_guidance, relationship, enhancement, general |
| `confidence` | string | No | high, medium, low |
| `entity_urns` | array | No | DataHub URNs (max 10) |
| `metadata` | object | No | Arbitrary metadata |
| `limit` | integer | No | Page size for list (default 20, max 100); also caps review_duplicates pairs (which may hold fewer when the byte budget hits, more_pairs=true) |
| `offset` | integer | No | Pagination offset for list (not used by review_duplicates) |

### Reading memory back

Reading memory back moved to the universal `search` tool (see Knowledge Tools): relevance ranking, entity-keyed lookup, and DataHub lineage/graph traversal are all served there. The memory toolkit retains `memory_manage` for the write path.

The ranking it uses fuses embedding cosine similarity with a lexical full-text match signal, blended as `0.6 * semantic + 0.4 * lexical`, which improves recall on identifier-heavy content (entity URNs, column names, error codes) over pure vector search. The vector arm uses an `hnsw` ANN index on `memory_records.embedding`; the lexical arm a GIN index on `to_tsvector('english', content)`. When no embedding provider is available it degrades to lexical-only instead of erroring, which also surfaces rows whose embedding is NULL (saved during an outage).

Memory registers as a consumer of the shared index-jobs framework (`source_kind = memory`). The synchronous embed-on-write is preserved so a just-saved memory stays immediately recallable; a periodic reconciler backfills embeddings that were missed during an embedder outage (`embedding IS NULL`) or invalidated by a provider model swap (`embedding_model` differs from the current model), and the `memory` kind appears on the admin Indexing dashboard with an indexed/expected coverage ratio. Migration `000054_memory_hybrid_search` adds the `embedding_model` and `embedding_text_hash` breadcrumb columns plus the hnsw and GIN indexes.

## Portal Tools

The portal toolkit persists AI-generated assets (JSX dashboards, HTML reports, SVG charts) to S3 with PostgreSQL metadata. Requires `portal.enabled: true`.

### save_asset

Save AI-generated content to the asset portal as a versioned asset. Captures the calls the asset was built from: by default every data-access call the session made since its last save or export, or exactly the calls named in `sources`.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `name` | string | Yes | - | Display name (max 255 chars) |
| `content` | string | Yes | - | Artifact content |
| `content_type` | string | Yes | - | MIME type the asset is stored under. One of `application/json`, `application/octet-stream`, `application/sql`, `application/x-ndjson`, `application/xml`, `application/yaml`, `image/svg+xml`, `text/css`, `text/csv`, `text/html`, `text/javascript`, `text/jsx`, `text/markdown`, `text/plain`, `text/tab-separated-values`, `text/x-python`; anything else is refused |
| `description` | string | No | "" | Description (max 2000 chars) |
| `tags` | array | No | [] | Tags for categorization (max 20 tags, each max 100 chars) |
| `sources` | array | No | session window | The calls this asset was built from, as the `call_id` (or `mcp:call:<id>` reference) each query and API invocation returns in its own result. Replaces the default window; only the caller's own calls resolve (max 100) |

Response includes asset_id, portal_url (if public_base_url configured), provenance_captured status, and calls_recorded count. Content stored at `{s3_prefix}{user_id}/{asset_id}/content.{ext}`.

### manage_asset

List, retrieve, update, delete, share, or relevance-search saved assets (and manage collections). Ownership is enforced: users can only modify, share, or find their own assets. Human feedback on assets is handled by the separate `manage_feedback` tool.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `action` | string | Yes | - | list, get, update, delete, search, patch, locate, get_content, outline, stats, diff, share, list_shares, revoke_share |
| `asset_id` | string | Conditional | - | Required for get, update, delete, share, list_shares, and every content action |
| `content` | string | No | - | New content (for update — replaces S3 object) |
| `name` | string | No | - | New name (for update) |
| `description` | string | No | - | New description (for update) |
| `tags` | array | No | - | New tags (for update) |
| `max_versions` | integer \| null | No | - | How many versions this asset keeps (update). Omit to leave the setting alone, `null` to go back to the deployment's `portal.max_versions`, `0` to keep every version, `N` to keep the newest N. Negative is refused |
| `content_type` | string | No | - | New content type (for update, only when replacing content) |
| `sources` | array | No | session window | The calls behind a content edit (update and patch), as `call_id` values or `mcp:call:<id>` references. Recorded as a new capture beside the ones earlier versions carry |
| `query` | string | Conditional | - | Free-text relevance query (required for search) |
| `limit` | integer | No | 50 | Max results for list (max 200); ranked search defaults to 20 (max 100) |
| `recipient` | string | No | - | Who to share with (share): an email address, or a name resolved against the known-users directory. Omit for a link share |
| `permission` | string | No | viewer | `viewer` or `editor` (share). A link share is always viewer |
| `access_mode` | string | No | authenticated | Who a link share admits (share, no recipient): `authenticated` or `public` |
| `expires_in` | string | Conditional | - | Duration bounding a public link (`24h`). Required for `access_mode: public`, refused for every other share |
| `share_id` | string | Conditional | - | Share to end (required for revoke_share) |

The `search` action ranks the caller's own assets by relevance to `query` using the shared hybrid (vector + lexical) ranking — weighted hybrid when an embedding provider is configured, automatic lexical-only fallback otherwise — and returns each match with a `score` plus a `ranking` field (`hybrid` or `lexical`). It is scoped server-side to the caller's own assets by `owner_id` — the same ownership key the asset library list and the update/delete checks use, so search returns exactly the assets the caller can list (note: `owner_email` is a secondary display field that can diverge from `owner_id` across API-key vs OIDC identities, so it is deliberately NOT the scope key) — and fails closed when the caller has no identity. The Portal exposes the same ranking over HTTP at `GET /api/v1/portal/assets/search?q=...` and `GET /api/v1/portal/collections/search?q=...` (each row carries a `score`), mirroring the Knowledge & Memory and prompt search endpoints.

### Version retention (max_versions)

Every write to an asset records a version, and the version trail was the one history on the platform with no bound: `portal_asset_versions` had INSERT and SELECT and nothing else, no sweeper, and an asset delete that is a soft delete, so the `ON DELETE CASCADE` never fired. A managed script refreshing one dashboard hourly writes 24 versions a day, each with its own stored object, indefinitely.

An asset keeps its newest `portal.max_versions` versions (100 unless configured; `0` keeps every version, and a negative default is refused at startup by `portalcfg.MaxVersionsError`), and may override that with its own `max_versions`: NULL inherits the deployment value, `0` keeps every version, `N` keeps the newest N. Three surfaces write the same column — the portal's asset metadata form, `manage_asset action=update`, and `PUT /api/v1/admin/assets/{id}` — and all three refuse a negative value where it was entered; the column carries a non-negative CHECK behind them.

Retention is owner-or-admin on every surface. The portal's asset update route otherwise admits an editor-share recipient, and this one field is excluded from that grant: lowering the cap deletes the owner's history and its stored content at the next write, which is not a decision an edit grant carries. The portal form omits the field entirely from an update that did not move it, so an editor saving a rename never sends it.

The prune runs inside the transaction `portalversions.CreateVersion` already takes the asset row lock in, which is the single door every asset write passes through (portal handlers, admin routes, the asset toolkit, managed-script output, and both export adapters), so the cap converges at the write and needs no schedule. It is skipped entirely when the asset's history has not reached the cap, which is the common case on every write. A pruned version's stored object is deleted with its row, along with the thumbnails derived beside it — no row records a thumbnail, so a prune that removed only the content would cap the table while the bucket kept growing. The object deletes happen after the commit and are best-effort: the version the caller asked for has already committed, and an object that could not be removed is logged rather than turned into a failed save.

**A version's key is not private to it, and the cleanup must not assume it is.** `scriptexec` keys a managed script's output object by RUN (`internal/platform/scriptexec/export.go`), so a run that exports a document and then publishes data into the same asset writes two version rows at one key, and every version of such an asset shares one directory — which is also where its thumbnails are derived. The cleanup therefore reads the keys the surviving versions still own (`survivingObjectKeys`, inside the same transaction, after the delete) and skips every one of them. It reads the asset row's own thumbnail pointers into the same set, because since #1431 a capture outlives the version it was taken from: the image an asset is serving can sit in a directory this prune is deleting, and no surviving row derives its key. Deleting by key shape alone would take the live content object or the current thumbnail with the pruned row.

The prune also never deletes the key the asset row points at. Version 1 of an asset created before it had a second version carries the flat content key the asset itself still names, so the DELETE joins `portal_assets` and excludes `v.s3_key = a.s3_key`, exactly as `resource.PruneVersions` does.

Retention applies at the next write, not on a migration: an asset already carrying more history than the cap is untouched until it is written again, then trimmed. Knowledge page versions (`portal_knowledge_page_versions`) are the same shape and are not yet bounded.

### Sharing an asset from the session (share, list_shares, revoke_share)

"Share this with John" is one call rather than a trip to the portal. `manage_asset action=share` takes the same three decisions the portal's share dialog takes — who, what they may do, how it ends — and builds the share row through the same constructor the REST share routes use (`portaldomain.BuildShare`), so a share created from a session is indistinguishable from one created in the UI.

`recipient` accepts an email address or a person's name. A name is resolved against the known-users directory (the same directory the portal's share picker reads), matching case-insensitively against email, first name and last name; a full name is resolved by querying the most selective word and narrowing to the entries every word appears in. The lookup refuses to guess: a name that matches nobody, or more than one person, comes back as an error naming the candidates, and no share is created. An email address is taken as written, so a person who has never signed in can still be shared with; a deployment without a database has no directory, and a name there is refused with a pointer to give an address instead.

A share addressed to a person is `restricted` — it opens only for the named recipient and its creator, whoever else holds the URL — is `viewer` unless `permission: editor` is asked for, never expires (it ends when it is revoked), and sends the recipient the same "shared with you" email the portal sends, subject to their own notification preferences. The response reports `notified`, so an agent claims someone was emailed only when one was actually sent.

Omitting `recipient` mints a link. `access_mode: authenticated` (the default) opens for any signed-in platform user and lasts until revoked; `access_mode: public` opens for anyone holding the URL without signing in and therefore requires `expires_in`, which is refused for every other shape (#1279). A link is always `viewer`: it admits an audience the creator never enumerated, so it can never carry write access.

`list_shares` returns the shares that currently grant access to an asset — a revoked or expired share is not one of them, decided by the same liveness rule the public viewer gate applies — each with its recipient, permission, access mode, view URL, expiry and access count. `revoke_share` ends one by `share_id`, and its token stops opening the asset immediately. All three are owner authority (an editor share never carries the right to hand access on), an unauthenticated caller is refused outright rather than matched against the shared "anonymous" owner sentinel, and an admin is unrestricted as everywhere else.

### Editing content in place (patch, locate, get_content, outline, stats, diff)

Regenerating a whole document to change one sentence costs output tokens proportional to the size of the document rather than the size of the change, and every regeneration is a chance to silently drop an unrelated paragraph. `manage_asset` and `manage_prompt` therefore share one content-editing grammar, implemented once in `pkg/textpatch` (pure text in, text out, with no knowledge of assets, prompts, S3, or MCP) and rendered to the protocol by `pkg/textpatch/patchmcp`. Both tools splice the identical JSON Schema fragment, so the grammar an agent reads on one tool is literally the grammar it reads on the other.

Verbs: `outline` (heading tree with levels, line numbers, per-section byte size, plus `landmarks` on an HTML/JSX/SVG asset), `get_content` (whole body, one `section` or `selector`-addressed element, or a `line_start`/`line_end` range), `stats` (size, line count, current version, content type, sha256 body hash, no body), `locate` (literal `find` or regex `pattern` matches with count, line, byte offset, enclosing section, and a copyable context window), `patch` (an ordered list of anchored edits), and `diff` (two stored versions, or a pending prompt draft against the approved snapshot). The intended loop for a large document is outline or locate to decide where, then patch that place; the body crosses the wire in neither direction.

Patch operations, each an entry in the `edits` array: `replace` (the default; literal `find` swapped for `replace`, an empty `replace` deletes), `insert_before` / `insert_after` (`text` positioned relative to the anchor, anchor kept), `replace_section` (a region, either a heading through the next heading of the same or higher level or a `selector`-addressed element's balanced subtree, replaced by `text`), `replace_content` (a `selector`-addressed element's interior replaced by `text`, its own tags kept byte for byte — the data-island operation; selector-only, and a void or self-closing element is refused because it has no interior), `move_section` (relocate a region `before` or `after` another heading, or with `position` set to `start` or `end`), and `append` / `prepend` (no anchor). `section` and `selector` also scope the anchor search on the anchored operations, which is how a repeated phrase becomes unambiguous without quoting a long anchor; a `Report > Methodology` path disambiguates repeated headings. The payload key is REQUIRED and is the only one: an edit writes from `replace` on `op: replace` and from `text` on every other writing op, and an edit that omits the key its op reads is refused with `PATCH_BAD_EDIT` naming its index, as is one carrying a key the grammar does not declare, which is named along with the accepted set. Deleting is therefore always asked for -- `"replace": ""` written out. Omission used to be read as an empty replacement, so an edit whose text rode on a misspelled key (`text` on a replace, `replacement`, `with`) had its payload ignored, replaced the anchor with NOTHING, saved a new version and reported success; the damage was a deletion nobody asked for in a source that usually still parsed, so `validate` passed too and a scheduled run was the first thing to notice (#1804). Every tool that adopts the grammar splices the same schema fragment from `pkg/textpatch`, so its edit items declare exactly these keys and are closed to the rest and an unknown key is refused by the argument validator before the handler runs; `manage_script` published a hand-written paraphrase whose edit items were a bare `{"type": "object"}`, which is how the keys inside an edit came to be validated nowhere, and `TestPatchGrammarIsNeverHandCopied` now fails any package outside `pkg/textpatch` that declares the property itself.

Naming a region is syntax-aware, derived from the content type through `pkg/contenttype` (never guessed from the bytes). Markdown (`text/markdown`, `text/plain`) names regions by ATX heading. HTML, JSX, SVG and XML (`text/html`, `text/jsx`, `image/svg+xml`, `application/xml`) name them by `<h1>`..`<h6>` heading (resolved as markdown resolves `#`) or by CSS `selector`, whose region is the addressed element's balanced subtree, running from its start tag through its matching end tag, so a structural edit can never cut a tag in half. The element tree is built from a tokenizer with byte offsets recorded per node; markup that cannot be resolved into a balanced tree is refused with `PATCH_UNRESOLVED_MARKUP` rather than guessed at. Supported selector forms: type (`section`, `Card`), `#id`, `.class` (also matches a JSX `className`), `[attr]`, `[attr=value]`, joined by descendant (space) or child (`>`) combinators; HTML/SVG tags match case-insensitively while a JSX component name is case-sensitive. `selector` on a markdown document, or `section`/`selector` on a structureless one (JSON, CSV, SQL), is refused with a message naming the right alternative (`PATCH_NO_STRUCTURE` for the structureless case). On a headingless dashboard, `outline` returns `landmarks`: every element carrying an `id` or `data-*` marker, each with tag, a copyable selector, line, and byte size, so an agent finds where to patch without reading the body.

Matching rules: an anchor or `selector` must resolve to exactly one span. Zero and more than one are both errors that report the count; `occurrence` (`first`, `last`, `all` for text anchors, or a 1-based integer) is the explicit opt-in, and an `occurrence: "all"` edit reports how many spans it changed. A region is a single element, so a `selector` uses `first`/`last`/index, not `all`. Matching is exact first, with exactly one retry that normalizes CRLF to LF and ignores trailing whitespace per line (reported as `normalized: true`); there is no fuzzy, similarity, or semantic matching, because a plausible-but-wrong edit applied silently is worse than a rejection the agent can correct. `pattern` is the regex alternative on both `locate` and `patch`, with `$1`-style capture expansion in `replace`; Go's RE2 gives linear match time so a pathological pattern cannot hang the server, and a pattern-length cap, a match cap, and the all-or-nothing failure rule still apply. Edits apply in order against the evolving body, so a later edit can anchor on text an earlier edit introduced.

Atomicity: every edit resolves against an in-memory copy and the first failure aborts the whole call, writing nothing; the error names the failing edit by index. `base_version` is optional and checked when supplied, refusing a stale patch with the current version in the error. The response never echoes the new body: it returns the new version, the new size and line count, a per-edit report (operation, normalized flag, spans touched, landing line), and a unified diff of the changed hunks only. Unified diff is deliberately an output format and never an input format: as input it would make correctness depend on line numbers and context lines the model must reproduce exactly. `dry_run: true` returns exactly that report without writing. Line numbers are read output only and are never accepted as an edit anchor.

Integration: an asset patch goes through the same version write path as a full content update, so it produces an ordinary new version and `revert` still works, and the version's change summary is the caller's `change_summary` or a generated one ("3 edits via patch") instead of a fixed constant. A prompt patch goes through `prompt.ApplyEdit`, so patching an approved global or persona prompt produces a pending draft version and the approved snapshot keeps being served. Ownership, persona authorization, and audit are untouched; the read-only asset verbs additionally accept a share grant (owner, direct share, or collection share), while `patch` remains owner-only like `update`.

Every verb is text-only: a PDF, image, or other binary target is refused with `PATCH_NOT_TEXT` via `contenttype.IsTextual`, the platform's single detection seam. Error codes, all carrying the standard `{code, category, message, hint}` envelope: `PATCH_NO_MATCH`, `PATCH_AMBIGUOUS`, `PATCH_STALE_BASE`, `PATCH_NOT_TEXT`, `PATCH_TOO_LARGE`, `PATCH_SECTION_NOT_FOUND`, `PATCH_BAD_PATTERN`, `PATCH_BAD_EDIT`, `PATCH_BAD_SELECTOR` (the selector does not parse), `PATCH_NO_STRUCTURE` (the content type has no addressable structure), and `PATCH_UNRESOLVED_MARKUP` (the markup could not be resolved into a reliable element tree).

### manage_resource (create, replace_content, get, list, delete, extract)

Manages a file in the managed resource library. A managed resource is the only kind of file an asset can reference, and creating and revising one was reachable only from the portal UI and the REST surface behind it, both of which need a person with a browser. That left the data half of a referencing asset unrefreshable by the platform: an agent could rewrite a report on a schedule and could not rewrite the CSV the report reads.

| Parameter | Type | Required | Default | Description |
|-----------|------|----------|---------|-------------|
| `action` | string | Yes | - | `create`, `replace_content`, `get`, `list`, `delete` or `extract` |
| `reference` | string | Conditional | - | `mcp:resource:<id>`, the reference a search hit, a fetch document, or a create carries. Required for `replace_content`, and for `extract`, where it names the archive; for `get` and `delete` it is the alternative to naming the address |
| `if_exists` | string | No | `fail` | What a `create` or an `extract` does when an address already holds a file: `fail` refuses it, `replace` records the next version of what is there |
| `members` | string | No | every file | For `extract`: a glob. A pattern with no slash (`*.csv`) matches a member's file name wherever it is filed; one with a slash matches the member's whole cleaned path |
| `force` | boolean | No | false | Delete a file even though something still points at it |
| `limit`, `offset` | integer | No | 100, 0 | Page a `list`; 100 is also the largest page |
| `content` | string | Conditional | - | The file as text |
| `content_base64` | string | Conditional | - | The file as base64-encoded bytes, for a binary file. Exactly one of the two |
| `content_type` | string | Conditional | - | Media type the bytes are (required for `create`). Not detected: SVG, HTML, JSX and Markdown all read as plain text to a byte sniffer. `replace_content` keeps the type the resource already carries when it is omitted; a file stored under a generic type is re-detected from its bytes |
| `filename` | string | Conditional | - | Name of the file (required for `create`), normalized to lowercase; ignored by `replace_content` |
| `display_name`, `category`, `description` | string | Conditional | - | Required for `create` |
| `tags` | array | No | [] | Tags for filtering in the library |
| `scope`, `scope_id` | string | No | the caller's own user scope | `user`, `persona`, or `global` |
| `change_summary` | string | No | generated | Why the content changed, shown in the version history beside the revision |

Content arrives in one of two fields rather than one field with an encoding flag, so a caller cannot declare an encoding the bytes are not in. Both are capped at `portal.max_content_size` (10 MB by default), the same cap `save_asset` applies: one number governs everything an agent writes through a tool call, and a file larger than that goes through the portal's upload form.

`create` reports the resource's `mcp://` URI (which is what `save_asset`'s `references` argument takes) and its `mcp:resource:` reference (which is what every other tool takes). It reports no `version`: version 1 is recorded only where the deployment keeps a version trail, and a number the history may not hold is worse than none. `replace_content` reports the `version` the content was recorded as and, when a table is registered over the file, `tables`: one sentence per table saying it followed onto the new version or is pinned and now behind it, appended to the message as well (#1536); a scheduled script that replaces a followed resource needs no second call, and the sentences reach its run log.

A replacement goes through `resource.ReviseContent`, the same call the portal's replace-content route makes, so the resource keeps its id, its canonical `mcp://` URI and its filename — which is what keeps every `mcp:resource:<id>` citation, every prompt attachment and every asset reference pointing at it resolving across the write, with no asset re-saved — the new bytes land on a fresh per-revision key so the prior version stays independently readable and restorable, the revision is recorded with its author and change summary, and the resource is re-registered so `resources/list_changed` reaches clients that would otherwise keep serving the bytes it moved off.

Creating is scope authority, checked with the same `resource.CanWriteScope` the REST surface applies: the caller's own user scope (the default, and the one place every signed-in caller may write), a persona they administer, or the global scope as a platform administrator. The refusal names the SCOPE and not the file, because where it was filed is what the caller has to change. Replacing is `resource.CanModifyResource` — the uploader, or an administrator of the scope — and a resource the caller cannot see is answered as absent rather than as forbidden, so a caller who should not learn it exists cannot tell the two apart. An unauthenticated session is refused outright: a resource records who uploaded it and decides its scope on that.

A deployment with no managed-resource layer (no database, or no S3 connection for resource storage) refuses the call with that reason and says nothing was saved, rather than reporting a write that did not happen; a store with no version trail still creates and refuses only `replace_content`, naming the missing history. The seam is `internal/platform/resourcewrite`, built by the platform because the re-registration callback is the platform's, and bound onto the toolkit through `portalstore.BindResourceWriter`; the create path itself is `resource.CreateResource`, exported for the same reason `ReviseContent` was — creating a managed resource is not the browser's alone — and the REST upload route goes through it too, so both surfaces write one record shape and one version-1 row.

The lifecycle around those two writes is #1665. `get` resolves an address (`scope` + `path` + `filename`) or a reference to the record a `fetch` returns without the bytes, `list` reports a folder and everything beneath it, and `delete` removes a file and its version trail. Without them a caller had to remember the id of everything it wrote: `create` refuses an occupied address, `replace_content` needs a reference, and `search` is relevance-ranked and capped, so a file it does not return is not a file that is not there -- a script whose state was cleared could never reach its own rolling file again, and a file a draft run or a test made was permanent from the agent's side. `Writer.Locate` is the one address resolution the lander already performed, reached by both: it validates through the managed-resource layer's own `ValidateScope`/`ValidatePath`/`SanitizeFilename`, composes the canonical URI with `BuildURI`, follows the alias trail a move left behind, and answers "nothing is filed there" as a value rather than an error, which is what lets one lookup serve a `create` deciding whether to revise and a `get` reporting an empty address; a read that FAILED stays an error, because creating on it would file a second file at an occupied address. The reference branch answers the same way (#1690): `Writer.Get` folds absent and forbidden into `resource.ErrNoSuchResource` -- the sentinel lives in the domain package so the tool surface can read it without importing the writer -- and `handleGetResource` renders that one error as `found: false` beside the `reference` it looked up, the shape `fetch` already answers a dangling reference in, while any other error stays a tool error. Before this, a script holding a reference in its state read a deleted file as a failed call. `create` with `if_exists=replace` routes into the same `replaceResourceContent` an explicit `replace_content` reaches, so the two cannot disagree about what a revision does to the file's type, its history or the tables over it, and `toolkit.ResourceAddress` sits beside `ResourceDestination` in `pkg/toolkit` so an export's destination and a tool's lookup are one definition of where a file is.

A delete is `resource.DeleteResource`, extracted from the REST route's handler so both doors remove the blobs (head first, then every superseded version's, best-effort past the head) and then the record, and `Writer.Delete` adds the `Unregistered` callback that takes the file out of the MCP resource list. Authority is `CanModifyResource`, the same rule a replacement meets: deleting is not stronger than overwriting, since a replacement already leaves nothing of the previous content beyond the trail the delete takes. What still points at the file is gathered BEFORE the delete by `internal/platform/resourceholds`, two reverse lookups behind one call -- `assetrefs.Store.ListByTarget` for assets whose content references it and `prompt.AttachmentStore.ListByResource` for prompts that attach it -- plus `TableRegistrar.Tables` for query-engine tables. A knowledge page is deliberately NOT a third: `ParseCitableRef` refuses an `mcp:resource:` citation on a shared page, because a resource is visibility-scoped and the citation would be broken for every reader outside that scope, and every writer of `knowledge_page_entity_refs` (the manual picker, `ScanBodyRefs`, and promotion through `parsePageReferences`/`promotedRefsFromURNs`) goes through it, so a page cannot hold a resource row and a lookup for one would be a count that is always zero. Neither of the two IS a foreign key: migration 000088 states the rule for attachments, and deleting the file must leave the row so the prompt reports the material as missing rather than erasing the evidence that the procedure is now incomplete. Nothing in the database therefore stops the delete, which is exactly why the tool does, and why a failed lookup refuses rather than counting zero -- an empty count reads as "nothing depends on this file", the one wrong answer somebody about to delete must not act on. The counts are counts and not names because each of those records carries an audience of its own (an asset has shares, a prompt has a scope) and the caller is not necessarily in it; the portal's used-by panel resolves those audiences and names what its reader may open. A table IS named, since `manage_table` already names it to the same caller. `force=true` deletes anyway, records what it broke, and calls `DropResourceTables`, the resource-kind sibling of `DropAssetTables` -- the REST route reaches the same `UnregisterAllForSource` through the store's delete hook, and a tool call does not cross that route. `resourceholds` returns its own `Holds` type and the composition root adapts it in `portalstore.holdReader`, so the checker depends on no toolkit and the toolkit depends on no reverse lookup.

`extract` (#1879) writes the members of a stored archive out as managed resources: the only way a delivery like a monthly `.zip` holding one large CSV becomes a file a table can be registered over, since `manage_table register` reads CSV, JSON lines and Parquet and Starlark has no decompression and could not hold a multi-hundred-MB member anyway. `internal/unarchive` opens the archive (format by magic bytes: `PK` for zip; `1f 8b` for gzip, read as a tar when the decompressed head carries the `ustar` magic at offset 257) and answers which members are selected, whether each name is safe, and whether the archive is inside `resources.managed.extract`; `resourcewrite.Extractor` plans every selected member's address -- the member's directories become folders beneath `path` through `resource.FolderSegment`, the Go twin of the bulk uploader's `folderSegment` (`Q3 2026` files as `q3-2026`, and `2026` keeps its name since #1886), or with `filename` the one selected member lands at `<path>/<filename>` -- and checks each through `Lander.plan` (scope authority, the authority to replace what is there, `if_exists`) before the first write, then streams each member through `Lander.landPlanned`, so a member is created, or recorded as the next version with the tables over it followed, exactly as an export landing at that address would be. Refused before anything is written: a name that is absolute, carries a drive letter or has a `..` segment (anywhere in the archive, selected or not; a backslash is read as a separator), an encrypted member (flag bit 0, or WinZip AES method 99), a zip method other than stored or deflate, a selection matching nothing (the refusal lists what the archive holds), `filename` with several members selected, two members at one address, an occupied address without `if_exists=replace`, a folder name with no usable characters, a folder chain past 8 levels, a denied extension, and any limit. Links and devices are skipped and listed as `skipped`. A failure while a member streams (CRC mismatch, truncated stream, a gzip past a limit) abandons that member's multipart upload and returns the members written before it alongside the error, which the tool names with each reference and version. The result is `{archive, format, members: [{member, resource_id, reference, uri, path, filename, content_type, size_bytes, version, created, table_changes}], skipped, message}`. The toolkit holds a `ResourceExtractor` interface over its own request and result types; `portalstore.BindResourceWrites` assembles the writer, the export lander and the extractor over the managed-resource layer in one place (moved out of `pkg/platform`), builds the extractor only when the blob client implements `resourcewrite.RangeReader`, and binds it through an adapter that carries the fields across, so the toolkit depends on no writer. A tar's global PAX header, which `git archive` writes first, is metadata and is neither extracted nor listed as skipped. `extract` is classified as a write in `internal/toolwrite`, so a script draft reaches it only with `allow_writes`.

A managed script reaches the tool through `platform.call` like any other, under its version author's permissions, which is what makes a scheduled refresh of a referenced file possible with nobody in the loop.

A run authenticates as `script:<name>`, a principal that owns no file and is in nobody's library, so the resource rules read the address it acts for the same way the asset toolkit's ownership check does (#1419). `resource.Claims.OnBehalfOf`, built by `BuildClaimsFor` and empty for every human caller, is read by four rules: `VisibleScopes` adds the library of the person acted for, `CanWriteScope` admits that person's user scope, `CanAccessResource` admits a file they uploaded INTO THEIR OWN LIBRARY, and `CanModifyResource` matches the address the resource records as its uploader wherever the file now sits. Those two are deliberately different predicates (#1576): a move rewrites a resource's library, folder and `mcp://` address and never its uploader columns, so the person who uploaded a file keeps modify authority over the row after somebody files it elsewhere, and a run acting for them has to keep it with them -- holding the run to the library instead meant a person who moved a CSV into a persona they merely BELONG to (which `CanMoveToLibrary` deliberately permits) silently broke the schedule refreshing it, with nothing warning at the move and whether it broke at all turning on whether the author happened to be a platform administrator. The modify arm is ONE rule for the person and for what acts for them (`uploadedByCaller`) rather than a person's rule beside a stand-in written for the run: a stand-in diverges, and this one would have -- a resource a RUN filed records the principal as its subject and the author as its address, so the person's own subject does not match it, and once such a file leaves their library every OTHER script that person writes could have rewritten what the person cannot touch. That rule reads the recorded ADDRESS for an unattended caller and the recorded SUBJECT only for a caller acting as themselves, because `script:<name>` is unique only within an owner (`idx_scripts_name_owner`, migration 000119) and two people's same-named scripts would otherwise share one identity. Visibility is unchanged: `CanAccessResource` keeps the scope-bound arm, and the libraries a move can reach are readable on their own terms (a persona by its members, the global one by everyone) -- though "may change" is a wider set than "may see", so a surface authorizing on `CanModifyResource` alone (registering a resource as a Trino table) grants the run that wider set exactly as it grants the person. The decay the uploader arm is known for now falls on the person and their script identically rather than on one of them: an administrator who uploaded into somebody else's scope and later lost the role keeps modify authority over that row, and so does their script -- one grant, one holder. Without it a run could neither read nor replace a file its author uploaded through the portal — which files a person's own resources under their SUBJECT, an identifier a run does not have — and a create with no scope named would land under `{user, script:<name>}`: visible to the run and absent from the author's Resources page, which is where the person who scheduled it looks. The create default is the address the run acts for, not the principal. A scope id is compared EXACTLY, because VisibleScopes emits it verbatim and the store's listing predicate is built from that -- folding case in the write check alone would make a resource modifiable and deletable by somebody it never appears in a listing for; the uploader address IS folded, since it is read off the row rather than listed by and so has no predicate to disagree with. An author whose address changes in the identity provider loses their scripts' reach over resources uploaded under the old one (the uploader address is frozen on the row, the run carries the current one), and the run is answered as for a deleted file, because telling the two apart would disclose that a file the caller may not see exists. Both sides of every identity comparison must be non-empty, so an absent address never matches an absent owner or an absent scope id; the same tightening now applies to the plain subject arm of `CanWriteScope` and `CanModifyResource`, which previously let an empty subject match an empty scope id (unreachable through the handlers, since `ValidateScope` refuses an empty user scope id, but not a rule worth leaving written the other way beside its new sibling).

---

# Personas

Personas control which tools a user can access. Map OIDC roles to personas for role-based access control.

## Configuration

```yaml
personas:
  analyst:
    display_name: "Data Analyst"
    roles: ["analyst", "data_engineer"]
    tools:
      allow: ["*"]
      deny: ["*_delete_*"]
  admin:
    display_name: "Administrator"
    roles: ["admin"]
    tools:
      allow: ["*"]
```

A caller whose roles match no persona is unmapped and reaches nothing: MCP tool
calls resolve to the built-in deny-all persona and are refused, the portal
answers 403 with a branded page naming the refused account, and the
managed-resources API refuses the request. There is no fallback persona -
granting access means granting a role one of the personas lists. A deployment
with auth.allow_anonymous, or with no authenticators configured, gives its
callers the role "anonymous", so open access is something a persona opts into by
listing that role. personas.default_persona was removed and a config that still
sets it is refused at startup.

## Tool Filtering

Patterns support wildcards:
- `*` matches any sequence of characters
- `trino_*` matches all Trino tools
- `*_delete_*` matches any delete tool

Evaluation order: deny patterns are checked first, then allow patterns.

Prefer `allow: ["*"]` with a targeted `deny`. An enumerated allow-list has to name every tool the persona will ever hold, so it silently loses each tool a later upgrade adds; the persona keeps working with less of the platform than the operator thinks it has.

## Some tools are a unit

A few tools produce work that only another tool can consume, so granting one without the other is a configuration error rather than a tightening: `search` needs `fetch` (search returns navigational pointers carrying a `reference` and fetch is the only tool that dereferences one, so without it the persona discovers that an answer exists and can never read it, and cannot follow a knowledge page's outbound references); `memory_capture` needs `search` (search is the retrieval front door for captured memory, so without it the persona writes knowledge nobody, including itself, can retrieve); `apply_knowledge` needs `search` (the review workflow's documented first step is discovery).

The platform evaluates these pairs against the tools it actually registered — at startup for every persona, and again on each persona write through the admin API — and logs a warning naming the persona, the granted tool, the missing tool, why the pair matters, and the fix (`persona.CheckCoherence`, `pkg/persona/coherence.go`). It is advisory, never a gate: a restricted persona may be intended, and a deployment that never registered the missing tool is never warned about it. The warning exists because the loss is otherwise silent — the instruction baseline names a tool only when the caller can reach it, so a persona missing `fetch` is simply never told that reading a result in full is possible, and the only other symptom is an `unauthorized` audit row if an agent guesses the tool name unprompted.

Relatedly, the search-first gate refuses `trino_query` and `trino_execute` until `search` has been called in the session, so a persona granted query tools but not `search` cannot open that gate and is refused permanently.

## Connection Access Control

Personas can restrict which toolkit connections a user may access. A tool call must pass both the tool pattern check and the connection check, and the discovery surfaces show only the connections a persona is granted.

The connection is the platform's authorization boundary rather than the end user, and connection rules are the load-bearing half of a persona. Several connections may front the same downstream system under different credentials at different permission levels (a read-only Trino account and a write-capable one on the same cluster are two connections), so the permission level a caller gets is the permission level of the credential bound to the connection they were granted. Tighten access by adding a connection with a narrower downstream account, not by adding a role that lands on the same connection. The platform does not impersonate the caller downstream, so warehouse row policies and column masks that key off the end user do not follow a caller through, and per-user attribution comes from the audit trail (user_id, user_email, persona, connection on every call). Full rationale, trade-offs, and boundaries: the Authorization Model section above.

```yaml
personas:
  analyst:
    connections:
      allow: ["prod-*"]
      deny: ["prod-admin-*"]
```

Connections are deny-by-default: if the `connections` block is omitted or `allow` is empty, the persona reaches no connection at all. Connection patterns use the same wildcard syntax as tool patterns. A pattern matches the name a call binds the connection by, which `list_connections` reports as `connection`: the `instances:` key of the toolkit instance, or the `connection_name` a DataHub or S3 instance sets in its place. A `connection_name` on a Trino instance has no effect (the server warns about it at startup): Trino routes by the `instances:` key, so that key is its connection's name everywhere. Before #1396 an unqualified Trino call bound the `connection_name` instead, so a persona that lists one must be changed to list the `instances:` key.

Discovery applies the same rules through the same predicate, so what a caller can find and what a caller can call never disagree. `search` omits catalog datasets, connections, and API endpoints belonging to connections the persona is not granted; `fetch` returns `found: false` for such a reference, so a citation cannot read around what search omitted; `list_connections` enumerates only the granted connections; the portal search shares the router and behaves identically. A dataset is attributed to a connection through its DataHub platform name — a dataset that maps to no configured connection stays visible, and one reachable through any granted connection stays visible.

Nothing is dropped silently: `search` coverage carries a per-source `withheld` count plus a `withheld_notice` naming the persona and the remedy, and `list_connections` returns `withheld` and `notice`. This is a metadata boundary — it hides names, descriptions, and inventory; the data behind them was already gated at the tool call.

## API Endpoint Rules

A persona that reaches an `api` connection can call every operation it exposes. `api_routes` narrows that to specific methods and paths, which is how one API is split into read-only and read-write access without two connections and two credentials.

```yaml
personas:
  analyst:
    connections:
      allow: ["crm-*"]
    api_routes:
      - connection: "crm-*"
        methods: ["GET", "HEAD"]
      - connection: "crm-*"
        methods: ["DELETE"]
        paths: ["/v1/orders/{id}"]
        action: deny
```

`connection` is a glob and is required. `methods` and `paths` are glob lists where empty means any. `action` is `allow` (the default) or `deny`. Evaluation against one `(connection, method, path)`: entries whose connection glob does not match are skipped; among the rest a matching deny refuses the call; otherwise a matching allow is required; and if no entry names the connection at all the check is a no-op, so adding a rule for one connection does not close the others and a persona written before `api_routes` existed behaves as it did.

A path glob is matched against the path the call reaches (`/v1/orders/42`) and the catalog path the operation declares (`/v1/orders/{id}`). Naming the declared path is the precise way to govern one operation; a glob such as `/v1/orders/*` is a different rule that also covers the sibling operations at that depth. Wildcards do not cross a `/` and there is no recursive form — `**` is two stars and stops at a separator like one does.

Rules are authored in the config file or in the portal at Settings > Personas > Permissions > API endpoints, which lists each api-kind connection and the operations its catalog declares with the persona's decision on each one and the rule that produced it. Selecting an operation writes that operation's own method and declared path, so a portal-written rule and a file-written rule are the same rule; a hand-written glob is displayed as the glob it was typed as and is not rewritten on save. `POST /api/v1/admin/personas/{name}/test-access` answers a `(connection, method, path)` question with the decision and the matched rule.

`api_routes` does not make an API connection deny-by-default. A connection no rule names stays fully reachable by any persona granted it.

---

# Go Library

Import the platform as a library for custom MCP servers.

## When to Use

Use the library when you need to:
- Add tools - Your domain-specific operations
- Swap providers - Different semantic layer, different query engine
- Write middleware - Custom auth, logging, rate limiting
- Embed it - MCP inside a larger application

## Packages

| Package | What it does |
|---------|--------------|
| `pkg/platform` | Main entry point, orchestration |
| `pkg/toolkits/*` | Trino, DataHub, S3, Knowledge adapters |
| `pkg/semantic` | Semantic provider interface |
| `pkg/query` | Query provider interface |
| `pkg/middleware` | Request/response processing |
| `pkg/mcpcontext` | MCP session/progress context helpers |
| `pkg/persona` | Role-based tool filtering |
| `pkg/auth` | OIDC and API key validation |
| `pkg/admin` | Admin REST API for knowledge management |

## Minimal Example

```go
package main

import (
    "log"
    "os"
    "github.com/txn2/mcp-data-platform/pkg/platform"
    "github.com/txn2/mcp-data-platform/pkg/toolkits/datahub"
)

func main() {
    p, err := platform.New(
        platform.WithServerName("my-data-platform"),
        platform.WithDataHubToolkit("primary", datahub.Config{
            URL:   os.Getenv("DATAHUB_URL"),
            Token: os.Getenv("DATAHUB_TOKEN"),
        }),
    )
    if err != nil {
        log.Fatal(err)
    }
    defer p.Close()
    p.Run()
}
```

## API Stability

The module is `github.com/txn2/mcp-data-platform` at major version 1 with no `/vN` suffix. This is the promise the project keeps across minor releases.

Supported import surface (breaking changes only in a major release):

| Package | What you import it for |
|---------|------------------------|
| `pkg/platform` | The facade: construct, configure, run (options and lifecycle) |
| `pkg/toolkit` | Shared types every toolkit implements |
| `pkg/registry` | Register and manage toolkits |
| `pkg/semantic` | Semantic provider interface |
| `pkg/query` | Query execution provider interface |
| `pkg/middleware` | Request/response middleware contracts |
| `pkg/toolkits/*` | Toolkit adapters' exported config types |

Other exported packages under `pkg/` are importable but are implementation packages, not a committed integration surface; their API may change in a minor release with a release-note callout. The set is bounded by a build gate (`TestPublicSurfacePolicy`, `pkg_stability_policy_test.go`) that fails when a package is added under `pkg/` outside the supported table with a single first-party importer, since that shape is an implementation seam and belongs under `internal/`; the remaining exemptions are the reference store and provider implementations a consumer passes to `platform.WithSessionStore`/`WithQueryProvider`/`WithStorageProvider` and their siblings, plus `pkg/admin` (a mountable router) and `pkg/database/migrate` (the embedded schema migrations). Facade-internal seams live under `internal/platform/`, the HTTP adapters the server mounts under `internal/httpserver/`, and the portal's own seams under `internal/portal/` (its domain types and store contracts, PostgreSQL and no-database stores, authorization core, feedback surface, public-viewer templates, rate limiter and share cache — every moved name aliased back so `portal.Asset`, `portal.Collection`, `portal.User` and the store constructors are spelled as before); all are unimportable from outside the module by Go's internal rule and were never a supported surface. What a connection kind does to reach an HTTP upstream is one of these seams too: `internal/upstreamauth` holds the outbound authenticator and every auth mode (`none`, `bearer`, `api_key`, `basic`, `oauth`, `mtls`), the operator-owned static headers and the header names a model may not claim, the connect/call timeouts, the response read cap and the TLS material, shared by the HTTP-based kinds rather than copied into each; `internal/cfgmap` holds the typed readers over a stored connection's `map[string]any`. `apigateway.Authenticator`, `apigateway.NewAuthenticator`, `apigateway.ErrNeedsReauth` and the `AuthMode*`, `CredentialPlacement*` and `OAuth2AuthStyle*` constants are aliased back, so the toolkit's API, its configuration keys and its error messages are unchanged.

Configuration-file compatibility is handled more conservatively: additive keys ship in minor releases, and breaking renames or removals are called out in that version's release notes with an admonition (old key, new key, runtime effect). Precedent: `workflow.require_search` replaced the former `workflow.require_discovery_before_query` as a hard, non-aliased rename documented in the release notes. Set `config.strict: true` to turn an unaccounted-for rename into a hard startup error instead of a silent no-op.

---

# Knowledge Capture

Domain knowledge shared during AI sessions -- column meanings, data quality issues, business rules -- is captured, reviewed through a governance workflow, and written back to DataHub with changeset tracking and rollback.

## The Problem

Every organization has tribal knowledge about its data that lives in the heads of experienced team members. AI-assisted data exploration surfaces this knowledge in conversations, but without capture, it's lost when the session ends. Knowledge capture persists these insights for admin review and catalog write-back.

## How It Works

The system has four MCP tools and an Admin REST API:

- `search`: The universal, topology-free discovery entry point, fanning one query (text and/or entity_urns) across the technical catalog, its context documents, knowledge pages, the caller's memory, captured insights, the caller's feedback, saved assets, uploaded reference material (managed resources, indexed over their file contents), prompts, managed scripts, the caller's own recorded calls, the caller's own sessions, API endpoints, and connections, returning a balanced, grouped-by-source, per-user-scoped result set with a coverage summary.
- `fetch`: The companion read verb to `search`. Dereferences any `reference` a search hit carries (knowledge page, context document, catalog dataset, saved asset, prompt, or connection) back to its full content, honoring the same per-user scope so it never reads what search could not surface. A stale or out-of-scope reference returns a structured not-found, not an error. Fetching a knowledge page also returns its outbound `references` (the pages and entities it links to, each a citable reference plus its type), so an agent can deep-crawl a graph of focused, cross-linked pages deliberately (fetch the index, follow into the relevant branch) instead of re-parsing links from the markdown body (#705).
- `memory_capture`: The one way to record knowledge (memory toolkit). Routed by sink-class; reviewed classes create insights with status `pending`. (To read knowledge back, use `search`.)
- `apply_knowledge`: Admin-only sink router for reviewing, approving, synthesizing, applying, and rolling back insights. `apply` promotes by `sink`: `datahub` (schema_entity → catalog) or `knowledge_page` (business_knowledge/operational_rule → a canonical portal knowledge page, found-or-created by slug). Both record reversible changesets. Actions: `bulk_review`, `review`, `synthesize`, `apply`, `approve`, `reject`, `rollback`, `list_changesets`, `bulk_untag` (remove a tag from every entity carrying it).
- Admin REST API: HTTP endpoints for managing insights and changesets outside the MCP protocol.

## Reflexive Capture

Reflexive capture removes the operator from the loop for the highest-signal case: a query error the same session later fixes. When a `trino_query`/`trino_execute` call fails with a data-model misunderstanding (unknown column/table, ambiguous reference, type mismatch, `GROUP BY` mistake) and a later related query on the same connection succeeds in the same session, the platform auto-mints one "misconception + fix" correction memory with no tool call by the agent and without blocking the tool response (the mint is asynchronous). The record carries `source: automation` and `category: correction`, and is a reviewed sink-class (`schema_entity`) so it enters review as a `pending` insight rather than mutating catalog state. Pairing is conservative: the error must match an allowlist of misconception signatures (infra, policy, timeout, and permission errors are ignored), and the success must run on the same connection and be a near-variant of the failing query (scored on shared identifiers, and requiring a novel corrected identifier) so an unrelated query on the same table, or an identical retry, is not treated as a fix. The correction is best-effort entity-keyed to the successful query's DataHub dataset URNs, and the whole path is gated by the caller persona's `memory_capture` grant. Enabled by default when the memory subsystem is available; disable with `knowledge.reflexive_capture.enabled: false` (#635).

## Configuration

```yaml
knowledge:
  enabled: true
  apply:
    enabled: true
    datahub_connection: primary
    require_confirmation: true
  reflexive_capture:
    enabled: true
```

| Field | Type | Default | Description |
|-------|------|---------|-------------|
| `knowledge.enabled` | bool | `true` (when a database is available) | Enable the knowledge review and write-back toolkit (`apply_knowledge`). Knowledge capture lives in the memory toolkit (`memory_capture`) and is enabled with the memory layer, not this flag |
| `knowledge.apply.enabled` | bool | `true` (when a database is available) | Enable the `apply_knowledge` tool for admin review and catalog write-back. Set `false` to disable. |
| `knowledge.apply.datahub_connection` | string | - | DataHub instance name for write-back operations |
| `knowledge.apply.require_confirmation` | bool | `false` | When true, the `apply` action requires `confirm: true` in the request |
| `knowledge.pages.dedup_threshold` | float | `0.85` | Cosine similarity in `[0,1]` at or above which creating a knowledge page is blocked as a near-duplicate. Acts only when a real embedding provider is configured. A non-positive value selects the default |
| `knowledge.pages.dedup_disabled` | bool | `false` | Turn the duplicate gate off entirely, so "no gate" is never confused with "left at default" |
| `knowledge.pages.oversize_bytes` | int | `16384` | Body size in bytes at or above which a page write returns a non-blocking split suggestion. Negative disables this arm. An editorial nudge, not a bound on search reach: page content is embedded as chunks sized to the provider input budget, so a page of any size is semantically searchable end to end (#1242) |
| `knowledge.pages.oversize_sections` | int | `12` | Markdown heading count at or above which the split suggestion fires. Negative disables this arm |
| `knowledge.reflexive_capture.enabled` | bool | `true` | Auto-capture a "misconception + fix" correction (source `automation`, reviewed sink-class) when a Trino query errors and a later same-session query over the same table(s) succeeds (#635). Default-on when the memory subsystem is available; set `false` to disable |
| `knowledge.verifiable_insights` | bool | `true` | Deliver a `verifiable` block ({urn, query_table, connection}) on an insight whose linked catalog entity resolves through the query provider to an available table, on every delivery surface: `search` insight hits, the record `fetch` returns for `mcp:insight:<id>`, and the insight entries of the `memory_context` enrichment block (#1220). Additive and absent whenever nothing resolves, so a deployment with no query provider is unchanged. Set `false` to deliver insights with no marker |
| `knowledge.search_provider_timeout` | duration | `5s` | Per-provider deadline for the `search` fan-out arms, so one slow knowledge source drops out as a collected error rather than stalling the whole search. Set a negative duration to disable the bound. |
| `knowledge.search_embed_timeout` | duration | `5s` | Deadline for the serial intent-embedding step in `search`, independent of `search_provider_timeout`. A slow or unreachable embedder degrades to lexical ranking rather than stalling the search; this knob lets you give a slow (cold or CPU-only) embedder headroom to preserve `hybrid` ranking without loosening the fan-out bound. Set a negative duration to disable the bound. |

Prerequisites: `database.dsn` must be configured. The `apply_knowledge` tool requires the admin persona.

## How much of a knowledge page is searchable

All of it, at any size. A page's content (title, body, tags) is embedded as a SET of chunks rather than one vector, each chunk sized to the embedding provider's per-text input budget (`memory.embedding.ollama.max_input_bytes`, default 6,000 bytes) and split on the page's own markdown section boundaries where it has them, falling back to paragraph and then rune-boundary cuts. Every chunk is composed through the same title+body+tags composition a query is embedded over, so each carries the page's identity. Both vector readers -- the hybrid search arm and the create-time near-duplicate gate -- rank chunks and score each page by its best-matching chunk, so results stay page-granular while a fact at the end of a long runbook ranks its page exactly as one in the opening paragraph would, and a large candidate page is compared for duplication in full rather than by its head. Before #1242 a page carried one vector over text trimmed at the provider cap, so anything past roughly 6 KB reached the lexical arm only. The byte cap is a proxy for the model's token budget and cannot be an exact one (token density varies by close to an order of magnitude between prose and dense content such as CSV, JSON, or source), so when the model refuses an input anyway the provider halves what it sends and retries down to a 256-byte floor; that is what keeps a dense document from failing identically on every attempt forever (#1350). The refusal is recognized by the wording of the response body whatever status carries it, because the same Ollama answers the condition with a 400 from its batch endpoint and a 500 from its single-input endpoint, body identical; a status-gated match classified the batch refusal and missed the per-input one that follows it, so the bound was never reduced (#1385). The converged vector covers a prefix of that chunk rather than all of it (the lexical arm still matches the whole text), and the shrink is logged with the model and the refused size, because a model that cannot hold the configured cap for ordinary content is a misconfiguration to correct at `max_input_bytes` rather than something to absorb silently. The bound is deliberately not remembered between calls: density is a property of the text, not the model, so one dense document learning a small bound would truncate every prose document with it. Vectors live in `portal_knowledge_page_embedding_chunks`, replaced atomically by the indexjobs consumer; a content edit deletes the page's chunks in the same transaction as the write, so no vector outlives the text it was computed from, and the same write enqueues the page's own index job (trigger `write`, #1256) so the new content enters ranked search within one embed rather than on the next reconciler sweep.

## Insight Categories

| Category | Description | Example |
|----------|-------------|---------|
| `correction` | Fixes wrong metadata | "The amount column is gross margin, not revenue" |
| `business_context` | Business meaning not in metadata | "MRR counts active subscriptions only, not trials" |
| `data_quality` | Quality issues or limitations | "Timestamps before March 2024 are UTC; after that, America/Chicago" |
| `usage_guidance` | Tips for querying correctly | "Always filter status='active' to avoid soft-delete duplicates" |
| `relationship` | Dataset connections not in lineage | "customer_id in orders joins to the legacy CRM export" |
| `enhancement` | Suggested metadata improvements | "Tag sales_daily with its 6 AM CT refresh schedule" |

## memory_capture Tool

Records domain knowledge shared during a session (memory toolkit).

**Parameters:**

| Parameter | Type | Required | Description |
|-----------|------|----------|-------------|
| `type` | string | Yes | Sink-class. Live: `personal_preference`, `episodic_event`. Reviewed (creates a pending insight): `business_knowledge`, `schema_entity`, `operational_rule` |
| `content` | string | Yes | The knowledge to record (10-4000 characters) |
| `confidence` | string | No | `high`, `medium` (default), or `low` |
| `source` | string | No | `user` (default), `agent_discovery`, or `enrichment_gap` |
| `entity_urns` | array | No | DataHub URNs this knowledge relates to (max 10) |
| `related_columns` | array | No | Columns related to this knowledge (max 20) |
| `suggested_actions` | array | No | Proposed catalog changes (max 5) |
| `thread_ids` | array | No | Feedback thread IDs this insight resolves; each is linked (insight_id set, `insight_linked` event, status `resolved`) |

**Suggested action types:** `update_description`, `add_tag`, `remove_tag`, `add_glossary_term`, `flag_quality_issue`, `add_documentation`, `add_curated_query`, `set_structured_property`, `remove_structured_property`, `raise_incident`, `resolve_incident`, `add_context_document`, `update_context_document`, `remove_context_document`, `delete_tag`, `set_custom_property`, `remove_custom_property`

**Target field:** Use `column:<fieldPath>` (e.g., `column:location_type_id`) for column-level descriptions. For `add_documentation`, the target is the URL. For other types, leave empty or omit.

**Detail field:** Description text, tag name or URN (e.g., `pii` or `urn:li:tag:pii`), glossary term name or URN, quality issue description, link description, or query name (for `add_curated_query`). Short tag/term names are auto-normalized to full URNs.

`flag_quality_issue` adds a fixed `QualityIssue` tag and, on entity types DataHub supports incidents for (datasets, dashboards, charts, data flows/jobs, containers, data products) with a detail, raises a "Data quality issue" incident whose description is the detail, so the issue is verifiable on the entity (the detail is not encoded in the tag name). Re-applying the same detail reuses the existing open incident rather than creating a duplicate. The created incident URN is recorded in the changeset's `new_value.created_urns`, and rolling the changeset back removes the tag and resolves any incident it created. On entity types that do not support incidents, only the tag is set.

`add_curated_query` creates a reusable Query entity in DataHub linked to the target dataset. Requires `query_sql` (the SQL statement) and optionally `query_description`. The `detail` field serves as the query name.

**Response:**

```json
{
  "insight_id": "a1b2c3d4e5f67890a1b2c3d4e5f67890",
  "status": "pending",
  "message": "Insight captured. It will be reviewed by a data catalog administrator."
}
```

When `thread_ids` is supplied, the response also includes `linked_thread_count` and `unlinked_thread_ids` (the IDs that matched no open thread, or that the caller is not authorized to resolve); both are omitted when `thread_ids` is absent. Linking is gated by the same owns-or-edit check as `manage_feedback resolve` (it is not a back door around the access model), is best-effort, and never fails the capture.

The tool also registers an MCP prompt `knowledge_capture_guidance` that provides AI agents with guidance on when to capture insights.

## Feedback Bridge

Human feedback threads on portal assets connect to the knowledge loop: an agent resolves a thread by capturing the insight it represents, and the chain stays visible end to end (**thread → insight → changeset → `target_urn`**).

- **`memory_capture` `thread_ids`** links threads to the captured insight (see above).
- **`manage_feedback` tool** (the dedicated, discoverable home for agent feedback): `list`, `get`, `reply`, `resolve`, `request_validation`, `respond_validation`. **`list` with no target is the cross-target entry point**: it returns the caller's pending feedback across the assets and collections they own or can edit AND the shared general channel (unresolved threads they did not author, plus any awaiting their validation), so an agent can "review and act on any pending feedback" in one call; `list` with an `asset_id`/`collection_id`/`prompt_id` or `target_type=standalone` scopes to one target (filter by status, `requires_resolution`, `validation_state`). Scoped to the assets and collections the caller owns or can edit; admins see all; standalone/general threads are readable by any authenticated caller and resolved by the thread author or an admin. (These six actions previously lived on `manage_asset`; they were moved to `manage_feedback` so agents discover feedback by tool name.)
- **@-mentions (#627)**: a thread comment body may address people with `@local(domain)` tokens (for example `@marcus.johnson(example.com)`). The parenthesized domain terminates the token, so a mention written before sentence punctuation is not swallowed by it, and the stored form is an address rather than a display name. On write (portal REST and `manage_feedback reply` alike) the named addresses are filtered to the target's **audience** -- the people who can already open it: an asset's owner plus recipients of active direct and collection shares, a collection's owner plus share recipients, everyone for a persona- or global-scoped prompt and owner-plus-shares for a personal one, and the known-users directory for knowledge pages and the standalone channel. Eligible mentions are recorded on the event as `metadata.mentions` (a lower-cased address array, GIN-indexed) and notified in the `mention` category, which has its own per-user preference so muting comment activity still leaves a person reachable by name. A name outside the audience -- and the author's own address -- is kept as plain text: nothing is recorded, the timeline renders it unchipped (the client chips only what `metadata.mentions` records), and nothing is delivered, so a comment can never mail an item's title and an excerpt to someone who cannot open it. A share that names its recipient only by user id grants access but records no address anywhere, so such a person is outside the audience by construction. `GET /api/v1/portal/mention-candidates?target_type=&target_id=&q=` lists the audience for the composer's type-ahead (the caller must belong to it, or be an admin); `GET /api/v1/portal/worklist/mentions` is the caller's mentions inbox, re-checked against present-day access so a revoked share stops surfacing the item's threads.
- **`GET /api/v1/portal/threads/{id}/chain`** returns a thread's resolved chain: its `insight_id` and the changesets that insight produced (each with `target_urn`, `change_type`, rollback state). The portal feedback panel renders this as a "Knowledge chain" section.

### Validation & sign-off (Phase 3)

- **Validation response**: after `request_validation` routes a request to the feedback author, that author confirms or disputes it via `manage_feedback action=respond_validation` (`validation_result` = `validated`|`disputed`, optional `validation_reason`) or `POST /api/v1/portal/threads/{id}/validation`. Both append a `validation_result` event and set `validation_state`; **disputing re-opens the thread**. Only the author (or an admin) may respond.
- **Worklists**: `GET /api/v1/portal/worklist/practitioner` (open resolution-required threads across assets and collections the caller owns/can edit), `GET /api/v1/portal/worklist/sme` (threads awaiting the caller's validation), and `GET /api/v1/portal/worklist/mentions` (threads where a comment @-mentioned the caller).
- **Activity feed**: `GET /api/v1/portal/feedback/activity` returns every feedback thread across the assets, collections, and prompts the caller can view (owned or shared, any permission), most recent activity first, each row enriched with the target's display label (`target_label`) so the client links straight back to the item. It is scoped server-side to the caller's reachable set and never discloses feedback on items they cannot see; global-prompt threads the caller did not author are excluded so the feed stays personal. With no push-notification system, this feed (plus a sidebar badge counting the caller's open worklist) is how a user discovers new feedback.
- **Sign-off**: `GET /api/v1/portal/assets/{id}/signoff` and `.../collections/{id}/signoff` return `signed_off` of `stakeholders` (N = distinct approvers; M = owner + active share grantees), rendered as "signed off by N of M".

## apply_knowledge Tool

Admin-only tool for reviewing and applying captured insights. Requires the `admin` persona.

### Actions

**bulk_review**: Counts of all pending insights (total_pending, by_entity, by_category, by_confidence) plus a review-queue staleness rollup (oldest_pending_at, oldest_pending_age_days, pending_over_30d, omitted when the queue is empty) so aging review debt is visible; the same rollup appears in platform_info under features.knowledge_apply.review_queue. An hourly operator alert emails a digest of the same rollup when the queue crosses an admin-configured threshold (see Review queue alerts). Pass itemize:true to enumerate the queue itself, paginated by offset/limit, each insight carrying its full insight_text body, captured_by (author), sink_class and suggested_actions_count (full suggested_actions omitted, fetch for it); the relevance-ranked search tool cannot list it completely. The response is bounded so it stays under the output limit: page_size_capped:true flags a short insights page (continue with next_offset) and by_entity_truncated:true flags a capped by_entity.

```json
{"action": "bulk_review"}
```

**review**: Insights for a specific entity with current DataHub metadata.

```json
{"action": "review", "entity_urn": "urn:li:dataset:(urn:li:dataPlatform:trino,hive.sales.orders,PROD)"}
```

**approve/reject**: Transition insight status with optional review notes.

```json
{
  "action": "approve",
  "insight_ids": ["a1b2c3d4e5f67890", "f6e5d4c3b2a10987"],
  "review_notes": "Verified with data engineering team"
}
```

**synthesize**: Structured change proposals from approved insights.

```json
{"action": "synthesize", "entity_urn": "urn:li:dataset:(urn:li:dataPlatform:trino,hive.sales.orders,PROD)"}
```

**apply**: Write changes to DataHub with changeset tracking.

```json
{
  "action": "apply",
  "entity_urn": "urn:li:dataset:(urn:li:dataPlatform:trino,hive.sales.orders,PROD)",
  "changes": [
    {"change_type": "update_description", "target": "", "detail": "Order records with gross margin amounts (before returns)"},
    {"change_type": "add_tag", "target": "", "detail": "gross-margin"},
    {"change_type": "remove_tag", "target": "", "detail": "urn:li:tag:outdated"}
  ],
  "insight_ids": ["a1b2c3d4e5f67890"],
  "confirm": true
}
```

## Insight Lifecycle

```
pending -> approved -> applied
pending -> rejected
pending -> superseded
applied -> pending (changeset rolled back)
```

| Status | Description |
|--------|-------------|
| `pending` | Newly captured, awaiting review |
| `approved` | Reviewed and approved for application |
| `rejected` | Reviewed and rejected |
| `applied` | Changes written to the canonical sink (DataHub catalog or a knowledge page); also the point the insight becomes organization-wide |
| `superseded` | Replaced by newer insights for the same entity |

A changeset rollback returns its source insights to `pending` rather than to a terminal state, so the review queue surfaces them again; they keep `applied_by`, `applied_at` and `changeset_ref`, and their `review_notes` names the rolled-back changeset. `reject` remains the way to discard an insight.

## Insight Visibility

Applying an insight is what turns one person's capture into knowledge the organization holds, so `applied` is also the visibility boundary.

| Status | Who can find it with `search` and read it with `fetch` |
|--------|--------------------------------------------------------|
| `pending`, `approved` | Only the capturer |
| `applied` | Every identified caller, attributed through the hit's `captured_by` |
| `rejected`, `superseded` | Only the capturer, and only under an explicit status query |

No insight is public: reaching applied insights requires an identified caller, so an anonymous visitor to a shared portal link never sees them. Discovery also does not depend on the sink, so a fact applied to the DataHub catalog is findable by search even when no tool result names the dataset it hangs off.

## Checkable Delivered Insights

An insight is a claim, and a delivered claim removes the consumer's reason to check it. When the platform can query the subject of a claim for itself it says so (#1220): an insight whose linked catalog entity resolves through the query provider to an available table is delivered carrying `verifiable: {urn, query_table, connection}` -- the table one query would settle the claim against, and the connection it lives on. Default on (`knowledge.verifiable_insights`, nil = enabled).

The marker rides the payload on every surface that delivers an insight, because the platform cannot detect what model is consuming it: `search` insight hits, the record `fetch` returns for `mcp:insight:<id>`, and the insight entries of the `memory_context` enrichment block pushed onto tool results. It asserts that the claim's subject is one query away, NOT that the claim was checked or that it is true; an insight linked to several entities names the first one that resolves, since a claim needing more than one table to settle is not the checkable shape the marker is for.

It is additive: an insight linked to no entity, an entity the query provider cannot see, a noop query provider, and the operator opt-out all deliver exactly the payload they always did -- absent, never an empty block. A plain memory record never carries it, since a note is not a claim about the warehouse.

The block is topology, so it honors the persona connection boundary (#1108): an insight about a dataset on a connection the caller's persona may not reach is still delivered -- an insight is not connection-scoped -- but carries no `verifiable` block, and that connection is never probed on the caller's behalf. The hit itself is not counted as withheld, since nothing about the record was hidden.

Cost: each delivery surface resolves through a bounded TTL cache (`internal/tableavail`, the same implementation behind the review path's observed-entity block), so a page of insights about the same entity costs one lookup rather than one per hit, and concurrent sessions asking about the same cold entity share a single in-flight lookup. Delivery resolves through `query.LocationResolver` (where an entity is queryable) rather than `GetTableAvailability` (which also runs `COUNT(*)` when `enrichment.estimate_row_counts` is on), so the marker never costs a full scan. The pass is bounded at 2s, so a slow warehouse costs a delivery its marker rather than the record or the search arm's budget. A negative answer is remembered for 30s rather than the positive 5 minutes, because a query adapter reports a transiently failed describe and a genuinely absent table in exactly the same shape; an answer computed after the budget expired is not remembered at all.

The `memory_context` block's insight entries carry `mcp:insight:<id>` as their reference rather than `mcp:memory:<id>`, which `fetch` always declined for a knowledge-dimension record. Because that push path is persona-scoped rather than caller-scoped, only an APPLIED insight is given a reference: fetch serves an insight to its capturer or, once applied, to everyone, and the push path cannot know which case it is in, so anything else is delivered with no reference rather than one that answers not-found.

## Governance Workflow

1. User shares domain knowledge during an AI session
2. AI calls `memory_capture` to record it (status: pending)
3. Admin uses `bulk_review` to see pending insights
4. Admin reviews and approves/rejects via `approve`/`reject`
5. Admin calls `synthesize` to generate change proposals from approved insights
6. Admin calls `apply` to write changes to DataHub with changeset tracking
7. Source insights are marked as applied with a changeset reference
8. Changeset stores previous values for rollback if needed

This is a human-in-the-loop metadata curation workflow. Every change to the catalog goes through admin review and is tracked for auditability.

## Changeset Tracking

Every `apply` action creates a changeset record:

| Field | Description |
|-------|-------------|
| `id` | Unique changeset identifier |
| `target_urn` | The DataHub entity that was modified |
| `change_type` | Summary of change types applied |
| `previous_value` | Entity metadata before changes (for rollback) |
| `new_value` | Changes that were applied |
| `source_insight_ids` | Insights that produced this changeset |
| `applied_by` | User who applied the changes |
| `rolled_back` | Whether this changeset has been reverted |

The `apply` response also returns `resulting_state`, a fresh read-back of the entity's `description`, `tags`, `glossary_terms`, and `owners` (the same shape stored as the changeset's `previous_value` before-image). Tags and glossary terms are read from the authoritative aspects: REST-exposed types (datasets, dashboards, charts, data flows, data jobs, containers, data products) from the REST aspects, and three GraphQL-only types (`domain`, `glossaryTerm`, `glossaryNode`) through the entity query (mcp-datahub v1.10.2+). Some fields cannot be read for some types and come back empty even when set: `tags`/`glossary_terms` are empty for `document` entities, and `description`/`owners` are empty for types with no dedicated read and no dataset/dashboard entity-query fragment (`domain`, `glossaryNode`, `container`, `chart`, `dataFlow`, `dataJob`). For those types treat an empty value as "not read" rather than "not set"; because the before-image shares these gaps, rollback cannot restore a field it could not read.

## Persona Integration

Control access through persona tool filtering:

```yaml
personas:
  analyst:
    tools:
      allow: ["search", "trino_*", "datahub_*", "memory_capture"]
      deny: ["apply_knowledge"]
  admin:
    tools:
      allow: ["*"]
  etl_service:
    # "search" is required so this query-capable persona can satisfy the
    # search-first gate (on by default); otherwise trino_* calls are refused.
    tools:
      allow: ["search", "trino_*"]
      deny: ["memory_capture", "apply_knowledge"]
```

## Admin REST API

HTTP endpoints for managing insights and changesets. All endpoints require admin authentication via API key (X-API-Key or Authorization: Bearer header) with the admin role.

### Insight Endpoints

| Method | Path | Description |
|--------|------|-------------|
| GET | `/api/v1/admin/knowledge/insights` | List insights with filtering and pagination |
| GET | `/api/v1/admin/knowledge/insights/{id}` | Get a single insight by ID |
| PUT | `/api/v1/admin/knowledge/insights/{id}` | Update insight text, category, or confidence |
| PUT | `/api/v1/admin/knowledge/insights/{id}/status` | Approve or reject an insight |
| GET | `/api/v1/admin/knowledge/insights/stats` | Get insight statistics |

**Query parameters for listing:**

| Parameter | Type | Description |
|-----------|------|-------------|
| `status` | string | Filter by status (pending, approved, rejected, applied, superseded) |
| `category` | string | Filter by category |
| `entity_urn` | string | Filter by entity URN |
| `captured_by` | string | Filter by user who captured |
| `confidence` | string | Filter by confidence level |
| `source` | string | Filter by source (user, agent_discovery, enrichment_gap) |
| `since` | RFC3339 | Filter by creation time (after) |
| `until` | RFC3339 | Filter by creation time (before) |
| `page` | integer | Page number (1-based) |
| `per_page` | integer | Results per page (default 20, max 100) |

**Observed warehouse state:** a pending insight whose entity URN resolves through the configured query provider to an available table carries `observed_entities` on both the list and detail payloads: one entry per resolved URN with `urn`, `query_table`, `connection`, `estimated_rows` (only when `enrichment.estimate_row_counts` is on), and an advisory `conflict` (`claimed_rows`, `observed_rows`, `message`) when the claim text states an integer the table disagrees with. When a claim states several numbers, the one nearest the estimate is compared. The marker never blocks promotion — approve and reject are unchanged. The field is omitted entirely, leaving the payload byte-identical to before, for a decided insight, an unresolvable URN, an unavailable table, a lookup that outran the read path's budget, and a deployment with no (or a noop) query provider.

### Changeset Endpoints

| Method | Path | Description |
|--------|------|-------------|
| GET | `/api/v1/admin/knowledge/changesets` | List changesets with filtering and pagination |
| GET | `/api/v1/admin/knowledge/changesets/{id}` | Get a single changeset by ID |
| POST | `/api/v1/admin/knowledge/changesets/{id}/rollback` | Rollback a changeset (restores previous metadata) |

**Query parameters for listing changesets:**

| Parameter | Type | Description |
|-----------|------|-------------|
| `entity_urn` | string | Filter by target entity URN |
| `applied_by` | string | Filter by user who applied |
| `rolled_back` | boolean | Filter by rollback status |
| `since` | RFC3339 | Filter by creation time (after) |
| `until` | RFC3339 | Filter by creation time (before) |
| `page` | integer | Page number (1-based) |
| `per_page` | integer | Results per page (default 20, max 100) |

### Memory Record Endpoints

| Method | Path | Description |
|--------|------|-------------|
| GET | `/api/v1/admin/memory/records` | List every user's memory records: `created_by`, `dimension`, `sink_class`, `category`, `status`, `source` filters, `limit`/`offset` paging (the portal route lists only the caller's own) |

## Database Schema

Knowledge capture uses two PostgreSQL tables (migrations 000006, 000007, 000008).

**knowledge_insights:**

| Column | Type | Description |
|--------|------|-------------|
| `id` | TEXT | Primary key (cryptographic random hex) |
| `created_at` | TIMESTAMPTZ | When the insight was captured |
| `session_id` | TEXT | MCP session that produced it |
| `captured_by` | TEXT | User who shared the knowledge |
| `persona` | TEXT | Active persona at capture time |
| `source` | TEXT | Where the knowledge came from: `user`, `agent_discovery`, `enrichment_gap` |
| `category` | TEXT | Insight category |
| `insight_text` | TEXT | The domain knowledge |
| `confidence` | TEXT | Confidence level |
| `entity_urns` | JSONB | Related DataHub entities |
| `related_columns` | JSONB | Related columns |
| `suggested_actions` | JSONB | Proposed catalog changes |
| `status` | TEXT | Current lifecycle status |
| `reviewed_by` | TEXT | Who reviewed the insight |
| `reviewed_at` | TIMESTAMPTZ | When it was reviewed |
| `review_notes` | TEXT | Reviewer comments |
| `applied_by` | TEXT | Who applied the insight |
| `applied_at` | TIMESTAMPTZ | When it was applied |
| `changeset_ref` | TEXT | Link to the changeset |

**knowledge_changesets:**

| Column | Type | Description |
|--------|------|-------------|
| `id` | TEXT | Primary key (cryptographic random hex) |
| `created_at` | TIMESTAMPTZ | When changes were applied |
| `target_urn` | TEXT | DataHub entity that was modified |
| `change_type` | TEXT | Type of changes applied |
| `previous_value` | JSONB | Metadata before changes |
| `new_value` | JSONB | Changes applied |
| `source_insight_ids` | JSONB | Insights that produced this |
| `approved_by` | TEXT | Who approved the changes |
| `applied_by` | TEXT | Who applied the changes |
| `rolled_back` | BOOLEAN | Whether changes were reverted |
| `rolled_back_by` | TEXT | Who reverted the changes |
| `rolled_back_at` | TIMESTAMPTZ | When changes were reverted |

---

# Portal: Your Work

The user portal is the day-to-day interface for analysts, engineers, and data consumers. On the site it is the `portal` rail section's Your Work group, one page per surface, in the order the portal's sidebar lists them: Assets (https://mcp-data-platform.txn2.com/portal/assets/), Collections (https://mcp-data-platform.txn2.com/portal/collections/), Shared With Me (https://mcp-data-platform.txn2.com/portal/shared/), Resources (https://mcp-data-platform.txn2.com/portal/resources/), Activity (https://mcp-data-platform.txn2.com/portal/activity/), APIs (https://mcp-data-platform.txn2.com/portal/apis/), Automations (https://mcp-data-platform.txn2.com/portal/scripts/), Inbox (https://mcp-data-platform.txn2.com/portal/feedback/), Knowledge and Memory (https://mcp-data-platform.txn2.com/portal/knowledge/), Prompts (https://mcp-data-platform.txn2.com/portal/prompts/), Scratch Tables (https://mcp-data-platform.txn2.com/portal/scratch-tables/) and Settings (https://mcp-data-platform.txn2.com/portal/settings/). The sidebar opens with Assets and Resources, set apart from the other sections, which follow in alphabetical order; Collections and Shared With Me are reached from Assets. The tour's entry page is https://mcp-data-platform.txn2.com/portal/, which also covers enabling and branding the portal.

Every page is addressable. Path recognition is one table (`ui/src/lib/portalRoutes.ts`) rather than a final `else` in the shell's switch, because the switch is spread across five section components that each own their own matching and none can see whether another matched: a retired or guessed name (`/assets`, `/shared`, `/knowledge-pages`) redirects to the surface it meant, a trailing slash is dropped when what is left is a real route, and a path with no page renders a not-found page naming the address and offering the way back. The old behaviour was the failure mode the platform refuses everywhere else — the shell rendered its chrome, a section title, and an empty content area, having requested nothing, which is indistinguishable from being told you own none of what you asked for (#1359). A drift gate reads the route literals and detail patterns back out of the shell and the section components and asserts the table knows each one, so a route added to the shell cannot silently become a refusal.

## Section intros
Every reader-facing section opens with a description of itself: a compact row carrying the section's own nav icon and one sentence, and a disclosure holding the fuller paragraph (`ui/src/components/patterns/SectionIntro.tsx`, copy in `sectionIntros.tsx`). Both fields describe the section; neither instructs the reader, who is finding out what they are looking at rather than being told where to file something. It is composed once in the shell (`AppShell`) over the section a route belongs to, so it heads a section's listing surfaces and no detail page beneath them, and a test fails when a nav entry has no intro. A reader who has never opened a section sees it expanded; closing it leaves the one-line summary and is remembered per section in `localStorage`, wrapped in try/catch so a private window still renders. Knowledge draws its compact row rather than writing it -- the Memory to Insight to Knowledge pipeline is its summary -- and keeps the storage key its former standalone header used, so a reader who had already closed it does not get it back. Settings has no intro: it is controls, not a place things live. This is what a person is told; `instructions.ResourcePositioning` is what the agent is told, still rendered verbatim in the resources empty state and the upload dialog. The two agree because both are drawn from the content model, not because one renders the other (#1570).

## Activity
Three tabs at `/activity`, `/activity/sessions`, and `/activity/calls`.

Overview: personal tool usage analytics with configurable time ranges (1h, 6h, 24h, 7d). Summary cards (total calls, avg duration, tools used), activity timeseries chart, and top tools bar chart.

My Sessions: the caller's own sessions, listed and openable, over `GET /api/v1/portal/sessions` and `GET /api/v1/portal/sessions/{id}`. Same audit-derived read model the admin Sessions surface reads (`internal/platform/sessionview`), scoped to the caller: the scope is a predicate inside the rollup over `audit_logs`, so another user's session id groups to no rows and is answered 404 — indistinguishable from an id that was never used. There is no `user_id` parameter (one sent is overwritten with the authenticated caller) and no user column or facet. Facets: time window (24h/7d/30d/all, default 7d, a visible control because the rollup reads every event in range), session kind (agent/portal/script/transport), has-failures, has-assets. The detail carries the summary, the assets and insights the session produced, and a page of its call timeline with the purpose stated for each call; the timeline is scoped to the caller too, so it cannot carry calls they did not make. Rows are plain rather than clickable: the admin event drawer is the only surface that opens a single audit event. Addressable at `/activity/sessions/{id}`. The same read model is what an agent recalls through `search` (the `sessions` source) and `fetch mcp:session:<id>` (#1322), under the same predicate, so what an agent recalls and what its owner opens are one session.

My Calls: the caller's own recorded data-access calls, over `GET /api/v1/portal/calls` and `GET /api/v1/portal/calls/{id}`, with `POST .../promote` and `POST .../reject`. Same catalog the admin Calls surface reads (`internal/platform/callrecord`), scoped the way My Sessions is: the caller is a predicate inside the query, so another user's record id is 404 and there is no user facet or column. A record is one query (`trino_query`, `trino_execute`, `trino_export`) or one API invocation (`api_invoke_endpoint`, `api_export`), written by an audit-store decorator on the audit writer's own goroutine, carrying the statement or request line, the targets (dataset URNs parsed from the SQL; the endpoint identity for API), the stated purpose, the session, the duration and the response size. Outcome is DERIVED on read, never stored: `failed` (the call errored), `satisfied` (an asset's provenance or a capture's `sources` names it), `superseded` (a later successful call in the same session over the same connection and targets, with nothing built from this one), else `ran`. Reuse counts later sessions that fetched the record and then ran what it holds; a session re-running its own query, or an identical query written without reading the record, counts for nothing. Facets: kind, connection, outcome, target, session, free text over purpose and statement, and `queue=promotable` (satisfied, undecided, ordered by reuse). Promotion writes a SQL record to DataHub as a Query entity over every dataset it reads (`knowledge.DataHubWriter.CreateCuratedQuery`) and an API record as a saved example on its endpoint; a rejection is recorded so the queue stops offering it. Retention (`calls.retention_days`, default 90) sweeps by what a record came to rather than by age: a record cited by an asset or a capture, promoted, declined, or re-run by another session is never swept; a query that came to nothing ages out. `calls.exclude_personas` names the personas that are machinery -- an automated system driving ingestion through the same tools people use writes a record per fetch that nobody re-runs, and every one is embedded -- and a call made under one of them is audited exactly as before and never cataloged, so no surface can return it and the indexer has no unit to embed; the records such a persona wrote before it was named are swept whatever their age, with the evidence clauses standing, and a name matching no known persona is warned about at startup. A managed script run's calls are excluded the same way and take no setting: the platform's own scheduler is the automated system, a run is by construction the re-run, and a run presents the persona of the person who wrote the script, so `exclude_personas` could never name it -- the rule reads the audit event's `source` instead, the sweep clears the rows written under a `script:` principal before this existed, and a person's call in the same persona is cataloged exactly as before. The audit row, its retention and the API gateway metrics are untouched; what is also withheld is the `mcp:call:<id>` reference a data call normally comes back with, since an id the catalog declined to record resolves to nothing and citing it would store a citation that can never be satisfied. The sweep runs at startup and then daily under a PostgreSQL advisory lock so replicas delete once. Addressable at `/activity/calls/{id}`.

The asset viewer walks the other way: the metadata sidebar's Session row and the provenance panel's Open session action both open the session that produced the asset (`/activity/sessions/{id}` for an owner, `/admin/sessions/{id}` for an operator). Both are omitted for a reader who is neither, since the session would answer them 404.

## Assets
AI-generated dashboards, reports, and visualizations saved via `save_asset`. Grid/table view with search, content type filter (HTML, JSX, SVG, Markdown, CSV), and tag filter. Ordering is a control, not a constant: a column select (updated, created, name, size) plus a direction toggle, mirrored by the sortable table headers, sent to the server as `sort` and `dir` on `GET /api/v1/portal/assets`. It defaults to `updated_at DESC` so the most recently revised work is first, resolves the column against an allowlist before it reaches an ORDER BY, and appends `id` in the same direction so LIMIT/OFFSET pagination cannot repeat or drop a row that ties. The date each card and row displays is the one the list is ordered by. Asset viewer renders each type natively: HTML/JSX as interactive components, SVG as vector graphics, Markdown with GFM+mermaid, CSV as sortable tables. Preview/Source toggle, Delete/Download/Share actions. The version picker beside the toggle lists every version the asset keeps with the time it was written, since a number alone does not identify a version of an asset written on a schedule; the trigger shows the version alone. The knowledge pages that reference the asset are a Referenced by button to the right of the version picker, carrying their count and opening a modal that lists them, each opening its page; it replaced a full-width card above the toolbar that pushed the document down to name one page (#1792). The metadata sidebar's References panel is the person's end of the asset reference mechanism (#1475, #1488), over `GET/POST /api/v1/portal/assets/{id}/references` and `DELETE .../{kind}/{targetID}`: each reference's name and type, a scope for a managed resource and an owner for a referenced asset, a thumbnail where it is an image (loaded through the reference's own /portal/refs URL rather than the target's own route, so it renders for a reader who was only ever shown the asset), a link to the target where the reader could open it themselves, and the reference string with a copy control -- adding a reference never touches the content, so the markup still has to name it. The owner, a shared editor and an administrator add through a picker with a tab per kind (the resources they can read, the assets they can open) and remove; the picker names what the reference gives away and the asset's CURRENT audience before it is confirmed, a removal whose reference the stored content still writes warns with those line numbers first, a reference whose target was deleted is flagged rather than dropped, and a reader without edit authority gets the list and no controls. The reverse edge is on both ends (`GET /api/v1/portal/resources/{id}/used-by` and `GET /api/v1/portal/assets/{id}/used-by`): a Used by section on the resource beside the prompts one and on the asset viewer's own sidebar, flagging any referencing asset that carries a public link and counting without naming the ones the reader cannot open, and the resource delete dialog carries the same list. A presentation is an HTML asset built on the slide runtime the platform serves (reveal.js, pinned in the portal build and answered at `/portal/vendor/reveal/` by the same origin as the page, so a deck renders on a network that admits only this deployment and inside the share viewer's existing script policy). The HTML frame is an `srcdoc` document rather than a blob: URL, so a root-relative path in the artifact resolves against the page. Every HTML asset carries, on the row that holds the Preview/Source toggle and the version picker, a Present control that fullscreens the frame and moves the keyboard into it (arrow keys and space advance, Esc returns) and an Export PDF control that frames a second, print-stepped copy of the document under a modals grant and opens the browser's print dialog on it, one slide per page for a deck. That copy is always rendered light, whatever the document's own theme: a browser prints backgrounds only when the reader turns "Background graphics" on, so a dark deck printed as light text on white. The copy gives the page a white background and moves every colour that would not read on white to the other side of the scale keeping its hue, so an accent prints as a deep version of itself and a panel filled near-black prints as a pale tint; a document already written light prints as it did (#1772). A document on the served runtime also carries Overview, the runtime's grid of every slide, asked for by message. The frame fills the page under that row, ending at the page's padding. The share viewer carries the same three controls in its header and fills the space under its notices, and the thumbnail is the title slide. An agent asked for a presentation, a deck or slides is pointed at `mcp:knowledge_page:platform-presentations`, which carries the served paths, the skeleton, what the controls do, and what the sandboxed frame does not allow: the runtime's speaker view, a second window (#1767, #1769).

**What gets a tile.** If the viewer renders it as a document, it gets a tile: the drawable set is derived from the renderer registry (`ui/src/components/renderers/registry.ts`) rather than restated beside it, which is how YAML, XML, SQL, Python, JavaScript, CSS and TSV -- every one of them laid out by the viewer -- kept a content-type icon forever (#1754). The bridge is a table over the registry's own renderer kinds, exhaustive by type, in which the families the viewer renders and the tile page deliberately does not draw (PDF, audio, video) carry the reason they do not; a test derives the expectation from it for every content type the registry names, and the Go/browser parity test holds the two languages to each other. A type is matched whole -- its canonical media type with parameters removed and aliases settled, or a `+json`/`+xml` structured suffix, with any other `text/` type drawn as plain text -- rather than as a fragment found anywhere in the type (#1882): every Office Open XML type contains "xml", so an Excel workbook, a Word document and a presentation were offered as XML and drawn as the text of their zip bytes. They, OpenDocument files and archives keep their content-type icon, and migration 000162 cleared the tiles already drawn for a type nothing draws.

**Thumbnails and how they stay current.** The platform draws every tile itself (#1787), in a headless Chrome (`chromedp/headless-shell`, pinned by digest) that runs as a second container in the platform's pod and is reached over the DevTools protocol at `thumbnails.renderer_url` (default `http://127.0.0.1:9222`). A worker in every replica claims owed rows with a lease (`FOR UPDATE SKIP LOCKED`, `thumbnail_claimed_until`), opens a tile page -- the content-viewer bundle's second entry, `ui/src/tile-entry.tsx`, carrying the document as JSON -- and screenshots it once the page reports itself drawn: HTML and JSX in the viewer's own frame at 1280x960, painted at full size and reduced to 800x600 with a Catmull-Rom filter (a browser asked for the small picture paints body text a few pixels tall, #1789), everything else laid out at 400x300 and painted at twice the density, scrollbars hidden. A PDF's tile is rasterized rather than borrowed from the viewer's frame (#1794): the tile page rasterizes page one with pdf.js onto a canvas sized in device pixels and scaled to cover, in a chunk loaded by a PDF tile and nothing else (the worker file is served from the viewer's own chunk directory, which is why `.mjs` is servable there). A PDF reaches the page by URL like a raster image rather than through the page's JSON payload (`isBinary`) and is drawn once for both schemes. It is one of the families past the 1 MB source bound: `thumbtypes.LargeSourceLimit` holds a PDF, a CSV and a TSV to 32 MB, because each is drawn from part of the file rather than all of it -- page one of a PDF, and for a table the head of the file, which the worker cuts at a record boundary (`headFor`, 64 records or 256 KiB) before the tile page is built, since the tile is the header row and ten rows and the rest used to be parsed and discarded (#1802). Both stores ask the bound as a CASE over the content type (`thumbtypes.SourceLimitExpr`) so the raise reaches those families alone, and `thumbtypes.LargeSourceFamilies` and `HeadDrawnFamilies` are the one Go statement of which they are, held to the browser's copy by a parity test. A document that is password-protected, damaged or not a PDF settles with its reason and keeps its icon. Every family except SVG, PDF and raster images is drawn once per colour scheme with `prefers-color-scheme` emulated: the families laid out on the portal's surface take the scheme's background, and HTML and JSX answer it through their own stylesheets as they do in the viewer's frame (#1789). Raising the renderer generation (`thumbnail_renderer`) redraws every stored tile once in the background, and the old tile serves until its replacement lands; the generation is on the asset, resource and collection-item JSON (`thumbnail_renderer`, `asset_thumbnail_renderer`) and the portal puts it in the tile URL (`&r=`) beside the version, because a redraw keeps the version and the thumbnail route is cacheable for an hour. A collection's mosaic is composed twice, light from its members' light tiles and dark from their dark tiles, the dark one stored beside the light one at `portal/collections/<id>/thumbnail_dark.png` and served by `GET /api/v1/portal/collections/{id}/thumbnail?variant=dark` (the light mosaic when a collection has no dark one yet); the mosaic's source signature carries each member's dark version, so a member's new dark tile recomposes it. The portal root sets `color-scheme` from its resolved theme (`:root { color-scheme: light }`, `:root.dark { color-scheme: dark }` in `ui/src/theme.css`, shared by the portal and the public share viewer), which every document frame inherits and which a framed document's `prefers-color-scheme` resolves against, so an HTML or JSX asset follows the reader's theme choice rather than the operating system (#1789). The browser paints it, so the tile is what the viewer shows; it replaced html2canvas 1.4.1, a JavaScript reimplementation of CSS painting that could not draw a CSS transform (a reveal.js deck became a sliver in the corner, #1784) and painted an inset box-shadow as a solid fill (#1787). The renderer never reaches the network: every page runs in its own browser context behind a dead proxy, every request of the page, its frames and its workers is intercepted and answered by the platform (its own routes in-process, a public URL fetched through `internal/egressguard`), and the socket constructors are removed before any script runs. A row is owed a tile when it has none, has one of an older version (an asset's `thumbnail_version`, a resource's `thumbnail_captured_at` against `updated_at`), one drawn by an older renderer generation (`thumbnail_renderer`), or one under a legacy filename; a version write leaves the stored tile serving until the new one lands. A document that cannot be drawn is recorded (`thumbnail_failure`, and the version or file time it was tried at) and not tried again until it changes or its owner clears the tile; the reason is on the file's Thumbnail panel, where Recapture (`DELETE .../thumbnail`) discards the tile and the failure and the panel polls the row until the new tile arrives; the control is disabled while the row says a tile is owed, since the draw it would ask for is already under way, and presses never queue work (a clear resets columns on the row, the worker leases the row, and the clear does not touch the lease). The light variant is drawn and stored first, so a failure can stand beside a stored tile, and the panel names the part that failed from the tile's own stamp: a light tile of the current version means only the dark tile failed, a light tile of an earlier version means this version could not be drawn and the reader is looking at the previous one, no tile means the file has none (#1791). The worker's in-process requests to the platform's own routes carry `viewerlimit.InProcess` on their context and are admitted by the public viewer's rate limiter without being counted: every one presents the loopback address, so counted they shared one bucket and a document drawn after it ran dry was recorded as not drawable although every file it names exists (#1791). A collection's mosaic of its first four drawn members is composed by the same renderer and redrawn when that signature (`thumbnail_source`) changes. The browser capture routes (`PUT` asset, admin asset, collection and resource thumbnail, and both `thumbnails/pending` lists) and the in-tab capture queue are gone.

## Collections
Curated groups of assets organized into ordered sections with markdown descriptions. Grid/table list view with search, and the same ordering control as Assets minus size, which a collection does not have (`sort`, `dir` on `GET /api/v1/portal/collections`; same `updated_at DESC` default and same id tie-breaker). Collection viewer renders sections with markdown, asset cards with thumbnails, each in the reader's color mode: a collection item carries both capture keys and both capture versions out of the same join that populates its name and content type, so a tile asks for `variant=dark` exactly when a dark capture exists and carries the capture's version as its cache key, the same as an assets-grid card (#1468). The admin asset thumbnail route accepts the same `variant` parameter, since an administrator reading a collection someone else owns gets its tiles from there (#1292). The public share viewer's tiles are still light-only. Configurable thumbnail sizes (Large/Medium/Small/None). The knowledge pages that reference a collection are a Referenced by button to the left of Feedback in its header, with their count and a modal listing them (#1792); the prompt viewer, knowledge page detail, and glossary, domain and tag views keep the card. Sharing via links (token-based, time-limited, opening for signed-in users or, by explicit choice, for anyone) and user shares (email, Viewer/Editor permission, restricted to that recipient).

Collection API endpoints:
- CRUD: POST/GET/PUT/DELETE `/api/v1/portal/collections[/{id}]`
- Sections: PUT `/api/v1/portal/collections/{id}/sections` (full replace)
- Config: PUT `/api/v1/portal/collections/{id}/config`
- Sharing: POST/GET `/api/v1/portal/collections/{id}/shares`
- Shared with me: GET `/api/v1/portal/shared-collections`
- Public view: GET `/portal/view/{token}` (branches on collection vs asset)

Database tables: portal_collections, portal_collection_sections, portal_collection_items. portal_shares extended with collection_id (CHECK constraint: exactly one of asset_id/collection_id set).

## Resources
Human-uploaded reference materials (SQL templates, runbooks, checklists) accessible to AI via MCP resources/read. The page is a FILE MANAGER (#1872): a folder tree on the left whose top-level folders are My Resources, Global and each persona the caller can read (a platform administrator adds People, one folder per person from `GET /api/v1/resources/people`), a path bar with Back/Forward/Up and clickable segments that take a drop, one listing of the folder in view (folders first, then files; Name/Modified/Size sort on the server through `sort=name|name_desc|updated|updated_asc|size|size_desc`; rows or tiles), a preview pane (one file: thumbnail, description, location as `/Global/data/weekly`, size, modified, uploader, URI, tags, with Open/Download/Copy URI; several: count, total size, bulk actions; nothing: the folder), and a status line. A folder opens on click and a file is selected and previewed, opening on double-click or Enter (the stated exception to the portal's row-click rule). Click, Shift-click, Cmd/Ctrl-click, Cmd/Ctrl-A and Escape select; the selection bar offers Move to, Tag and Delete, each reporting per file; rows dragged onto a folder row, a tree node or a path segment move there with Undo on the toast; files dropped from the desktop upload there; right-click offers Open, Rename, Move to, Tag, Copy URI, Delete, and New folder and Upload here on empty space; F2 renames inline; the tree follows the WAI-ARIA tree pattern. The page never says "library". A file's tile is a PNG the platform's renderer drew and stored beside it, in the reader's own colour scheme where the family carries two; a type nothing draws keeps its content-type icon. A file's own page carries a Thumbnail panel showing that tile, or the reason the renderer gave for not drawing it, with a Recapture control, for whoever may change the file.

The scope tab in view is the upload's destination, and the dialog states it before a file is chosen (#1444). `ui/src/pages/resources/scopes.ts` mirrors `resource.CanWriteScope` in the browser — own user scope always, persona scope with a `persona-admin:{name}` role (the name cut out of the role, so any identity-provider prefix works) — and the page renders Upload, in the filter bar and in the empty state, only on a tab that check passes. Where it does not, a note in the control's place names who publishes that library. On All, both upload dialogs name the persona library the caller belongs to and cannot upload into (`withheldUploadPersonas`), the `persona-admin:{name}` role that grants it, and the move (Edit details > Library) that membership does allow (#1866). A persona admin therefore adds to their persona's library from the user page, not only from the admin page, which keeps its scope picker and its persona and per-user fan-out.

The platform-admin override is the identity's, not a section's (#1527). `canWriteScope`, `moveTargets` and `libraryOptions` read the caller and nothing else: a platform admin is offered Upload on every scope tab and every move target — their own library, every persona, Global, and a named person's library — on the portal Resources page and in Admin > Resources alike, which is what `CanWriteScope` and `CanMoveToLibrary` have always granted the same caller whatever route the request arrived on. The `Surface = "portal" | "admin"` parameter those three took, and the `adminReachNote` sentence that pointed an administrator at Admin > Resources in the withheld control's place, are removed rather than defaulted: a parameter whose only remaining value is `"admin"` is a seam the behavior grows back on. The portal enables the `usePersonas` fetch for a platform administrator so the picker offers the deployment's personas and not only the ones their own claims name; everyone else's personas stay claims-derived, and `moveTargets` offers a non-admin nothing out of a list it is handed. `ResourceViewerPage` reads Edit and Delete the same way — `CanModifyResource`'s two arms, the uploader and `canWriteScope` over the library the file is in — instead of the section it was mounted in, which had left an administrator on the portal with no Edit button on a file they did not upload. A blob store that refuses the write now answers 503 with what happened and that nothing was saved, rather than the fragment "storing file" that `writeError`'s colon truncation left of the internal chain — a storage outage read, to the person who hit it, as a refusal of permission.

Resources are also discoverable through the universal `search` (#1012). An indexjobs consumer (`internal/platform/resourceindex`, source_kind `resources`) embeds each resource's metadata plus a bounded text prefix (32 KiB) read from its S3 blob for text-family MIME types (classified by `pkg/contenttype`); the vector, the model/text-hash breadcrumbs, and the extracted `content_text` live inline on the `resources` row (migration 000091), so a DELETE removes the index entry with the resource. The request-path metadata Update clears the vector and enqueues the row's own index job (trigger `write`, #1256), as do the create and content-revision paths, so a just-uploaded or just-edited resource is re-indexed within one embed rather than on the next sweep; the reconciler stays the backstop for what a write could not produce. `content_indexed_at` is a second, independent gap signal: the consumer stamps it only once it has extracted the text or established there is nothing to extract, so a transient blob failure (which still yields a valid metadata-only embedding) leaves the row in `FindGaps` and the next sweep retries, instead of the job succeeding and the file's contents never being indexed while coverage read 100%. Coverage counts a resource indexed only when both signals are settled. Content extraction needs the index queue (database + configured embedding provider); an object over 8 MB is indexed on metadata alone, since the blob API has no range read and the whole object would be held in memory to keep 32 KiB of it. A transient blob failure keeps the previously extracted text; a confirmed missing object clears it and indexes metadata only (the row is pruned only by the read path's orphan self-heal). The by-id REST handlers (get, content, patch, delete) gate on `resource.CanAccessResource` (visible scopes OR current write authority over the scope — deliberately not `CanModifyResource`, whose uploader clause would outlive the role that authorized the upload), so a platform admin can manage a persona resource they uploaded rather than being 404'd on material `CanWriteScope` let them create; ENUMERATION stays membership-scoped via `CanReadResource` (a listing and the ranked search hand a caller material they did not name), while a file the caller NAMES is answered on `CanAccessResource` wherever it is named -- the by-id routes, an asset reference declaration, and `fetch` on an `mcp:resource:` reference (#1584). `resources/read` remains membership-scoped. Search hits carry `mcp:resource:<id>` plus an MCP resource link with the canonical `mcp://` URI; `fetch` on that reference returns text contents inline (at or under the 1 MB threshold shared with the MCP read path) or metadata with the URI, MIME type, and size.

A RESOURCE'S SOURCE IS EDITABLE IN THE PORTAL (#1775). A text family the viewer can show as source carries a Preview/Source switch and a Save above its content for a reader who may change the file -- `ViewModeToggle`, `SaveControls` and `SourceEditor` imported from the asset viewer rather than rewritten, so the two surfaces cannot drift. A save posts the edited text through the SAME `POST /api/v1/resources/{id}/content` route the file picker uses, so it records a version, keeps the id, URI and filename, and carries the server-side validation; it sends `change_summary: "edited in the portal"`, which that route now reads (it had the field on `RevisionUpload` for the correction a table registration saves and sent an empty one whatever the caller passed), so Version history tells an edit from a file picked off disk. The switch is withheld from a reader who may not change the file, from a family that is not text, and from a file past the inline preview limit. Until this a resource was VIEW-ONLY in the portal and the only control named Edit opened the metadata dialog, so somebody looking at their own CSV pressed it and could not touch the contents; that button is now named Edit details.

A resource's content is revisable in place (#1014). `POST /api/v1/resources/{id}/content` uploads a replacement that keeps the resource's id, canonical URI, and filename — so every `mcp:resource:<id>` citation and prompt attachment keeps resolving, which delete-plus-re-upload breaks by minting a new id — and the uploaded file's own name is therefore ignored. Each revision writes a `resource_versions` row (version number assigned inside the insert transaction, uploader, timestamp, size, MIME type, and its own per-revision S3 key) and moves the head in the same transaction, so the trail and the live content can never disagree; migration 000092 backfills every existing resource as version 1 pointing at its current blob, copying nothing. `GET .../versions` lists the trail with the version the head serves, `GET .../versions/{n}/content` downloads one, and `POST .../versions/{n}/restore` re-promotes a prior version's exact bytes as a NEW head revision (recorded with `restored_from`) rather than rewinding, so history stays append-only. A revision written on the uploader's behalf records why (`change_summary`, migration 000123, #1450) and the version panel renders it beneath the row -- the same thing `portal_asset_versions.change_summary` has carried since 000022, so a corrected managed resource and a corrected portal asset say the same thing in their histories rather than the resource's correction reading as an upload. Retention is bounded by `resources.managed.max_versions` (default 10): a revision past the cap deletes the oldest version rows and their S3 objects, never the key the head points at. A revision clears `content_text`, `content_indexed_at`, and the embedding and enqueues the resource's index job so the new bytes are re-indexed within one embed rather than on the next sweep, and fires `resources/list_changed`. Deleting a resource removes every version's blob (the rows cascade).

Reads are audited (#1014). Content served through any of the four doors — MCP `resources/read` (`mcp_read`), a `search` fetch of an `mcp:resource:<id>` reference (`fetch`), a REST content download (`rest_download`), or the portal drawing an image tile in the resource library (`portal_preview`) — writes one `resource_read` audit event (`resource_id`, `resource_uri`, `surface`, and `version` when a specific revision was named) through one shared recorder (`internal/platform/resourceaudit`), gated by `audit.enabled`; listing is not a read, and a read that could not be served records nothing. The same recorder stamps `resources.last_read_at`, which is what `GET /api/v1/resources?sort=last_read` orders on (`DESC NULLS LAST`, so never-read material sorts last) and what outlives the audit retention window. `GET /api/v1/resources/{id}` carries a `usage` object aggregated from those events by the audit store (`ResourceUsage`, backed by a partial expression index on `(parameters->>'resource_id') WHERE event_kind = 'resource_read'`): reads over 30 and 90 days, a per-surface 30-day breakdown, and the last read the rollup saw. The portal detail view renders version history (with download and restore), the usage panel, and the prompts attaching the resource; the admin table adds a Last read column, a Recently read sort, and a flag on anything never read since it was uploaded over 30 days ago.

A resource can be MOVED to another library after it is uploaded (#1502). The library was chosen once on the upload form and never again, so the only route from a personal library to a shared one was to upload the file a second time -- a second id, a second URI, a second blob, two separate version trails, and every asset and prompt still referencing the first. `Library` on the Edit dialog rewrites `scope`, `scope_id` and `uri` on the row instead; the S3 key is a stored column that nothing recomputes on read, so the blob stays where it is and the version trail, table registration and read history are untouched. `resource.CanMoveToLibrary` is the destination check and is looser than `CanWriteScope` on exactly ONE arm: a persona you BELONG to accepts a file you already own, while uploading into it still takes `persona-admin:{name}`. Widening `CanWriteScope` itself would have been the smaller diff and the wrong one -- it is what the upload route checks and what the Resources page derives its Upload control and scope tabs from, so every member of every persona would have gained an upload door nobody asked to open. Global and another person's library stay `CanWriteScope`'s, so an administrator reaches every target and nobody else reaches those two. The move travels on the update route (`PATCH /api/v1/resources/{id}` carrying `scope`/`scope_id`) because the permission check it needs -- `CanModifyResource` on the resource itself -- is the one that route already runs, and it runs BEFORE any metadata edit in the same request so a refused move cannot leave half of it committed. A move to the library the file is already in writes nothing and audits nothing, so an idempotent PATCH neither fails nor reads as a refile.

THE URI FOLLOWS THE FILE, AND THE ADDRESS IT LEAVES KEEPS RESOLVING. `RelocatedURI` composes the target library's prefix, the target folder path and the filename off the stored URI, so a rewrite that changes one half leaves the other exactly as it was; a URI that will not parse falls back to the row's own filename. A folder edit takes the same path as a library move (#1528): `Update.Path` is not metadata, `buildUpdate` writes no `path` column, and `applyMove` in the handler resolves whichever halves the request named against the row's current values into ONE `Destination`, so a library and a folder changed in the same PATCH produce one URI carrying both, one alias for the one address vacated, and one audit event. The no-op test compares the composed URI as well as the two columns, so a row whose stored URI disagreed with its folder -- which every category edit before #1528 produced -- is repaired by saving the folder it already shows. `resource_uri_aliases` (migration 000129, `uri` PRIMARY KEY, `resource_id` ON DELETE CASCADE) records every address a resource has vacated, written in the SAME transaction as the row -- the new address and the alias for the old one are one fact stated twice, and a commit carrying only the first leaves every citation dangling. `GetByURI` tries the live column first and the alias table second, so an alias can never shadow whichever resource occupies that address NOW: a person who moves a file out and uploads another under the same name reaches the new one by its own URI. The move also DELETES any alias equal to the address it is taking, so moving a file back where it came from leaves no row claiming its own current URI. The collision pre-check tells a live occupant from an alias by comparing the hit's own `URI` against the address asked for, and names the occupant in the refusal (the mover cannot see the other library, so "that address is taken" alone is an error they can do nothing about); the store additionally maps a UNIQUE violation, at exec or at commit, to `ErrURIConflict`, since the pre-check and the write are not atomic. Reference and attachment rows are NOT touched: `portal_asset_resource_refs` and `prompt_resource_attachments` key on `resource_id`, and the serve-time rewrite matches the URI string recorded on the reference row -- which is what the author wrote and stays what they wrote -- so a declaring asset and an attaching prompt both keep serving the file, proven end to end over the real PATCH route, the real reference-serving route and the real attachment resolver rather than by asserting a store call. The MCP registry IS keyed on the URI, so a move withdraws the old one and registers the new. Each move writes a `resource_move` audit event (`resource_id`, `display_name`, and `from_scope`/`from_scope_id`/`from_path`/`from_uri`/`to_scope`/`to_scope_id`/`to_path`/`to_uri`) through the same `internal/platform/resourceaudit` recorder reads go through; the "from" side is snapshotted BEFORE the write, since a store is free to hand back the same `*Resource` it then mutates. `ui/src/pages/resources/scopes.ts` mirrors the rule in the browser through `moveTargets`/`libraryOptions`, withholding the platform-admin override on the portal surface exactly as `canWriteScope` does, and offering the current library even when the caller could not move it there, because leaving the file where it is has to be expressible; the field is absent entirely when there is nowhere else to put it. There is no MCP action: `manage_resource` creates and replaces content, and deciding that a file becomes a persona's or the platform's is a human act.

A FOLDER VIEW LISTS ITS OWN LEVEL (#1837). The listing for a folder used to be the folder and everything beneath it, paged in the view's sort order, with the page then filtering to the files whose path is exactly this folder; the footer reported the fetched row count and the subtree total. A `brand` folder whose four subfolders held 2,528 files showed four folder tiles, no files, a Load more, and `Showing 146 of 2541 resources`: 146 poster rows fetched, none on screen, and Load more fetching more rows to hide. `GET /api/v1/resources?path=<folder>&direct=true` (`Filter.Direct`, a single `path = $n` equality in place of the `path = $n OR path LIKE $n+1` subtree arm) lists only the files filed at that path, and the portal's folder view always asks for it (`listingFor` in `useResourceLibrary.ts`), so the page, its total and Load more all describe the files on screen; the subfolders and their counts come from the facets as before. A search and a tag at a library root still list flat across the library. Without `direct` the route answers the subtree, which is what a folder move plans over.

RENAMING A FOLDER IS ONE TRANSACTION OVER THE WHOLE SUBTREE (#1529). `Store.Move` takes a BATCH of relocations rather than one, because a half-renamed folder is not a state anyone should be able to observe; `MoveFolder` (`pkg/resource/folder.go`) lists every resource beneath the prefix, refuses the whole move if the caller may not modify one of them (403) or if a destination address is occupied (409, naming the occupant) or if the target is inside the folder being moved, and hands the store one batch. A multi-row batch VACATES every address before taking any: renaming `a/b` to `a` hands one resource an address another in the same batch still holds, and the UNIQUE constraint on `uri` is not deferrable, so the rows are parked on a sentinel derived from their own primary key first and the statement order stops mattering. One rename covers at most `MaxFolderMoveResources` (500) resources and is REFUSED with its true count past that rather than moving part of the subtree. `POST /api/v1/resources/folders/move` is the route; the MCP registry is keyed on the URI, so every vacated address is withdrawn and every new one registered from the row as stored. The folder path is part of `IndexText`, so `applyMove` drops the embedding only for a row whose path actually changed -- compared by the database against the stored value, since a library-only move changes no indexed text. `scopeLabel` now takes the reader, because a user library is only "My Resources" to the person it belongs to and the admin section lists every library at once -- a library keyed by address is named by that address, one keyed by a subject identifier is described rather than printed.

## Shared With Me
Items shared by other users appear in the corresponding pages, filtered by ownership scope: the Assets and Collections pages each have a Mine / Shared / All scope control (shared items show sharer email and permission level), and the Prompts page has a Shared tab (shared prompts are runnable over MCP as shared-<name>).

## Share landing page and guest links (#1001)
A refused `/portal/view/{token}` browser navigation renders a branded landing page instead of bare text: sign-in with a return path for account holders; for a signed-in caller who is not the recipient, a page naming the signed-in account with a sign-out-and-switch action; a branded 410 for revoked/expired. When the share names an email address, the landing page also offers "Email me a one-time view link" for recipients with no platform account: POST `/portal/view/{token}/request-link` (uniform response regardless of share state, per-IP rate limit plus a per-share hourly cap) emails a single-use link only to the stored `shared_with_email`. The link (GET `/portal/view/{token}/guest?otk=...`) is claimed atomically (SHA-256 hash at rest in `portal_share_guest_links`, 15-minute expiry, dead after first use, so forwarding transfers nothing) and opens a view-only guest session: an HMAC-signed browser-session cookie scoped to that one share (key domain-separated from the browser-session signing key), rendered in the public viewer with a "Viewing as guest" indicator, download allowed, no editing even on Editor shares, no auto-promotion, no portal or API access. Revoking the share refuses existing guest sessions immediately. Same flow covers collection shares and their item routes. Subresource and API-style refusals stay plain-status. (`pkg/portal/shareguest`)

## Feedback
Structured feedback from reviewers (including non-agent subject-matter experts and stakeholders) on shared work, replacing email round-trips. Built on a generic thread substrate (`portal_threads` + `portal_thread_events`) where a comment is one event type among many, so later thread kinds slot in with no schema churn. A thread targets one asset, collection, prompt, or knowledge page (mirroring the `portal_shares` 1-of-N polymorphism, enforced by a CHECK constraint) or lives on a standalone channel; it carries a kind (comment, question, correction, rating, approval, rejection, suggestion), a status (open, answered, resolved, wont_fix, acknowledged), an optional requires_resolution flag, an optional inline anchor (JSONB, e.g. a W3C-style text-quote or a collection section) plus the target_version it was raised against, and a typed event timeline (comment, status_change, resolution, rating, approval, rejection, and the Phase 2 knowledge-link events). A status change records a timeline event in the same transaction so the timeline never shows a status with no event. Standalone threads are visible to any authenticated user; object threads follow the target's existing view access; knowledge pages are org-shared, so any authenticated user can read and add feedback on them. Moderation (status change, delete) is allowed for the thread author, the target owner/editor (for knowledge pages, an apply_knowledge holder), or an admin. When an anonymous viewer opens a public share link, they see a "Sign in to leave feedback" prompt; signing in, when they have no prior share for the item, auto-creates a viewer share (`origin=public_link_login`) so the item appears in their portal, and never downgrades an existing editor. A signed-in platform user who can open the target in the portal (its owner, a holder of an active share, or one the derived share just granted) AND whom the portal application would actually serve (the deployment carries a frontend build and the shell gate admits their roles, both answered by the composition root through one seam so the redirect and the page it lands on cannot disagree) is sent there instead of being shown the public page, carrying the token so the portal viewer offers a Shared page action back to the page its recipient sees; a promotion that failed, an account no persona claims, a deployment serving no portal application, a guest holding a one-time link, and an anonymous reader all get the public page unchanged. REST under `/api/v1/portal/threads` (GET list scoped to one target, POST create, GET/PATCH/DELETE by id, GET/POST `/events`) plus `GET /api/v1/portal/threads/counts` for list-page open-thread badges, owner-scoped for assets/collections (a non-admin receives counts only for objects they own) but open to all for org-shared knowledge pages, and `POST /api/v1/portal/threads/{id}/insight` to capture a correction/suggestion thread as a pending, knowledge-dimension insight (requires apply_knowledge) that enters the review queue and links/resolves the source thread via the existing insight bridge. List rows carry timeline aggregates (event_count, last_event_at, last_event_type) so the panel renders activity without an N+1 fan-out. Humans author and triage feedback through a portal panel: a non-modal slide-out drawer mounted in the asset, collection, prompt, and knowledge-page viewers, plus a full-width Inbox page in the sidebar (named Feedback until #1798). The drawer lists threads with kind/status/activity, supports a new-thread form (kind, requires_resolution, rating, optional text-quote anchor captured from a markdown/plain-text selection), threaded replies, and owner/editor/admin status changes and deletion. The Inbox has four views: Recent (the cross-target activity feed, each row linking to its item and opening the thread in a right slide-over with a "Go to item" link), Worklist (the practitioner, SME and mentions worklists, selected by chips rather than a second tab strip), General (the standalone channel), and Notifications (the history of what the platform has emailed the caller, read from `GET /api/v1/portal/notifications`, which used to be a card on /settings beside the preferences that decide what gets sent; the preferences stay there). The page returns flowing content with no scroll region of its own -- `main` is the one scroll container, as on every other portal page (#1798). The sidebar's Inbox item carries a badge of the caller's open worklist count as a stand-in for push notifications. My Assets and Collections show an owner-scoped open-thread badge per item.

## Knowledge
The single home for the Memory to Insight to Knowledge lifecycle (formerly the separate Knowledge Pages, Knowledge & Memory, and admin Knowledge & Memory routes, which now redirect here). A header teaches the model: everything learned is a Memory; a memory others would benefit from becomes an Insight (a proposal awaiting review); whoever holds the `apply_knowledge` capability promotes good insights into Knowledge (business/domain facts become knowledge pages, technical/entity facts go to the DataHub catalog). Three tabs, with review/promote affordances gated on the `apply_knowledge` tool (a capability, not an admin role). Knowledge (default): unified search across every accessible source grouped by source with a coverage summary (the same federation as the `search` tool, over `GET /api/v1/portal/search`); with an empty query, browse of canonical knowledge pages (create/edit/remove for `apply_knowledge` holders; the platform's own built-in pages sit in the same corpus — reconciled from the binary at startup, badged Built-in, read-only where people edit, hidable per deployment with the hide respected across upgrades and reversible via the Knowledge list's Restore built-in / POST /api/v1/portal/knowledge-pages/restore-builtin) in either of two layouts, a card list or an interactive graph of the corpus (#1162) where every page and every entity a page references is a typed node and every stored reference is a directed edge. The graph is exploratory rather than a whole-corpus hairball: it opens on the corpus's strongest bridge and its neighborhood, with a hops control to widen it and a whole-corpus overview on demand. It is analysed, not merely drawn - Louvain community detection partitions the corpus into clusters (tinted as regions in the overview, with a clustering force separating them, and reported with the partition's modularity) and betweenness centrality scores each node for how much of the graph it bridges, which sets node size and is reported per node with its percentile. Clicking a node opens a side inspector (references in each direction, bridge score and rank, cluster, the reference's URN, selectable neighbor lists) rather than navigating away; its actions are focus, expand, shortest-path tracing between any two nodes, and open. Selecting a catalog node resolves it against the DataHub catalog and reports what is there (description, domain, owners, tags) naming the connection queried, or states plainly that the cited dataset is not in that catalog - a page citing a dataset the catalog does not have is surfaced as the gap it is. Which of the two it is comes from the read itself, a 404 for a URN the catalog has never ingested (#1610), so a dataset the catalog holds and nobody has documented reads as held rather than as missing. Catalog entities are URL-addressable at /knowledge/catalog?urn=..., which is where a catalog reference links from anywhere in the portal. Plus hover neighborhood highlighting, node drag with pinning, pan/zoom, type and tag filters, and search-as-focus; the graph read is `GET /api/v1/portal/knowledge-pages/graph`, access-filtered so an entity the viewer cannot see has neither node nor edge, and explicitly reporting any node/page cap instead of truncating silently; and for `apply_knowledge` holders the changesets (the record of insights promoted into knowledge, with rollback) since a changeset is created only at apply time and belongs with the promoted knowledge, not the unpromoted insights. Insights: the review pipeline only (insights are the one memory type that crosses between users, and they cross when applied) - your captured insights with status and relevance search, plus for `apply_knowledge` holders the full review queue (approve/reject); a pending-review count is badged on the sidebar Knowledge item and the Insights tab. Memory: personal, scoped to your own records (the cross-user unit is the insight), classified by lifecycle class (`sink_class`: Preference, Event, Business knowledge, Operational rule, Schema/entity). The Knowledge tab also has a Catalog sub-tab (#719/#720/#1156/#1157/#1158/#1194), a first-class route at `/knowledge/catalog`, which holds every DataHub-backed surface in the portal: the rule is that everything under Catalog is DataHub and anything the portal's own database backs (knowledge pages, changesets) stays outside it, which is what keeps the Knowledge sub-tab row at four (Search All, Knowledge Pages, Catalog, Changesets) while the catalog surfaces grow underneath. Catalog's own inner tabs are Tables, Context Docs, Tags, Domains, and Glossary - the described things first, the vocabularies that describe them second. The DataHub connection is picked once for the whole section and applies to every inner tab; changing it returns each tab to its list, since an open table, document, tag, domain, or glossary entity belongs to the connection it was read from. The inner tab is carried in the hash (`/knowledge/catalog#tags`, `#domains`, `#glossary`, `#context-docs`, `#tables`) rather than its own route, so the selection survives a refresh and back/forward without unmounting the container that holds the shared connection; there is no `/knowledge/tags` or `/knowledge/context-docs` (they were removed outright in #1194, with no redirect, since only the tab bar itself produced those URLs). Tables: browse/search the tables the connection catalogs, open one to see description, tags, owners, glossary terms, domain, and columns, and edit each facet inline when the persona grants `datahub_update` and the connection is writable (no table create/delete since tables originate in source systems). Context Docs: browse/search and full create/edit/delete of markdown context documents through a markdown editor, gated on `datahub_create`/`datahub_update`/`datahub_delete`; a document attaches only to Dataset/GlossaryTerm/GlossaryNode/Container. Tags: the tag vocabulary itself rather than one table's tags - list and name-filter a connection's tags, open one to see its description and the tables carrying it (each linking into the Tables entity editor), and create, describe, or retire a tag when the persona grants the matching datahub tool on a writable connection; the delete confirmation states how many tables carry the tag first. A tag description is plain text, not markdown - the one deliberate exception among the Catalog vocabularies (#1200), because DataHub's own tag page renders the field as plain text and formatting authored in the portal would show as raw source everywhere else in the catalog. Domains: the business areas the catalog is grouped into rather than one table's domain - list and name-filter a connection's domains, open one to see its description and the tables in it (each linking into the Tables entity editor), create, describe, or retire a domain, and move tables in and out of it, each gated on the matching datahub tool on a writable connection; the delete confirmation states how many tables are in the domain and that deleting leaves them without one, since it touches no table. A domain description is markdown (#1200): the domain view renders it formatted and the editor is the split source/preview markdown editor the Tables tab and Context Docs use. Glossary (#1158): the business vocabulary itself - a tree, so it is walked one branch at a time (the root shows the nodes and terms with no parent; opening a node shows what is inside it, and a node's browse view IS its detail view, carrying its definition, its attached context documents, and its children on one screen). A term shows its definition, a breadcrumb built from DataHub's parent chain (so it is the same wherever the term was reached from), the context documents attached to it, and the tables annotated with it, each linking into the Tables entity editor and marked when a COLUMN rather than the table carries the term. Create a term or a node, edit either's definition, and retire a term or an EMPTY node, each gated on the matching datahub tool on a writable connection; a new entity lands in the open branch and the form names it. A node that still holds entries is not offered a delete at all, because DataHub takes the node without taking what is inside it - the surface says to empty it first rather than showing a confirmation that cannot state the outcome. Term and node definitions are markdown (#1200), rendered formatted and edited through the same split source/preview markdown editor, so a definition can carry a heading, a list of the cases it includes and excludes, and a worked example; the field itself was never the constraint, since all of these kinds write through the one `PUT catalog/entity/description` route and markdown survives it byte-for-byte. One glossary backs both surfaces: a term defined here is immediately what the Tables tab's glossary picker offers. All five are backed by the portal DataHub REST API at `/api/v1/portal/datahub/{connection}/...` (`GET .../connections` lists connections with a writable flag): reads require DataHub access on the persona; a write requires the matching MCP tool grant AND a write-enabled connection (`read_only: false`), both enforced server-side and recorded in the audit log. Tag and glossary-term edits use batched add/remove sets (the clobber-safe write path, #721/#729). The same REST surface exposes the business glossary as a tree rather than only a flat name search (#1155, requires mcp-datahub v1.15.0): `GET catalog/glossary/roots` returns the nodes and terms with no parent (each paged with its own total, since DataHub pages the two independently), `GET catalog/glossary/children?urn=` returns one page of the nodes and terms directly under a node (DataHub pages a node's children as one mixed collection, so start/count/total describe the combined page rather than either slice), `GET catalog/glossary/parents?urn=` returns the ancestor nodes of a term or node direct-parent-first for a breadcrumb, `POST catalog/glossary/nodes` creates a node from {name, definition, parent_node} and returns its URN (empty parent_node creates it at the root), gated on `datahub_create` plus a write-enabled connection, and #1158 adds the rest of the editor: `POST catalog/glossary/terms` creates a term from the same body through the same handler (a term and a node differ only in which upstream call runs), `DELETE catalog/glossary/entity?urn=` retires either kind through the one route because upstream is one call (`datahub_delete`), and `GET catalog/entity/documents?urn=` returns the context documents attached to one entity - the one document read the corpus-wide browse and search cannot express. Two things the glossary needs are not routes of their own: a definition is edited with `PUT catalog/entity/description` (DataHub stores a glossary entity's text in the glossaryTermInfo/glossaryNodeInfo aspect's `definition` field, and the platform routes the write there by entity type), and the tables a term is applied to come from the catalog search's glossary filters, `GET catalog/search?q=*&glossary_term=<urn>` for every annotated table and `&column_glossary_term=<urn>` for those where a column carries it - two reads because DataHub's `glossaryTerms` index folds column-level annotations into the table's and only `fieldGlossaryTerms` isolates them, with no table-level-only field. Deleting a node does not delete what is inside it and deleting a term does not remove the term from the tables annotated with it, since upstream DeleteGlossaryEntity touches only the entity named, which is why the portal shows a node's children and a term's usage before offering the delete. Every node read carries `terms_count`/`nodes_count`, DataHub's own tally of its direct children, so a branch renders as expandable without first fetching it. A URN of the wrong kind is a 400 (children hang off a node only; a parent chain exists for either kind) and an unknown node is a 404, not a 503. Children are served from DataHub's asynchronously populated graph index, so a just-created entity may not appear under its parent yet; the parent chain reads the entity itself and is immediately consistent. Tag governance (#1156) adds only two routes, because its reads already exist: `POST catalog/tags` creates a tag from {name, description} and returns the URN DataHub assigned it (201, gated on `datahub_create`), and `DELETE catalog/tags?urn=` retires one (gated on `datahub_delete`); listing and name-filtering tags is `GET catalog/lookup/tags` (the picker's read), the datasets carrying a tag are `GET catalog/search?q=*&tags=<urn>` through the catalog search's tag filter, and a tag's description is edited with `PUT catalog/entity/description`, which takes any entity URN. A URN that is not a tag is a 400 before the call reaches DataHub, and a newly created tag is not immediately listable because the list read is served from DataHub's asynchronously populated search index. Domain governance (#1157) adds the same two routes for the same reason: `POST catalog/domains` creates a domain from {name, description} and returns its URN (201, `datahub_create`), `DELETE catalog/domains?urn=` retires one (`datahub_delete`), while listing domains is `GET catalog/lookup/domains`, the tables in a domain are `GET catalog/search?q=*&domain=<urn>`, the description edit is `PUT catalog/entity/description`, and membership is edited with `PUT catalog/entity/domain` aimed at the table rather than at the domain. Two limits the surface states rather than hides: the domain list is capped at 100 by DataHub's own `listDomains` query (the lookup route takes no limit), and a table has at most one domain, so adding a table already in another domain moves it. Knowledge pages link to governance entities as first-class references (#1159): the manual-reference picker searches glossary terms, tags, and domains by display name through the existing lookup routes (asking which DataHub connection to search, and offering the three catalog types only when a connection exists), and stores the entity's own URN. A stored governance reference renders with the name the catalog reports rather than the key inside its URN - DataHub generates a UUID key for anything created without an explicit id, so a chip built from the URN alone would read as `8f3c1a94` where the page meant `Net Revenue`. Names are resolved server-side in one batch on the existing refs/resolve path (an optional `CatalogLabeler` over the same DataHub bridge the REST surface uses, wired only when a connection is configured), because a tag and a domain have NO by-URN read upstream and are resolved by listing their vocabulary once per request and matching; only a glossary term has one, `GET catalog/glossary/term?urn=`, added by #1159 and also what opens a cited term in the portal. An unresolved or unreachable name falls back to the URN-derived label rather than failing the page, and resolution is gated on catalog access - the one rule `portal.HasCatalogAccess` now holds for both the DataHub REST surface and the labels (any `datahub_*` tool on the persona, or admin), so a persona denied the Catalog tab does not learn a governance entity's name through the reference list instead. A catalog reference links to the Catalog inner tab that manages that kind of entity (`/knowledge/catalog?urn=...#glossary|#tags|#domains`, everything else `#tables`); each inner tab claims only its own URN kinds, so a stale link opens the list rather than a read that cannot succeed, and going back drops the `?urn=`. A tag, domain, or term the connection does not list says so instead of opening a detail view assembled from the URN alone, and a failed read is reported as a failure rather than as a missing entity. Each governance detail view lists the knowledge pages that reference its entity, through the existing `GET /api/v1/portal/knowledge-pages/backlinks?urn=` reverse lookup keyed on `entity_urn` (no schema change). The MCP side already accepted these refs via ParseCitableRef, so `apply_knowledge` attaches them with no new input.

## Prompts
The organization's prompt library, presented as two buckets (#1010, #1124): My Prompts (every prompt the caller owns at any scope — shared scopes carry a scope badge — plus prompts shared with them, each attributed to its sharer; never another owner's personal prompt) and Library (the approved shared prompts visible to the caller). Both buckets are browsed grouped by collection, each group headed by the collection name, its prompt count, and the collection description beneath it; only search results are flat, holding their relevance order. The scope taxonomy (global/persona/personal) appears only inside the promote and admin flows, never in the user-facing library. Collections are named groups organizing the library by team, domain, or workflow (`prompt_collections` table; a prompt belongs to at most one via `prompts.collection_id`, released to the default General group when its collection is deleted). Any user creates collections; renaming/deleting is creator-or-admin; assignment follows the prompt's own mutation rule (owner for personal, admin for shared) and is organizational metadata: it never versions or triggers review. REST: `GET/POST /api/v1/portal/prompt-collections`, `PUT/DELETE /api/v1/portal/prompt-collections/{id}`, `PUT /api/v1/portal/prompts/{id}/collection` (admin-prefixed equivalents exist); agents assign with `manage_prompt` `collection_id` on create/update (an id from the `collections` array list returns; empty string clears; collections themselves are created in the portal). Facets narrow by collection, tag, status (My Prompts), owner (Library), and usage; rows show run count and last-run age with sorts by name, runs, and last run, and dead prompts carry a badge naming the exact usage condition ("never run" or "unused 60d+"; a prompt created within the last week carries no flag). Search ranks prompts by relevance to a phrase (semantic vector similarity when an embedding provider is configured, keyword fallback otherwise) across the caller's full visibility, best-first; browse mode keeps sortable columns. The same ranking backs the MCP tool `manage_prompt list query=...`. Every enabled prompt is embedded off the request path by the shared index-jobs framework (source_kind `prompts`) regardless of lifecycle status, since visibility is decided at query time; editing a prompt's title, description, body, or tags clears its vector so it re-embeds against the new text. Visibility is applied before ranking, so a prompt you cannot read is never returned, and it is the same rule browsing applies (#1124), so a prompt visible in a list is findable by search: an admin ranks across everything (any owner, any status); a non-admin ranks over approved global prompts, approved matching-persona prompts, and every prompt they own at any scope and status — the publication gate applies to other people's shared prompts, never to your own work. Sortable columns, expandable rows with full content and copy button. Scope badges: Personal, Global, Persona, System. Lifecycle status badge (draft, approved, deprecated, superseded) and comma-separated tags on create/edit. An admin creating a global or persona prompt is its approver: the prompt lands approved with the stamp set (#1124), on both the admin REST API and `manage_prompt`; personal creates stay draft and publish through the promotion flow. Request Promotion on a personal prompt asks an admin to promote it to a chosen persona or to global; the prompt stays personal with a "Promotion requested" badge until an admin approves or rejects it. Share sends a personal prompt directly to another user by email (owner-initiated, no approval): the recipient gets a real runnable prompt, not a markdown snapshot, and it appears in the Prompts page's My Prompts bucket with a shared-by attribution. Personal prompt names are unique per owner; when served over MCP, names carry a scope prefix computed at serve time: `personal-<name>`, `<persona>-<name>` (one per persona), `global-<name>`, or `shared-<name>` for prompts shared with the caller, keeping the surface collision-free by construction; every descriptor carries a `title` from display_name. Users never need machine names: agents resolve any handle (stored name, display name, `mcp:prompt:<id>`, or free text) to a ready-to-run prompt with the `manage_prompt` `use` command, which returns rendered content, argument specs, and provenance (including version, approver, and approval time), or ranked candidates when ambiguous. Every database prompt is versioned (#1009): each mutation of content, display name, description, arguments, or tags snapshots an immutable `prompt_versions` row with its author, and approval stamps bind to the specific version approved (approving v5 never alters v4's recorded approval). Editing the content or arguments of an approved global or persona prompt does not change what is served: the edit lands as a pending draft version (manage_prompt update returns status pending_approval; a gated content edit cannot be combined with scope/status changes in one call) and the approved snapshot keeps serving until an admin approves the draft via `POST /api/v1/admin/prompts/{id}/versions/{version}/approve` (reject with `.../reject`; full history with author and content via `GET /api/v1/admin/prompts/{id}/versions`; `GET /api/v1/portal/prompts/{id}/versions` serves history to any caller who can view the prompt (own personal prompts and enabled shared prompts), since history is the library's verification surface; non-admin viewers of a shared prompt get the served history only: applied snapshots in full, the pending draft as a content-redacted stub, and rejected/superseded drafts omitted). The portal prompt page renders this history with per-version approval provenance, flags a pending draft (readers keep being served the approved version), and diffs any version against the current content as a line diff; it also shows point-of-use invocation help (a copyable natural-language invocation built from the stable name and required arguments, resolved by agents via `manage_prompt use`). Metadata-only edits (tags, category, description, display name) apply directly; personal prompts version silently. Served prompts carry provenance: `prompts/get` stamps `prompt_version` / `prompt_approved_by` / `prompt_approved_at` / `prompt_reference` into `_meta`. Usage stats are aggregated from `prompt_serve` audit events (emitted on every database-prompt `prompts/get` and resolved `use`, within the audit retention window): `manage_prompt get` and `list` report `run_count` and `last_run_at` per prompt (`list` also includes the shared `collections` list when the store supports collections, so MCP clients see the same organization model as the portal), and `GET /api/v1/admin/prompts/usage` / `GET /api/v1/portal/prompts/usage` return the per-prompt rollup (portal scoped to the caller's visible prompts, including prompts shared person-to-person with them) for library curation and dead-prompt detection. A prompt also carries the reference material its procedure depends on (#1013): attachments are ordered links from `prompt_resource_attachments` to managed resources, stored by resource id so editing the uploaded file updates every prompt that attaches it, and deliberately without a foreign key so deleting a resource leaves the link visible as broken rather than silently erasing the evidence that the SOP is incomplete. A resolved prompt delivers them after the prompt text: textual material at or below a 64 KiB threshold as an MCP `EmbeddedResource` carrying the contents, anything binary or larger as a `ResourceLink` the client reads on demand, with a framing message stating the material is authoritative (fill an attached template, follow an attached checklist); `manage_prompt use` additionally lists them in an `attachments` array (uri, media type, size, availability, and inline content for embedded items) so the agent can state what it received. An attachment must be at least as widely visible as the prompt that carries it: global resources attach anywhere, a persona resource to personal prompts and to persona prompts scoped to exactly that persona, and a user-scoped resource only to its owner's own personal prompts. The rule is enforced at attach time (`POST /api/v1/portal/prompts/{id}/attachments`, admin-prefixed equivalent exists; `PUT` reorders, `DELETE .../{resourceId}` detaches, `GET /api/v1/portal/resources/{id}/prompts` answers what depends on a resource) and again on every prompt write through the shared store wrapper, so a scope edit or promotion request that would strand an attachment is refused with a message naming the resource, whichever surface it arrives from. At serve time it is re-checked per caller with `resource.CanReadResource`: an unreadable or deleted attachment is reported only as an aggregate count of undelivered materials, never by name and never by contents, and the prompt still serves. A prompt also references the managed scripts its procedure depends on (#1289): ordered links in `prompt_script_attachments` storing the canonical `mcp:script:<id>` reference rather than a bare id, because that reference is the platform's one way to name a script from outside `pkg/script` (the same string `search` emits and `fetch` dereferences), and again deliberately without a foreign key so deleting a script leaves the reference visible as broken. `manage_prompt attach_script` / `detach_script` take `script` as either that reference or a bare script id and normalize to the reference. Serving a prompt delivers each referenced script's contract plus the instruction to call `run_script` for fresh output, and `manage_prompt use` lists them in a `scripts` array; serving NEVER executes a script, because a prompt read is a read path and running code from it would blur audit attribution and turn every read into a potential asset write. A reference resolves for the script's owner and for nobody else; referencing is allowed from any prompt and the response names who it will resolve for when the prompt serves anybody else, and only a script the caller can see may be referenced, with the per-caller check applied again at serve time, where an unreadable or deleted reference is reported only as an aggregate count with the prompt still serving. The rule itself is one implementation for both kinds (`prompt.CheckAttachScope` over an `AttachmentScope` carrying a SET of audience ids), because a script may serve several personas where a resource serves exactly one, and a single id cannot answer whether the prompt's audience is contained in the material's. In MCP Apps-capable hosts, the built-in `prompt-browser` app (#1011) is bound to the presentation-only `show_prompts` tool: a host renders an app on every call to the tool it is bound to, so the browser lives on a tool whose only job is to open it for the human rather than on `manage_prompt`, which the agent calls throughout its own work. `show_prompts` opens the browser for the human (search, buckets, collection/tag facets, argument forms, preview with provenance, and a Run action that resolves via `use` and inserts the rendered prompt when the host supports `ui/message`); the rendered app populates itself from its own `manage_prompt` calls. `manage_prompt` carries no app and renders nothing, and its JSON results stand alone in non-app clients.

## Automations

The portal's view of automations, for the people who own them (#1290, #1307). An automation is work that runs on a schedule or when someone asks; an agent builds one as a managed script, and every automation is a script today (#1912). The section lives at /portal/automations (the administrator's at /portal/admin/automations); the old /portal/scripts paths, including a run link mailed with a failed run before the rename, redirect to the same page with the rest of the path, query and hash intact. The detail page of one script is titled Script and backs out to Automations. The listing's first column is Automation and a Kind column states each row's kind as Script. A Grid/Table control switches the listing to a grid (#1909), remembered per browser and opening on the table: each card's tile is the script's flow diagram, drawn light and dark by the thumbnail worker (internal/platform/thumbworker, a scripts kind over internal/platform/scripttiles and the script_tiles table, migration 000164) whenever a new version is saved, recorded as not drawable with the parse error when the latest version does not parse, and removed with a deleted script; served at GET /api/v1/portal/scripts/{id}/thumbnail (?variant=dark) to everyone who can read the script. The failed-run email reads "A scheduled automation failed". The section is three tabs: Automations, Schedules (#1891) and Runs (#1405). The Schedules tab draws when every visible schedule fires, one row per script with a mark at each fire, on one of three axes chosen by how often it fires over four weeks: Intraday (the viewer's today, more than 28 fires), Multi-day (the viewer's Monday-to-Sunday week, at least four) and Long-term (three months from the first of the viewer's month, fewer); a row is colored by its rhythm (the median gap between fires: minutes, hours, days, weeks, months) from the --chart-1..5 theme tokens, a paused row is muted and keeps its marks, a row over 100 fires draws 1.3px ticks and one with at most 12 irregular fires draws dots, hover or focus states a fire in the viewer's and the schedule's zones, and a click opens the script; with nothing scheduled the tab shows an empty state that returns to the Scripts tab. Every fire is expanded server-side by GET /api/v1/portal/scripts/fires?tz=<zone> (internal/httpserver/scripthttp/fireshttp over script.Cron.FiresBetween, the scheduler's own robfig/cron parse), never in the browser. The listing shows every script the caller may see with, per row, what the script is called (with a badge beside the name where it will execute nothing — disabled, or its lifecycle status — because the exception is what a listing is scanned for, while the version a run executes is true of every healthy script and belongs on the script's own page (#1407)), the cadence and next fire of its schedule, stated in words ALWAYS (#1405) — the sentence the schedule editor states, plus the step cadences an agent writes that the builder has no control for ("Every 30 minutes, UTC"), and a named custom cadence for what is left — because this is the column an owner scans to answer what is running and when; a cron expression appears only in the schedule editor, where somebody is editing one, and both the listing and the editor read what the schedule is doing off one `scheduleState` so neither can call a paused schedule due (a paused schedule says so rather than showing a fire that will not happen; a script without one runs on demand) (#1358), and the state of its most recent run. Above the table is ONE health line (#1795): the total, the scheduled count and the failed-last-run count, of which only the last is a control (pressing it filters to those scripts, `aria-pressed` while it does) because it is the one number a person opens the page for. The total and the scheduled count are the SERVER'S, counted over the predicate with no limit (`Store.Count`, `Store.CountScheduled`), and the line says "showing N" beside the total when the cap truncated the page; the failed count is the page's, and is the one that cannot be otherwise, since a run is attached per page and only for rows the caller owns. They replaced three bordered tiles that counted the ROWS ALREADY LOADED while the chips beside them counted the platform, so past the 200-row cap the two disagreed and nothing said the page was truncated (`total` was `len(rows)`). Under the line is ONE filter bar in the shape the assets browser established: Mine/All scope tabs (a non-admin's; an administrator's listing is every script by construction), free text, and a `FilterSelect` each for author, category, tag and status — tag being a select rather than a chip cloud because it is the least discriminating axis and the only one with no bound on its vocabulary, its options ordered by count exactly as the chips were. Every axis is a SERVER predicate (`GET /api/v1/portal/scripts?search=&category=&tag=&owner=&status=&enabled=&scope=`, over `script.ListFilter`) so the narrowing covers every script the caller can see rather than the capped page the listing holds. ORDERING is the server's too: `sort` and `dir` carry a whitelisted column (`script.SortColumn`: name, display_name, owner_email, created_at, updated_at) into `ORDER BY` ahead of `LIMIT`, with an id tie-breaker, and an unrecognised column falls back to updated_at DESC rather than erroring, because the caller of a listing is a page of a portal and a typo in a query string is not worth a blank screen. The Script, Author and Updated headers sort through `SortableHead`; LAST RUN does not and carries no affordance, because it is attached per page after the query and an ordering over it would be an ordering over the page. A script is VISIBLE to everyone and READABLE as before: a non-admin's `scope=all` lists every script, with `reportableScript` blanking the source of a row the caller does not own, `attachLastRuns` filling runs only for owned rows, and no action offered on one; `scope=mine` is the default and is what the listing did unconditionally before. The same widening reaches `manage_script command=list`, deliberately — an agent asked to write a weekly report should be able to see that one already exists — and the tool's projection carries no source and must keep not carrying it. The second tab is every run across the caller's own scripts, newest first (`GET /api/v1/portal/scripts/runs`, bound to the caller's owned script ids through `script.RunFilter.ScriptIDs` — an empty, non-nil set for a caller who owns nothing, so a listing across every run on the platform is unreachable; an administrator reads every run, the same reach their script listing has): what triggered each run, how it ended, how long it took, and the reason a failure failed, wrapped in the row rather than behind it, with the row linking to the run and the script's name linking to the script. A run has an address of its own, `/scripts/{id}/runs/{runID}`, which opens that run in its script's history rather than in a page of its own. Opening a script shows its Details — owner, the version that runs, the schedule, next fire, lifecycle status, and the typed parameters a run binds against, read in that one section rather than in a card apart from it (#1406) — which is the SAME document `search`, `fetch` on an `mcp:script:<id>` reference, and a prompt that references the script serve, so the page and an agent describe a script identically; where the run gate would refuse a run requested now, the page carries the gate's own refusal (`script.RefuseRun`) rather than re-deriving runnability from a status and an enabled flag. The page is ordered for the person debugging a script (#1406): Details, the schedule controls, About (the description as the markdown document it is), the Source, and the run history directly beneath the Source, because an error in the history is answered by the text above it. The code is one card with two tabs, Flow and Source, and Flow opens first for owners and readers alike (#1906). Flow is a diagram the platform derives from the saved version's source (internal/platform/scriptflow, served by GET /api/v1/portal/scripts/{id}/versions/{version}/graph under the source's read rule and GET /api/v1/admin/scripts/{id}/versions/{version}/graph; no MCP tool carries it, an agent reads the code): every platform.* call except platform.progress is a card, colored by role (input, reads, writes, output); edges follow values, a value carrying the set of steps and run.state it was computed from, with the pure functions it passed through named on the edge, and an edge a longer path implies removed; a user function is expanded at each call site (except inside its own expansion: the interpreter refuses a recursive call when it is made), a function holding steps is one box per enclosing box captioned with its author's comment, and a function whose only effect is one platform call is folded into its step as a chip; module constants are shown by value and string building rendered with each computed part as {its source}, and a connection, destination, tool, table or target the source computes marks its card dashed and is never guessed; parameters are listed beside the diagram with the steps whose arguments their value reaches, and a parameter a condition reads is marked as deciding which steps run; run.state is an input node and platform.save_state draws a dashed edge back to it. Selecting a card fills a side panel (what it reaches, what feeds it, what it feeds, the loop, the helper and its call site, a source excerpt); double-clicking it opens Source with its lines marked, and lines selected in Source light the cards they produce. For the owner and administrators the Flow tab opens on the latest run drawn on the diagram (#1907), with a Run menu choosing any run in the history or none: each card states its calls and their duration, the rows it exported, or that it ran; a card the run never reached is dimmed; the card a failed run stopped at is drawn in the error color with the run's cause and error in the side panel (when the run failed in the script's own code just after a call failed, such as fail() on the line after a call answered 500, the card that made that call from the same function is the one marked, #1933); a card's run chip counts its failed calls and shows the failed call's message on hover; a call that failed in a run that went on counts on its card as a failed call without marking the card failed; and the panel counts the run's calls, those on cards, and those no card made, which it lists. A run of an older version is drawn on that version's diagram. The attribution is exact rather than guessed: the script host records the call site of every host call (internal/scriptcallsite: the line:col of each script frame on the Starlark stack, outermost first, at the call's opening parenthesis), sends it on the in-process MCP request's _meta under mcp-data-platform/call_site, and the audit middleware stores it in audit_logs.call_site (migration 000163) for calls whose source is script; the output writer stores it on each run output, and a failed run's backtrace names the site it stopped at. The graph records the same positions on each card, so a helper called from two places is two cards with their own calls. Served at GET /api/v1/portal/scripts/{id}/runs/{runID}/flow to the owner and administrators, reading at most 10,000 of the run's calls and answering calls_truncated past that. Each version older than the running one has a Compare with button (#1908) that opens the pair as a diagram and a text diff: the running version's graph marked against the older one (?compare=<version> on the graph route), a step matched by what it reads, writes or produces rather than by its line, so a card is added, changed (carrying what it was) or removed (the older version's card, dashed, with its edges), and moved code is not a change. A version that no longer parses answers ok false with its findings and empty lists; a script with no platform calls answers nodes, edges, groups and params as [] rather than null; the graph is cached per distinct source. The same constant resolution now feeds manage_script validate: a name the module binds once to a literal, or string building over literals and such names, is read as its value into connections, tools, destinations and refresh_targets without setting the matching dynamic_ flag. Two of those sections fold (#1407): the schedule is FOLDED by default and states what the script runs in its header ("Runs: Every weekday at 7:00 AM, America/Los_Angeles", or "Not scheduled", or the same sentence prefixed "Paused" for a schedule firing nothing), with the builder and the fire bindings behind the reveal and pause/resume on the header either way, because a cadence is set once and read constantly and a form nobody is filling in pushes the code and the run history off the screen; About starts OPEN, because what a script is for is what a reader came for, and folds to the first line of prose in its document when the description has grown into one. Details states the cadence in the same words the listing and the folded header do rather than as the expression the platform stores, so the page carries a cron expression in exactly one place: the editor. For a script the caller owns, Run and Dry run sit side by side above the editor over ONE parameter form they both bind — Run executes the saved version, a dry run executes the text on screen — and a script the run gate would refuse carries no Run control at all rather than a button that fails when pressed; the version history folds into that same section behind a collapsed reveal, each version with its author and the roles that author held at the save, which are the roles a run of that version presents, and nothing opens by default because the version that runs is the text in the editor directly above it. The run history carries each run with its trigger, the version it executed, its duration, its outputs, its failure reason where it failed, and, opened in place, its bound parameters, its cost in steps and queries and exports, its outputs, and the bounded log it printed. The run rows are composed rather than laid out flat: how a run ended and when it ran are one fact and are set as one, while the trigger (a short enumeration) and the version (the same number down the whole column) qualify it from underneath rather than each holding a column open, which is what lets the history fit the width the page has instead of scrolling sideways (#1362); a failure message wraps to as many lines as it needs, because a table cell does not wrap by default and a Starlark traceback on one line put the page's own reason for existing off the edge of it (#1406). A `skipped_overlap` row is shown as a distinct state, neither success nor failure: it names a fire that never executed because the previous run was still going, which is exactly what a report that stopped producing has to be able to show. An output written to the portal links to the asset version it produced — a recurring script writes new versions of ONE asset keyed `script:<id>:<name>`, so that asset's version history is the script's refresh history — while an object delivered to a configured bucket destination names its bucket and key and is deliberately not a link, because the platform wrote those bytes and does not hold them. Visibility rules, shared with every other surface: a script's definition -- its contract, its source and its version history -- to everyone signed in (#1866), because a script is how a resource or an asset was produced, with the roles each version's author held withheld from all but the owner and administrators (`manage_script` `get`, `get_content`, `outline`, `stats`, `locate`, `diff` and `versions` address another person's script with `owner_email`; `get` leaves out `live_runs`); the run history, the produced list, the state and every action to that owner and to administrators, because a log is free text the script printed while presenting its author's captured roles and may echo rows the reader has no access to of their own; and one particular run additionally to whoever requested it, since the result was handed to them when they asked for it. A command that acts on another person's script is refused naming that rule. The pages write six things and none of them is an authority beyond the author's own (#1307, #1363, #1364, #1369, #1575). The schedule: on a script they own, at any scope, a caller sets or replaces the cron expression, the timezone it is read in, and the values every fire binds, and pauses or resumes it — over `GET`/`PUT /api/v1/portal/scripts/{id}/schedule` and the enable/disable pair, inside the portal's own authentication and CSRF handling, restricted to the owner and to administrators and refusing a non-owner exactly as it refuses a caller who may not see the script. Pausing is its own route rather than a field of the schedule, because re-sending the schedule to turn it off would re-base the fire it resumes on. A schedule saved against a disabled or retired script saves and stays inert, and the page says so in the gate's own words. The schedule is asked for in the terms a person has it in — hourly, daily, weekdays, chosen days of the week, a day of the month, a time, a timezone — and the cron expression is DERIVED and shown rather than asked for, because cron is a precise notation for people who already know it and the owner of a report is not required to be one of them; a Custom field keeps every expression reachable, and an expression an agent wrote through manage_script opens there as itself rather than being rewritten into something near it. The SOURCE: the code is editable in the portal's own editor with Starlark highlighted as the Python dialect it is, and the edit crosses `script.ApplyEdit` — the one gate every mutation surface crosses — landing on the live row as the version that runs and recording the roles its editor held, which is what a run of it presents; the source is parsed before anything is stored. A RUN of the latest saved version (`POST /api/v1/portal/scripts/{id}/runs`), which queues exactly what `run_script` queues under the same gate, worker and principal and records `portal` as the trigger. A DRAFT run of an edit, executed as the caller with the draft limits, persisting nothing it produced and refusing every write it reaches through `platform.call` (#1664) — a "Write for real" control beside the Dry run button lifts that for one run, and the report then lists what the run persisted. And what the script SAYS about itself — display name, markdown description, category, tags — which is not an input to any decision the platform makes and is still captured as a version. And the DELETE (#1575), at the bottom of a script the caller owns and of every script for an administrator: `DELETE /api/v1/portal/scripts/{id}`, authorized the way the tool's delete is — the owner, or an administrator over any script, with a caller who may not see the script answered as one who named a script that does not exist — and calling the same `Store.Delete` `manage_script command=delete` calls, so the cascade cannot come out differently depending on which surface asked. The confirmation names what goes with the script rather than asking whether the reader is sure: every saved version, the schedule stated in words, the whole run history, and the state it carried, with the schedule and the state named only where the script has them. It names what stays for the opposite reason — the assets and resources the script wrote remain, owned by whoever owns them, and the producer records naming it as their writer remain with them (#1569) — because "delete the script" reads to a lot of people as "delete the reports it wrote". The delete is recorded in the audit log as a `script_delete` event of kind `admin` naming the script and the owner who lost it, whether the removal succeeded or failed at the write, the way the owner transfer is; a caller the route refuses is answered before the script is read and no event is written, which is what every other portal script route does, and a delete made through the tool is already in the log as the `manage_script` call it was, refusals included. After it lands the reader is back at the script listing. `show_scripts` opens the pages for a human and performs no data work, following the `show_prompts` split (#1040) so an agent's own `manage_script` calls never render UI; it returns a confirmation and, where the deployment is configured with its public address, a link.

---

# Registered Tables

A file already in object storage -- a CSV somebody uploaded as a managed resource, or one `trino_export` wrote as a portal asset -- is made readable as a Trino table by pointing an external table at the directory the file already sits in (#1327). Nothing is copied and nothing is ingested: a `CREATE TABLE ... WITH (external_location = 's3://.../', format = 'CSV', skip_header_line_count = 1)` over the object's own directory, so the table reads whatever the object holds now and a vendor drop that overwrites its file changes the query result on the next run with no further action. The join was always the easy part; loading was the gap, because Trino has no client-side upload and an agent emitting `INSERT ... VALUES` batches for a fifty-thousand-row CSV is slow, error-prone, and burns context.

## What an operator configures

A Trino connection gains `scratch: {catalog, schema}` naming where registrations are written. A connection with no such block cannot hold one and the surfaces do not offer it there; both keys are required, and a block naming only one is ignored with a warning rather than half-honored, because a registration built on it would fail at the DDL. Per-instance `Config` is discarded after `NewMulti`, so the target is kept per connection on the toolkit the way the read-only flag is, maintained by `AddConnection` and `RemoveConnection` so a connection added through the admin API can hold a table from its first call rather than from the next restart.

Behind the connection is a Hive catalog over the same object store the platform's resources and assets live in -- a file metastore is enough (`hive.metastore=file`, `hive.metastore.catalog.dir=s3://<bucket>/trino-metastore/`, `fs.native-s3.enabled=true`, `s3.path-style-access=true` for an S3-compatible store), with Trino's OWN credentials to the bucket, separate from the platform's S3 connection. `hive.recursive-directories` stays false. Dynamic catalog management (`CREATE CATALOG`) is deliberately not the mechanism: it only creates connector instances, it is marked experimental, and the full statement including credentials is logged. One properties file is the whole setup.

## Two consequences a reader has to know

A CSV TABLE'S COLUMNS ARE ALL `VARCHAR`. That is the Hive CSV storage format's rule and not a platform choice -- declaring a CSV table with any other type is refused by Trino itself ("Hive CSV storage format only supports VARCHAR (unbounded)") -- so a join to a typed warehouse column needs a `CAST`. Search hits, fetch documents, the portal panel and `manage_table action=list` all carry a sample statement showing it, because a reader who writes the obvious join otherwise gets a type error that explains nothing. A JSON-lines or Parquet table declares each column's type (#1833), its sample is the plain join, and every surface carries the declared types as `column_types`.

A DIRECTORY HOLDING A SIBLING TRINO WOULD READ IS REFUSED, BY NAME. Trino reads every non-hidden object under an external location and parses it as CSV WITHOUT ERRORING: a `thumbnail.png` beside a `content.csv` comes back as rows of PNG bytes. There is no error to lean on, so the platform's refusal is the only protection, and it names the object in the way rather than reporting a generic failure. Hive skips names beginning with `.` or `_`, which is why the portal's derived thumbnails are written as `.thumbnail.png` / `.thumbnail_dark.png`, and the sibling check applies that same rule so a CSV asset registers with its thumbnails in place; nothing derives a thumbnail key to READ one (an asset row records the key that was written), so thumbnails captured under the older names keep serving, and the prune path covers both spellings. An asset still carrying a legacy thumbnail IS refused, because those names are ones Trino reads; the thumbnail refresh queue offers such an asset for capture (a recorded key under a legacy name is pending however current its version) and the capture endpoint deletes the object it supersedes once the row points at the new key, so having the portal open anywhere clears it with no bucket work. The same rule applies to the file being registered: a source object whose own name begins with `.` or `_` is REFUSED, because the table would be created, recorded and queried without error and return nothing. A subdirectory is invisible while `hive.recursive-directories` is false, so a resource's revisions under `v/` and an asset's versions under `<headDir>/<versionID>/` are never read.

## CSV or JSON lines (#1820)

A registration reads its file through the reader its FORMAT names, decided by the file's content type or, when that is generic, its extension (`.csv`; `.jsonl` or `.ndjson`) -- and a `.jsonl`/`.ndjson` file typed `application/json` is JSON lines too, since one record on one line is also a JSON document and detection names it that -- and recorded on the registration as `format` (`table_registrations.format`, migration 000151, default `csv` for every row written before it) because the CREATE TABLE is written again at every follow and when a failed follow puts a table back. A JSON-lines table is `format = 'JSON'` with no header to skip and no escape to declare; its columns were declared `VARCHAR` like a CSV's until #1833, which types them from the values (see "Typed JSON lines and Parquet (#1833)" below). What each format carries was established against Trino 453: a JSON-lines table returns every string byte for byte -- LF, CR, CRLF, NUL, backslash, quotes, a leading `=`/`+`/`-`/`@`, tabs, emoji, private-use code points, U+2028 -- and a null or missing key as NULL, where a CSV table cannot carry a line break inside a value (refused unless `repair`, which joins the value's lines with single spaces after trimming each and dropping blank ones) and reads a null back as an empty string; no other character is altered by either. Columns are every key any record carries, lowercased, in first-seen order, because the JSON reader matches a key to its column without regard to case; a key cannot be renamed the way a CSV heading is, since a renamed column is one the reader never fills. `internal/platform/tablejsonl` refuses, by line number, the shapes the reader fails on or reads wrongly: a blank line anywhere but the end, a line that is not exactly one JSON object (a second object on one line is DROPPED without error, the one silent case), a key repeated in a record exactly or by case, a key whose values cannot be one type (#1833; until then every nested object or list value was refused, because the reader fails a query on one in a `VARCHAR` column), and a key that is non-ASCII (its column reads empty on every row), holds a comma, or has leading or trailing space (both refused by Hive as column names). A JSON-lines file is never corrected: `repair` does nothing for it. A file with no records declares no columns and is refused at registration, while a later version of a following registration's file that holds none moves the table and keeps its columns, so a day with nothing in it reads as zero rows. `trino_export` and a script's `platform.export` write the format as `format=jsonl` through one formatter (`pkg/toolkits/trino/format.go`): one object per row, keys in column order, a list or dict value written as its JSON text so the reader can read it and `json_parse` restores it, and a string that is not valid UTF-8 refused rather than replaced, since encoding/json would substitute U+FFFD silently. The portal's Query as a table panel appears for a JSON-lines file as it does for a CSV.

## Typed JSON lines and Parquet (#1833)

A JSON-lines registration types each column from EVERY record (`internal/tabletype.Inferrer`, shared with a script's Parquet export so the two cannot disagree): all booleans `BOOLEAN`; all integers within int64 `BIGINT`; any number with a fraction or exponent, alone or among integers, `DOUBLE` (not `DECIMAL`: a JSON number carries no scale and one wide value would widen the whole column); all strings, a mix of scalar kinds, an integer too large for `BIGINT`, or nothing but nulls `VARCHAR` (verified on Trino 453: the JSON reader returns a number's or a boolean's text in a `VARCHAR` column, and a mixed list reads as `ARRAY(VARCHAR)`); an object a `ROW` over the union of its keys, lowercased, each inferred the same way; a list an `ARRAY` of its elements' type. A key holding an object on one line and a list or scalar on another is refused naming the key and line. No date or timestamp is inferred from a string. A NESTED field name may hold only letters, digits, `_`, `.`, `$` and spaces (`tabletype.CheckFieldName`): the metastore keeps a nested type as Hive type text, and on Trino 453 a table whose ROW declares `b-c`, `a:b` or a quote is CREATED and then fails every statement with "Could not read table schema", its DROP included, so the refusal is what keeps an undroppable table out of the scratch schema. A JSON-lines registration made before this change carries `all_varchar` (`table_registrations.all_varchar`, migration 000152, set true for every existing `jsonl` row): a follow keeps its columns `VARCHAR` and refuses a nested value as before, so a query written against it keeps working; registering the file again under the same name makes a typed registration.

A PARQUET registration (`format = 'PARQUET'`, `internal/tableparquet`) reads the file's footer and nothing else: a ranged read learns the size, the last eight bytes give the footer length, a second ranged read fetches it (`ObjectReader.GetObjectRange`, over `GetObjectRange` added to txn2/mcp-s3 for this), so a Parquet file larger than the cap a CSV is read whole under still registers. Detection: content type `application/vnd.apache.parquet` (aliases `application/x-parquet`, `application/parquet`) or the `.parquet` extension; `pkg/contenttype` names Parquet from the `PAR1` magic, and with the whole payload in hand holds it to `PAR1` at both ends; a file named `.parquet` whose bytes are not Parquet is refused as not Parquet. Mapping: BOOLEAN -> BOOLEAN; INT32 (none / INT(32,signed)) -> INTEGER; INT(8|16,signed) -> TINYINT/SMALLINT; INT64 (none / INT(64,signed)) -> BIGINT; INT(8|16,unsigned) -> SMALLINT/INTEGER; FLOAT/DOUBLE -> REAL/DOUBLE; DECIMAL(p,s) on any physical type -> DECIMAL(p,s) (p <= 38); BYTE_ARRAY STRING/ENUM/JSON -> VARCHAR; plain BYTE_ARRAY -> VARBINARY; INT32 DATE -> DATE; INT64 TIMESTAMP (any unit) and INT96 -> TIMESTAMP(6); LIST -> ARRAY (the format's backward-compatible element rules), MAP -> MAP, group -> ROW; converted-type-only files map the same. TIME, UUID, FLOAT16, INTERVAL, plain FIXED_LEN_BYTE_ARRAY, an unsigned 32- or 64-bit integer (a UINT32 declared BIGINT reads 4294967295 as -1; a UINT64 declared DECIMAL(20,0) is refused by the connector), a repeated field outside LIST/MAP are refused naming the column and its Parquet type -- no partial registration -- as are two columns one apart by case (the reader matches by name case-insensitively) and a nested name the metastore cannot store. Every timestamp column is declared `TIMESTAMP(6)`: a Hive table's timestamps are read at the catalog's `hive.timestamp-precision` and the connector refuses a column declared at any other ("Incorrect timestamp precision for timestamp(3); the configured precision is MICROSECONDS"), so the scratch catalogs set `hive.timestamp-precision=MICROSECONDS`, which is also what reads a microsecond Parquet timestamp back exactly (the default MILLISECONDS truncates it). `repair` has nothing to correct in a Parquet file and is accepted as a no-op. A follow re-reads the new version's footer; an added, removed or retyped column is reported in the write's `table_changes` (`FollowOutcome.ColumnChanges`: "added region VARCHAR; id is now DOUBLE (was BIGINT)"), and a version that is not Parquet or holds a refused type sets `follow_error` and leaves the table where it was.

EXPORT: `trino_export format=parquet` and a script's `platform.export(format="parquet")` write Parquet through one formatter (`pkg/toolkits/trino`, `tableparquet.Write`, parquet-go, ZSTD). `trino_export` reads the query with `QueryOptions.RawValues` and the column precision and scale `ColumnInfo` now reports (both added to txn2/mcp-trino, whose `convertValue` also stopped cutting fractional seconds with RFC3339), so a `TIMESTAMP(6)` keeps its microseconds, a `VARBINARY` its bytes and a `DECIMAL` its declared precision. Trino to Parquet: BOOLEAN; TINYINT/SMALLINT/INTEGER -> INT32 INT(8|16|32,signed); BIGINT -> INT64; REAL/DOUBLE -> FLOAT/DOUBLE; DECIMAL(p,s) -> DECIMAL(p,s) on INT32 (p<=9), INT64 (p<=18) or FIXED_LEN_BYTE_ARRAY; VARCHAR/CHAR/JSON/UUID/IPADDRESS/TIME/interval -> STRING; VARBINARY -> BYTE_ARRAY; DATE -> DATE; TIMESTAMP(p) -> TIMESTAMP MILLIS (p<=3) or MICROS, not adjusted to UTC; TIMESTAMP WITH TIME ZONE -> TIMESTAMP MICROS adjusted to UTC, the zone not kept and the result message saying so; ARRAY/MAP/ROW -> LIST/MAP/group, fields lowercased. The result message names every column that lost something, at any depth (a nested value loses what a top-level one does): a zone, a timestamp finer than the microsecond a table reads, or a type written as text. A `TIME WITH TIME ZONE` written as text keeps its offset, which the CSV and JSON exports of the same query also keep. `Write` holds every column and ROW field name to the rules `Columns` reads them by -- non-empty, no comma, no edging whitespace, ASCII, no two one apart by case, a ROW field only letters, digits, `_`, `.`, `$` and spaces -- and refuses before writing, so an export cannot store a file the registration it was made for would refuse. Columns and ROW fields are written in declaration order through an ordered group node, because parquet-go's `Group` is a map that lists fields alphabetically. Every file written registers. A script's rows carry no types, so the Parquet formatter infers them with the JSON-lines rules (a dict a ROW, a list an ARRAY), and a key that cannot be one type -- two of a dict, a list and a scalar across rows -- fails the export naming it; `register=` accepts `parquet`. The script's export record carries `column_types` as `{"name", "type"}` entries, the shape every other surface uses.

VIEWERS: a JSON-lines asset opens on a Table view (the records' key union, first-seen order and spelling, the CSV viewer's search, sort and row dialog, a nested value as one-line JSON in the cell and indented in the dialog) when every line is an object of scalars or flat values and the records carry most of the columns, and on the Records view otherwise; a toggle switches. A Parquet asset or resource, and a public share of one, opens in a viewer that reads the file by HTTP range through hyparquet: a schema panel with each column's Parquet type, nullability and the Trino type a registration declares (the same mapping, in TypeScript), the row count, row groups, size, codecs and writer, and one row group's rows at a time. The tile is the content-type icon.

## A CSV a line-based reader cannot read (#1441)

A BACKSLASH IS AN ORDINARY CHARACTER (#1819). Registration declares `csv_escape = U&'\0000'`, the value that tells the Hive CSV reader to escape nothing; its default escapes with a backslash, a dialect no writer of these files uses, under which `back\slash` read as `backslash` and a quoted cell holding a backslash beside a quote read as an empty string, row and column counts intact. `csv_escape='"'` does not help. NUL cannot collide with a value because a file with a NUL byte is refused before registration. A table registered before this keeps the reader's default until its file gets a version it follows or it is registered again.

Trino's Hive CSV reader is LINE-BASED: the text input format splits records on `\n` before the quote-aware serde sees them, so a quoted line break TEARS one record into fragments, the first fragment ends on an unbalanced quote, and every field after it lands in the wrong column. This is the ordinary shape of a spreadsheet export -- a multi-line address in one cell -- and it parses without error under Go's `encoding/csv` and Python's `csv`, so nothing about the file looks wrong. The table is created, the row is recorded, the query returns rows, and the rows are fragments: there is no error anywhere on the path, which is why the file is inspected before the DDL rather than trusted. A real 34 KB file with 157 records of 27 fields each, 153 of them carrying an embedded line break, registered as a 530-row table -- 530 being exactly the count of `\n` bytes after the header -- with only the 2 records containing no line break correct.

The bytes to see it were already in hand: `columnsFor` did a full `GetObject` and handed the whole body to `ReadHeaderColumns`, which read one record and discarded the rest. It now returns the body (`contentFor`) and the body is INSPECTED. Three conditions refuse, all wrapping `ErrNeedsRepair` as well as `ErrRefused`: lines that end in a bare `\r` and never in `\n`, any field containing `\r` or `\n`, and bytes that are not valid UTF-8 (which reach every cell as replacement marks). The refusal names how many rows carry a break and which columns they are in, using the SAME column names `ReadHeaderColumns` would declare. A parse failure part-way through STOPS the scan rather than failing it -- what was read is still evidence, and a file this reader cannot finish is one that registers today, so refusing it on that ground alone would take away a working registration.

A LEADING UTF-8 BYTE-ORDER MARK IS DROPPED BEFORE ANYTHING READS A RECORD (#1774). It is not an encoding declaration to any reader here -- it is the first bytes of the first field -- so `encoding/csv` reads `<mark>"Post ID"` as an UNQUOTED field carrying a bare quote and fails the parse on line 1. A header whose first field is unquoted parsed and had the mark trimmed off the NAME afterwards (`tablecsv.ColumnsFrom`), which is why the defect was invisible for as long as it was: it only appears when the export QUOTES its strings, which Facebook Insights, Excel's "CSV UTF-8" and Google Ads all do, so no export from those surfaces could be registered at all. `ReadHeaderColumns` turned every read error into `ErrEmptyHeader`, so the answer was "the file has no header row, so the table has no column names" about a header the file has and the portal's own CSV viewer renders; `tablecsv.Inspect` runs first and died on the same line, recorded it as `Unreadable` and returned NIL, so the defects BEHIND the mark -- the line breaks inside cells the correction exists for -- were never found and `repair: true` did nothing. `tablecsv.TrimBOM` now drops it at the top of `Inspect` and `ReadHeaderColumns`, and only for bytes already established as UTF-8, because a code page assigns those three bytes three of its own characters and `decodeToUTF8` carries them rather than dropping them. A first line the reader cannot parse AT ALL is now a defect in its own right rather than nil: such a file never registered -- the registration meets the same line three lines later -- and the refusal carries the parse error ("this file's first line cannot be read as a CSV header (...)"), never "no header row". `ErrEmptyHeader` is kept for its one true case, a file with nothing in it.

A COMMA IN A HEADING DOES NOT SURVIVE INTO THE COLUMN NAME (#1774). The Hive metastore stores a table's column list comma-separated, so the connector refuses a name holding one whatever the quoting -- `USER_ERROR: Hive column names must not contain commas` -- and a Facebook Insights export names a column `Reactions, Comments and Shares`, so no export from that surface could be registered even once the byte-order mark stopped hiding its header. It failed at the DDL, which is AFTER the correction of the file's rows had already been saved. `tablecsv.withoutCommas` turns each comma into a space and collapses the runs, in `ColumnsFrom`, on the same footing as the blank name filled in positionally and the repeat that is suffixed. Which characters Hive refuses was established against Trino 476 rather than assumed: a space, a dot, a colon, a parenthesis, a slash, a semicolon, a percent, an equals and a hyphen were each created without complaint, and the comma alone was refused, which is why it alone is dropped.

A RECORD SHORT OF THE HEADER IS NOT A DEFECT (#1779). A CSV record may end before the header does -- an exporter writing nothing rather than trailing commas -- and those columns are ABSENT rather than wrong. Every reader supplies them: Go's encoding/csv and Python's csv return the short record, PapaParse fills the missing keys, Excel and Preview open the file, and the Hive CSV reader a registered table is served by reads the values it finds and leaves the rest null. Refusing such a file put this platform alone among them: a Facebook Insights export omitting two trailing columns for one post type (21 of 178 records) could not be registered, AND the short records withdrew the offer to correct the 148 torn rows it also had (Correctable() is false while Ragged is non-empty, #1449), so the file had no path forward at all. `noteRagged` and `checkFieldCounts` now flag only a record carrying MORE fields than the header, where dropping one to fit would lose a value the file holds; the correction writes a short record back as it came. `padding a short record invents data` was the stated reason and does not hold -- an absent trailing value is not invented by being read as empty.

THE CSV VIEWER SHOWED THE PARSER'S ROWS AND DISCARDED ITS ERRORS (#1779). `CsvRenderer` parsed with `Papa.parse(content, {header: true})` and read `parsed.data` alone. PapaParse reports a record that does not have the header's fields as a FieldMismatch in `parsed.errors` and maps the row onto the header keys regardless, leaving the missing ones unset, which the table drew as empty cells. So a file whose exporter omits trailing fields looked COMPLETE in the portal while a registration over it named the same records -- which is why the viewer "rendering fine" was not evidence the file was rectangular. A note under the table now says how many rows end before the last column and how many carry more fields than the header names.

A REFUSED REGISTRATION IS READ IN A DIALOG (#1780). The register form lives in the viewer's details column (`ViewerLayout`, `lg:w-80` = 320px) and a CSV refusal is several sentences naming rows and columns; wrapped into that column it was small red type below the fold, and #1617 had already had to make the repair control wrap like a paragraph (`h-auto max-w-full whitespace-normal`) to fit. `RefusalDialog` (extracted, since TablesPanel was at its 600-line budget) puts the reason at a readable measure on `ModalShell` and the action in a footer as an ordinary button; a refusal with no next step reads the same way with only a dismiss. THE DESTRUCTIVE TOKEN WAS A FILL BEING USED AS TEXT: `--destructive` was 0 84.2% 60.2% light and 0 62.8% 30.6% dark -- DARKER for dark mode, which is right for a fill and backwards for text -- putting every refusal in the product at 3.76:1 on a card in light and 1.78:1 in dark, both under WCAG AA. It is now the on-surface text colour (0 72% 42% light, 0 70% 65% dark; 6.47:1 and 5.39:1 on a card), with the old values under `--destructive-fill` for the button, badge and status dot. That inversion was chosen by counting: 100 places set text with it against 2 that fill a shape, so the common case gets the plain name and only the fills were migrated.

A TABULAR CELL THAT DOES NOT FIT IS READ BY OPENING THE ROW (#1781). Cells are capped at 200px and the only way to see a longer value was the browser's `title` tooltip, which cannot be selected or copied, does not wrap for a paragraph, does not exist on touch and vanishes when the pointer moves; the columns cannot be resized and the row had no click, so on a 35-column export with paragraphs in Title and Description the table was a wall of ellipses over data that was all there. It was also the ONE list in the portal that did not open on row click. A row is now a focusable control (Enter and Space open it) and `RowDetailDialog` shows the record as key/value pairs on `ModalShell`, values wrapping and selectable, with copy per value and for the whole row and previous/next so reading several records is not close-and-re-aim. A field ABSENT from the record reads "not in this row" rather than as an empty value, because after #1779 the two are different facts. THE PATTERN THIS SETTLES, written into docs/server/content-viewers.md for the next dense table: truncate in the table, open the record in a dialog, never leave something a person needs to read reachable only through `title`.

A REGISTRATION THAT FAILS SAYS WHERE IT STOPPED, AND IS ALWAYS AUDITED (#1775). A 500 answered "the registration could not be completed" -- no field, no reason, no code -- because the wrapped store or driver text may carry topology a caller should not see. The stage is the half that can be told: the registrar wraps each platform failure with `failedf(stage, err)` (reading the file, checking the table name, saving a corrected version of the file, listing the file's directory, registering the table, recording the registration, replacing the previous registration, removing the registration), `tableregister.StageOf` reads it back, and the detail becomes "the registration could not be completed while reading the file". The whole error still goes only to the log and the audit event. And it now GOES to an audit event: a registration that failed before it had corrected the caller's file was not audited at all, so a deployment where registrations were failing had a `table_register` history in which every attempt had succeeded -- CloudSent's was 14 events, every one `success: true`, with neither of the two failures that prompted the ticket among them. `plan` fills the connection, source kind and source id onto the registration before the first check, so an attempt refused at the connection boundary still names what it was for, and `ErrConnectionDenied` is recorded `authorized: false` rather than `true`.

THE READ IS NOT BOUNDED BY A FLAT 30 SECONDS (#1773). `contentFor` reads the WHOLE object by design -- the header row and the line-break scan both need the full bytes -- and neither `createPortalS3Client` nor the managed-resource layer passed a `Timeout`, so mcp-s3 filled in its 30 s default around the request AND the `io.ReadAll` of the body. On a full-object read that is a THROUGHPUT BUDGET rather than a timeout: a 126 MB CSV registered in 2.05 s on a healthy path and failed at exactly 30.0 s, 116 MB in, the day one deployment's DNS sent the pod to the other site's ingress and back over the WAN at ~3.8 MB/s -- so whether a customer's file could be registered turned on which A record the resolver handed out. The instance's documented `timeout` key reached the `s3_*` tools and nothing else, so an operator could not raise it either. `toolkitcfg.S3Config` now reads `timeout`, and `toolkitcfg.BlobReadTimeout(configured, maxObjectBytes)` is what both clients are built with: a configured value wins outright, and with none the deadline is the deployment's own object ceiling -- the larger of `resources.managed.max_upload_bytes` and `portal.max_content_size`, the same ceiling #1634 pinned the registration's own cap to -- at a floor of 1 MB/s, never below the 30 s mcp-s3 would have used. A 250 MB ceiling is a 250 s deadline; a deployment storing small objects keeps exactly the client it had.

CARRIAGE-RETURN LINE ENDINGS ARE THE SAME FAILURE ONE LAYER EARLIER (#1445). Go's `encoding/csv` splits records on `\n` only, exactly as the Hive text input format does, so a CSV whose lines end in a bare `\r` -- the classic Mac ending some spreadsheet exports still write -- is ONE record to both. Before #1445 that made `InspectCSV` report `Rows: 1` with column names assembled out of a header line joined to the line after it (`address101`, `12 Mill Rd102`), and the offered correction, once taken, joined those lines with a space and saved a two-column header row with NO data rows as a new version of the person's file. `withLineFeeds` now rewrites every LONE `\r` -- one that is not part of a `\r\n`, which every reader here already folds -- as `\n` before anything is counted, in `InspectCSV` and again in `NormalizeCSV`. WHICH CARRIAGE RETURNS WERE LINE ENDINGS IS DECIDED FROM THE PARSE, NOT FROM THE BYTES, BY TWO MEASURES FOR TWO REGIMES (`recoversRecords`). A body with NO `\n` ANYWHERE is one record to every reader on this path and no ordinary CSV is written that way, so nothing in it is ambiguous and the PLAIN count (`recordCount`) decides. A body that ALREADY HAS `\n`-delimited records is ambiguous -- splitting a cell on a `\r` adds a record exactly as a real line ending does -- so records carrying the HEADER'S FIELD COUNT (`wellFormedRecords`) decide, since a record recovered from a line ending has the columns the header declares and a fragment torn out of one cell does not. THREE WEAKER RULES WERE TRIED AND ARE EACH WRONG, in ways only measurement found. (1) A byte test ("no `\n` anywhere") disqualifies an otherwise CR-delimited file over one stray line feed -- a tool that appended one, a cell pasted in from a Windows source -- and restores the exact pre-#1445 behaviour the correction then reports as a repair. (2) A plain record count EVERYWHERE is wrong for an UNQUOTED in-cell `\r` (`store_id,address\n101,12 Mill Rd\r9 Oak St\n102,4 Elm Ave\n`): the translation splits that cell, the count rises by one just as a recovered boundary would, and the file gets a false "lines end in a carriage return" plus a field-count refusal over the one-field fragment -- a REGRESSION against pre-#1445, which diagnosed the in-cell break and corrected it. The quoted shape cannot catch this: a `\r` inside quotes stays inside them, so the count cannot move and both rules agree, which is why the test carries an UNQUOTED case. (3) The header-width measure EVERYWHERE is wrong in the DESTRUCTIVE direction, and this is the #1445 failure itself: for an untranslated Mac file the parse returns the whole file as one record and `want` is taken from THAT record, so the original side always scores >= 1 and the translated side must reach >= 2 -- the header PLUS a data row of the header's width. A Mac file with no such row (an unquoted comma in an address, a header wider than its rows, rows wider than the header) never clears the bar, the translation is discarded, and the file is merged into one row and saved that way as a new version. Enumerating classic-Mac CSVs of 2-5 rows x 1-4 header widths x 1-4 row widths, 41 of those 64 shapes are silently flattened by the header-width measure and 0 by the shipped rule (at the narrower 2-4 x 1-3, 14 of 27 against 0); `TestNormalizeCSV_NeverMergesTheRecordsOfACarriageReturnFile` pins the property. Kept, those same files are refused HONESTLY by the field-count rule naming the record. ONE SHAPE STAYS AMBIGUOUS AND IS DECIDED RATHER THAN KNOWN: a SINGLE-COLUMN file with an unquoted `\r` in a value, where every fragment is also the header's width, so field counts cannot separate the readings and the `\r` is taken as a line ending; the quoted form a spreadsheet actually writes is correctly left alone. `wellFormedRecords` and `recordCount` both set `ReuseRecord` and read only `len(record)`, the header's width captured before the loop since the reused slice is the one the header came back in. The parse test also picks up the case a byte test cannot see at all: a CR-delimited file with a quoted `\n` in a cell fails `csv.Read` on its FIRST record, so the record scan returned nothing, `InspectCSV` returned nil, and the file reached `ReadHeaderColumns` to be refused with `ErrEmptyHeader` -- telling the person the one thing that was not wrong with it and offering no repair. A lone `\r` that recovers no record is left alone and stays what it is, a line break INSIDE A CELL, so nothing untrue is said about the file's lines. Translation is applied only to bytes the platform reads as single-byte text (`CSVDefect.convertibleEncoding`, the encoding half of `Correctable`), so a wide encoding is still named rather than rewritten. `CSVDefect.LineEndings` carries the finding, `Correctable` stays true (one line ending maps to another without guessing at anything), and the correction goes through the SAME version trail: `NormalizeReport.FromLineEndings` reports it and `repairSummary` renders it as "rewrote the carriage return line endings as newlines". Because the endings are settled first, the record scan sees the records the file actually holds, so a ragged CR file is refused by record number -- at the inspection, with no correction offered (#1449) -- and a CR file with a real break inside a cell names the column it is in.

THE CORRECTION IS A NEW VERSION OF THE FILE ITSELF, not a derived copy beside it. `manage_table` gains `repair` (and the REST body `"repair": true`, and the portal a control on the refusal); with it, `NormalizeCSV` decodes as UTF-8 or, on invalid UTF-8, as windows-1252 (`golang.org/x/text/encoding/charmap`), drops a leading BOM, replaces each run of `\r`/`\n` inside a field with one space and trims it, and re-emits with `csv.Writer` as UTF-8 with no BOM. The bytes then go through the version mechanism the source kind ALREADY has -- `resource.ReviseContent` (extracted from the replace-content route so a revision written on somebody's behalf is the same revision, same trail, same retention) for a managed resource, `portaldomain.VersionStore.CreateVersion` for a portal asset -- both of which write a fresh per-version directory and move the head in one transaction, which is exactly the one-file directory `locationFor` requires. The uploaded bytes stay as the version before the correction, revertible from the same panel every other version is. The registrar reaches these through a `Reviser` port keyed by source kind, the way `ObjectReader` is, since the two kinds keep their trails in different tables; a kind with no reviser refuses instead of correcting.

A record whose field count differs from the header is REFUSED, naming the records: padding a short record invents data and truncating a long one discards it, and neither is a correction the platform can make on somebody's behalf. THE OFFER IS TESTED AGAINST WHAT THE CORRECTION WILL ACTUALLY DO (#1449). Until then `Correctable` inspected the ENCODING AND NOTHING ELSE, while `NormalizeCSV` applied two further rules the offer knew nothing about, so a caller told to "register it again asking for the file to be corrected" asked for it and got a SECOND, DIFFERENT refusal naming a problem the first had not mentioned -- `a,b\n1,"x\ny"\n2\n` was offered the correction for its torn row and then refused by `checkFieldCounts` for its one-field record. The single reader pass in `scanRecords` (was `scanEmbeddedBreaks`) already visits every record, so it now also records the RAGGED RECORDS (`CSVDefect.HeaderFields`/`Ragged`, numbered as `checkFieldCounts` numbers them) and the PARSE ERROR that ended the read early (`CSVDefect.Unreadable`), and `Correctable` is the conjunction of all three: a convertible encoding, no ragged record, no truncated parse. THE HEADER READ IS THE SAME READ: a scan whose FIRST `Read` fails returns claiming nothing, and if that early return does not record the parse error the offer stands for exactly the files whose defect was settled BEFORE the scan -- a windows-1252 file or one with carriage-return endings, whose header carries a bare quote -- and the correction refuses them on its own `ReadAll`, which is the same two-answer bug one layer up. `raggedClause` and `unreadableClause` are the ONE wording for each condition, shared by the inspection that declines to offer and by `NormalizeCSV` which still refuses independently. NEITHER NEW CONDITION MAKES A DEFECT BY ITSELF: `InspectCSV` still returns nil unless there are torn rows, an encoding or carriage-return endings, so a ragged CSV with newline endings and no in-cell break registers exactly as it did, as does one the reader gives up on partway -- which is the invariant `scanEmbeddedBreaks` tolerated a parse failure for in the first place. The remedy sentence is now the defect's (`CSVDefect.remedy`) rather than the registrar's fixed "Re-export it as UTF-8 CSV and upload that", which was wrong advice for a file whose bytes were fine and whose records were not. A file that is already line-safe and valid UTF-8 behaves exactly as before -- no version, no rewrite, nothing extra in the result -- including when the repair was asked for unnecessarily.

ONLY A SINGLE-BYTE CODE PAGE IS CONVERTED. `charmap.Windows1252` maps all 256 bytes and never errors, so "not valid UTF-8 therefore windows-1252" would decode a spreadsheet's UTF-16 "Unicode Text" export into one character per byte and write that mojibake back as a corrected version of somebody's file, reporting it as a repair. `sourceEncoding` checks the byte-order marks widest-first (a UTF-32LE mark begins with a UTF-16LE one) and treats a NUL byte in a markless file as "not single-byte text either"; `CSVDefect.convertibleEncoding` is false for all of those, and `correct` refuses them outright with `refusedf` rather than `needsRepairf`, so the surfaces offer nothing the platform cannot honestly do. WINDOWS-1252 IS THE LAST ANSWER, NOT THE FALLBACK FOR EVERYTHING THAT IS NOT UTF-8 (#1448): the code page leaves five byte values undefined (0x81, 0x8D, 0x8F, 0x90, 0x9D) and `charmap.Windows1252.NewDecoder().Bytes` emits U+FFFD for each of them and returns a NIL ERROR, so a file holding one converted "cleanly" into a file with replacement marks in it, was saved as a new version of the person's resource or asset, and reported `FromEncoding: windows-1252`. `sourceEncoding` therefore checks the source for those five bytes AFTER `utf8.Valid` (a valid UTF-8 file may carry any of them as a continuation byte) and returns `encodingUnidentified` for them, which `Correctable` is false for, so the refusal comes before anything is written and no correction is offered either way. The refusal names the evidence rather than an encoding -- "a byte windows-1252 does not define" -- because that byte rules the code page out without saying what the file is instead. THE NUL TEST RUNS BEFORE `utf8.Valid`, NOT AFTER IT (#1447): a NUL is valid UTF-8, so a markless UTF-16/UTF-32 export of ASCII content passes a UTF-8 test with a NUL beside every character and would register as a table whose column names and cells carry them. The refusal therefore never says its bytes "are not UTF-8", which that file's bytes are: an encoding identified by a byte-order mark is named ("look like UTF-16"), and one identified by the NUL alone is reported as the NUL, since nothing else about those bytes says what they are. `InspectCSV` returns on the encoding alone for anything `convertibleEncoding` is false for, claiming NOTHING about line endings, torn rows or column names: every reader below that point is a single-byte one, so over a UTF-16 file with Windows endings it takes the `\r` beside a NUL for a line break inside a cell and names the column out of the NUL-laden bytes around it, which would put a count nobody can check and a NUL into the refusal text and the audit event.

A CORRECTION IS A WRITE THAT MAY OUTLIVE ITS REGISTRATION. It is written before `locationFor` and `ReadHeaderColumns` have run and before the DDL, so a refusal, an unreachable coordinator, or a store that cannot record the registration can arrive after the person's file has already changed. `plan` therefore returns its partial plan alongside the error, `Register` audits the correction on that path too (the record is filled in by `claim`, before the file is read, so the event names the table it was for rather than an empty one), and ALL FOUR failure paths wrap the error in `repaired`, whose message leads with what changed and which `RepairOf` reads back. The two store writes did not wrap it until #1446: a file that was corrected, whose DDL then succeeded and whose row then failed to insert, produced a 500 reading "the registration could not be completed" while the new version sat on the file unmentioned. `tablehttp.detailFor` uses that to keep the sentence about the file while still replacing a 500's own text, which is a wrapped driver error carrying topology the caller should not see. A revision of a managed resource additionally re-registers it with the MCP server and fires `resources/list_changed`, the same call the replace-content route makes, so a client that already read the resource does not keep serving the bytes the correction replaced.

`plan` was reordered so nothing is written before every refusal that can happen has: the connection boundary, the derived table name and the name-ownership check are settled FIRST, then the body is read, then the correction, then the location and columns. A caller who may not take the name never gets their file rewritten on the way to being told so. The correction is recorded on the registration's own audit event (`repaired`), because it rewrote somebody's file for them. The REST refusal carries an RFC 9457 problem TYPE (`urn:mcp-data-platform:problem:csv-needs-repair`, via `httpjson.WriteErrorCode`) so the portal keys its offer on the code rather than on prose; the detail stays the sentence a person reads. An answer whose registration failed AFTER a correction carries `urn:mcp-data-platform:problem:file-corrected` on the same mechanism (#1450): the correction is written before the last checks and before the DDL, so a refusal or a platform failure can follow one and the file stays changed, and the version trail and the file's own record on the page are then behind what is stored. The register mutation refreshes them from `onError` on that type as it does from `onSuccess` on `repaired`.

## The surfaces

One registrar serves both kinds, because a registration says the same thing about either: this object's directory is readable as this table, on this connection, with these columns. It checks the caller's persona against the connection (`connscope.Scope.AllowConnection`, the same predicate the authorizer applies to a tool call), reads the object's header line for column names, refuses a directory with a sibling, and issues `CREATE SCHEMA IF NOT EXISTS`, a `DROP TABLE IF EXISTS` only when replacing a registration the caller is entitled to replace, and the `CREATE TABLE`. Everything that can refuse does so BEFORE any statement runs, so a refused registration leaves nothing behind in Trino, and the record is written LAST, because a row naming a table that was never created is a lie a search hit repeats. EITHER STORE WRITE THAT FAILS AFTER THE `CREATE` TAKES THE TABLE BACK OUT (`rollBackTable`, #1446), because neither can be undone by the one that follows it. Unrecorded, the table stands with nothing naming it: no surface lists it, `BuildDDL` issues a `DROP` only when replacing a registration that EXISTS, and the second attempt the answer tells the caller to make would therefore meet it in Trino and fail. Recorded still by the registration it was REPLACING -- the `Delete` of that row failed after the DDL re-pointed the table -- it is worse, and is the case nothing else reports: the table now reads the new file with the new columns while the surviving row keeps advertising the old ones through `toolView` and `SampleJoinSQL`, and `IsStale` does not catch it, because a replacement registering the SAME key to pick up a changed header leaves the location it compares identical. What the rollback leaves instead is no row and no table. THE DDL FAILURE IN A REPLACEMENT IS THE SAME RECONCILIATION FROM THE OTHER SIDE (`forgetDroppedRegistration`): a replacement runs DROP then CREATE, so a coordinator that refuses the CREATE leaves the previous table gone with its row the only thing still claiming it exists -- and not marked stale either, for the same reason. That row is removed. When the DROP is the statement that FAILED, nothing ran that changed anything and the row is still accurate, so it is left alone; distinguishing the two is why `runDDL` returns the statements it got through rather than only an error, and why `dropTableStatement` is one function the three callers that must agree on its text all use. The drop runs on a `context.WithoutCancel` bounded by a 30s timeout, since taking cancellation off leaves nothing else between a wedged pool and a request goroutine that never returns (the audit write goes to the same pool the write that failed came from) -- the store write may have failed BECAUSE the caller disconnected, which is exactly when cleanup has to outlive the request -- and is audited, since the event written a moment earlier says the table was created; a drop that fails is logged and joined onto the cause on that event rather than replacing the store error the caller can act on.

The DDL goes through `Toolkit.Exec(ctx, connection, sql)`, the platform's one write path into Trino, which runs the SAME `ReadOnlyInterceptor` the MCP tools run -- a `read_only: true` connection refuses registration exactly as `trino_execute` would. The two pre-existing direct callers of `Manager().Client(name).Query` (the HTTP query func and `trino_export`) run SELECT and reach the client without that check, and are not a model for a statement that writes.

Three surfaces: the portal's *Query as a table* panel in both the managed-resource viewer's sidebar and the asset viewer's (ONE component, since the action is identical on either), the REST routes (`GET /api/v1/table-connections`, and `GET`/`POST /api/v1/{resources,portal/assets}/{id}/tables` with `DELETE .../{registrationID}`), and the `manage_table` tool, actions `register` / `list` / `unregister`. The connection picker is filled from the connections the caller reaches that CARRY a scratch target AND accept writes, so every option it offers is one the registrar accepts. Both halves are required: a scratch target is a destination and grants nothing, so a `read_only: true` connection naming one still refuses the CREATE TABLE, and offering it produced a 500 "the registration could not be completed" the moment it was chosen. Naming a read-only connection directly is refused by the registrar itself with `ErrConnectionReadOnly` -> 400, the sibling of `ErrNoScratchTarget`, before any DDL is attempted; reaching the DDL instead produced an error the layer could not classify and the surface could only call a 500. `tableregister.Executor.AcceptsWrites` is on the port rather than discovered by type assertion precisely so the picker cannot forget to ask; it mirrors `checkExecWritable` including its two asymmetries (no interceptor at all means a single-connection toolkit that was not configured read-only, so writes are ALLOWED; an empty connection name resolves to the default).

THE TOOL IS KEYED BY REFERENCE, NOT BY AN ID (#1428). `manage_table` takes the `reference` a `search` hit emits and `fetch` dereferences -- `mcp:resource:<id>` or `mcp:asset:<id>` -- so ONE action serves every kind of stored file: the kind travels inside the reference, there is no per-kind argument, and there is no second tool. `register` takes `follow` (default true; false pins the table) and reports which rule the table was made under; `list` reports `follow` and `follow_error` beside `stale` (#1536). An id-keyed action could only ever reach one kind, which is why `manage_asset` carries no table actions: the case registration exists for is a file a PERSON uploaded, and the agent is on the other side of that upload. `manage_table` is registered by the portal toolkit rather than a new one, because `pkg/platform` sits at its LOC cap and a new toolkit would need construction there; the wiring is `wireTableToolRegistrar` in `internal/httpserver/tablemounts.go`, which hands the toolkit a `tableregister.ToolAdapter`.

A reference the caller may not register -- missing, deleted, or somebody else's -- is answered as `ErrNoSuchFile` in every case, the way `fetch` answers one outside its reach, so the tool cannot enumerate what exists. A well-formed reference to something that is not a stored file (`mcp:knowledge_page:...`) is refused as `ErrBadReference` NAMING what was passed, because the caller can see that difference themselves.

## Registration is owner authority

Registering puts a file's contents in a schema everyone granted the connection can read, so it is the owner's call, like sharing. The rule is AUTHORITY TO CHANGE THE FILE, NOT AUTHORITY TO READ IT, and it is the same for both kinds (#1428): an asset by its owner or an administrator (`assetVisibleTo`), a resource by `resource.CanModifyResource` -- its uploader, a platform administrator, or an administrator of the scope it lives in, which is exactly the rule for updating or deleting it. A read rule would not do: resource scopes are NOT carried into Trino, so anyone who could see a persona-scoped file could publish it to everyone with the connection. Both surfaces resolve a file through ONE `tableregister.Subject` per kind (`tableSubjects`, `internal/httpserver/tablemounts.go`), so the rule cannot drift between the REST routes and the tool. Table names are persona-prefixed (`<persona>_<slug>`, the slug from the filename by default) and the name is CLAIMED in a unique index on (connection, catalog, schema, table): re-registering your own name replaces it, taking someone else's is refused naming who holds it, and administrators are unrestricted. The prefix is collision avoidance and legibility, NOT a boundary -- resource scopes and asset ownership are not carried into Trino, and the scratch schema is a shared workspace by design.

What keeps a registration off the warehouse is the Trino identity the scratch connection authenticates as. The platform's `read_only` is a statement-prefix denylist evaluated per connection name; nothing in the toolkit restricts a catalog or a schema, and `catalog`/`schema` on a connection are session defaults rather than bounds.

## Finding what is registered: the Scratch Tables section (#1472)

Every read of `table_registrations` was keyed by ONE source -- by id, by qualified name, by source -- and the only surface a person saw was `TablesPanel`, rendered inside one asset's or one resource's detail view. So the way to find out what a deployment had registered was to open every asset and every resource in turn. That is worse than an inconvenience, because the scratch schema is a shared workspace: the unique index on (connection, catalog, schema, table) exists precisely because everyone granted the connection sees every table in it, which left a reader able to query a table through Trino that the portal gave them no way to find, no way to identify the source of, and no way to tell was current.

`/scratch-tables` is the listing, in the portal's own section list rather than the administrator's, and `/scratch-tables/{registrationID}` is one registration at an address of its own. Behind them are `Store.List(ctx, Filter)` -- the first read here that spans sources -- and `GET /api/v1/tables` / `GET /api/v1/tables/{regID}`, in `internal/httpserver/tablehttp/listing.go`.

VISIBILITY FOLLOWS THE CONNECTION, which is the boundary Trino itself applies and the same predicate `Unregister` already applied: a caller sees the registrations on connections their persona is granted (`connreach.ForPersona` narrowed to the Trino kind, which delegates to `connscope`), and an administrator sees all of them. It is deliberately NOT the register form's connection list, which narrows further to connections carrying a scratch target that also accept writes: a connection turned read-only after a registration would otherwise hide that table from the person who made it while Trino went on answering queries against it. The boundary is a `Filter.AllConnections` FLAG rather than the emptiness of a `Connections` slice, because the two states that separates are opposites and both are reachable -- a persona granted no connection reaches nothing, an administrator reaches everything -- and reading one as the other would either hide every table from an operator or show every table to a persona that may query none of them. The count is taken under the same predicate as the page, so a pager states the rows the caller may see rather than the rows that exist, and a caller who names a `connection` facet outside their reach is answered with an EMPTY PAGE rather than a 403, since the parameter is a facet of a listing they may read and a refusal would confirm the connection exists. `GET /api/v1/tables/{regID}` answers a registration outside the persona as 404, the same answer an id that never existed gets.

The two things a cross-source read has to add to a stored registration are which file it is and whether the table is still reading that file's current contents, and neither is on the row: staleness needs the source's head key NOW. `tableregister.Sources` is the bulk form of `Subject` -- one read per kind for a page, `portal.AssetStore.GetByIDs` and a new `resource.Store.GetByIDs` -- and it returns a `SourceRef` carrying the name, the bucket, the head key, and `CanModify`. AUTHORITY IS A FIELD ON THE ANSWER RATHER THAN THE ANSWER ITSELF, which is what separates it from `Subject`: the listing shows what a caller may QUERY, decided by connection, while the unregister action needs authority over the SOURCE, and a row is shown either way. A source id absent from the map is a record that is gone, rendered as "Source deleted" rather than as an ordinary row -- deleting a file unregisters its tables, so it is the residue of a cleanup that did not complete. A source store that could not answer degrades the page (no name, no action) rather than failing it.

UNREGISTER GOES THROUGH THE SOURCE'S OWN ROUTE (`DELETE /api/v1/{resources,portal/assets}/{sourceID}/tables/{regID}`), so the rule for who may drop a registration stays written once: authority over the source AND having registered the table or being an administrator. `can_unregister` on each row is that rule evaluated from the same bulk lookup that supplies the name and the head key, so the control is absent rather than present and refusing. REGISTERING STAYS ON THE FILE'S OWN PAGE, because it needs the file: the platform reads the header row to learn the columns.

The listing is driven by the STORED ROWS rather than by what can be registered, so it carries no content-type gate of its own -- `TablesPanel` renders only for a content type containing "csv", which belongs to the register action, and a source of another Hive-readable format appears in the listing as soon as one exists with no change here. The empty state says what a scratch table is and where registration happens rather than showing an empty table, and it is distinguished from a filter that matched nothing, which keeps the table and the facets in front of the reader.

## The lifecycle, staleness, and deleting

A FILE IS NOT LIMITED TO ONE REGISTRATION, which is why the portal panel keeps its Register control available after the first one and lists registrations as a set. Registering under a different name, or onto a different connection, ADDS a table and leaves the existing ones alone; registering the SAME name on the same connection REPLACES that registration, which is the repair for staleness below and the reason the panel's warning says to register again rather than to unregister first. Unregistering is a separate act with a separate consequence, and neither it nor a replacement touches the file. THE DISCOVERY LAYER REPORTS THE SET, NOT ITS NEWEST MEMBER (#1627): `fetch` on a stored file carries `tables`, one entry per registration over it, newest first, each with `registration_id`, `query_table`, `columns`, `sample_sql`, `stale`, `follow`, `repair` and `follow_error` -- the same facts and the same order `manage_table action=list` reports under `table_registrations`, so the document and the managing tool cannot disagree. THE TWO KEYS ARE THE TWO USES (#1666): `tables` is the query view a `fetch` document and a `search` hit carry, the DOCUMENT's rows carrying the columns so the query can be written without a second call while the hit's one `table` carries none (`preferredTable` clears them: a hit is a pointer chosen from a ranked page of them, and a page each carrying a wide table's column list is a large answer to the question of which record to read), and `table_registrations` is the maintenance view `manage_table action=list` and the per-file REST route answer with, carrying the `registration_id` an unregister takes. It answered under `registrations` until #1666 while every other surface said `tables`, so a script checking for an existing registration read `tables`, found nothing, and re-registered on every run -- the table survived, since registering the same name replaces the registration, but the registration id changed every time. A `search` hit still carries ONE `table`, because a hit is a pointer to somewhere the data can be queried rather than an inventory, and it is the newest registration whose `follow_error` is EMPTY; a file whose every registration is disowned gets no `table` on its hit and its reasons are in the document. `TableLookup.TablesFor` therefore returns the whole per-file list and the hit path chooses from it. Before that, both paths took `regs[0]` -- the newest -- whatever state it was in, so a file with a healthy table and a newer one a follow had reported gone was described by the gone one: the document named a table that did not exist, offered `sample_sql` over it, and said nothing about the table that would have answered.

A new revision of a resource or a new version of an asset writes to a new directory and moves the head key, and a registration FOLLOWS that move by default (#1536): `Registrar.FollowSource` is called from the one place each kind's head moves -- `resource.ReviseContent` behind the replace-content and restore routes and behind `manage_resource replace_content`, every asset version writer (the portal's PUT content and revert, the admin console's, `manage_asset` update/patch/revert through `uploadContentUpdate`, and a script's `platform.export` / `publish_data` through scriptexec's `storeVersion`), and the correction path inside `Register` for every OTHER registration over the file -- through `TableSourceHooks.AssetRevised` / `ResourceRevised` beside the delete hooks, with a caller-less `tablesource.Locator` resolving the record. For each following registration it re-reads the new head exactly as a registration does (whole body, CSV inspection, the sibling check, the header), runs the same DDL a re-registration runs (DROP, CREATE at the new directory with the new columns), relocates the row, and audits a `table_follow` event under the REGISTRANT naming the version followed to and the columns before/after when the header changed. The follow NEVER fails the write: a refused CREATE after the DROP ran puts the old table back, the reason is kept on the row as `follow_error`, and the registration is behind the file exactly as a pinned one is, with the reason shown. After any write that ran DROP TABLE (a follow, an unregister, a replacing registration) the registrar asks the connection whether every OTHER registration on it still holds its table (`Executor.TableExists`, an information_schema lookup), records the ones that do not (`follow_error` opening with *The table no longer exists*) and reports them -- in the follow's `table_changes` sentences, in a replacing registration's result, and on the rows an unregister leaves (#1546): Trino's file metastore drops a table by listing its metadata directory, and an object store whose prefix listing does not stop at a directory boundary (SeaweedFS before 4.17) answers `x/` with `x_pinned/...` too, so one drop took a name-prefix sibling with it while every surface kept listing the sibling as registered. A lookup the connection cannot answer is logged, not reported. The write's result names every table over the file under `table_changes`, a change report rather than the `tables` a caller queries (#1666; `script.RunOutput` is STORED as well as served -- `script_runs.outputs` is a JSONB array unmarshalled back into it -- so migration 000141 rewrites the key on every recorded output, preserving call order, or every run already in the history would read back with no table report at all) -- on `replace_content` and the asset content results (also appended to the message), the portal/admin `statusResponse`, the export record and `RunOutput`, and the run log as `table_changes: <output>: <sentence>` (echoed by the host for `platform.call` results too) -- so the write that leaves a table behind says so. `follow=false` / `"follow": false` / the unticked portal box PINS a table to the version it was registered over, stored on the row (`table_registrations.follow`, default true for existing rows too); a pinned table keeps serving the directory it was registered against and the recorded location compared against the head's directory NOW reports it stale on the portal panel, in Scratch Tables (Pinned / Follows the file / Behind the file, with the failure on the badge), on a `search` hit, on a `fetch` document and in `manage_table action=list`. Re-registering the name targets the current head and re-sets the choice. A pin holds only because every version has a directory of its own, since a table reads every file in the directory it points at (#1851): a script's `platform.export` / `publish_data` versions are written as `scripts/<script>/<asset>/<run>/content.<ext>` (they were `<asset>/<run>.<ext>`, every version of an output in one directory, so the next run's file landed beside the one a pinned table read, the table returned both versions' rows, and `followOne` reported the pinned table as followed because the head's directory matched the recorded one), and `api_export` / `graphql_export` write `<prefix>/api_export|graphql_export/<user>/<asset>/content.<ext>` rather than one directory per user. `Registrar.pinned` answers a pinned registration: behind the file (`Pinned`) when the head is elsewhere; reading the version (`Followed`) only when the head was replaced in place and is the only file in the directory; and, when the directory holds more than one file Trino reads -- an output written in the old layout -- `Pinned` with a `Reason` whose sentence says the table does not read its version alone and to register it again, the files being left where they are. It lists that directory only when the head is at or beneath it. Old versions keep the keys their rows recorded, and a legacy output whose head still shares its directory is refused

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.