agentleFS
Sign inSign up

aldus-palace

heymi/aldus-palace/llms-full.txt

GENERATED FILE — do not edit. Regenerate: pnpm gen:llms Source: README.md, EVALUATION.md, docs/ARCHITECTURE.md, docs/DOMAIN-SCHEMA.md, docs/CAPABILITIES.md, docs/INTELLIGENCE.md, docs/EVAL.md, docs/POSITIONING.md, docs/MAP.md, docs/GLOSSARY.md, docs/RETRIEVER.md, docs/BENCHMARKS.md, docs/BENCHMARKS-LLM.md, ROADMAP.md [](https://github.com/heymi/aldus-palace/actions/workflows/ci.yml) [](https://www.npmjs.com/package/@aldus-palace/core) [](https://www.npmjs.com/package/@aldus-palace/mcp) [](https://www.npmjs.com/package/@aldus-palace/client) [](LICENSE) A trustworthy context layer for AI that remembers, plans and acts. Free text goes in; typed commitments, decisions and memories come out — each with the evidence behind it, in a SQLite file you own. Evaluate it in five minutes: EVALUATION.md maps every claim to the command that proves it. 中文版…

llms.txt17 starsChanged 12 days ago
  • Reads credentials
  • Installs packages
  • Sends data out
# Aldus Palace — full documentation

GENERATED FILE — do not edit.
Regenerate: pnpm gen:llms
Source: README.md, EVALUATION.md, docs/ARCHITECTURE.md, docs/DOMAIN-SCHEMA.md, docs/CAPABILITIES.md, docs/INTELLIGENCE.md, docs/EVAL.md, docs/POSITIONING.md, docs/MAP.md, docs/GLOSSARY.md, docs/RETRIEVER.md, docs/BENCHMARKS.md, docs/BENCHMARKS-LLM.md, ROADMAP.md


<!-- README.md -->

# Aldus Palace

[![CI](https://github.com/heymi/aldus-palace/actions/workflows/ci.yml/badge.svg)](https://github.com/heymi/aldus-palace/actions/workflows/ci.yml)
[![npm core](https://img.shields.io/npm/v/%40aldus-palace%2Fcore?label=core)](https://www.npmjs.com/package/@aldus-palace/core)
[![npm mcp](https://img.shields.io/npm/v/%40aldus-palace%2Fmcp?label=mcp)](https://www.npmjs.com/package/@aldus-palace/mcp)
[![npm client](https://img.shields.io/npm/v/%40aldus-palace%2Fclient?label=client)](https://www.npmjs.com/package/@aldus-palace/client)
[![license](https://img.shields.io/badge/license-Apache--2.0-blue.svg)](LICENSE)

![A capture session](docs/assets/capture-session.svg)

**A trustworthy context layer for AI that remembers, plans and acts.** Free text
goes in; typed commitments, decisions and memories come out — each with the
evidence behind it, in a SQLite file you own.

**Evaluate it in five minutes:** [`EVALUATION.md`](EVALUATION.md) maps every
claim to the command that proves it. [中文版 README](README.zh.md) ·
[中文评估指南](EVALUATION.zh.md).

---

A to-do list asks you to file, tag, date and prioritize every item. The list
grows. The managing becomes the work.

An AI assistant remembers things about you. That memory sits in a closed box. You
read it through one app, and you move it nowhere.

Aldus Palace holds one place for what you need to do. You write a sentence. The
system does the filing.

```
$ pnpm demo

$ input  Ship the onboarding page next week
Captured · 1 commitment
  commitment  Ship the onboarding page next week  ·  window 2026-09-19 → 2026-09-26

$ input  I prefer simple tools
Captured · remembered 1
  memory      Prefers simple products and interfaces; avoids complexity  ·  active
```

That output comes from the deterministic provider, offline; dates resolve on the
day you run it. `pnpm demo` runs the same pipeline with no API key.

**Try it in your assistant.** MCP inside Claude, Cursor or any MCP client:

```bash
npm install -g @aldus-palace/mcp
claude mcp add aldus-palace -- node "$(npm root -g)/@aldus-palace/mcp/dist/index.js"
```

## What it does

**It reads a sentence and sorts it.** Write "Ship the onboarding page next week"
and you get one commitment, with the date worked out. No project picker, no
priority field, no due-date calendar. Paste a paragraph and you get several
objects: thoughts, commitments, decisions, memory candidates.

**It catches repeats.** Repeat yourself and the system points at the first record.

**It remembers the rules you state.** "I prefer simple tools" lands as a rule the
system follows from the next capture onward. The memory list archives it.

**It shows its work.** Every memory carries the sentence you said, the
confidence behind it, and a note that says whether you stated it or the system
inferred it.

**It follows you.** One file, one API. Claude, Cursor, your own frontend, a script
you write.

## Who it's for and when

**Who it's for**

| You are | What you actually get | Start at |
|---|---|---|
| A developer building an AI product | A user writes a sentence; the system sorts it into a task, a decision or a preference worth remembering, stores it, and schedules it when needed. You do not design those rules yourself. | [INTEGRATION.md](docs/INTEGRATION.md) |
| An individual who wants their context to be theirs | Your memory and to-dos sit in one file you own; Claude, Cursor and your terminal all read the same one, every memory shows where it came from, and you can delete any of it. | [Capabilities 5, 6, 9](docs/CAPABILITIES.md) |
| A team that must audit AI-written data | Every action the system takes for you is recorded, reviewable and revocable; names and amounts are stripped before anything reaches a model; deleting really deletes. | [SECURITY.md](SECURITY.md) · [ADR 0005](docs/adr/0005-action-gate-with-published-risk.md) |

**What developers get out of the box**

| Capability | What it lets you ship |
|---|---|
| **Text → typed objects** | A user writes a sentence and the product already has the task, the decision and the preference — dates resolved, duplicates skipped, model failures handled. When the rules are unsure, it asks instead of guessing. No form, no classification rules of your own. |
| **A memory layer** | An assistant that remembers a person over time, where every memory explains itself, can be corrected and can be undone — no "personality from one remark". |
| **A planning engine** | A "what now" answer instead of a wall of red items: work ordered by dependency, a daily buffer, and slipped work that moves itself. |
| **Local first + background enrichment** | Instant feedback, then the model catches up; retries, concurrency and crashes never lose or duplicate data. |
| **The Action Gate + privacy** | Let the agent act: high-risk steps ask first, autonomy grows as trust is earned, every step is revocable; cloud calls are redacted, users can revoke permissions and truly delete. |
| **Replaceable, offline-testable, four surfaces** | SQLite and an offline model today; Postgres or another model later without touching your code; `pnpm verify` green with no key; library / HTTP / MCP / typed client. |

**When to use it**

| Your situation | What the system does |
|---|---|
| "My memory does not follow me between clients." | One file behind MCP, read by Claude, Cursor and your terminal. |
| "My assistant invented a preference and I cannot delete it." | The gate waits for your confirmation, every memory shows its evidence, and any of them can be archived. |
| "Twelve red items and I freeze." | One scored "now", missed dates as risks instead of failures, and a quarter of the day left free. |
| "We must show how an AI-written record came to be." | Every change writes an action log with a reason; memories keep their evidence and versions. |
| "A risky action needs a human in the loop." | The Action Gate grades every proposed agent action; high risk waits for one approval, critical (permanent deletion) for two, and the log is revocable. |
| "Sensitive data cannot go to a model as-is." | The Privacy Gateway redacts by data level; level 4 stays local, permissions are scopes, Memory is private by default. |

## Sentences it understands

A real run of the offline provider. Relative dates resolve at capture time.

| You write | It files |
|---|---|
| Ship the onboarding page next week | commitment · window 2026-09-19 → 2026-09-26 |
| I prefer simple tools | memory · active, because you stated it |
| Keep Mac only, no Windows version | thought + decision + memory waiting for confirmation |
| Fix the notification bug tomorrow, twice | commitment · the repeat points at the first record |
| 下周把 onboarding 做完 | 要做 · 时间窗 2026-09-19 → 2026-09-26 |
| 以后产品不要做太复杂,保持克制 | 记忆 · 已生效,因为你亲口说了 |
| 今晚先不做评审,改排到周五 | 要做 · 截止 2026-09-25 |
| The search box does not work | question · defect / work / note; the answer becomes the fix |

## How this compares

| The dimension | What you use today | Aldus Palace |
|---|---|---|
| **What you give it** | a form: project, due date, priority, tags | a sentence in your own words; a paragraph yields several records |
| **Who decides the shape** | you classify, prioritize and schedule | the runtime files the thought, the commitment, the decision and the memory |
| **How time works** | one due date, and a red label when it passes | four kinds held apart — deadline, availability window, suggested slot, unscheduled — and a missed date becomes a risk you can move |
| **How memory behaves** | the assistant infers and stores inside that app | confident memories take effect on capture, and each one carries the sentence it came from plus a note saying why it is active |
| **Where your words live** | summarized into a task or a chat log | kept as you wrote them, with the system's own reading beside them |
| **What the daily view answers** | a list of everything | what to do now, what is at risk, what is unscheduled; a full day gets a rest suggestion |
| **How many stores you have** | one per app | one record, reached by MCP, HTTP, a library and a typed client: Claude, Cursor, your own frontend, a script |
| **Where the record sits** | a vendor cloud | a SQLite file you own, or a Cloudflare Worker; copy it, back it up, hand it on |
| **How you verify it** | by using it | a deterministic provider runs the pipeline with no network and no API key; 30 suites and 11 fixtures replay each run |

## How you run it

- **One person, one instance.** One database and one bearer token; run one
  instance per person.
- **A SQLite file you own**, or one Cloudflare Durable Object. Copy it, back it
  up, hand it on.
- **Four surfaces over the same data:** MCP, HTTP, the library and a typed
  client. You bring the screen.
- **You decide what acts.** The runtime records intent and plans; acting on the
  world is your call.
- **Offline mode** runs the whole pipeline with a deterministic provider; connect
  a model for general understanding, and the receipts show which provider
  produced them.

## Build with it

Four ways in, one schema, open source under Apache-2.0. Offline mode runs with
no API key.

**Library**

```ts
import { DevLLMProvider, ensureDevUser, formatActionCard, initialize, newId,
         nowIso, processRawInput } from "@aldus-palace/core";
import { openSqliteDatabase } from "@aldus-palace/core/db/sqlite";

const db = await openSqliteDatabase("./aldus.db");
await initialize(db);                              // canonical DDL + migrations
const user = await ensureDevUser(db, { name: "Me", timezone: "UTC", language: "en" });

const id = newId("inp");
const text = "Ship the onboarding page next week";
await db.prepare(`INSERT INTO raw_inputs
  (id, user_id, content, source, processing_status, created_at, updated_at)
  VALUES (?, ?, ?, 'text', 'pending', ?, ?)`)
  .run(id, user.id, text, nowIso(), nowIso());

const card = await processRawInput(db, new DevLLMProvider(), user, id, "local");
console.log(formatActionCard(card));
// Captured · 1 commitment
//   commitment  Ship the onboarding page next week  ·  window 2026-09-19 → 2026-09-26
```

**MCP** — inside Claude, Cursor or any MCP client

```bash
npm install -g @aldus-palace/mcp
claude mcp add aldus-palace -- node "$(npm root -g)/@aldus-palace/mcp/dist/index.js"
```

**HTTP** — self-hosted, SQLite in a volume

```bash
cp .env.example .env && docker compose up --build
curl -X POST localhost:8787/v1/inputs \
  -H 'Authorization: Bearer dev-local-token' -H 'Content-Type: application/json' \
  -d '{"content":"Ship the onboarding page next week","mode":"sync"}'
```

**Client** — a typed client over the same HTTP API

```ts
import { createClient } from "@aldus-palace/client";

const aldus = createClient({ baseUrl: "http://localhost:8787", token: "dev-local-token" });
const card = await aldus.capture("Ship the onboarding page next week", "local");
console.log(card.action_card.summary);        // Captured · 1 commitment
console.log(await aldus.actions("proposed")); // what the Action Gate is holding
```

---

## Scale

`packages/core` is runtime-agnostic and dependency-light: `zod`, `dayjs` and
`nanoid`. **Requirements:** Node 20 or newer. `better-sqlite3` ships prebuilds
for common platforms; other platforms need a C toolchain.

## The system behind it

The runtime runs as a pipeline, and the Core Intelligence Layer organises it into
four engines. Understanding — the capture front end in `agent/understand.ts` —
turns a sentence into typed objects; the engines decide what is kept, what
happens next, what the system may do on its own, and which model is used. A model
can contribute to understanding and to conflict detection, and every engine has a
path that runs without one. Each engine ships today and has a designed
extension — the shipped parts name the file they live in, and the full design is
mapped in [`docs/INTELLIGENCE.md`](docs/INTELLIGENCE.md).

```
user / environment
      |
input          raw text, stored as written
      |
understanding  intent, typed objects, resolved dates, gates
      |
memory         candidates, evidence, confirmation, conflicts, versions
      |
planning       four kinds of time, today, risk, adaptive limits
      |
context        active memories and projects feed the next capture
```

| Engine | Shipped today | Designed next |
|---|---|---|
| **Memory** | extraction, a pollution gate, an activation gate, evidence on every row, duplicate collapse, conflict detection, versioned supersede, concepts and their links, graded levels with decay and a value score, full-text retrieval into the next capture | more kinds and extraction signals, an identity level, a memory graph and its exploration UI |
| **Planning** | four kinds of time, constraint handling (deadline, window, preference, blocked-by dependencies), priority scoring, slot search, duration estimation from project history, a daily buffer, migration for slipped flexible work, a scored Now with context match, a morning core/optional/deferred plan, replanning after a change, Today, risk, adaptive limits, a learned behaviour model | soft-constraint trade-offs, a blended priority score, a rhythm-aware Now, an events surface for meetings |
| **Trust & autonomy** | an Action Gate (low and medium run, high waits, critical needs two) with durable, idempotent execution; a trust score with levels 0–4; permission evolution capped by a user ceiling | proactive rules |
| **Model orchestration** | one `LLMProvider` interface and three implementations, configuration resolved by the caller | routing by task — fast classification, reasoning, embeddings, a local model for sensitive input |

The runtime lives in `packages/core/src`: `agent/understand.ts` for capture, `services/` for planning, memory, the Action Gate and privacy, `lib/` for the pure helpers, and `providers/` for the model interface. The capability-by-capability map is in [`docs/MAP.md`](docs/MAP.md), and the full design, with each engine's shipped and planned parts in depth, is in [`docs/INTELLIGENCE.md`](docs/INTELLIGENCE.md).

### The memory pipeline

```
capture -> extraction -> candidate -> evaluation -> conflict check -> storage -> activation -> retrieval
```

- **Extraction** reads two signals: durability markers ("from now on", "as a rule") and repeated behaviour. Rules run with no model; a model adds general understanding.
- **Evaluation** drops what should not be remembered: a temporary state, a one-off creative fragment, a low-confidence guess.
- **Activation** follows one published rule: `confidence >= 0.8` and `importance >= 0.8`. Everything below that waits as a candidate.
- **Conflict check** compares a candidate against active memories on the same dimension and reports the contradiction instead of storing both.
- **Retrieval** injects active memories into the next capture, so understanding improves with use. Each injection lands in the action log.

### Three rules that keep memory honest

1. **A mood does not become a profile entry.** "I'm tired today" is dropped before storage.
2. **One inference does not make a principle.** A principle the system inferred waits for confirmation, whatever its score.
3. **Every memory carries evidence.** The sentence it came from, a confidence value, and a note that says whether you stated it or the system inferred it.

### Privacy is the architecture

*The system touches work, decisions, relationships and habits; privacy is the
shape of it, not a feature on top.*

**Shipped today** — a SQLite file you own (or one Cloudflare Durable Object), a
single-user runtime, immutable `raw_inputs`, model output validated and gated, an
`action_log` entry on every mutation, and a memory gate that confident, stated
memories pass on capture while an inference waits for you. A Privacy Gateway
prepares a cloud call by data level (names, money, emails and phone numbers are
replaced, level 4 stays local), permissions are progressive with Memory private
by default, and `purgeUserData` deletes every row the user owns in one
transaction — a critical action, so it takes two approvals. [`SECURITY.md`](SECURITY.md) records the posture and the current
threat model.

Cloud providers are only built together with that guard, so a call cannot leave
the device unredacted.

**Designed next** — local encrypted storage (Keychain / Secure Enclave).

## Where the difficulty lives

**Models return text. Code needs records.** A model writes `"0.8"` where you
asked for a number, invents a date, or stores one task under two titles. The
runtime validates each answer, skips duplicates and resolves dates on the server.
A failed answer leaves the previous record in place.

**The slow path breaks products.** A user waits on a model and loses interest. A
background model meets retries, two clients writing at once, and records stuck
mid-update after a crash.

**Memory is a liability.** Store the first inference and the user owns a
personality from one remark. Replace the old record and the history disappears.

## What the code enforces

| Capability | The point | Where |
|---|---|---|
| Schema & domain | `overdue` has no state to occupy; a deadline, a window and a slot are three fields | `db/schema.ts` |
| Providers | a deterministic provider shares the interface with the paid ones, so agent logic runs in CI | `providers/dev.ts` |
| Understanding | the model proposes; the server decides (mode, duplicates, dates, fallback), and a grey-zone sentence asks instead of guessing | `agent/understand.ts`, `lib/objectAmbiguity.ts` |
| Progressive capture | a lease and a generation id make "local first, model second" idempotent | `services/enrichmentLease.ts` |
| Memory | a confident memory takes effect, explains itself, and versions instead of deleting | `lib/memoryActivation.ts`, `services/memoryEvolution.ts` |
| Today & planning | a day with no plan gets suggestions; a full day gets a rest suggestion; a quarter of the day stays free, slipped flexible work moves forward, a blocked commitment is never scheduled, and Now is scored rather than first-in-line | `services/today.ts`, `services/workMigration.ts`, `services/dependencies.ts`, `lib/nowScore.ts` |
| Work streams | grouping is a rebuildable projection; the records stay as they are | `services/workStreams.ts` |
| HTTP API | one schema, two runtimes: a local SQLite file and a Cloudflare Durable Object | `apps/server` |
| MCP server | runs with no server process, against the same local file | `packages/mcp` |
| Action Gate | an unclassified agent action waits; a critical one needs two approvals; every decision is logged and revocable, and trust is the approval rate of those decisions | `services/actionGate.ts` |
| Privacy | a cloud call is redacted by data level and level 4 stays local; Memory is private until a scope is granted; deletion empties every table for the user | `services/privacyGateway.ts`, `services/permissions.ts`, `services/dataLifecycle.ts` |

The pipeline runs offline: a deterministic provider implements the same interface
as the model-backed ones, so 30 test suites and 11 acceptance fixtures replay
with no key.

## See it run

```bash
git clone https://github.com/heymi/aldus-palace.git && cd aldus-palace
pnpm install

pnpm test     # 30 suites — deterministic, offline, no API key
pnpm eval     # 11 acceptance fixtures — the behaviour this project promises

pnpm --filter @aldus-palace/example-understanding-only start
pnpm --filter @aldus-palace/example-memory-gate-only start
pnpm --filter @aldus-palace/example-today-only start
```

Each example prints the records it stored and the reasoning behind them, with no
API key and no network. To use the macOS-shaped reference client in a browser,
start the server and the client — capture a sentence, then watch it land in Home,
Memory and the decision queue:

```bash
pnpm --filter @aldus-palace/server start                  # API on :8787
pnpm --filter @aldus-palace/example-reference-client start # client on :5173
```

## Install

| Surface | Command |
|---|---|
| MCP | `npm install -g @aldus-palace/mcp` |
| HTTP | `cp .env.example .env && docker compose up --build` |
| Library | `npm install @aldus-palace/core` |
| Client | `npm install @aldus-palace/client` |

## Documentation

| Doc | Contents |
|---|---|
| [INTEGRATION.md](docs/INTEGRATION.md) | the three levels, with copy-paste configs |
| [CAPABILITIES.md](docs/CAPABILITIES.md) | the nine capabilities and their contracts |
| [USE-CASES.md](docs/USE-CASES.md) | five things people build with this |
| [POSITIONING.md](docs/POSITIONING.md) | differentiation, and the current scope |
| [ARCHITECTURE.md](docs/ARCHITECTURE.md) | the runtime, module boundaries, storage port |
| [INTELLIGENCE.md](docs/INTELLIGENCE.md) | the four engines and the privacy design, shipped vs planned |
| [BENCHMARKS.md](docs/BENCHMARKS.md) | measured rule-layer accuracy, with the command to reproduce it |
| [BENCHMARKS-LLM.md](docs/BENCHMARKS-LLM.md) | the same pipeline through a live model, key-gated |
| [MAP.md](docs/MAP.md) | capability → code → doc → test |
| [GLOSSARY.md](docs/GLOSSARY.md) | the shared vocabulary |
| [EVALUATION.md](EVALUATION.md) | every claim mapped to the command that proves it |
| [DOMAIN-SCHEMA.md](docs/DOMAIN-SCHEMA.md) | objects, invariants, memory evolution |
| [DEPLOYMENT.md](docs/DEPLOYMENT.md) | local, edge, embedded, backups |
| [EVAL.md](docs/EVAL.md) | the acceptance fixtures, and how to add one |
| [PROGRESSIVE-CAPTURE.md](docs/PROGRESSIVE-CAPTURE.md) | the enrichment lease, written to be copied |
| [packages/client/README.md](packages/client/README.md) | the typed HTTP client |
| [adr/](docs/adr) | decisions already made, and why |

## Repository layout

```
packages/core          domain, agent runtime, storage port, migrations, providers
packages/mcp           MCP server (stdio) — profiles, local and HTTP backends
packages/client        typed HTTP client, one method per route
apps/server            Hono reference server (local SQLite and Cloudflare Durable Object)
examples/              six runnable examples, one per capability group (the last is a browser client)
spec/schema.sql        generated, readable schema (CI-checked)
eval/fixtures          acceptance scenarios
docs/                  everything above
```

## Status

`0.x` — usable and tested; the API may change between minor versions. See
[SECURITY.md](SECURITY.md) before exposing an instance.

## Contributing

Fixtures, docs and focused fixes are welcome. See [CONTRIBUTING.md](CONTRIBUTING.md).
Commits carry a DCO sign-off (`git commit -s`).

## License

[Apache-2.0](LICENSE).


<!-- EVALUATION.md -->

# Evaluate this project

A checklist for a person or an agent. Every public claim maps to a command or a
file that proves it.

> 中文版见 [`EVALUATION.zh.md`](EVALUATION.zh.md)。

## Five minutes, no API key

```bash
git clone https://github.com/heymi/aldus-palace.git && cd aldus-palace
pnpm install
pnpm verify     # typecheck · suites · fixtures · benchmark · build · spec · openapi · llms · docs
pnpm demo       # the offline capture demo
```

A passing `pnpm verify` ends with fragments like:

```
packages/core test: All 24 suites passed.
packages/mcp test: mcp tool tests passed.
apps/server test: All 4 suites passed.
All fixtures passed.
spec/schema.sql is up to date.
```

`pnpm demo` prints the card shown at the top of the README. Nothing needs a key,
a network or a running server.

If you cannot run the repository, the real output is committed:
[`examples/output/`](examples/output) (demo, suites, fixtures, benchmark) and
[`docs/assets/demo.cast`](docs/assets/demo.cast) (asciinema).

## Claim → proof

| Claim | Prove it | Expect |
|---|---|---|
| Free text becomes typed objects | `pnpm demo` | `Captured · 1 commitment`, with the window resolved |
| The pipeline runs offline | `pnpm verify` | 30 suites and 11 fixtures pass with no key |
| `overdue` has no state to occupy | `rg overdue packages/core/src/db/schema.ts` | no matches; statuses are `captured/planned/scheduled/completed/risk/cancelled` |
| Memory activates by one published rule | `sed -n '19,21p' packages/core/src/lib/memoryActivation.ts` | `MEMORY_ACTIVE_THRESHOLD = 0.8` |
| A mood never becomes a memory | `pnpm eval` | fixture `S04` passes |
| A repeat points at the first record | `pnpm demo` | `Already tracked, so this repeat was skipped` |
| Dates resolve on the server | `pnpm --filter @aldus-palace/example-capture-cli start "Ship the onboarding page next week"` | a window, not the literal phrase |
| Every memory carries evidence | `packages/core/test/memoryEvolution.test.ts` | evidence and activation notes asserted |
| A contradiction is flagged, not stored twice | `packages/core/test/memoryEvolution.test.ts` | supersede and the version chain |
| Background enrichment survives retries | `packages/core/test/enrichmentLease.test.ts` | lease and generation assertions |
| The model proposes, the server decides | `packages/core/test/inputObjectClassification.test.ts` | one object mode per capture |
| A missing duration is estimated from history | `packages/core/test/adaptivePlanning.test.ts` | the slot follows the project median |
| The planner keeps a quarter of the day free | `packages/core/test/adaptivePlanning.test.ts` | nine hours plan as 6.75; a full day is not over-filled |
| Slipped flexible work moves forward, deadlines do not | `packages/core/test/workMigration.test.ts` | slot cleared and deferral counted; a deadline stays a risk; three deferrals ask for a decision |
| A blocked commitment is never scheduled | `packages/core/test/dependencies.test.ts` | the planner skips it; completing the blocker releases it; cycles are refused |
| Now is the best current action, not the first in line | `packages/core/test/planningIntelligence.test.ts` | the context match wins, and the reason travels with it |
| The day is classified into core, optional and deferred | `packages/core/test/planningIntelligence.test.ts` | a risk item is core, work that fits is optional, the rest is deferred |
| Finishing early refills the day | `packages/core/test/planningIntelligence.test.ts` | a completion replans and the next candidate is scheduled |
| A fresh principle outranks an old experience | `packages/core/test/memoryValue.test.ts` | levels, decay and the value score rank retrieval |
| A cloud call can be redacted before it leaves | `packages/core/test/privacy.test.ts` | names, money and emails become placeholders; level 4 stays local |
| A cloud provider cannot be built without the guard | `packages/core/test/providers.test.ts` | `createLLMProvider` and the provider classes refuse a cloud kind without a guard |
| An approved action runs once, and deletion runs through the gate | `packages/core/test/actionGate.test.ts`, `apps/server/test/actionGate.test.ts` | the executor runs once; purge proposes and approval deletes |
| A live model's dates are resolved on the server | `packages/core/test/modelDateNormalization.test.ts` | free text becomes ISO; a past date for a future phrase is dropped |
| Memory retrieval is full-text, and CJK works | `packages/core/test/retriever.test.ts` | FTS5 with bm25; a Chinese substring is found |
| Retrieval quality is measured | `pnpm bench:retrieval` | Recall@1/3/5 and MRR on a labeled corpus |
| The system writes in your language | `packages/core/test/language.test.ts` | the deterministic provider follows the input; memory wording follows the user's language |
| An ambiguous capture asks instead of guessing | `packages/core/test/inputObjectClassification.test.ts`, `apps/server/test/reclassify.test.ts` | a grey-zone input asks defect / work / note; the answer creates the object |
| A correction is remembered | `packages/core/test/inputObjectClassification.test.ts` | a similar sentence is classified without asking again |
| Memory is private until a scope is granted | `packages/core/test/privacy.test.ts` | the default scopes exclude memory; revoking works |
| Deletion is real | `packages/core/test/privacy.test.ts` | the purge needs `confirm` and empties every table for the user |
| A high-risk agent action waits for a decision | `packages/core/test/actionGate.test.ts` | critical needs two approvals; revocation is final |
| The Action Gate is reachable from MCP | `packages/mcp/test/tools.test.ts` | the `actions` profile exposes list, decide and revoke |
| Trust grows from decisions, not silence | `packages/core/test/trustScore.test.ts` | level 0 with no evidence; six decisions at 0.75 reach level 2 |
| Autonomy widens only as far as the user's ceiling | `packages/core/test/trustScore.test.ts` | default ceiling 2; raising it to 3 runs high risk only after the record reaches level 3, and rejections lower it again |
| The schema is the single source of truth | `pnpm spec:check` | `spec/schema.sql` matches `db/schema.ts` |
| One record, three surfaces | `packages/mcp`, `apps/server`, `packages/core` | the same schema and migrations everywhere |
| MCP tools return a readable card | `pnpm --filter @aldus-palace/mcp test` | card text asserted, payload in `structuredContent` |
| The rule layer has measured numbers | `pnpm bench` | see [`docs/BENCHMARKS.md`](docs/BENCHMARKS.md) |

## Layout for agents

| File | Purpose |
|---|---|
| [`AGENTS.md`](AGENTS.md) | the development contract |
| [`CLAUDE.md`](CLAUDE.md) | Claude Code entry, imports AGENTS.md |
| [`.cursor/rules/aldus.mdc`](.cursor/rules/aldus.mdc) | Cursor rules |
| [`llms.txt`](llms.txt) | curated index for LLM tools |
| [`llms-full.txt`](llms-full.txt) | the key docs concatenated |
| [`docs/MAP.md`](docs/MAP.md) | capability → code → doc → test |
| [`docs/GLOSSARY.md`](docs/GLOSSARY.md) | the shared vocabulary |


<!-- docs/ARCHITECTURE.md -->

# Architecture

## The loop

```
raw input ──► understanding ──► structured objects ──► ActionCard
                                        │
                                        ├─► Today (planning projection)
                                        ├─► Work streams (rebuildable projection)
                                        └─► Memories (active through the gate, or candidates)
```

Every step is a pure function over the storage port plus an `LLMProvider`, so the
whole loop is testable offline.

## Progressive capture

The most distinctive part of the runtime. A single capture produces **two
passes**:

1. **Local pass** (`mode: "local"`) — deterministic rules write Thoughts,
   Commitments, Decisions and Memory candidates immediately. The client can
   render a real ActionCard within milliseconds.
2. **AI pass** (`mode: "full"`) — the model re-parses the *original* input and
   replaces the local derivatives.

The hand-off is coordinated by an **enrichment lease**:

| Column | Purpose |
|---|---|
| `processing_generation_id` | identifies the current writer; a stale worker cannot overwrite a newer one |
| `processing_lease_until` | a crashed worker's lease expires and becomes reclaimable |
| `result_generation_id` | records which generation produced the stored result |

`claimEnrichment` returns `null` when another generation holds a live lease, so
concurrent enrich calls are safe by construction. See
[ADR 0003](adr/0003-progressive-capture-with-enrichment-leases.md) and
`packages/core/src/services/enrichmentLease.ts`.

## Module boundaries

```
packages/core/src
├── domain/types.ts      object types + enums (no I/O)
├── db/port.ts           the storage port (async, no globals)
├── db/schema.ts         canonical DDL (single source of truth)
├── db/migrate.ts        applySchema / migrate / initialize (forward-only)
├── providers/           LLMProvider implementations + config resolution
├── lib/                 pure helpers (time, titles, matching, memory filters,
│                        now score, day plan, memory value, capacity, redaction,
│                        object ambiguity, classification signals)
├── services/            domain services (planning, classification, today, work
│                        streams, memory lifecycle & evolution, enrichment leases,
│                        dependencies, work migration, replanning, action gate,
│                        permissions, privacy gateway, data lifecycle, reclassify)
├── repos/               thin data access (users, action log)
└── agent/understand.ts  the Understanding Agent: input → objects → ActionCard
```

`packages/core` never imports Express/Hono/Workers APIs and never reads
`process.env`. Configuration and I/O are owned by the caller.

### The storage port

```ts
interface SqlDatabase {
  prepare(query: string): SqlStatement;       // all/get/run → Promise
  exec(query: string): Promise<void>;
  transaction<T>(cb: () => Promise<T>): () => Promise<T>;
}
```

It is deliberately **async** so an asynchronous driver (Postgres) can implement
it later without a breaking change. Both bundled adapters wrap synchronous
engines and therefore resolve immediately:

- `apps/server/src/db/local.ts` — `better-sqlite3`
- `apps/server/src/db/durableObject.ts` — Cloudflare Durable Object SQLite

## Services are independent

Every service in `packages/core/src/services` depends only on the storage port,
the domain types and pure helpers — **there are no dependencies between the
services themselves**:

```
services/*  ←  domain/types · db/port · lib/* · repos/actionLogs
```

That is what makes "use just one capability" true rather than aspirational, and
it is what keeps the module boundaries honest as the project grows.

## Agent responsibilities

`understand.ts` is the only place that talks to a model, and it treats the model
as a *proposer*, not an authority:

1. Build context: registered projects (name + description + aliases + brief) and
   active memories.
2. Ask for strict JSON and validate it with `zod`. Malformed numbers are coerced;
   malformed structure falls back to the local rule engine.
3. Enforce invariants the model cannot be trusted with:
   - one object mode per capture (thought / commitment / mixed)
   - near-duplicate commitments are skipped
   - relative dates (`明天`, “next Friday”) are resolved server-side, and
     ambiguous early-morning phrases produce a clarification instead of a guess
   - a grey-zone capture (a work signal too weak to act on) asks
     defect / work / note, and the correction is remembered (ADR 0012)
   - a memory becomes active only through the published gate; below it, and for
     an inferred principle, it waits as a candidate
4. Write an `action_log` entry for every mutation so behaviour stays explainable.

## Runtime targets

| Target | Entry point | Storage |
|---|---|---|
| Node | `apps/server/src/index.ts` | local SQLite file |
| Cloudflare Workers | `apps/server/src/worker.ts` | one named Durable Object (SQLite) |

Both call the same `createApp({ db, llm, config })` factory, so the HTTP surface
is identical. The Worker migrates inside `blockConcurrencyWhile` on first use.
The single Durable Object is intentional for a single-user product and is not
horizontally partitioned — see [SECURITY.md](../SECURITY.md).


<!-- docs/DOMAIN-SCHEMA.md -->

# Domain schema

The canonical DDL lives in [`packages/core/src/db/schema.ts`](../packages/core/src/db/schema.ts);
[`spec/schema.sql`](../spec/schema.sql) is generated from it (`pnpm gen:spec`).

## Objects

| Object | Table | Role |
|---|---|---|
| RawInput | `raw_inputs` | immutable audit record of what the user actually said |
| Thought | `thoughts` | an idea, insight, observation, research note or decision candidate |
| Commitment | `commitments` | something the user intends to get done |
| Decision | `decisions` | a choice, with a reason, that can be superseded or retracted |
| Memory | `memories` | durable personal context — active through the gate, or a candidate waiting for the user |
| Concept | `concepts` | a reusable cognitive node (e.g. “simplicity”, “privacy”) |
| Project | `projects` | context container with a free-form `brief` used for grounding |
| Event | `events` | fixed external calendar entries and AI work blocks (read-mostly) |
| ActionLog | `action_logs` | every user/agent mutation, with a reason |

### Agent state

| Table | Role |
|---|---|
| `commitment_dependencies` | a commitment waits on another; the planner skips it while a blocker is open |
| `action_proposals` | the Action Gate: risk, status, decision and revocations for agent actions |
| `autonomy_settings` | how far earned trust may widen autonomy (the ceiling) |
| `permission_grants` | granted permission scopes; absence means not granted |
| `classification_signals` | terms from inputs the user corrected, and the mode they chose; applied before asking (ADR 0012) |
| `clarifications` | a pending question about a record (a relative day, or what an input is) |

`commitments` also carries `deferral_count` and `migration_surfaced_at` for
slipped flexible work (see [ADR 0008](adr/0008-daily-buffer-and-work-migration.md)).

### Projections (never truth sources)

| Table | Rebuildable from |
|---|---|
| `today_assignments` | commitments + planning decisions |
| `commitment_classifications` | commitments + projects + active memories |
| `planning_day_states`, `planning_profiles`, `planning_feedback_episodes` | planner behaviour over time |
| `memory_search` (FTS5) | `memories.search_text`, written by the runtime |

Deleting a projection must never change what the user committed to — see
[ADR 0002](adr/0002-keep-work-classification-as-rebuildable-projection.md).

## Key invariants

1. **`raw_inputs` is sacred.** Model output only ever writes fields *derived*
   from it; the original text is never rewritten.
2. **Memory passes a published gate.** A high-confidence rule the user states is
   inserted as `active` on capture; an inferred principle and anything below the
   gate is inserted as `candidate` and waits. Only `active` memories are injected
   into the understanding context. `evidence` and `confidence` are required, and
   near-duplicates are rejected.
3. **Commitments are deduplicated.** Re-capturing a near-identical intention
   skips the insert and surfaces the existing commitment instead.
4. **Dates are resolved server-side.** `deadline`, `window_start`/`window_end`
   and `ai_slot_start`/`ai_slot_end` are distinct fields: a deadline is not a
   schedule, and an AI suggestion is not a calendar event.
5. **Time is stored as ISO-8601 UTC**, with the user's timezone applied when
   computing “today”.
6. **Statuses are constrained** by `CHECK` in the schema and mirrored in
   `domain/types.ts`:

| Field | Values |
|---|---|
| `thoughts.status` | `captured`, `exploring`, `converted`, `archived` |
| `thoughts.type` | `idea`, `insight`, `observation`, `research`, `decision_candidate` |
| `commitments.status` | `captured`, `planned`, `scheduled`, `completed`, `cancelled`, `risk` |
| `memories.status` | `candidate`, `active`, `archived` (plus the derived `superseded`) |
| `memories.type` | `preference`, `project_context`, `principle`, `decision`, `experience` |
| `raw_inputs.processing_status` | `pending`, `local`, `enriching`, `processed`, `failed` |
| `memories.source` | `user_explicit`, `ai_inferred`, `decision_promote` |
| `action_proposals.risk` | `low`, `medium`, `high`, `critical` |
| `action_proposals.status` | `approved`, `notified`, `proposed`, `pending_second`, `rejected`, `revoked` |
| `action_proposals.actor` | `user`, `agent` |
| `autonomy_settings.ceiling` | `2`, `3`, `4` |

## Memory evolution

A memory is never overwritten. These columns carry the history:

| Column | Meaning |
|---|---|
| `supersedes_id` | the older memory this one replaced |
| `superseded_by_id` | the newer memory that replaced this one |
| `supersede_reason` | why the user accepted the replacement |
| `conflicts_with_id` | the confirmed memory this candidate disagrees with |
| `conflict_reason` | the detected contradiction |

`memoryState()` derives one of `candidate` / `active` / `superseded` / `archived`
from `status` plus `superseded_by_id`, so the `CHECK` constraint stays intact and
existing databases need no table rebuild.

`activation` distinguishes how a memory became active without another column:
`confirmed_at` set means the user confirmed it, `confirmed_at` null means the
capture activated it (`confidence >= 0.8` and `importance >= 0.8`). See
`lib/memoryActivation.ts`.

## Migrations

`migrate()` is forward-only and records every version in `schema_migrations`.
Never edit a shipped migration — add a new one. The migration set is part of the
public contract, because it runs against databases other people already own.

```ts
await applySchema(db);          // idempotent DDL
await migrate(db);              // pending migrations only
await initialize(db);           // both, in order — what servers call
```


<!-- docs/CAPABILITIES.md -->

# Capabilities

Nine capabilities, three ways in. Use one, or all of them.

| # | Capability | You don't have to build | MCP profile | HTTP | Library |
|---|---|---|---|---|---|
| 1 | [Schema & domain](capabilities/01-schema-and-domain.md) | the object model for human intent, its constraints and migrations | — | — | `core/domain` `core/db/*` |
| 2 | [Providers](capabilities/02-providers.md) | model abstraction + a deterministic provider so agents are testable | — | — | `core/providers` |
| 3 | [Understanding Agent](capabilities/03-understanding.md) | prompt engineering, output validation, fallbacks, dedupe, date resolution | `capture` | `POST /v1/inputs` | `core` |
| 4 | [Progressive capture](capabilities/04-progressive-capture.md) | concurrency, retry and crash-safety around background LLM enrichment | `capture` | `POST /v1/inputs/:id/enrich` | `core` |
| 5 | [Memory](capabilities/05-memory.md) | a memory store that explains itself, with levels, decay and revisions | `memory` | `/v1/memories*` | `core` |
| 6 | [Today & planning](capabilities/06-today-and-planning.md) | scheduling heuristics, risk detection, adaptive limits, dependencies, buffer and migration | `today` | `/v1/today` `/v1/plan/today` `/v1/plan/migrate` | `core` |
| 7 | [Work streams](capabilities/07-work-streams.md) | grouping that is rebuildable and never touches the source of truth | `workstreams` | `GET /v1/work-streams` | `core` |
| 8 | [HTTP API](capabilities/08-http-api.md) | the REST layer, storage adapters and deployment | — | all 62 routes | `apps/server` |
| 9 | [MCP server](capabilities/09-mcp.md) | the Model Context Protocol surface and tool selection | — | — | `packages/mcp` |

## Modules are decoupled in code

Every capability service depends only on the storage port, the domain types and
pure helpers — **there are zero dependencies between the services themselves**:

```
services/*  ←  domain/types · db/port · lib/* · repos/actionLogs
```

That is what makes "use just this one" real, and it is enforced by the module
graph rather than by convention.

## One shared contract

Capabilities 1–7 read and write the same schema. "Bring one module" therefore
means **bring our migrations**:

```ts
import { initialize } from "@aldus-palace/core";
await initialize(db);   // canonical DDL + forward-only migrations
```

Migrations are forward-only and versioned in `schema_migrations`, so a database
created by an older release upgrades on startup. `spec/schema.sql` is generated
from the canonical schema and checked in CI (`pnpm spec:check`).

## Three ways in

| Level | You write | Best for |
|---|---|---|
| **MCP** | a JSON config block | using it inside Claude, Cursor or any MCP client |
| **HTTP** | `fetch` / `curl` / any language | your own frontend, mobile app or service |
| **Library** | TypeScript | embedding the runtime in your product |

In TypeScript, `@aldus-palace/client` is a typed wrapper over the HTTP API: one
method per route, no runtime dependencies.

Start at [`INTEGRATION.md`](INTEGRATION.md).


<!-- docs/INTELLIGENCE.md -->

# Core intelligence & privacy design

> **This is design, not a promise about the current release.** It records the
> system Aldus Palace is being built toward. Every section is split into what
> **ships today** and what is **designed**, so the two cannot be confused.

Aldus Palace is not a chat model with a database behind it. It is a system that
runs continuously: it reads input, decides what to remember, decides what it may
do, arranges real work, and calls a model only where a model helps.

## How to read this

| Marker | Meaning |
|---|---|
| **Shipped today** | runs in the current `0.x` release; the file that implements it is named |
| **Designed** | the direction the system is being built toward; tracked in [`ROADMAP.md`](../ROADMAP.md) |

Where this document disagrees with the code, `spec/schema.sql` or an
[ADR](adr), the code and the ADR win — the same rule as the
[design archive](internal/design-archive/README.md). The archive holds the
original Chinese essays these sections condense.

## The loop that ships

```
user / environment
      |
input          raw text, stored as written
      |
understanding  intent, typed objects, resolved dates, gates
      |
memory         candidates, evidence, confirmation, conflicts, versions
      |
planning       four kinds of time, today, risk, adaptive limits
      |
context        active memories and projects feed the next capture
```

The **Core Intelligence Layer** organises the middle of that loop into four
engines. Understanding — the capture front end in
[`agent/understand.ts`](../packages/core/src/agent/understand.ts) — is what
turns a sentence into typed objects today; the engines below decide what is kept,
what happens next, what the system may do on its own, and which model is used.

| Engine | Goal | Shipped today | Designed |
|---|---|---|---|
| **Memory** | understand a person over years, not store a chat log | extraction, pollution gate, activation rule, evidence, dedupe, conflict, versioning, FTS5 retrieval, graded levels, decay, value score | more kinds and signals |
| **Planning** | keep what matters happening while the environment changes | four kinds of time, priority scoring, slot search, Today, risk, adaptive limits, feedback model, constraints, duration estimate, scored Now, morning plan, buffer, migration | blended priority, rhythm-aware Now, schedule optimization |
| **Trust & autonomy** | widen what the system may do on its own, safely | the risk table, autonomy levels, trust score, permission evolution, durable execution | proactive rules |
| **Model orchestration** | use the right model for each job | one provider interface, three implementations, the privacy guard at the boundary | routing by task |

---

## Memory Intelligence Engine

*The goal is to understand one person over years — not to archive a chat log.*

### Shipped today

- **Extraction** reads durability markers ("from now on", "as a rule") and
  repeated behaviour. Rules run with no model; a model adds general
  understanding.
- **Evaluation** drops what should not be remembered: a temporary state, a
  one-off creative fragment, a low-confidence guess.
- **Activation** follows one published rule — `confidence >= 0.8` and
  `importance >= 0.8` — everything below waits as a candidate.
- **Evidence on every row**: the excerpt it came from, a confidence value and the
  input id.
- **Duplicates collapse** to a normalised key (for example
  `preference|prefer_simplicity`).
- **Conflict detection** compares a candidate against active memories on the same
  dimension and reports the contradiction instead of storing both.
- **Versioning** marks a replaced belief `superseded` with a reason, and keeps it
  readable. Nothing is deleted.
- **Retrieval** injects active memories into the next capture, and every
  injection lands in the action log.
- **Five kinds** ship: preference, project context, principle, decision,
  experience.
- **A graded value model** (`lib/memoryValue.ts`): the kinds map onto levels
  0–3, each carries a decay half-life, and a value score (explicitness,
  frequency, impact, scope, future relevance) weighs retrieval. A fresh
  principle outranks an old experience on the same topic.

`lib/memoryExtract.ts` · `lib/memoryActivation.ts` ·
`services/memoryLifecycle.ts` · `services/memoryEvolution.ts`

### Designed

- **An identity level.** Levels 0–3 ship; identity (4) needs an identity kind
  before it can exist.
- **The full formation pipeline** — user experience → extraction → candidate →
  evaluation → conflict check → storage → activation → retrieval. Most stages
  ship; the candidate lifecycle is the part that keeps growing.
- **Four extraction signals** — a long-term phrase, repeated behaviour, impact on
  future decisions, and reach across projects. The first two ship; impact and
  scope extend the same extractor.
- **More kinds** — goal, relationship, knowledge, habit and episode, each with
  its own lifetime and evidence rules.
- **Decision memory keeps the *why*** — What, Why, When, Status — not only what
  was chosen.
- **A three-layer store, a memory graph, richer context assembly and a user
  memory control centre.**

---

## Planning Intelligence Engine

*This is the core that turns understanding into things happening: it decides when
and in what order work occurs in the real world. The goal is not a pretty
calendar — it is important work done with the least cognitive load while the
environment changes. Where a traditional calendar holds fixed blocks you adjust
by hand, this plans from goals, constraints and resources, and keeps adjusting.*

### Shipped today

- **Inputs are Commitments, not Tasks.** Planning reads commitments, the calendar
  (`events` of kind `fixed_external`) and a learned behaviour model.
- **Four kinds of time held apart** — deadline, availability window, suggested
  slot, unscheduled. There is no `overdue` state to occupy.
- **Constraints, concretely** — a deadline is a hard boundary and sets the risk
  tiers; an availability window bounds when work is eligible; a project
  preference is learned from behaviour.
- **Priority scoring** (`scoreCandidate`): risk and a deadline within 24 h or
  72 h form tiers; then recent-project continuity, an actionable title, a short
  duration, importance, title overlap with recent completions, learned project
  weights, and recent self-defined / short-next-step preferences.
- **Time-window generation and conflict avoidance** (`findSlot`): walk 15-minute
  steps from now to the end of the day, skipping AI slots and fixed external
  events, and take the next window that fits.
- **Duration estimation** (`estimateDurationMinutes`): a stated estimate keeps
  the larger weight and is calibrated against the median of completed work in
  the same project; two samples or more fill a missing estimate.
- **Dependency constraints.** A commitment can wait on another; the planner
  skips it while a blocker is open, a finished blocker releases it, and cycles
  are refused (`services/dependencies.ts`).
- **A daily buffer.** Capacity is the 09:00–18:00 window minus 25%; auto-fill
  counts the minutes already on the day and stops before the day is full
  (`planCapacityMinutes`).
- **Migration for slipped flexible work.** An open, unstarted commitment with no
  deadline whose slot or window ended moves forward: the slot is cleared and the
  deferral is counted. At three deferrals it surfaces for a decision instead
  (`services/workMigration.ts`).
- **Scheduling** writes `ai_slot_start/end`, a `today_assignments` row carrying a
  human-readable reason, and an `action_log` entry.
- **Execution monitoring** (`observePlanningOutcome`): record what the user did
  next — started, completed, scheduled today, created a commitment — as a
  feedback episode (project switch, same-project switch, self-defined task,
  stopped working).
- **Replanning** is a reconcile pass keyed on a plan version, triggered by the
  plan endpoint and when a new commitment is arranged for today; it can add up to
  three items to an empty or light day, and stall detection pauses auto-fill.
- **A scored Now** (`lib/nowScore.ts`): urgency, importance, whether the work
  fits the time left and whether it continues the current context decide the
  current action, and the reason travels with it.
- **A morning plan** (`lib/dayPlan.ts`): the day is classified into core,
  optional and deferred.
- **Replanning on change** (`services/replan.ts`): finishing something
  re-derives the day; a new task and a removal already did. A planning failure
  never fails the change that triggered it.
- **The day view**: Now (exactly one thing), timeline, risks (what replaces
  overdue), unscheduled, and a rest suggestion when the day is full.
- **Adaptive limits**: automatic additions stop at 5, or 10 after a deliberate
  add; a stalled queue of 1–3 items with no completion for 24 h pauses auto-fill.
- **Light triage**: an empty day is filled preferring concrete bugs and small
  executable work, and deprioritizing research or long epics.
- **A learned behaviour model** over a rolling 15-day window, kept as reversible
  planning state — never memory, and never overriding a deadline you set
  ([ADR 0001](adr/0001-keep-adaptive-planning-state-outside-memory.md)).

`services/today.ts` · `services/planToday.ts` · `services/adaptivePlanning.ts`

### Designed

- **A fuller pipeline** — constraint analysis → priority calculation →
  time-window generation → schedule optimization → conflict resolution →
  execution monitoring → replanning, with each stage carrying more of the model.
- **Soft constraints** — preferences that trade off against each other, not only
  rules that hold or fail.
- **A blended priority score** — impact and goal alignment join urgency,
  dependencies, context and time fit.
- **Complexity in duration estimation** — read task complexity alongside the
  history, not only the stated estimate and the project median.
- **A rhythm-aware Now** — energy match and a user rhythm join priority,
  available time and context match.
- **Event-driven replanning** — react to a postponed meeting, a new task,
  finishing early, or a change in state.
- **Buffer management by user rhythm** — keep a share of the day free that
  follows energy patterns, not only a fixed ratio.
- **Migration with a richer rule set** — dependencies and goal alignment decide
  what moves, not only the slot and the window.

---

## Trust & Autonomy Engine

*Widen what the system may do on its own — safely, and only as far as it has
earned.*

### Shipped today

- **One fixed rule.** A capture lands on its own; a principle the user states
  takes effect; everything else waits for the user. The runtime writes an
  `action_log` entry for every mutation.
- **An action gate** (`services/actionGate.ts`). Every proposed agent action is
  graded against a published table: low and medium run (medium is recorded as a
  notification), high waits for one approval, critical needs two, and an
  unclassified action waits. `action_proposals` holds the trail; a revocation is
  a status change, and every step writes an `action_log` entry.
- **Durable execution.** An approved action carries an executable descriptor, an
  idempotency key and an execution status; a registered executor runs it under a
  lease, the outcome is recorded, and a succeeded action never runs twice
  (`executeApprovedAction`). Deletion is the first real one: `POST /v1/me/purge`
  proposes `user_data_purge` — permanent, so it is critical and takes two
  approvals, which then run it.
- **A trust score and autonomy levels 0–4.** Trust is the Laplace-smoothed
  approval rate of the decisions the user made in the last 90 days; automatic
  runs do not count, so trust grows from decisions. The level follows thresholds
  with minimum samples.
- **Permission evolution with a user ceiling.** The effective level is
  `min(max(earned, 2), ceiling)`: the published rule is the floor, the user's
  ceiling (default 2, up to 4) is the consent, and the gate applies the result
  by default. Trust rises and falls with the record; high-risk autonomy needs the
  ceiling raised. `GET /v1/autonomy` and the MCP card expose the state.

### Designed

- **Proactive rules**, judged on evidence, pattern and value.
- **Routing the capture and planning flows through the gate.** Their action
  types are graded low today, so the gate is available without changing them.

---

## Model Orchestration Engine

*Use the right model for each job, instead of one model for every call.*

### Shipped today

- **One `LLMProvider` interface** and three implementations — dev
  (deterministic, offline), OpenAI-compatible and Anthropic. Configuration is
  resolved by the caller (`resolveProviderConfig`) and passed in explicitly; the
  runtime reads no environment variables.

`providers/`

### Designed

- **Routing by task**: a fast model for classification, a reasoning model for
  planning and conflict, an embedding model for memory retrieval, and a local
  model for sensitive input.

---

## Privacy & Security Architecture

*The system touches a person's work, decisions, relationships and habits. Privacy
is the shape of it, not a feature on top.*

### Shipped today

- A **SQLite file you own**, or one Cloudflare Durable Object — no vendor cloud.
- **Single-user** runtime: one static bearer token guards the API.
- `raw_inputs` is **immutable**; the AI pass writes only derived fields.
- Model output is **validated and gated** before it reaches storage.
- Every mutation writes an **`action_log`** entry.
- A memory becomes active only through the **published gate**
  (`confidence >= 0.8` and `importance >= 0.8`); an inferred principle waits for
  the user.
- **A Privacy Gateway** (`services/privacyGateway.ts`): a cloud call is prepared
  by data level — project names, money, emails and phone numbers are replaced,
  level 4 stays local, and each call is logged without its content.
- **The gateway is a boundary** (`providers/guard.ts`): a cloud provider is only
  constructed with a message guard, which redacts user-role messages and refuses
  level 4, so a call cannot leave unredacted.
- **Progressive, fine-grained permissions** (`services/permissions.ts`): calendar
  read by default; mail, files and memory scopes on request; Memory is private
  until a memory scope is granted.
- **True deletion** (`services/dataLifecycle.ts`): `purgeUserData` deletes every
  row the user owns in one transaction, after an explicit confirmation.

`SECURITY.md` records the current posture and the threat model.

### The privacy model

**Three principles**

1. **User owns the context.** Memory, Thoughts, Decisions and project context
   belong to the user, not the platform.
2. **Minimum Data Exposure.** Send only what the task needs; never the whole
   database.
3. **Local First.** What can be processed locally is processed locally.

**Five data levels**

| Level | Data | Handling |
|---|---|---|
| 0 | public data | no protection needed |
| 1 | personal preferences | ordinary |
| 2 | work context | high |
| 3 | sensitive work data | very high |
| 4 | private cognitive data — unshared ideas, decision process, business plans, long-term Memory | highest |

**Local intelligence layer, cloud AI layer**

- *Local*: input parsing, simple classification, Memory indexing, sensitive
  detection, calendar reads, basic planning.
- *Cloud*: deep reasoning, long-text analysis, complex planning, high-quality
  generation.

**Privacy Gateway.** Every cloud call passes through:
`user input → sensitive detection → redaction → permission check → cloud AI`.
The real mapping stays local.

> "Discuss Orvia funding with Zhang tomorrow" leaves as
> "discuss [business] funding with [contact] tomorrow".

**Memory storage.** Memory is the most sensitive data, so the design stores it
encrypted on device — Keychain plus an encrypted database on macOS, Secure Enclave
plus encrypted storage on iOS. The cloud receives temporary context by default,
never the full Memory.

**Permissions are progressive and fine-grained.** Calendar first, then mail, then
advanced files. A permission is split by scope (read / create / modify events).
Memory is private by default, with optional Sync or AI Assist.

**The Action Gate.** Every action follows
`propose → risk assessment → permission check → execute → record`, with risk
graded:

| Risk | Behaviour |
|---|---|
| low | automatic |
| medium | notify |
| high | confirm |
| critical | confirm again |

The audit log is viewable and revocable.

**Data lifecycle.** Create → process → store → use → archive → delete. When the
user deletes, the deletion is real: the local database, the full-text memory
index and every derived row are all covered.

### Designed

- **Local encrypted storage** for Memory — Keychain plus an encrypted database on
  macOS, Secure Enclave plus encrypted storage on iOS.
- **A level per call** — today the guard uses one level per provider
  (`PRIVACY_LEVEL`, default 2); a capture about money could ask for a higher one
  at the call site.

## Action items

- [x] A **Local Intelligence Layer** — input parsing, classification, Memory
      indexing and planning run on the device.
- [x] A **Privacy Gateway** and redaction flow, enforced at the provider
      boundary.
- [x] **Progressive and fine-grained permissions**, with Memory private by
      default.
- [x] An **Action Gate** with risk levels, and a viewable, revocable audit log.
- [x] A **delete policy** that reaches the local database, the full-text memory
      index and every derived row.
- [ ] **Local encrypted storage** for Memory (Keychain / Secure Enclave plus an
      encrypted database).

---

## Sources

The design above condenses four essays in the
[design archive](internal/design-archive/README.md):

| Archive document | Subject |
|---|---|
| [`13__Core_Intelligence_Specification…`](internal/design-archive/13__Core_Intelligence_Specification_核心智能系统设计_.md) | the four engines |
| [`7__Memory_System_Technical_Design…`](internal/design-archive/7__Memory_System_Technical_Design_长期记忆系统技术设计_.md) | memory types, scoring, decay |
| [`8__AI_Planning_Engine_Technical_Design…`](internal/design-archive/8__AI_Planning_Engine_Technical_Design_智能规划引擎设计_.md) | constraints, priority, replanning |
| [`9__Security___Privacy_Architecture…`](internal/design-archive/9__Security___Privacy_Architecture_安全与隐私架构设计_.md) | principles, gateway, action gate |

See also [`ARCHITECTURE.md`](ARCHITECTURE.md) for what runs today and
[`ROADMAP.md`](../ROADMAP.md) for the status of each designed part.


<!-- docs/EVAL.md -->

# Eval and testing

Two tiers, deliberately separated.

## 1. Deterministic suites — `pnpm test`

Assertion scripts under `packages/*/test` and `apps/server/test`, executed
offline by `tsx`. They cover the parts where a regression would be silent:

| Suite | Locks down |
|---|---|
| `relativeDay` | relative date resolution and its timezone edge cases |
| `modelDateNormalization` | a model's free-text or hallucinated dates are resolved or dropped |
| `clarificationReply` | short replies to time clarifications |
| `thoughtTitle` | summary quality gates, transfer grounding |
| `projectMatch` | project matching by name/alias/description |
| `inputObjectClassification` | thought vs commitment vs mixed |
| `adaptivePlanning` | Today planning, adaptive caps, feedback episodes, the daily buffer |
| `workMigration` | slipped flexible work, deferral counting, the confirmation threshold |
| `dependencies` | blocked-by edges, cycle rejection, planner eligibility, the audit trail |
| `planningIntelligence` | the Now score, the core/optional/deferred plan, context match |
| `memoryValue` | levels, decay half-lives, the value score |
| `memoryActivation` | which memories activate for a capture, and the evidence |
| `retriever` | FTS5 search, CJK segmentation, backfill, re-index and removal |
| `privacy` | redaction by data level, the gateway, permission scopes, true deletion |
| `client` (packages/client) | every method's path, method and body, plus error mapping |
| `enrichmentLease` | concurrent enrichment and supersede semantics |
| `commitmentOriginalInput` | optimized content vs original input |
| `commitmentClassification` | work-stream projection, `/v1` API, lease races |
| `memoryEvolution` | rule-based conflict detection, supersede, version chains |
| `actionableWork` | English imperatives, Chinese build verbs, non-work cases |
| `actionGate` | the published risk table, decisions, two-step critical approval, revocation, audit |
| `actionGate` (apps/server) | the gate over HTTP: propose, decide, execute, purge, races |
| `trustScore` | the Laplace-smoothed score, level thresholds, what each level runs, the decision window |
| `durationEstimate` | stated/history blending, the median, bounds, invalid input |
| `format` | the locale-aware capture card and Today text projection |
| `providers` | request shaping for Anthropic and OpenAI-compatible providers |
| `language` | the dev provider follows the input script; rule memories follow the user's language |
| `reclassify` (apps/server) | the reclassify API: ask, answer, correct, and learn |
| `llm-config` | wrangler and code agree on the default model |
| `mcp tools` (packages/mcp) | the tool surface and profiles run on the offline provider |

```bash
pnpm test                       # every package
pnpm --filter @aldus-palace/core test
```

No API key, no network, no shared state: `createTestDb()` gives every suite its
own in-memory database, migrated exactly like production.

## 2. Acceptance fixtures — `pnpm eval`

`eval/fixtures/*.json` are user-visible behavioural requirements. Each fixture is
an input plus the properties the result must have:

```json
{
  "id": "S04",
  "input": "最近觉得 AI 产品都太吵了,干扰太多",
  "expect": {
    "thoughts_min": 1,
    "commitments_max": 0,
    "thought_types_allowed": ["idea", "insight", "observation", "research", "decision_candidate"],
    "forbidden": ["thought_type:principle", "recurrence"]
  }
}
```

Supported expectations: `thoughts_min`/`thoughts_max`,
`commitments_min`/`commitments_max`, `memory_candidates_min`/`memory_candidates_max`,
`memory_active_min`/`memory_active_max`, `memory_pending_min`/`memory_pending_max`,
`thought_types_allowed`, `memory_types_allowed`, `must_include_thought`, and
`forbidden` markers (`thought_type:*`, `commitment:*`, `recurrence`, …). Every
fixture key is enforced; an unknown key would not fail on its own, so keep to
this list.

They run against the **deterministic** provider, so results are stable and the
suite is safe to require on every pull request.

```bash
pnpm eval
# PASS S04 … PASS S31
# All fixtures passed.
```

### Adding a fixture

Contributing a fixture is the best first contribution — it is data, not code:

1. Pick the next `S<n>` id (`S32`, …).
2. Copy the shape above; write the input in any language.
3. State only properties that must hold, never the exact model output.
4. Run `pnpm eval` and open a PR explaining the user behaviour you are pinning.

If a fixture documents behaviour that the current rules get wrong, open the PR
with the fixture and mark it `xfail` in the description — that is a valuable bug
report.

## The offline provider's reach

`LLM_PROVIDER=dev` recognises a defined set of patterns: commands ("Ship the
onboarding page next week"), stated rules ("I prefer simple tools"), platform
decisions ("Stay Mac-only, skip Windows"), temporary states, one-off creative
work, and a few date forms, in English and Chinese. It resolves relative dates,
skips near-duplicates and gates memory candidates.

Connect a model (`anthropic`, `deepseek`, `openai-compatible`) for general
understanding: arbitrary phrasing, multi-paragraph pastes, and inferences the
rules do not cover. The capture receipt shows which provider produced the
result, and the deterministic path stays available as the fallback.

## What is *not* tested offline

Live-model quality. `LLM_PROVIDER=deepseek` (or any OpenAI-compatible endpoint)
is exercised manually; there is no CI job that depends on a paid key. When you
change the understanding prompt, re-run the acceptance fixtures **and** capture
a few real inputs with a live provider before merging.


<!-- docs/POSITIONING.md -->

# Positioning

## The one-liner

> **A trustworthy context layer for AI that remembers, plans and acts.**
>
> A self-hosted runtime that turns free-form input into typed, explainable
> objects — usable over MCP, HTTP, a library or a typed client.

## What it is

Aldus Palace turns free-form input into structured objects a program can act on —
and keeps them honest over time.

You speak or type. The runtime decides what it is (a thought, a commitment, a
decision, something worth remembering), when it is due, whether you have said it
before, and whether it contradicts something it already believes about you. It
records all of it in a database you own.

It is not a todo app, a calendar client, a note editor or a chatbot wrapper.
Tasks and calendar entries are *projections* of deeper objects.

## The three things worth telling someone about

### 1. Delete "overdue" from your vocabulary

Time is not an attribute of a task; it is the **result of scheduling**.

Most tools model one field — `dueDate` — and then punish you with a red badge
when it passes. Aldus Palace separates four different kinds of time:

| Kind | Field | Meaning |
|---|---|---|
| Deadline | `deadline` | the world imposes this |
| Availability window | `window_start` / `window_end` | it can happen any time in here |
| AI-suggested slot | `ai_slot_start` / `ai_slot_end` | a suggestion, not a promise |
| Nothing yet | — | unscheduled work, still visible |

Statuses are `captured → planned → scheduled → completed`, with `risk` and
`cancelled`. **There is no `overdue` state in the schema, because there is no
such thing.** A missed date becomes a risk you can see and rearrange.

*Proof:* `packages/core/src/db/schema.ts` (`commitments`), `services/today.ts`.

### 2. Memory you can audit — and that knows when you changed your mind

Assistant memory is usually a black box: something gets written, nobody knows
why, and it stays as written.

Here, every memory can explain itself and can be taken back:

- **High-confidence memories take effect on capture.** Rules you state, and
  inferences the system is confident about, start working at once. Everything
  else waits as a candidate until you confirm it.
- **Every row says why it is active.** `activation_note` distinguishes "stated by
  you", "confirmed by you" and "inferred during a capture", with the confidence
  and importance that decided it.
- **Evidence on every row.** Each memory carries the excerpt it came from, a
  confidence value, and the input id. Archive any of them, including one the
  system stored on its own.
- **Temporary states are rejected.** “I'm tired today” stays a mood; the gate
  drops it before it reaches your profile.
- **Duplicates collapse.** Different phrasings of the same principle normalise to
  one key and merge.
- **Contradictions surface.** Say the opposite of a confirmed belief and the
  candidate is flagged with the memory it conflicts with.
- **Replacing keeps history.** Confirming a replacement marks the old memory
  `superseded` with a pointer and a reason. Nothing is deleted, so “why do you
  think that about me?” has an answer.

*Proof:* `services/memoryLifecycle.ts`, `services/memoryEvolution.ts`,
`lib/memoryExtract.ts`, `test/memoryEvolution.test.ts`.

### 3. The raw input stays as written

Everything you say is stored verbatim in `raw_inputs`. Models
only fill *derived* fields, and every write passes server-side gates:

- one object mode per capture (a thought cannot become the same commitment twice)
- near-duplicate commitments are skipped, not duplicated
- relative dates (“next Wednesday”) are resolved on the server, not trusted to a prompt
- a failed model pass leaves the deterministic result in place instead of an empty record

That makes the AI layer **auditable**: you can diff what you said against
what the system stored.

*Proof:* `agent/understand.ts`, `eval/fixtures/`, `test/inputObjectClassification.test.ts`.

## Two more that developers feel immediately

**Background enrichment you can trust.** Waiting on a model is where products
lose people. The runtime answers in milliseconds with a deterministic result,
then lets the model replace it under a lease — so retries, concurrent clients and
crashed workers cannot corrupt or duplicate anything. See
[ADR 0003](adr/0003-progressive-capture-with-enrichment-leases.md).

**An agent you can run in CI.** A deterministic provider plus
replayable acceptance fixtures means the whole pipeline is testable offline, with
no API key. `pnpm verify` is green in a fresh clone.

## How it compares

| | Aldus Palace | Todoist / Things | Notion / PKM | Claude / ChatGPT memory | Motion / Reclaim |
|---|---|---|---|---|---|
| Who structures your input | the runtime decides | you do | you do | assistant, conversationally | partially |
| Time model | 4 kinds, held apart | one due date | free text | conversation only | calendar slots |
| Memory | a gate, candidates, evidence, versioning, levels and decay | saved views you maintain | documents you maintain | assistant memory, inside the app | learned preferences |
| Data location | your SQLite file or your Worker | vendor cloud | vendor cloud | vendor cloud | vendor cloud |
| Programmable | MCP · HTTP · library · typed client | API | API | in-app | calendar API |
| Replayable offline | 30 suites and 11 fixtures, no key | n/a | n/a | requires the service | requires the service |

## Who it is for

- **Developers building an AI product** who need commitments, memory or planning
  and would rather not invent a domain model, a memory gate and a scheduler.
- **People who want their context to be theirs** — one SQLite file, portable,
  inspectable, self-hosted, and shared across whatever assistant they happen to use.
- **Teams that need AI-written data to be auditable** — evidence, provenance and
  an action log on every change.

## How you run it

- **One person, one instance.** One database and one bearer token, per person.
- **Self-hosted.** Your SQLite file, or Cloudflare's edge. No telemetry.
- **You decide what acts.** The runtime records intent and plans; acting on the
  world is your call.
- **MCP-level clients.** The UI is yours to build. The integration surfaces are
  MCP, HTTP, the library and a typed client.


<!-- docs/MAP.md -->

# Map: capability → code → doc → test

Use this instead of searching. Every capability names the file that implements
it, the doc that explains it and the suite that pins it.

| Capability | Code | Doc | Test |
|---|---|---|---|
| Schema & domain | `packages/core/src/db/schema.ts` | [01](capabilities/01-schema-and-domain.md) | `pnpm spec:check` |
| Providers | `packages/core/src/providers/` | [02](capabilities/02-providers.md) | `packages/core/test/providers.test.ts` |
| Understanding | `packages/core/src/agent/understand.ts` | [03](capabilities/03-understanding.md) | `inputObjectClassification.test.ts`, `relativeDay.test.ts`, `modelDateNormalization.test.ts`, `thoughtTitle.test.ts`, `projectMatch.test.ts`, `language.test.ts` |
| Grey-zone classification | `lib/objectAmbiguity.ts`, `lib/classificationSignals.ts`, `services/reclassify.ts` | [ADR 0012](adr/0012-the-grey-zone-asks-and-learns.md) | `inputObjectClassification.test.ts`, `apps/server/test/reclassify.test.ts` |
| Progressive capture | `packages/core/src/services/enrichmentLease.ts` | [04](capabilities/04-progressive-capture.md) | `packages/core/test/enrichmentLease.test.ts` |
| Memory | `lib/memoryExtract.ts`, `lib/memoryActivation.ts`, `lib/memoryValue.ts`, `lib/retriever.ts`, `lib/search.ts`, `services/memoryLifecycle.ts`, `services/memoryEvolution.ts` | [05](capabilities/05-memory.md) | `memoryActivation.test.ts`, `memoryEvolution.test.ts`, `memoryValue.test.ts`, `retriever.test.ts` |
| Today & planning | `services/today.ts`, `services/planToday.ts`, `services/adaptivePlanning.ts`, `services/workMigration.ts`, `services/dependencies.ts`, `services/replan.ts`, `lib/nowScore.ts`, `lib/dayPlan.ts` | [06](capabilities/06-today-and-planning.md) | `adaptivePlanning.test.ts`, `workMigration.test.ts`, `dependencies.test.ts`, `planningIntelligence.test.ts` |
| Work streams | `services/workStreams.ts`, `services/commitmentClassification.ts` | [07](capabilities/07-work-streams.md) | `apps/server/test/commitmentClassification.test.ts` |
| HTTP API | `apps/server/src/` | [08](capabilities/08-http-api.md) | `apps/server/test/` |
| MCP server | `packages/mcp/src/` | [09](capabilities/09-mcp.md) | `packages/mcp/test/tools.test.ts` |
| HTTP client | `packages/client/src/` | [client README](../packages/client/README.md) | `packages/client/test/client.test.ts` |
| Text projections | `packages/core/src/lib/format.ts` | [EVALUATION](../EVALUATION.md) | `packages/core/test/format.test.ts` |
| Action Gate | `packages/core/src/services/actionGate.ts`, `apps/server/src/actions.ts` | [INTELLIGENCE](INTELLIGENCE.md) · [ADR 0005](adr/0005-action-gate-with-published-risk.md) | `packages/core/test/actionGate.test.ts`, `apps/server/test/actionGate.test.ts` |
| Privacy | `providers/guard.ts`, `services/privacyGateway.ts`, `services/permissions.ts`, `services/dataLifecycle.ts`, `lib/redaction.ts` | [INTELLIGENCE](INTELLIGENCE.md) · [SECURITY](../SECURITY.md) | `packages/core/test/privacy.test.ts` |
| Acceptance fixtures | `eval/fixtures/` | [EVAL](EVAL.md) | `pnpm eval` |
| The engine design | `docs/INTELLIGENCE.md` (design) | [INTELLIGENCE](INTELLIGENCE.md) | shipped parts above |

## Where to start reading

| Question | File |
|---|---|
| What does the product do? | [`README.md`](../README.md) |
| Who is it for, and when? | [`README.md`](../README.md#who-its-for-and-when) |
| How is the runtime put together? | [`ARCHITECTURE.md`](ARCHITECTURE.md) |
| What are the objects and invariants? | [`DOMAIN-SCHEMA.md`](DOMAIN-SCHEMA.md) |
| Why is a decision the way it is? | [`adr/`](adr) |
| What is planned, and what is not? | [`ROADMAP.md`](../ROADMAP.md) |
| How do I verify a claim? | [`EVALUATION.md`](../EVALUATION.md) |
| What does a word mean here? | [`GLOSSARY.md`](GLOSSARY.md) |
| How does retrieval scale? | [`RETRIEVER.md`](RETRIEVER.md) |

## The runtime loop

```
raw input → understanding → thoughts / commitments / decisions / memory
                │
                ├─ Today (projection)
                ├─ Work streams (projection)
                └─ Memory lifecycle (candidate → active → superseded)
```


<!-- docs/GLOSSARY.md -->

# Glossary

The shared vocabulary. Planning-specific terms live in
[`CONTEXT.md`](../CONTEXT.md).

## Objects

| Term | Meaning | Table |
|---|---|---|
| **Raw input** | the user's words, stored as written and never edited | `raw_inputs` |
| **Thought** | something worth keeping that is not yet work: an idea, insight, observation, research note or decision candidate | `thoughts` |
| **Commitment** | something the user intends to do; carries the four kinds of time and a status | `commitments` |
| **Decision** | a choice, with the reasoning attached | `decisions` |
| **Memory** | a durable belief about the user: preference, project context, principle, decision or experience | `memories` |
| **Project** | a durable container that commitments and memories can point at | `projects` |
| **Concept** | a reusable cognitive node (for example "simplicity") linked from memories | `concepts` |

## Time and status

| Term | Meaning |
|---|---|
| **Deadline** | the world imposes this time |
| **Availability window** | the work can happen anywhere inside `window_start` / `window_end` |
| **Suggested slot** | an agent proposal in `ai_slot_start` / `ai_slot_end`; not a promise |
| **Unscheduled** | open work with no time attached; still visible |
| **Risk** | the state that replaces `overdue`: a date that needs attention and can be moved |
| **Buffer** | the share of the daytime window the planner keeps free (a quarter by default) |
| **Deferral** | one automatic move of a slipped, flexible, unstarted commitment; counted on the row |
| **Migration** | clearing a slipped slot so the planner can place the work again; after three deferrals the item asks for a decision |
| **Dependency** | a commitment that waits on another; the planner skips it while a blocker is open, and cycles are refused |
| **Now score** | urgency, importance, time fit and context match decide the current action; the reason travels with it |
| **Core / optional / deferred** | the morning classification: must happen, fits the remaining capacity, or waits |
| **Replanning** | re-deriving the day after a change — a completion, a new task, a removal; a planning failure never fails the change |
| **Status** | `captured → planned → scheduled → completed`, plus `risk` and `cancelled` |

## Memory lifecycle

| Term | Meaning |
|---|---|
| **Candidate** | a proposed memory waiting for confirmation |
| **Active** | a memory in effect; `activation_note` says whether the user stated it, confirmed it, or the system inferred it |
| **Superseded** | a replaced belief, kept readable in the version chain |
| **Archived** | a memory the user removed from use |
| **Evidence** | the excerpt, confidence and input id behind a memory |
| **Activation gate** | `confidence >= 0.8` and `importance >= 0.8`; everything below waits as a candidate |
| **Memory level** | how much a kind shapes behaviour: experience 0, decision/project context 1, preference 2, principle 3 |
| **Memory decay** | the half-life of a kind — a month for an experience, years for a principle |
| **Value score** | explicitness, frequency, impact, scope and future relevance averaged into one weight |
| **Retriever** | the port that ranks memories for a query; the default uses FTS5 (bm25) blended with level, decay and value (`lib/retriever.ts`) |

## Runtime

| Term | Meaning |
|---|---|
| **Understanding** | the capture front end: a sentence into typed objects, dates resolved, duplicates skipped |
| **Progressive capture** | the local pass writes first; the model replaces derived fields under a lease |
| **Enrichment lease** | the mechanism that makes "local first, model second" idempotent under retries |
| **Action log** | an append-only row written on every mutation, with a reason |
| **Action Gate** | the published risk table that decides whether a proposed agent action runs or waits; a decision is logged and revocable (`services/actionGate.ts`) |
| **Data level** | 0 public, 1 preference, 2 work context, 3 sensitive work, 4 private cognitive; level 4 never leaves the device |
| **Privacy Gateway** | prepares a cloud call: sensitive detection, redaction, permission check, then the model; enforced at the provider boundary by the message guard |
| **Message guard** | the port a cloud provider calls before sending; redacts user-role messages by level, refuses level 4, and writes one audit row per call for every outcome including a pass-through (`providers/guard.ts`) |
| **Permission scope** | a grantable capability (`calendar.read`, `mail.read`, `memory.ai_assist`, …); absent means not granted |
| **Purge** | true deletion of every row the user owns (memory index and classification signals included), in one transaction; a critical action, so it takes two approvals |
| **Clarification** | a pending question about a record — a relative day, or what a capture is; answered by option or in words, never a guess (`clarifications`, ADR 0012) |
| **Classification signal** | the terms of an input the user corrected, with the mode they chose; applied before asking again (`classification_signals`, ADR 0012) |
| **Execution** | running an approved action under a lease, once, with the result or error recorded (`executeApprovedAction`) |
| **Executor** | the function registered per action type that performs the effect; only the composition root knows effects |
| **Trust score** | the Laplace-smoothed approval rate of the decisions the user made in the last 90 days; automatic runs do not count |
| **Autonomy level** | 0–4, derived from the trust score with minimum samples |
| **Autonomy ceiling** | how far earned trust may widen autonomy (default 2, up to 4); raising it is the user's explicit consent |
| **Effective level** | `min(max(earned, 2), ceiling)` — what the gate applies; the published rule is the floor |
| **Storage port** | the async `SqlDatabase` interface; the only I/O boundary in `packages/core` |
| **Projection** | a view rebuilt from the source of truth, never a second truth (Today, work streams) |

## Documentation markers

| Marker | Meaning |
|---|---|
| **Shipped today** | runs in the current `0.x` release |
| **Designed** | the direction the system is being built toward; tracked in [`ROADMAP.md`](../ROADMAP.md) |


<!-- docs/RETRIEVER.md -->

# Retriever

How the runtime finds the memories that matter, and the design for scaling it
from dozens to thousands.

## The question

Memory value already ships: levels, decay and a value score weigh what to keep
(`lib/memoryValue.ts`). Retrieval is the other half. FTS5 now ranks candidates
with bm25 and the value score re-ranks them, on both runtimes; keyword overlap
remains only as the fallback when the index has no match. The port leaves room
for more as the store grows, without giving up local-first.

## Spike findings (2026-09-19)

| Question | Result |
|---|---|
| FTS5 on the local adapter (`better-sqlite3`, SQLite 3.53.2) | works, with `bm25()` ranking |
| FTS5 on Cloudflare Durable Object SQLite | supported; Cloudflare lists the FTS5 module (including `fts5vocab`) among the supported extensions |
| CJK tokenization | `unicode61` keeps a CJK run as one token, and `trigram` misses two-character queries; separating CJK characters and matching a quoted phrase works for any length (`lib/search.ts`) |
| SQL triggers for CJK | a trigger cannot segment; the index stores text the runtime already segmented |
| Vector search (`sqlite-vec` or similar) | **not portable**: a native extension, absent from the Durable Object supported set and not bundled with `better-sqlite3` |
| Same schema on both runtimes | FTS5 keeps the "one schema, two runtimes" promise |

The decision follows from the spike: **full-text search is the default, and
embeddings are an optional adapter.**

## Design

```
query ──► Retriever port
               ├─ default: FTS5 (portable on both runtimes)
               │     bm25 + level + decay + value + importance
               └─ optional: embeddings + reranker (opt-in, never required)
```

- **A port, not a concrete dependency.** `Retriever` takes the storage port and
  returns ranked memory ids; callers (context assembly, the memory list) do not
  know how the ranking happened.
- **Segmented text the runtime writes.** `lib/search.ts` turns CJK into single
  characters (`segmentForSearch`) and a user query into an FTS5 phrase
  (`toMatchQuery`). A `search_text` column on `memories` holds the segmented
  form, and `indexMemory` writes it into `memory_search(memory_id UNINDEXED,
  search_text)` — the runtime does this because a trigger cannot segment. A
  query is capped at 512 characters and 24 terms so a long capture cannot
  become an unbounded OR chain. The
  backfill on first use repairs rows written before the column existed, or that
  lost their search row; a memory with nothing to index keeps an empty marker
  instead of a row.
- **The score blends both worlds.** `bm25` orders the match; the existing
  `memoryRetrievalScore` (level, decay, value, importance) breaks ties and keeps
  a fresh principle ahead of an old experience on the same topic.
- **A fallback stays in code.** If a runtime ever lacks FTS5, retrieval falls
  back to the current keyword scoring, so nothing breaks.
- **Embeddings are optional.** They sit behind the same port as a second
  implementation, and the default path never needs a model or a vector store.

## Slices

1. **FTS5 retriever — shipped.** The `memory_search` table ships in the schema;
   `lib/retriever.ts` indexes a memory (`indexMemory`), backfills on first search
   (`ensureMemoryIndex`), and answers with ranked ids (`retrieveMemoryIds`), and
   `retrieveActiveMemoriesForContext` uses it before the keyword fallback.
   `lib/search.ts` segments CJK. Locked by `packages/core/test/retriever.test.ts`.
2. **Ranking and measurement — shipped.** `bm25` ranks, `memoryRetrievalScore`
   re-ranks, and `pnpm bench:retrieval` reports Recall@K and MRR (1.000 on the
   labeled corpus; `docs/BENCHMARKS.md`).
3. **Optional embeddings.** A second `Retriever` behind the port, opt-in, with no
   change to the default behaviour.

Slices 2 and 3 are recorded in [`ROADMAP.md`](../ROADMAP.md).


<!-- docs/BENCHMARKS.md -->

# Benchmarks

Measured numbers for the **deterministic rule layer**. The corpus ships in this
repository (`packages/core/bench/run.ts`), so every number is reproducible
offline with no API key.

```bash
pnpm bench          # table
pnpm bench --json   # machine-readable
```

## Results (37 contract cases)

| Suite | Pass | Accuracy | Precision | Recall |
|---|---|---|---|---|
| Relative dates (en/zh, fixed clock) | 8/8 | 1.000 | — | — |
| Duplicate commitments | 6/6 | 1.000 | 1.000 | 1.000 |
| Memory conflicts | 5/5 | 1.000 | 1.000 | 1.000 |
| Memory activation | 6/6 | 1.000 | 1.000 | 1.000 |
| Actionable work | 7/7 | 1.000 | 1.000 | 1.000 |
| Object mode | 5/5 | 1.000 | — | — |

A case that fails exits non-zero, so a rule regression fails `pnpm verify`.

## Adding a case

Add an entry to the matching corpus in `packages/core/bench/run.ts`. Each case is
a promise that must hold; a regression is a red run.

## Retrieval

`pnpm bench:retrieval` runs a labeled corpus of twelve memories and ten queries
(English and Chinese) through the FTS5 retriever:

| Metric | Result |
|---|---|
| Recall@1 | 1.000 |
| Recall@3 | 1.000 |
| Recall@5 | 1.000 |
| MRR | 1.000 |

The corpus is small and authored alongside the retriever, so treat it as
regression evidence, not a leaderboard result.

## Live model

[`BENCHMARKS-LLM.md`](BENCHMARKS-LLM.md) runs the same pipeline through a live
model (`pnpm bench:llm`, key-gated) and reports object mode, memory behaviour,
date resolution and latency.


<!-- docs/BENCHMARKS-LLM.md -->

# LLM benchmark

Real captures through `deepseek-chat`, measured by `packages/core/bench/llm/run.ts`.

```bash
DEEPSEEK_API_KEY=… pnpm bench:llm
```

It needs a key and a network, so it is not part of `pnpm verify`. The corpus is
ten labeled captures, English and Chinese, each with an expected object mode,
commitment count, memory behaviour and resolved date.

## Results (deepseek-chat, 10 cases)

| Metric | Result |
|---|---|
| Object mode accuracy | 100% |
| Commitment count | 90% |
| Memory behaviour | 100% |
| Date resolution | 100% |
| Latency p50 / p95 | 1482 ms / 2915 ms |

Object mode covers thought, commitment and mixed. Memory behaviour checks that a
stated rule becomes active, that a mood is dropped, and that a platform decision
becomes a memory. Date resolution checks the resolved window or deadline against
the day the words point at.

## What the run improved

The first run showed date resolution at 25%: the model sometimes returned a
free-text date ("next week", "Friday") or a hallucinated past date, which the
runtime stored as written. `normalizeModelDate` in `agent/understand.ts` now
resolves free text with the same server rules the deterministic path uses, and
drops a past date when the words point at the future; resolution reached 100%.
Locked by `packages/core/test/modelDateNormalization.test.ts`.

The one remaining miss is "Keep Mac only, no Windows version": the model created
a commitment and two memories where the corpus allows none. It is reported, not
hidden.


<!-- ROADMAP.md -->

# Roadmap

This is a public summary. The design each item comes from — with what ships
today marked against what does not — is in
[`docs/INTELLIGENCE.md`](docs/INTELLIGENCE.md). Detailed design history lives in
[`docs/design-archive`](docs/internal/design-archive).

## v0.1 — public foundation (current)

- [x] `@aldus-palace/core`: domain model, agent runtime, storage port, migrations
- [x] Deterministic offline provider + provider-agnostic `LLMProvider`
- [x] Progressive capture with enrichment leases
- [x] Reference server: local SQLite **and** Cloudflare Durable Object adapters
- [x] Two-tier tests: deterministic suites + acceptance fixtures (offline)
- [x] `spec/schema.sql` generated from the canonical schema
- [x] Three published packages (`@aldus-palace/core`, `@aldus-palace/mcp`, `@aldus-palace/client`)
## v0.2 — ecosystem (current)

- [x] Anthropic provider (Messages API) alongside DeepSeek / OpenAI-compatible
- [x] `@aldus-palace/mcp` — capture, today, commitments and memory tools over stdio
- [x] Local SQLite adapter published as `@aldus-palace/core/db/sqlite`
- [x] [Technical write-up](docs/PROGRESSIVE-CAPTURE.md) on the enrichment lease
- [x] `claude mcp add` / Claude Desktop integration documented and smoke-tested
- [x] Published to npm with a release pipeline
- [ ] More acceptance fixtures contributed by users

## Memory, deepened

The shipped memory layer covers extraction, evaluation, activation, conflict
detection, versioning and retrieval. The design extends it:

- [x] **A graded memory model.** Levels 0–3 are derived from the shipped kinds
  (experience, decision/project context, preference, principle) and weigh
  retrieval (`lib/memoryValue.ts`). An identity level needs an identity kind
  first.
- [x] **Memory decay.** Each kind carries a half-life — an experience fades in a
  month, a principle holds for years — and retrieval weighs it.
- [x] **A value score.** Explicitness, frequency, impact, scope and future
  relevance average into one score that retrieval uses. Extraction still reads
  two signals, so frequency and scope stay at their defaults.
- [ ] **More memory kinds.** Goal, relationship, knowledge, habit and episode
  memories, each with its own lifetime and evidence rules.
- [ ] **More extraction signals.** Impact on future decisions, and scope across
  projects, alongside durability markers and repeated behaviour.

## Planning, deepened

The shipped planner covers four kinds of time, priority scoring, slot search,
Today, risk, adaptive limits and a learned behaviour model. The design extends
it:

- [x] **A constraint model** — hard (deadline), availability (window),
  preference (learned project weighting) and dependency (blocked-by, with cycle
  rejection) constraints. Soft constraints that trade off against each other are
  not modelled yet.
- [ ] **A blended priority score** — impact and goal alignment are not read yet;
  urgency, dependencies (as eligibility), context and time fit are.
- [x] **Duration estimation from history** — a stated estimate keeps the larger
  weight and is calibrated against the median of completed work in the same
  project (`lib/durationEstimate.ts`). Complexity is not read yet.
- [x] **Context-switch cost.** Continuing the project the user is already in
  scores higher and switching scores lower (`lib/nowScore.ts`); learned project
  weighting adds to it.
- [x] **An explicit morning plan.** The day is classified into core, optional
  and deferred in the Today payload (`lib/dayPlan.ts`).
- [x] **A scored Now.** Urgency, importance, whether the work fits the time left
  and the context match decide the current action; the reason travels with it.
- [ ] **A rhythm-aware Now** — energy match and a user rhythm are not read yet.
- [x] **Event-driven replanning.** Finishing something re-derives the day
  (`replanAfterChange`); a new task and a removal already did. A postponed
  meeting still needs an events surface.
- [x] **Buffer management** — a quarter of the daytime window stays free
  (`lib/planCapacity.ts`).
- [x] **Task migration** — flexible, unstarted work moves forward, and repeated
  deferrals surface for a decision (`services/workMigration.ts`).

## Retrieval

The design is in [`docs/RETRIEVER.md`](docs/RETRIEVER.md); FTS5 is confirmed on
both runtimes, and vector search is not portable.

- [x] **An FTS5 retriever** — `lib/retriever.ts` keeps `memory_search` in step
  with a segmented `memories.search_text`, backfills on first search, and is
  wired into context retrieval.
- [x] **Ranking and measurement** — bm25 ranks, the value score re-ranks, and
  `pnpm bench:retrieval` reports Recall@K and MRR on a labeled corpus
  (`docs/BENCHMARKS.md`).
- [ ] **Optional embeddings** — a second `Retriever` behind the port, opt-in,
  with no change to the default behaviour.

## Autonomy and models

- [x] **An Action Gate.** A published risk table grades every proposed agent
  action: low and medium run, high waits for one approval, critical needs two,
  and every decision is logged and revocable (`services/actionGate.ts`).
- [x] **A trust score and autonomy levels.** The Laplace-smoothed approval rate
  of decided actions, with level 0–4 thresholds; a caller can pass the level to
  the gate to widen what runs without asking (`services/actionGate.ts`).
- [x] **Permission evolution.** The effective level is
  `min(max(earned, 2), ceiling)`: the published rule is the floor, the user's
  ceiling (default 2, up to 4) is the consent, and the gate applies the result
  by default (`services/actionGate.ts`).
- [ ] **Model orchestration.** One provider interface serves every call today.
  The design routes work by task: a fast model for classification, a reasoning
  model for planning and conflict, embeddings for memory retrieval, and a local
  model for sensitive input.

## Privacy & local intelligence

The shipped runtime is local-first and single-user, and `SECURITY.md` records the
current posture. The privacy architecture takes shape in these pieces:

- [x] **A Local Intelligence Layer** — input parsing, classification, Memory
  indexing and planning all run in the runtime on the device.
- [x] **A Privacy Gateway** with redaction: `prepareCloudPayload` replaces
  project names, money, emails and phone numbers by data level, refuses level 4,
  and logs each call (`services/privacyGateway.ts`).
- [x] **Progressive, fine-grained permissions** — calendar read by default, mail,
  files and memory scopes on request, with Memory private until granted
  (`services/permissions.ts`).
- [x] **An Action Gate** with four risk levels and a viewable, revocable audit
  log (ADR 0005).
- [x] **A delete policy** — `purgeUserData` deletes every row the user owns in
  one transaction (`services/dataLifecycle.ts`).
- [ ] **Local encrypted storage** for Memory — Keychain plus an encrypted
  database on macOS, Secure Enclave plus encrypted storage on iOS.
- [x] **Routing every cloud call through the gateway** — a cloud provider is only
  built with the guard attached, so a call cannot leave unredacted
  (`providers/guard.ts`, ADR 0011).

## Clients

- [x] **A reference client** — a macOS-shaped browser app over the HTTP API
  (Home, Thoughts, Projects, Memory, Decisions, Activity, Input) with a
  quick-capture sheet; ships in the repo and is not hosted
  (`examples/reference-client`).
- [ ] A SwiftUI reference client (`clients/` is a placeholder today)

## Later

- [ ] Postgres adapter behind the existing async port (the port was designed for it)
- [ ] Cognitive-map exploration UI
- [ ] Planning engine beyond Today (horizon, dependencies, energy patterns)

See [AGENTS.md](AGENTS.md) for the product rules behind these.

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.