agentleFS
Sign inSign up

pydantic-ai / pydantic_clai2

pydantic/pydantic-ai/src/pydantic_clai2/AGENTS.md

Read this before touching pydantic-clai2/. The repository-level AGENTS.md still applies (no em-dashes, no Any, pyright strict, keyword-only arguments, 100% branch coverage). This file adds what is specific to the terminal shell. A thin terminal around any Pydantic AI agent. It reads a prompt, runs the agent, streams the answer, and repeats. Everything beyond that is a plugin, including the default coding tools. The three layers, and who owns what: If a change needs the agent loop to behave differently, propose…

AGENTS.md20k starsChanged 8 days ago
# CLAI 2 guide for AI code assistants

Read this before touching `pydantic-clai2/`. The repository-level `AGENTS.md`
still applies (no em-dashes, no `Any`, pyright strict, keyword-only arguments,
100% branch coverage). This file adds what is specific to the terminal shell.

## What CLAI is

A thin terminal around any Pydantic AI agent. It reads a prompt, runs the agent,
streams the answer, and repeats. Everything beyond that is a plugin, including
the default coding tools.

The three layers, and who owns what:

| Layer | Owns | Never does |
|---|---|---|
| Pydantic AI core | the agent loop, hooks, events, toolsets | know CLAI exists |
| `pydantic_ai_harness` | reusable capabilities (`Coder`, `Shell`, ...) | print to a terminal |
| `pydantic_clai2` | the prompt loop, rendering, `/commands`, plugin loading | reimplement a core hook |

If a change needs the agent loop to behave differently, propose it in core. If it
is a reusable behavior with no terminal in it, it belongs in harness. Only the
shell itself lives here.

## The plugin model

A plugin is a module with `activate(host: PluginHost)`, or a bare capability
class. `@host.on(name)` registers a handler for a lifecycle moment;
`@host.on(EventClass)` for a typed event; `host.commands.register` for a
`/command`; `host.add` for tools and instructions; `host.render` for custom
output; `host.settings(Model)` for validated config. `PLUGINS.md` is the user
contract. If code and `PLUGINS.md` disagree, fix one so they agree in the same PR.

## Rules for the plugin API

- **No global state.** Registration goes on a `PluginHost` instance someone
  constructed. No module-level dicts, no import-time side effects. Tests build a
  host and call `activate` directly.
- **Strings are keys, never payloads.** Hook names are a `Literal`; each name has
  an `@overload` binding one typed event dataclass. A handler never receives
  `*args`, `**kwargs`, `dict`, or `context: object = None`.
- **No string sub-dispatch.** A handler does not receive `event_type: str` and
  switch on it. Use the event class as the key.
- **One async spelling.** No `_sync` or `_async` pairs.
- **Observers return `None`.** Deciders mutate the event (`event.text = ...`,
  `event.cancel()`). Nothing collects a list of return values.
- **Fail closed, one way.** A raising handler on a decidable moment cancels the
  action and reports the error. No per-registration "fail open" flag.
- **Host hooks are `<subject>_<moment>`** (`turn_start`). Core hooks keep core's
  names verbatim (`before_tool_execute`). Do not rename a core hook for taste.
- **Payloads are kw-only dataclasses.** Not `BaseModel`. Pydantic is for the
  settings JSON boundary only.
- **No `getattr`/`hasattr` on a plugin** to discover what it supports. It
  registered the thing or it did not.
- **Only four host hooks.** `session_start`, `session_end`, `turn_start`,
  `turn_end`. Adding a fifth needs a use case that core cannot serve; say which
  core hook you checked and why it does not fit.

## Loading and unloading

Plugins load and unload while CLAI runs. The rules that make that safe:

- **One `PluginHost` per plugin.** The host is the ownership scope. Everything a
  plugin registers is recorded on its own host, so unloading is "discard this
  host". No `callback -> owner` map, no scanning registries for a plugin's name.
- **Load and unload only between turns.** `/commands` already run between turns,
  so this falls out for free; do not add a mid-run path.
- **Load fires `session_start` for that plugin; unload fires `session_end`.** A
  plugin cannot tell whether it was loaded at startup or later, and must not
  need to.
- **`reload` is unload, re-import, load.** Drop-in entry modules use fresh source.
  Installed modules use `importlib.reload`, which retains globals absent from the
  new source. Plugins must explicitly initialize their state on activation.
- **Instruction order is capability order.** Placement is a core
  `CapabilityOrdering` (`position`, `wraps`, `wrapped_by`), not a CLAI list.
- **Registration is idempotent per name.** A capability is bound per run
  (`agent.run(capabilities=...)`), so "active for the next prompt" is the
  natural unit; nothing rebuilds the agent.
- **Shipped plugins register first, in declared order.** The menu's alphabetical
  order is for scanning only. Registration order is the order instructions,
  renderers, and status segments are consulted in, so `coder`'s guidance leads
  the prompt. `customization_guide()` orders itself after the guidance plugins
  contribute and before harness `RepoContext`, so the CLAI hint never leads.
- **Built-ins are declarations, not code paths.** `DEFAULT_PLUGINS` in
  `_app.py` lists what CLAI ships enabled (`coder`, `ask_user`, `repo_context`,
  `compaction`, `persistence`, `logfire`). The loader treats them like drop-ins with the lowest
  precedence: a store declaration with the same id replaces one, `disable`
  persists an override, `remove` resets it. Do not special-case `Coder`
  anywhere else; the agent from `create_agent()` has no coding tools of its
  own. `coder` is declared with `repo_context: false` because `repo_context`
  binds harness `RepoContext` itself; keep it that way or `AGENTS.md` reaches
  the model twice.
- **Project declarations rank just above built-ins and start off.**
  `.clai/settings.json` (`project_settings.py`) may declare plugins; the loader
  takes them as `project=`, every one `enabled=False`, because a repository
  must not run code as the user on launch. `/plugins enable` is the approval
  and persists the approved declaration in the store. Precedence is store,
  drop-in folder, project, built-in. CLAI never writes the project file.
- **A load failure leaves the session as it was.** Import or `activate` errors
  are reported and the plugin stays unloaded; partial registrations from a
  failed `activate` are discarded with the host. Registered `session_end` handlers
  run with `reason='error'` under a shield with a five-second cooperative timeout
  per handler before a failed/cancelled load drops the host. Cleanup must tolerate
  incomplete `session_start`; a handler failure must not skip later cleanup.

Compaction registers harness `FallbackCompaction` directly with `max_fraction`
and `context_window` for both strategies. Harness owns the trigger; do not add
threshold math or an orchestrator in CLAI. Register the usage gauge after the
chain so yellow means the compacted request still exceeds the threshold.
`/compact` drives the same chain regardless of threshold. Only `ModelAPIError`,
`FallbackExceptionGroup`, and `UsageLimitExceeded` select truncation after a
summary failure; other exceptions propagate.

## Adding or changing a hook

1. Add the name to the `Literal`, the event dataclass, and the `@overload`.
2. Add the parity test entry: the `Literal` must equal `Hooks.on`'s attribute
   names plus the host names. Drift fails CI, not code review.
3. Fire it from exactly one place in the shell.
4. Document it in `PLUGINS.md` in the table it belongs to.

## The `/plugins`, `/set`, `/theme`, `/model`, and `/add_model` menus

Built on termflow's `MenuBuilder` (and `TextInputBuilder` for typed values),
exactly like Code Puppy's `/agent`, `/mcp`, `/set`, and `/model` menus:
alternate screen, a `.preview` panel on the right, `.on_key` for single-key
actions, `.footer_hint` for the key legend, `markdown_style()` for colours.

- Split it in two: a pure `build_plugins_menu(...)` that returns the menu (so
  tests drive it headless, no terminal), and a thin async runner that owns the
  screen and calls `menu.run` in a thread.
- Every key mutates immediately and `replace_items` redraws. No pending-changes
  state, no save/cancel pair.
- Nothing prints to the console while the menu is open; the alternate screen
  would hide it. Show empty states and errors inside the menu as disabled rows.
- Esc and Ctrl-C close cleanly. They are not errors.
- A widget opened mid-run (including the inline `ask_user` picker) goes inside
  `async with host.full_screen()`, which flushes streamed text and suspends the
  editor's input reader first, preserving its draft. Slash-command handlers
  already run with the editor suspended. Do not start a second input reader
  alongside the live editor.
- Adding a plugin is not in the menu. It needs free text, so it stays
  `/plugins add`.
- Anything that is "edit named, validated fields" uses `field_menu.py`: a
  `FieldSource` supplies rows, current values, validation, apply, and reset;
  `FieldMenu` builds the widgets; `run_flow` is the loop. `/set` and per-model
  settings are two sources, not two editors. Do not write a third editor.
- `/set` edits go through `CommandContext.set_setting` / `reset_setting`, the
  same path as the typed command, so validation lives in one place.
- Widget runners are a `Runners` value passed into the loops; tests pass
  scripted ones (`tests/menu_script.py`). Only the real `widget.run()`
  one-liners are `no cover`.
- Model sources live in `model_catalog.py`. To add one (models.dev, a provider
  API), write a function returning `CatalogModel`s and merge it in `catalog()`.
  The menu never talks to a source directly.
- Per-model settings are the editable subset of core's `ModelSettings`,
  declared once as `ModelSettingsForm` with descriptions and bounds. Extend the
  form, not the menu, to expose another setting.

## Rendering

`StreamRenderer` owns text and thinking. It knows nothing about any specific
tool. Tool-specific output (shell previews, diffs, grep) is registered through
`host.render` by the plugin that owns the event. Match on event classes, never
on `tool_name` strings. Always flush the stream before printing anything else;
the host does this for renderers, so do not call `console.print` from inside an
`on` handler when a renderer would do.

## Colours

`/theme` offers the unchanged `default` appearance and `termflow.themes.PALETTES`.
Do not define new palettes. Resolve brand roles with `theme.color(...)` for Rich;
`theme.sgr(...)` resolves raw ANSI itself. `theme.current()` returns a Termflow
palette or `None` for the original appearance. Termflow owns palette application
and reset; `theme.use(...)` leaves the terminal untouched in the default session.
Markdown keeps its original style by default and uses `to_render_style()` for a
selected palette. The preview renders a sample without OSC changes or persistence.
Heavy imports in `theme.py` stay lazy for the splash. Code uses the terminal
foreground and ANSI syntax colours through `theme.syntax_theme()`, shared by
streamed fences and theme previews. Default diff colours stay unchanged, while
bundled palettes use Termflow defaults.

## File map

| File | Holds |
|---|---|
| `_cli.py` | argument parsing, startup, `--agent` |
| `_app.py` | the prompt loop and built-in `/commands` |
| `_session.py` | conversation state, revision-checked saves, restore-only resume, per-run plugins |
| `sessions.py` | resume command and background namer ownership; built-in step capture |
| `forks.py` | `/fork` and `/forks`: history snapshot, background child sessions, deferred fork output |
| `session_browser.py` | project/session browser using Termflow layout and terminal primitives |
| `_rendering.py` | streaming Markdown and thinking |
| `plugins.py` | `PluginHost`, hook names, event dataclasses |
| `plugin_loader.py` | discovery, load, unload, reload; the `/plugins` subcommands |
| `plugin_menu.py` | the `/plugins` full-screen menu (`PluginMenu` plus its runner) |
| `ask_user_menu.py` | the built-in `ask_user` plugin: `QuestionMenu`, `TerminalAnswerer`, the transcript renderer |
| `screen.py` | `Screen`, what `host.full_screen()` binds to during a prompt |
| `field_menu.py` | the shared field editor (`FieldSource`, `FieldMenu`, `Runners`, `run_flow`) |
| `set_menu.py` | `/set`: `SettingsSource` over `CommandContext` |
| `model_menu.py` | `/add_model`: provider discovery, `ModelSettingsSource`, `run_model_flow` |
| `model_picker.py` | `/model`: selection and completion of saved models |
| `model_catalog.py` | model sources (genai-prices today) merged by `catalog()` |
| `model_settings.py` | `ModelSettingsForm`, the editable subset of `ModelSettings` |
| `logfire.py` | the default-enabled, locally configured Logfire plugin over core `Instrumentation` |
| `compaction.py` | the built-in `compaction` plugin: harness `FallbackCompaction([SummarizingCompaction, SlidingWindowCompaction])`, `/compact`, the context alert |
| `commands.py` | `Command`, the registry, completion |
| `usage_report.py` | `/usage`, `/cost`, and the footer cost, derived from `Session.messages` |
| `status.py` | the footer `Status` fields, `StatusSegment`, and the `StatusLine` row painter |
| `live_prompt.py` | pinned editor lifecycle, completion worker, submission queue and menu handoff |
| `prompt_surface.py` | scroll-region ownership, serialized transcript writes and changed-row painting |
| `prompt_transcript.py` | bounded styled transcript tail for viewport replay |
| `prompt_resize.py` | scoped resize notifications, without terminal IO in signal handlers |
| `prompt_buffer.py` | pure draft editing, history navigation, search and cell-width wrapping |
| `prompt_completion.py` | bounded daemon completion worker; no terminal ownership |
| `prompt_keys.py` | keyboard decoder attachment only; no prompt-toolkit Application or renderer |
| `config.py` | `Settings`, `PluginSettings` |
| `settings_store.py` | the SQLite store under `$XDG_CONFIG_HOME/pydantic-clai2/` |
| `project_settings.py` | `.clai/settings.json`: the walk-up to the git root, validation, `ProjectSettings` |
| `repo_context.py` | the built-in `repo_context` plugin over harness `RepoContext` |
| `speculation.py` | the `run.speculative_code_mode` switch, `Ctrl+X Ctrl+S` toggle, session counters and pinned row |
| `speculative_mode.py` | harness `CodeMode` wiring (native writes, read-only speculation allowlist, guidance), imported only while on |
| `eager_timing.py` | eager `run_code` latency measurement and the nested-call id pattern |
| `sandbox_calls.py` | events and ordering that render calls from inside `run_code` like direct calls; no harness imports |
| `theme.py` | Existing brand roles, opt-in Termflow palette scope, `color()`, `sgr()` |
| `theme_picker.py` | `/theme` picker over Termflow's bundled palettes |
| `spinners.py` | the working-animation catalogue: builtins, plugin `host.spinner`, the user's `spinners.json`, `Spinners` |
| `spinner_frames.py` | frame data for the Code Puppy cli-spinners pack |
| `spinner_picker.py` | `/spinner`: animated picker, by-name selection with speed, `init` |

Keep files concise - we don't need any 10,000 line files. Single responsibility.

## Testing

- When adding or modifying a CLAI2 CLI UX feature, test every affected UX feature
  with your changes in a fresh tmux window before opening a PR. Run CLAI2 with
  `uv run clai2`. Verify the actual terminal behavior against the intention of
  the user's request, not just the implementation or automated tests.
- `pytest-anyio`; real model calls are blocked globally.
- Drive the shell with `TestModel` and a `Console(file=StringIO())`.
- Test a hook by building a `PluginHost`, registering a handler, and firing the
  event from the shell path that owns it. Assert the handler's effect (the
  cancelled turn, the rewritten text), not a mock call count.
- Renderers get synthetic events. They must not need `Coder` installed.
- Cancellation: use a real `anyio` cancel scope, order with `Event`s, no sleeps.

### Settings and database compatibility

Before adding, removing, or renaming a setting, changing its type, meaning, or
default, or changing the SQL schema, consider upgrades, downgrades, and branch
switches. Different versions can share the same settings database.

- Add regression cases to `tests/test_settings_compatibility.py` and affected CLI
  tests using the previous stored format. Keep historical fixtures unchanged;
  do not regenerate them with current models or rewrite them to make a change pass.
- Verify older databases load with documented defaults for missing settings and
  preserve existing preferences, plugin declarations, and model settings. Test
  migrations and repeated initialization for data preservation and idempotence.
- Preserve unknown saved setting names and their values when reading, editing
  other settings, resetting, and reopening. Unknown does not mean obsolete.
  Keep validation strict for known values, new writes, and plugin declarations.
- For renamed fields or changed types, meanings, or defaults, define the migration
  or intentional behavior change and test it explicitly. Comparing only against
  the current `Settings()` defaults will not catch an unintended default change.
- Verify rejected values and unsupported schema versions leave stored data intact.
  If a change introduces a migration, test failure rollback as well as success.

## Local verification

Run from the repository root. CLAI shares the root `uv.lock`, `.venv`, and
Pyright configuration with Harness.

```bash
uv sync --locked --all-packages --group lint
uv run --no-sync ruff format --check .
uv run --no-sync ruff check .
PYRIGHT_PYTHON_IGNORE_WARNINGS=1 uv run --no-sync pyright pydantic-clai2/src pydantic-clai2/tests
uv run --no-sync pytest -p no:cacheprovider -c pydantic-clai2/pyproject.toml pydantic-clai2/tests
```

## Docs parity

`README.md` is the tour, `PLUGINS.md` is the plugin contract, this file is for
agents. A user-facing change updates the first two in the same PR. Plain
language, short sentences, no jargon a first-time plugin author would not know.

The resume browser is a focused exception to the single `MenuBuilder` convention:
it needs independently focused projects/sessions and two-line cards. Keep its
frame pure, inject keys/size/IO, use Termflow terminal ownership and layout, and
run through `menu_worker`. Metadata refreshes must preserve selection by ID.
History-changing plugins use `await conversation.commit_messages`, not the legacy
in-memory `replace_messages`, so exiting immediately after `/compact` is durable.

The interactive editor owns its layout explicitly. Do not reintroduce a
PromptSession renderer or mutate generated layout children. Transcript writes
go directly to the scroll region, never through an erase/redraw of the editor.
The hardware cursor stays hidden until release; the input cursor is a painted
reverse-video cell. Keep terminal mutations in `PromptSurface`, and detach the
key reader before a menu owns the screen. The remaining prompt-toolkit decoder
preserves paste and modified keys not yet exposed by Termflow's `read_key`.

Physical resize blanks the viewport and defers output until size notifications
have been quiet for 250 ms. Rebuild from `TranscriptBuffer`, not guessed old row
coordinates or cursor reports. Never send erase-scrollback (CSI 3 J). Keep editor
height changes separate from physical resize, preserve the draft, and close the
resize output spool on both normal handoff and failure. `SIGWINCH` only marks the
resize and schedules a paint; the signal handler must not perform terminal IO.

The `ask_user` picker is an inline exception to the full-screen menu convention.
It borrows the released `PromptSurface` while the editor is suspended, retaining
the shared transcript for resize replay. Keep its numbered choices and Enter
toggles; do not reintroduce alternate-screen switching or Space-to-toggle.

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.