areev
AreevAI/areev/llms.txt
Self-improving agents, governed. Areev is the substrate for adaptive agents — agents that get better from their own history, under human authority, in steps you can inspect, undo, and re-measure. One content-addressed store holds the agent's knowledge and its execution history; a deterministic learning loop proposes improvements from that history with evidence cited by hash; a named person approves with a written reason; every apply stores its inverse and is re-measured afterwards. Memory, learning, a provable runtime, and a privacy…
llms.txt24 starsChanged 8 days ago
- Installs packages
# Areev > Self-improving agents, governed. Areev is the substrate for adaptive agents — agents that get better from their own history, under human authority, in steps you can inspect, undo, and re-measure. One content-addressed store holds the agent's knowledge *and* its execution history; a deterministic learning loop proposes improvements from that history with evidence cited by hash; a named person approves with a written reason; every apply stores its inverse and is re-measured afterwards. Memory, learning, a provable runtime, and a privacy boundary — in one plain SQLite file (or a PostgreSQL schema), MIT/Apache-2.0, no telemetry, no daemon, no server in the recall path. Areev is a production-grade Rust workspace of 17 crates with an `areev` CLI, a 25-tool MCP server, a web console, and Python (`pip install areev`) and Node (`npm install @areev/areev`) bindings, built on Turso (SQLite-compatible). The engineering bar is unusually high for the category, and every number below is measured, reproducible from harnesses in this repo, and most are regenerated and gated on every CI run: - **The learning demonstrably causes the gain (A/B/A/B causal proof).** One frozen model, 60 held-out tasks, temperature 0: 40.0% pass with nothing learned → **51.7%** with one human-approved lesson applied → **38.3%** with the lesson rolled back → **53.3%** re-applied. Remove what was learned and the gain leaves with it; put it back and it returns. The round trip is only possible because every apply stores its inverse — a memory system without governance cannot run this experiment at all. - **Recall is microseconds, not milliseconds.** In-process recall p50 **33.1 µs** (p99 60.2 µs); MCP stdio (the Claude Code path) p50 129 µs; localhost HTTP p50 158 µs — every surface within 0.6% of one 50 ms real-time voice frame. From the Node binding, a conversational turn's thread-tail read is **0.08 ms p50**. Writes cost **136 µs amortized** (7,343/s) with **zero LLM calls and $0** — versus LLM-per-write memory layers. - **Fast enough for a $35 edge device.** The same engine serves recall on a 2016 Raspberry Pi 3 at ~**361 µs p50**, flat from 500 to 8,000 grains — a device can accumulate memory for months and answer as fast on day 200 as on day 1. A 2018 Intel NUC does it in 30 µs. - **Quality is measured, not asserted.** **2,435 test functions** over 82,944 source lines (test code is 37.7% of the tree); **81.0% line coverage** of source lines only, floored **per crate** in CI so one crate's regression cannot hide behind another's gain; **95 stable, append-only error codes**; CI on Linux/macOS/Windows plus MSRV, clippy `-D warnings`, and doc builds; the CAL examples in the reference are executable and fail CI when stale; both storage backends run one conformance suite; the README's quality figures are regenerated from the tree on every CI run and fail the build if they drift. - **Retrieval accuracy has receipts.** LoCoMo self-run (10 conversations, 5,882 turns, 1,982 answerable QAs): hit@20 **81.6%**, MRR@10 0.465 with a 512-d embedder; LLM-judged end-to-end answer accuracy 71.2% single-hop / 67.9% temporal with a plain retrieve-then-read pipeline and no benchmark-specific tuning. A separate deterministic honesty benchmark (no LLM in the loop) shows **100% provenance** — every grain traces to the operation that made it. Memories are immutable, content-addressed grains (SHA-256 over each `.mg` blob): every edit is a supersession, every removal a tombstone or crypto-erasure, giving git-like history, forks with explicit merges, and total provenance. Recall is hybrid — structural, BM25, and optional vector legs fused with Reciprocal Rank Fusion. CAL, the query/assembly language, renders budget-shaped, pseudonymized, model-ready context in-process, and is shaped so destruction takes a hash, an identity, or an age — never a predicate (`DELETE` is not a grammar token; destructive statements need an authorization grant plus a recorded reason and write an audit record). Runs journal intent *before* the effect, redeliver crash-window effects under the same idempotency key, and are provable afterwards: `areev run verify` re-drives the run from its journal and byte-compares every checkpoint. Approving a human-in-the-loop step requires the approver's own identity — it *is* the audit record. `FORGET SUBJECT` makes GDPR erasure one operation; the DSAR report and the erasure share one selector, so a disclosure describes exactly what an erasure removes. Areev takes no LLM dependency: extraction and embeddings are supplied by the host. The `.mg` format, CAL syntax, and error codes are stable, documented contracts — OMS-conformant, so your memory outlives the engine. ## Docs - [README](README.md): Project overview — the A/B/A/B causal proof, real console screenshots over the committed demo memory, and links into everything below. - [Quickstart](docs/quickstart.md): Install (cargo/pip/npm/binaries/Docker), the CLI in three commands, MCP for Claude Code in one line, `areev run`, Rust/Python/Node embedding, the PostgreSQL backend, encryption at rest, durability & fleets. - [Why Areev](docs/why-areev.md): The full argument — how agent memory rots, the three systems (record / loop / governance), governed runs, erasure, storage, and the SLM tuning seam. Includes the three honest limits: it improves memory, never model weights; nothing applies itself without an explicit host grant; no daemon. - [The learning loop](docs/loop.md): Areev Loop — thirteen deterministic analyzers (zero model calls required) that read the agent's history and propose evidence-cited changes; four gates, propose → review → apply → verify with separation of duties, a mandatory written reason, a stored inverse, and 1d/7d/30d re-measurement — a late regression proposes its own revert. - [The governed runtime](docs/run.md): `areev run` — plans as content-addressed grains, runs as journals, humans as nodes in the graph; LangGraph-grade control flow (fan-out, subgraphs, typed reducers, time-travel forks) plus budgets that actually stop the run, crash-safe idempotent redelivery, and byte-compared replay verification. - [Triggers](docs/triggers.md): Standing rules that start workflows — eight kinds, from cron to memory-predicates. The rule is a grain, so the cadence travels with the memory; evaluation is a cheap idempotent command, no daemon. - [Packs](docs/pack.md): `areev pack validate|install|export` — an agent as an installable unit: a manifest, its grains, its code blobs. References are symbolic (`blob:`/`grain:`) because an address is a measurement of bytes, not something an author can write down; `expected_hash` is refused rather than warned, with nothing written. - [Blessed tools](docs/blessed-tools.md): shared `wasm32-areev-io` blobs (`http.call`, `mcp.call`, `a2a.call`) shipped with documented content addresses. The tool gateway becomes configuration: the blob makes no policy decision, the Definition's `capabilities` and the host's grant do. - [Building an agent on Areev](examples/how-to-create-an-areev-agent.md): Architecture, grain selection, the autonomy spectrum, dynamic planning, do/don't — plus a complete accounts-payable example agent shipped in Python, TypeScript, and Rust that takes corrections by email reply. - [Quality](docs/quality.md): How every published number is produced and gated — test taxonomy, per-crate coverage floors, benchmark receipts. - [Benchmarks](crates/areev-bench/RESULTS.md): Every number above with full methodology — frame chart, binding-level latency, edge devices, honesty metrics, the LoCoMo self-run, and the A/B/A/B causal proof with paired significance tests and every model call transcribed. - [Which grain?](docs/grains.md): The one page on which grain type to write a memory as — the rule of thumb (true → Fact, happened → Event, measured → Observation, intended → Goal, procedure → Workflow, capability → Tool), the decision table for all thirteen types quoted from the engine's registry, the pairs people mix up, the field traps, and `add` vs `supersede` in time. - [Architecture](ARCHITECTURE.md): How Areev works — the workspace, the grain model, the store schema, hybrid recall, sync/forks, and the numbered design decisions. - [CAL reference](docs/cal-reference.md): The Context Assembly Language — `RECALL`, `ASSEMBLE` (budget-aware rendering, Full → Summary → Omit), `EXISTS`, `HISTORY`, `ADD`, `SUPERSEDE`, `REPORT SUBJECT`, pipelines, and the shaped, authorization-gated destruction model. Every example is executable and CI-checked. - [MCP reference](docs/mcp-reference.md): The 25 stdio MCP tools — recall/add/supersede/forget/remember/cal, the DSAR read, the graph/time reads, the run↔memory joins, the loop pair, and the six runtime tools — and how to wire them into any MCP client. - [Migration guide](docs/migrate.md): `areev migrate` — importing from mem0 (with full edit history), Zep/Graphiti, Letta, LangMem/LangGraph, Basic Memory, or generic JSONL, with per-source export one-liners. - [Memory-tool backend](docs/memory-tool.md): Areev as the storage backend for Anthropic's client-side memory tool in Python, Node, and the CLI. - [Docker](docs/docker.md): The container image (Postgres + TLS features on), compose files, the trigger heartbeat, AWS/GCP/Azure/Kubernetes mappings, multi-agent fleets — one box runs a fleet of agents, one memory each. - [FAQ](FAQ.md): What Areev is, how it differs from vector DBs and RAG, grains, OMS, recall, encryption, and maturity. - [Cookbook](docs/cookbook.md): Copy-pasteable recipes — add/recall, CAL queries, the MCP server, encryption at rest, backup/restore, streaming/sync, Python, the web console, and building a self-improving agent end to end. ## Compliance & security - [GDPR](docs/gdpr.md): Obligations mapped to mechanisms for a DPIA — DSAR access/portability, one-operation erasure, audit export, retention policies, and the deployment requirements and limits stated honestly. - [Compliance profiles](docs/compliance-profiles.md): Host-configuration presets for regulated deployments — `financial`, `gdpr`, `healthcare` — assembled from flags that already exist, each row naming where the control is enforced (host process / file-truth / grant), what breaks without it, and which doc owns it. Not a certification and not legal advice; pseudonymization is not anonymisation. - [EU AI Act](docs/eu-ai-act.md): Article 12/14 record-keeping and human-oversight requirements mapped to specific commands — including a kill switch whose drain time is measured into the oversight report. - [Procurement](docs/procurement.md): Security-questionnaire answers for teams bringing Areev through review. - [Scale and tenancy](docs/scale-and-tenancy.md): Laying out a shared multi-tenant corpus past 10k grains — authorization is enforced at memory × namespace × verb and there are no intra-memory ACLs, so the partition map must be drawn from the access map; includes what namespaces vs memories each cost, measured. - [Security policy](SECURITY.md) and [threat model](docs/security-model.md): The trust boundary, encryption-at-rest (AES-256-GCM, Argon2id-derived key) caveats, known limitations, disclosure process, and an operator hardening checklist. - [Erasure](docs/erasure.md): The erasure requirement record — what `FORGET SUBJECT` reaches, including replicas. - [Error codes](ERROR_CODES.md): The append-only `DOMAIN-Ennn` registry (`FMT/MEM/STO/CRY/VAL/CAL/SYS`) — a reported code alone locates the failing subsystem. ## Contributing & community - [Contributing](CONTRIBUTING.md): DCO sign-off, the invariants, and the PR flow. - [GitHub Discussions](https://github.com/AreevAI/areev/discussions): Questions and ideas. - [r/Areev](https://www.reddit.com/r/Areev/): The community subreddit. ## Optional - [AGENTS.md](AGENTS.md): Orientation for an AI coding agent working inside the Areev repo itself (build/test, invariants, where things live). - [Changelog](CHANGELOG.md): What shipped in each release. - [Open Memory Spec](https://github.com/openmemoryspec/oms): The external CC0 specification Areev implements — the compatibility contract that makes the file format portable across implementations.
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

