agent2
Artesiana/agent2/llms-full.txt
Agent2 turns domain experts into production AI agents. The framework is for building typed, hostable backend agents that work like professionals: they read documents, consult books, use tools, remember context, ask clarifying questions, propose side effects for approval, and return validated structured output. Product code stays in agent modules. Framework code stays in shared/. Make AI coding agents such as Claude Code, Codex, Cursor, and Gemini CLI able to open this repo and immediately understand how to build Agent2 agents.
llms.txt36 starsChanged 5 months ago
# Agent2
Agent2 turns domain experts into production AI agents. The framework is for
building typed, hostable backend agents that work like professionals: they read
documents, consult books, use tools, remember context, ask clarifying questions,
propose side effects for approval, and return validated structured output.
Product code stays in agent modules. Framework code stays in `shared/`.
## Design Goal
Make AI coding agents such as Claude Code, Codex, Cursor, and Gemini CLI able to
open this repo and immediately understand how to build Agent2 agents.
The target agent is not a chatbot and not a deterministic script. It is a
professional workspace:
- identity and professional mindset
- reference books and knowledge packages
- domain-specific tools
- memory and case history
- observable review process
- three mutually exclusive outcomes
- typed output contract
- sandboxed side effects
- host-controlled persistence and approval
## Agent2 Laws
1. The prompt teaches how to think and work; books contain what to know.
2. Add domain rules as knowledge documents before adding code tables.
3. Tools should be named in domain language, not generic abstractions.
4. Every serious domain agent has complete/clarify/reject outcomes.
5. A schema validator is part of the agent, not optional polish.
6. Side effects are pending actions until a host or human approves them.
7. Resume uses `message_history`; the agent continues the case, not a fresh run.
8. Evals must test real domain behavior.
## Public API
Every agent service exposes:
- `GET /health`
- `POST /tasks?mode=sync`
- `POST /tasks?mode=async`
- `GET /tasks/{task_id}`
- `POST /tasks/{task_id}/actions/execute`
`POST /tasks` accepts `{"input": {...}}`.
Input may include product fields such as `text`, `purchase_request`, IDs, or
`message_history`. Results may include typed output, `_message_history`, and
`pending_actions`. Errors use RFC 7807 `application/problem+json`.
## Core Modules
### `shared/runtime.py`
Responsibilities:
- load per-agent config
- resolve model IDs
- build OpenRouter-backed PydanticAI agents
- apply provider policy
- fetch optional Langfuse prompts
- expose `create_agent()`
Rules:
- use `instructions=` for new agents
- `system_prompt=` is compatibility only
- no LLM key triggers PydanticAI test model and API mock behavior
- pass MCP toolsets through `toolsets=[]` or per-run `_toolsets`
### `shared/api.py`
Responsibilities:
- expose the FastAPI app via `create_app()`
- validate request shape
- enforce auth and rate limiting
- load source-layout or Docker-layout agents
- run sync and async tasks
- deserialize `message_history`
- invoke `before_run` and `after_run`
- pass dynamic `_instructions` and per-run `_toolsets` into `Agent.run()`
- strip runtime control fields from the user prompt
- execute pending actions
- normalize failures as problem responses
Agent hooks:
- `before_run(input_data) -> dict`
- `after_run(input_data, output) -> None`
- `mock_result(input_data) -> dict`
- `execute_action(action) -> dict`
### `shared/message_history.py`
Serializes and deserializes PydanticAI messages so hosts can persist
conversation state anywhere.
### `shared/approval_workflow.py`
Executes approved pending actions through an injected action executor and updates
stored task state.
### `shared/tool_policies.py`
Provides composition helpers for request-scoped tool interception, including
collection scoping.
## Canonical Agent Anatomy
A full Agent2 agent directory contains:
- `schemas.py`: typed output models, status literals, validators
- `agent.py`: prompt, `create_agent`, tool registration, hooks, mock result
- `tools.py`: domain tools and sandbox tool implementations
- `config.yaml`: model, timeout, collections, capabilities
- `main.py`: `create_app("<agent-name>")`
- `Dockerfile`: service image
Full domain agents should also add:
- knowledge books in `knowledge/books/<collection>/`
- collection entries in `knowledge/collections.yaml`
- tests under `tests/test_agents/`
- Promptfoo evals under `tests/promptfoo/<agent>/`
## Canonical Examples
### Full Pattern
- `agents/procurement-compliance-officer`: the in-repo flagship. It demonstrates
Brain Clone prompt architecture, Knowledge MCP, request-scoped collection
filtering, per-run MCP toolsets from `before_run`, memory, sandbox approvals,
three outcomes, validators, mock mode, `after_run`, tests, and evals.
### Reference Pattern
- `docs/reference-agents/sachbearbeiter-pattern.md`: explains the
production-proven Sachbearbeiter architecture: one professional agent handles
one complete case from document intake to review, communication, and typed
output.
### Primitive Demos
- `approval-demo`: pending actions
- `resume-demo`: pause/resume
- `provider-policy-demo`: provider routing
- `scoped-tools-demo`: tool scoping
- `rag-test`: basic Knowledge MCP search
- `example-agent`, `support-ticket`, `code-review`, `invoice`: small reference
agents
## Knowledge
Knowledge is foundational. Domain experts rely on books, policies, regulations,
reference material, and institutional notes. Agent2 gives agents the same
pattern through R2R collections and Knowledge MCP.
Collection metadata lives in `knowledge/collections.yaml`. Source documents live
under `knowledge/books/<collection>/`.
The full Docker profile starts R2R and Knowledge MCP:
```bash
docker compose --profile full up -d
```
Ingest collections with:
```bash
python -m shared.ingest --all
```
## Build and Test
```bash
uv sync --extra dev
uv run pytest tests/ -v
docker compose up -d
docker compose --profile full up -d
```
Promptfoo eval example:
```bash
npx promptfoo eval -c tests/promptfoo/procurement-compliance-officer/eval.yaml
```
## AI Coding Agent Guidance
When creating a new domain agent:
1. Use `/brain-clone` first.
2. Study `agents/procurement-compliance-officer`.
3. Put domain facts in books.
4. Put the expert's Sachbearbeiter Chain-of-Thought in the prompt.
5. Use three outcomes and a validator.
6. Sandbox side effects with pending actions.
7. Add tests and evals that prove behavior.
When changing framework behavior:
- update docs and skills if the contract changes
- add unit tests under `tests/test_shared/`
- preserve source-layout and Docker-layout imports
- do not break mock mode
## Documentation Index
- [AGENTS.md](./AGENTS.md)
- [README](./README.md)
- [Brain Clone Pattern](./docs/brain-clone-pattern.md)
- [Sachbearbeiter Pattern](./docs/reference-agents/sachbearbeiter-pattern.md)
- [Architecture](./docs/architecture.md)
- [Creating Agents](./docs/creating-agents.md)
- [Capabilities](./docs/capabilities.md)
- [Knowledge Management](./docs/knowledge-management.md)
- [Approvals](./docs/approvals.md)
- [Resume and Conversations](./docs/resume-conversations.md)
- [Provider Policy](./docs/provider-policy.md)
- [Observability](./docs/observability.md)
- [Deployment and Scaling](./docs/deployment.md)
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

