session-orchestrator / site
Kanevry/session-orchestrator/site/llms-full.txt
Plan the work. Run it in checked waves. Pick up where you left off. Session Orchestrator adds a repeatable Plan, Go, Close workflow to AI coding sessions. It is a workflow layer that runs on top of the coding agent you already have (Claude Code, Codex CLI, Cursor IDE, or Pi), not a replacement for it. Maintained by one person, Bernhard Götzendorfer, Vienna; shipped as-is with best-effort maintenance. Not affiliated with or endorsed by Anthropic, OpenAI, or Cursor. MIT license,…
llms.txt52 starsChanged 11 days ago
- Deletes or force-pushes
- Installs packages
# session-orchestrator (full description for LLMs)
Plan the work. Run it in checked waves. Pick up where you left off. Session Orchestrator adds a repeatable Plan, Go, Close workflow to AI coding sessions. It is a workflow layer that runs on top of the coding agent you already have (Claude Code, Codex CLI, Cursor IDE, or Pi), not a replacement for it. Maintained by one person, Bernhard Götzendorfer, Vienna; shipped as-is with best-effort maintenance. Not affiliated with or endorsed by Anthropic, OpenAI, or Cursor. MIT license, local by default, no account required, telemetry strictly opt-in.
Version 5.2.0 · npm package: session-orchestrator · requires Node.js >= 24 · https://session-orchestrator.com
Cost: free, MIT licence. You still pay whatever your coding tool costs; that part is not affected.
Windows: native use is untested. The Node core uses portable paths, but shell hooks and the optional MCP server need WSL or Git Bash. macOS and Linux are covered by CI.
## The loop
Bootstrap once per project. In Codex, select `$session-orchestrator:bootstrap`, then `$session-orchestrator:session feature`, `$session-orchestrator:go`, and `$session-orchestrator:close` in the skill picker or invoke them in a message. The slash commands below describe the shared workflows on Claude Code, Cursor and Pi. Codex requires version 0.144.4 or later and a fresh task after plugin installation; review the hook bundle through `/hooks`.
- /session [housekeeping|feature|deep] - research and Q&A: inspects git state, open issues, CI status, prior-session records, then presents one structured summary with a recommendation. Scope is agreed with the operator before any code.
- /go - executes the agreed scope in a shape matched to the session type: one wave for housekeeping or five for deep (Discovery, Impl-Core, Impl-Polish, Quality, Finalization) with quality checks between waves and a full configured test, typecheck and lint gate before completion. Claude Code and Codex can dispatch parallel agents; Cursor and Pi execute sequentially. For work that outgrows five waves there is a named ultradeep profile - a profile over session-type deep, not a fourth enum value - running seven waves: Research plus Code-Discovery, a blocking coordinator Synthesis-Gate, Impl-Core, Impl-Polish, a read-only Review-Panel, Quality, Release.
- /close - verifies every planned item against evidence, runs the full gate one last time, files carryover issues for unfinished work, commits cleanly, mirrors to GitHub when configured, writes session records and learnings.
## Example and origin
The landing page shows an illustrative CSV-export session: read the code, issues and tests; plan separate interface and data work; combine the changes and verify download, formatting and error cases; close with evidence and record selectable columns as unfinished follow-up work. It is an example, not live execution data or a benchmark.
Bernhard describes starting with a Notion page containing roughly 20–30 prompts copied between projects. That routine gradually became Plan, Go, Close, which he uses on his Mac M4 and office M5. The story does not assign a date to the prompt library or claim that the public plugin has existed for years.
## Why waves and parallel sub-agents
Declared file scopes are checked for overlap before dispatch. Separate agent contexts and inter-wave reviews help contain mistakes, while the Quality wave simplifies code before adding tests. These measures reduce collisions and repeated errors; they do not provide filesystem isolation or guarantee correctness.
## Guardrails and workflow rules
- Destructive-command guard: the policy defines 11 blocking rules and 4 warning rules over git reset/checkout-discard/clean/stash, rm -rf, force-push and SQL DROP. An active hooks/pre-bash-destructive-guard.mjs applies them on Claude Code and through compatible Cursor or Pi bridges. Codex treats these rules as instructions only; the plugin does not enforce this guard there. Policy: .orchestrator/policy/blocked-commands.json.
- Scope enforcement: with an active compatible handler on Claude Code, Cursor or Pi, hooks/enforce-scope.mjs blocks supported out-of-scope edits in strict mode, reports them without denial in warn mode, and skips scope checking in off mode. Codex has no compatible handler; its scope rules are instructions only.
- Verification rule: the agent is instructed to support completion claims with fresh evidence (.claude/rules/verification-before-completion.md).
- Full Gate: the workflow requires typecheck, tests, lint and a debug-artifact scan at the Quality wave and session end.
- Session lock with heartbeat liveness per repo working copy (scripts/lib/session-lock.mjs)
- Subagents are instructed to leave git add/commit/stash/push to the coordinator (PSA-007). This is a workflow rule, not a sandbox that removes access to the git index.
- 10 autopilot kill-switches in a frozen enum; echo-stub detection catches a test command that is secretly a no-op
- Platform note: Claude Code runs guard hooks directly; Cursor bridges preToolUse and beforeShellExecution with additional post-edit checks; Pi uses a tool-call bridge. Cursor and Pi have documented event-coverage limits. Codex leaves PreToolUse handlers empty because its tool names and edit payloads do not match these guards. Both destructive-command and file-scope rules are behavioural only there. Quality checks run on all four. See docs/codex-setup.md, docs/cursor-setup.md and docs/pi-setup.md.
## Memory and self-improvement
Nothing is learned silently. Sessions append plain-text JSONL records (sessions, learnings with confidence scores and expiry). /evolve extracts patterns after 5+ sessions; a reconcile engine turns eligible learnings into PROPOSED rules that the operator approves one by one - it structurally cannot emit always-on rules. Failed approaches are recorded as "What Not To Retry" and force-read at the next session start.
## Cross-repo (single operator, many repos)
/portfolio aggregates issues/MRs/CI health across all registered repos; /dispatcher ranks free repos by backlog, staleness and readiness, then claims a lease atomically before launching; a vault live-status board shows in-progress/force-closed sessions across every repo on the host. GitLab and GitHub are both first-class (auto-detected; glab and gh drive the full issue/MR lifecycle).
## Eval standard
aiat-llm-eval v1: an open standard for honest LLM session evaluation. Pre-registered rubric, deterministic checks before any LLM judge, three-state verdicts with explicit abstention (cannot-determine instead of fake zeros), no global score (the validator rejects overall/total/mean fields), no superlatives in conforming reports, reproducibility as an executable proof (--verify). A session scoring itself is labelled a self-evaluation. Spec: docs/eval/aiat-llm-eval-v1.md in the repository.
## Numbers
<!-- census:start -->
Version 5.2.0 · counted 2026-09-19 at 8f6ac022 · skills 50 · commands 26 · agents 14 · hooks 28 · test files 692 · sessions 433 · learnings 253 · npm downloads (30d) 1728 · GitHub stars 52
<!-- census:end -->
Generated by `node scripts/site-numbers.mjs --write`; the machine-readable receipt is site/_census.json. Facts outside the census, measured 2026-09-06 at bc49301b: 18 ADRs, 45 modules under scripts/lib/validate/ (`find scripts/lib/validate -name '*.mjs' ! -name '*.test.mjs' | wc -l`), 15 destructive-command policy rules (11 marked blocking, 4 marked warning; applied only by a compatible active hook), 26 always-on rule files, 26 of the hook files plugin-wired across 10 event types, and 13,529 top-level it()/test() call sites (`rg -c '^\s*(it|test)\(' tests --glob '*.test.mjs'` summed per file; the runtime total is higher because of parameterised blocks that expand at run time). The v4.0.0 release REMOVES public surfaces, so counts from the preceding 3.24 line (49 skills / 28 commands / 16 agents / 61 rules) are stale by construction.
Repository session records describe development activity, not active users. npm downloads can include automated requests; neither count is a measure of adoption or productivity.
## Install
- Claude Code: /plugin marketplace add Kanevry/session-orchestrator then /plugin install session-orchestrator@kanevry - afterwards run npm install once inside the plugin directory (the runtime imports npm packages) and restart Claude Code.
- Codex CLI: git clone https://github.com/Kanevry/session-orchestrator.git ~/Projects/session-orchestrator && cd ~/Projects/session-orchestrator && npm install && node scripts/codex-install.mjs
- Cursor IDE: same clone, then node scripts/cursor-install.mjs /path/to/your/project
- Pi: pi install npm:session-orchestrator
Portable cross-harness surface: a root AGENTS.md generated byte-identical from CLAUDE.md, a .cursor-plugin/plugin.json following the agent-plugins.org 1.0.0 schema, and a .agents/skills/<name>/SKILL.md mirror of every skill carrying only spec-legal frontmatter plus a pointer body. All three are generated by scripts/generate-agents-skills.mjs and drift-checked in scripts/validate-plugin.mjs; none is hand-edited.
Upgrade: on Claude Code, /plugin update session-orchestrator@kanevry then restart. A Codex Git marketplace installation uses codex plugin marketplace upgrade kanevry followed by codex plugin add session-orchestrator@kanevry; a local clone uses git pull, npm install and node scripts/codex-install.mjs. Start a fresh task and reopen the skill picker after either Codex path. Cursor and the Pi clone fallback use git pull and their original installer; npm-installed Pi packages are managed through Pi. Session-start compares the version of the code that is RUNNING against the published npm version and fails silent rather than claiming "up to date" from a failed lookup. Uninstall removes the plugin from the harness; .orchestrator/ (bootstrap.lock, metrics JSONL, policy, steering), STATE.md and the Session Config block stay in the user's repo as plain text.
Minimum config: a "## Session Config" section in CLAUDE.md on Claude Code, or AGENTS.md on Codex CLI and Pi, declaring test-command, typecheck-command, lint-command, agents-per-wave, waves, persistence, enforcement. Cursor supports AGENTS.md; native CLAUDE.md pickup is unverified, though the orchestrator parser reads it. Run bootstrap to create the required marker as well. Additional features and defaults are documented in docs/session-config-reference.md.
## Common questions
- Is this itself an AI that writes code? No. It adds a workflow for planning, working and checking results. Your coding tool writes the code. Technical enforcement depends on the tool and guard configuration.
- Which tools does it work with? Claude Code runs the guard hooks directly. Cursor and Pi use bridges with documented limits. Codex CLI offers selectable workflow skills, but its destructive-command and file-scope rules are instructions only; the plugin does not enforce either guard there. Parallel agents are available on Claude Code and Codex; Cursor and Pi run tasks sequentially.
- What is different from just using Claude Code? It adds a repeatable session workflow: read the project, agree the scope, divide the work into steps with declared files, and check the results between steps. The session record carries progress and unfinished work into the next session.
- Does it send data anywhere? The plugin runs locally without its own account. Optional usage telemetry is off until you consent. By default, session-start checks the npm registry for updates. Successful results are cached for 24 hours per repository; a failed request or cache write may cause another request at the next session. SO_DISABLE_UPDATE_CHECK=1 or DO_NOT_TRACK=1 disables the check. Your coding tool and configured services still use their own network connections. This website uses cookieless Vercel Web Analytics, described in its privacy policy.
- Does it work on Windows? macOS and Linux are covered by CI. Native Windows use has not been tested. The Node core uses portable paths, but the shell-based hooks and optional MCP server need WSL or Git Bash. Treat Windows support as best-effort.
- What does it cost? Nothing. It is free and open source under the MIT licence. You still pay whatever your coding tool costs; that part is not affected.
- Can it break my project? It can still make mistakes. On Claude Code, an active destructive-command hook applies 11 blocking rules and 4 warning rules. Cursor and Pi use bridges with documented limits. Codex does not enforce this guard or the file-scope guard. For supported edits with an active scope hook, strict blocks violations, warn reports them, and off disables the check. Review changes and keep backups; the checks do not guarantee correct code.
- Who maintains it? One person, Bernhard Götzendorfer. Questions filed on GitHub are answered by him. It is shipped as-is, with best-effort maintenance and no service level agreement.
## Links
- Source, docs, issues: https://github.com/Kanevry/session-orchestrator
- npm: https://www.npmjs.com/package/session-orchestrator
- Methodology courses (optional, not required to use the plugin): https://agenticbuilders.at
- Optional support: https://paypal.me/Kanevry
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

