360-skills
shahabahreini/360-skills/llms-full.txt
This bundle contains the user guide, contributor instructions, and public skill instructions with supporting Markdown references. References are included for single-fetch access; normal skill loading should retrieve them only when relevant. [](LICENSE) [](https://agentskills.io) [](https://skills.sh) [](llms.txt) [](https://github.com/shahabahreini/360-skills/actions) 360-skills is an open collection of Agent Skills that give AI coding agents senior-level expertise for specific, high-stakes tasks. Most agents produce plausible work. Each skill in this repository packages the process, judgment, and quality gates of a senior specialist into an installable skill…
# 360-skills: Complete Agent Skills Catalog & Documentation
This bundle contains the user guide, contributor instructions, and public skill instructions with supporting Markdown references. References are included for single-fetch access; normal skill loading should retrieve them only when relevant.
## Table of Contents
- [Overview & Installation](#overview--installation)
- [Contributor Guide](#contributor-guide)
- [Contributing](#contributing)
- [Skills Catalog](#skills-catalog)
- [360-backend-audit](#skill-360-backend-audit)
- [360-blueprint](#skill-360-blueprint)
- [360-execute](#skill-360-execute)
- [360-expert-review](#skill-360-expert-review)
- [360-faculty](#skill-360-faculty)
- [360-optimize](#skill-360-optimize)
- [360-token-efficiency](#skill-360-token-efficiency)
---
## Overview & Installation
# 360-skills
[](LICENSE)
[](https://agentskills.io)
[](https://skills.sh)
[](llms.txt)
[](https://github.com/shahabahreini/360-skills/actions)
<p align="center">
<img src="assets/cover-zen-dark.png" alt="360-skills Workflow" width="100%">
</p>
**360-skills is an open collection of Agent Skills that give AI coding agents senior-level expertise for specific, high-stakes tasks.**
Most agents produce plausible work. Each skill in this repository packages the process, judgment, and quality gates of a senior specialist into an installable skill folder, to help agents produce work that can be reviewed against explicit checks. Skills follow the open [Agent Skills](https://agentskills.io) standard and install into Claude Code, Cursor, Codex, Copilot, Windsurf, Gemini CLI, and 70+ other agents through [skills.sh](https://skills.sh).
## Contents
| Topic | What you will find |
|---|---|
| [Install](#install) | Individual skills and optional dependencies |
| [Skills](#skills) | Descriptions and versions |
| [Routing](#which-skill-do-i-need) | Choose the right skill |
| [Workflow](#how-the-skills-work-together) | Shared plans and handovers |
| [Token efficiency](#optional-token-efficiency) | Consent, capabilities, accuracy and measurement |
| [Design principles](#design-principles) | Quality expectations |
| [Skill loading](#how-agent-skills-work) | Progressive disclosure |
| [Repository structure](#repository-structure) | Files and references |
| [Contributing](#contributing) | Validation commands |
| [FAQ](#faq) | Common questions |
| [License](#license) | MIT terms |
## Install
```bash
npx skills add shahabahreini/360-skills
```
The installer lists every skill in this repository and lets you choose which agents to install it for. To install a single skill non-interactively:
```bash
npx skills add shahabahreini/360-skills --skill 360-expert-review --agent claude-code
```
Each skill works individually. Review includes its own minimum plan contract; blueprint is not a required dependency. Install `360-token-efficiency` separately if you want the optional companion. A missing companion never blocks the primary task and is never installed automatically. Keep each skill's supporting `references/` directory with it.
## Skills
| Skill | Description | Version |
|---|---|---|
| [`360-backend-audit`](skills/360-backend-audit) | Deep-audit backend code, write the full report to a file, and brief the user in chat with bugs, updates, and dead weight. Use before or after significant backend work, or when inheriting, refactoring, or handing off services. | 2.2.0 |
| [`360-blueprint`](skills/360-blueprint) | Create an executable plan from a new objective with explicit tasks, constraints, and verification. Use when a goal exists but the path is unclear, or the request is "plan this". | 2.4.0 |
| [`360-execute`](skills/360-execute) | Execute a finalized plan task by task with a persisted coverage ledger, verify every item with evidence, and brief the user in chat. Use when a plan exists and work must begin, or when resuming a partial execution. | 2.2.0 |
| [`360-expert-review`](skills/360-expert-review) | Stress-test a draft plan, write the finalized executable plan back to the same file, and brief the user in chat. Use before executing any plan where a missed case could cause real damage. | 3.4.0 |
| [`360-faculty`](skills/360-faculty) | Seat a living, tailored expert team on a plan or task. Use when work must fit this developer's goals, taste, mindset, and strategy, when a plan needs the right expertise chosen for its complexity, depth, and nature, or when a named faculty team must be created, called, or updated. Recommends a short list, asks only the questions that still change the work, polishes immediately, and keeps upgradable memory. | 1.3.0 |
| [`360-optimize`](skills/360-optimize) | Audit working code for zero-cost speed, weight, and reliability gains — measure first, rank drop-in upgrades and restructures, and report conservative expected effects on this system. Use when existing code must run faster, lighter, or more robustly without changing what it does. | 2.2.0 |
| [`360-token-efficiency`](skills/360-token-efficiency) | Reduce avoidable context and tool-output overhead with capability-aware retrieval, reuse, and verified handovers. Use as an optional, session-approved companion when token cost or context growth matters, with strict preservation of task requirements by default. | 2.1.0 |
## Which Skill Do I Need?
| Your situation | Load | Not for |
|---|---|---|
| Backend correctness, dead weight, structure, or observability needs auditing | `360-backend-audit` | Not for planning work that does not exist yet (`360-blueprint`) or executing a plan (`360-execute`) |
| A goal exists, but no plan yet | `360-blueprint` | Not for hardening a plan that already exists (`360-expert-review`) or building one that is already final (`360-execute`) |
| An authorized plan needs implementation or resumption | `360-execute` | Not for creating plans (`360-blueprint`) or reviewing drafts (`360-expert-review`) |
| A draft plan needs adversarial review | `360-expert-review` | Not for drafting plans from scratch (`360-blueprint`) or executing a finalized plan (`360-execute`) |
| Choose expertise, tailor work to this developer, or manage a named team | `360-faculty` | Not for writing the plan itself (`360-blueprint`), attacking a finished draft (`360-expert-review`), or building one (`360-execute`) |
| Working code needs performance or weight recommendations | `360-optimize` | Not for greenfield design (`360-blueprint`), not for executing changes (`360-execute`), not for correctness or dead-weight hunts (`360-backend-audit`) |
| Context growth or avoidable token overhead matters; optional companion | `360-token-efficiency` | Not for creating plans (`360-blueprint`), executing the task itself (`360-execute`), or optimizing application performance (`360-optimize`) |
## How the Skills Work Together
On completion, each skill offers an interactive next-action recommendation based on the actual result and project stage. Choose the suggested step, stop, or write a custom direction outside the flow. Existing next-step instructions are honored without another prompt; the token-efficiency companion shares its parent’s prompt. Suggestions never authorize work by themselves.
Most skills hand off to one another sequentially on the same piece of work. Two are more flexible: `360-faculty` attaches wherever expertise is needed — before planning, after a draft, or ahead of review — and `360-token-efficiency` is an optional companion active only with session approval.
```mermaid
flowchart TD
Objective([Objective to plan]) --> Blueprint["360-blueprint: draft the plan"]
Blueprint --> Faculty["360-faculty: tailor the plan to this developer"]
Faculty --> Review["360-expert-review: stress-test and finalize the plan"]
Review -->|concrete blocker or revision| Blueprint
Review --> Execute["360-execute: run the finalized plan task by task"]
Execute --> Audit["360-backend-audit: audit the resulting backend code"]
Audit --> Optimize["360-optimize: audit for zero-cost speed and weight"]
Optimize -->|findings seed the next objective| Objective
Faculty -.->|house style before planning| Blueprint
Faculty -.->|which lenses this review needs| Review
Efficiency["360-token-efficiency: optional with session consent"] -.-> Blueprint
Efficiency -.-> Faculty
Efficiency -.-> Review
Efficiency -.-> Execute
Efficiency -.-> Audit
Efficiency -.-> Optimize
```
Plan producers and consumers carry the same local contract. Every task has ID, What, How, Where, Depends on, Skills (list or `None`), `Parallel: yes | no`, `Effort: S | M | L`, `Priority: must | should | could`, and an observable Done when. Only should/could work can sit below the cut line. Preserve checkpoints, change policy, replanning triggers, traceability, stable IDs and user decisions.
Readiness moves from `Draft` to `Ready for review` to `Reviewed and ready to execute`. Material edits invalidate prior review; execution progress stays in the coverage ledger. A direct user instruction to execute a supplied plan is sufficient authorization without forcing another skill review; record that basis honestly. Handover fields are Context, Decisions, State (done / pending / blocked), Remaining tasks (what, how, where), Verification, and Risks and how to detect them early.
- **`360-blueprint`** turns a vague goal into a complete, unambiguous plan written directly to a file, briefing the user in chat with key decisions, assumptions, and risks.
- **`360-faculty`** seats a short list of named experts fitted to this developer and to the plan's own complexity, depth, and nature, tailoring it surgically, remembering what it learns, and saving reusable teams that can be called by name later.
- **`360-expert-review`** tests that plan against relevant concrete failure cases, writing the finalized plan back to the file and briefing the user in chat.
- **`360-execute`** builds it, tracking every task in a coverage ledger written to disk, verifying each against its own acceptance check with evidence, and briefing progress in chat.
- **`360-backend-audit`** audits the resulting backend code for correctness, duplication, performance risks, and observability, writing the full report to a file and briefing findings in chat.
- **`360-optimize`** audits working code for zero-cost speed, weight, and reliability gains, ranking drop-in upgrades and restructures with conservative expected effects written to a file and briefing highlights in chat.
- **`360-token-efficiency`** runs alongside the active task with session consent, reducing avoidable overhead while preserving requirements and required checks.
Blueprint and expert review inspect the current project stage and completed work before asking questions. They clarify remaining gaps in interactive rounds with evidence-backed recommendations and a free-text note option, resolving material decisions while recording supported reversible defaults. When interactive tools are unavailable, they use equivalent choices and notes in chat.
Each skill also works standalone: ask `360-faculty` which expertise a plan needs without seating anyone, skip straight to `360-expert-review` for a plan someone else drafted, point `360-execute` at a plan someone else finalized, run `360-backend-audit` on existing code with no plan involved at all, audit working code for performance and weight with `360-optimize`, or apply `360-token-efficiency` to any task regardless of which other skills are in play.
## Optional Token Efficiency
Every sibling offers the companion once per session and reuses your approval or refusal across handoffs. Explicitly invoking it approves it for that session; a generated plan listing it does not. Unanswered offers leave it inactive. You can revoke consent immediately. A new session starts without approval, and consent does not authorize cross-session memory writes. Declining still permits ordinary efficient work.
The companion chooses techniques from capabilities the host actually exposes: search, selective reads, tool discovery, result processing, recoverable artifacts, context management, caching controls, delegation, and telemetry. Unknown capabilities stay unavailable. It does not rewrite agent configuration or installed instructions.
Strict accuracy is the default: never intentionally weaken requirements, exact values, evidence, uncertainty, or verification to save tokens. This is a working rule, not a guarantee that an AI cannot make mistakes. Potentially lossy compression requires separate approval of the technique, task-specific metric and tolerance. Missing or conflicting evidence triggers fuller retrieval or the ordinary workflow.
Caching can reduce processing cost without reducing context size. Savings reports are optional; absent telemetry is labeled `UNMEASURED`. Comparisons include skill loading, summarization, retries and delegation overhead, and use the same acceptance checks. See [evidence and limitations](skills/360-token-efficiency/references/evidence-and-techniques.md) and the [evaluation record](docs/validation/token-efficiency-redesign.md).
All skills prefer files where supported, honor explicit output requests, and provide complete in-session deliverables when files are unavailable. They label that fallback honestly.
## Design Principles
Understand the intended outcome, establish the relevant facts, choose the simplest supported action, and verify the result. When evidence is insufficient, preserve uncertainty and identify the next useful check.
1. **Evidence**: distinguish observed facts, supported defaults, assumptions and unknowns. Expert lenses advise; evidence validates.
2. **Proportionate coverage**: check relevant requirements, likely failures and severe plausible failures; disclose inspection limits.
3. **Observable gates**: an unknown check never becomes a pass. A completed assessment with limitations does not verify a proposed implementation.
4. **Simple decisions and recovery**: resolve routine choices within authorization, ask about material departures, and reconcile current state before resuming or retrying.
These rules reduce reliance on unstated judgment; they cannot guarantee flawless performance from every model. See the [logic review and scenario evidence](docs/validation/skill-logic-review.md).
## How Agent Skills Work
A skill is a folder containing a `SKILL.md` file with a `name`, a `description`, and step-by-step instructions. Agents load skills through progressive disclosure: they scan every skill's name and description at startup, then load the full instructions only when a task matches. This keeps many skills available at once without bloating the agent's context window. See the [Agent Skills specification](https://agentskills.io) for the full format.
## Repository Structure
```
360-skills/
├── README.md Project overview and install instructions
├── AGENTS.md Contributor guide for adding new skills
├── CONTRIBUTING.md Quick pointer to contributor guide
├── llms.txt Machine-readable index for AI engines
├── llms-full.txt Full compiled context for single-fetch LLM ingestion
├── LICENSE MIT license
├── docs/validation/ Redesign coverage and behavioral evaluation record
├── scripts/
│ ├── build-llms-full.mjs Compiles full documentation into llms-full.txt
│ ├── check-consistency.mjs Validates metadata, contracts, links and indexes
│ ├── lib/ Shared validator conventions
│ └── consistency.test.mjs Valid and invalid repository fixtures
├── .github/workflows/
│ └── consistency.yml Tests and validates on main pushes and PRs
└── skills/
├── 360-blueprint/
│ └── SKILL.md Skill definition and instructions
├── 360-faculty/
│ ├── SKILL.md Skill definition and instructions
│ └── references/ Optional seating roster
├── 360-expert-review/
│ └── SKILL.md Skill definition and instructions
├── 360-execute/
│ └── SKILL.md Skill definition and instructions
├── 360-backend-audit/
│ └── SKILL.md Skill definition and instructions
├── 360-optimize/
│ └── SKILL.md Skill definition and instructions
└── 360-token-efficiency/
├── SKILL.md Skill definition and instructions
└── references/ Evidence and optional techniques
```
Supporting references: faculty has an optional [seating roster](skills/360-faculty/references/seating-roster.md); token efficiency has optional evidence and technique guidance. Runtime faculty dossiers belong beside the user's plan and are not part of this catalog.
## Contributing
New skills must follow the structure, naming, and quality bar defined in [AGENTS.md](AGENTS.md). In short: one skill per directory under `skills/`, kebab-case names prefixed with `360-`, required frontmatter (`name`, `description`, `version`), and a fixed section order (Purpose, When to Use, Core Principle, Workflow, Output Format, Quality Gate).
Before opening a pull request, run the consistency validator:
```bash
node scripts/build-llms-full.mjs
node --test scripts/*.test.mjs
node scripts/check-consistency.mjs
```
The checker validates exact frontmatter fields, naming, section order, shared contracts, negative routing, local links, index registration, and compiled documentation (including references). Tests cover valid and invalid fixtures; they do not establish runtime accuracy. Behavioral scenarios and observed limitations live in the evaluation record.
## FAQ
**What is an Agent Skill?**
A portable, version-controlled folder that packages domain expertise and a repeatable workflow into instructions an AI agent can load on demand. See [agentskills.io](https://agentskills.io) for the open specification.
**Which AI agents can use these skills?**
Any agent supported by the [skills.sh](https://skills.sh) CLI, including Claude Code, Cursor, Codex, Windsurf, GitHub Copilot, OpenCode, and Gemini CLI.
**Why is it called 360-skills?**
Because the quality failures that matter most hide in the angles nobody checked. Every skill checks relevant requirements and failure cases, and states the limits of its inspection.
**How do I add a new skill?**
Read [AGENTS.md](AGENTS.md), create `skills/360-<name>/SKILL.md` following the required structure, then register it in this README's skills table.
## License
[MIT](LICENSE)
---
## Contributor Guide
# AGENTS.md: Contributor Guide for Agents
This file tells AI agents how to add new skills to this repository. Follow it exactly.
## Where Skills Live
Every skill is a directory under `skills/`, named in kebab-case, containing one `SKILL.md`:
```
skills/<kebab-case-name>/SKILL.md
```
Experimental or unproven skills go under `skills/.experimental/<kebab-case-name>/SKILL.md` instead. Promote them to `skills/` only once they're stable and ready to be listed publicly.
## Naming
- Every skill name carries the `360-` prefix (e.g. `360-expert-review`, `360-api-design`).
- The folder name and the frontmatter `name` must match exactly.
## Required Frontmatter
Every `SKILL.md` starts with YAML frontmatter containing exactly these fields:
```yaml
---
name: 360-example-skill
description: Perform a concrete action with a defined output. Use when a specific task needs that action.
version: 1.0.0
---
```
- `name`: must exactly match the folder name
- `description`: two or three sentences, action-oriented, covering both what the skill does and when to reach for it. This is the only text an agent sees before deciding to load the skill, so it is the trigger surface: make it specific enough to win the right tasks and lose the wrong ones.
- `version`: semantic version (`MAJOR.MINOR.PATCH`). Bump major when a skill's output shape changes, since other skills consume it; use minor for compatible workflow additions.
- Use single-line scalar values for these three fields; no extra keys, duplicate keys, YAML blocks, or implicit objects.
## Required Skill Body Structure
Every `SKILL.md` follows this section order:
1. **Purpose**: why this skill exists, in one or two sentences
2. **When to Use**: concrete triggers for reaching for this skill
3. **Core Principle**: the single idea the skill is built around
4. **Workflow**: the ordered steps the agent executes
5. **Output Format**: the exact shape of the final deliverable. May carry a domain-specific heading instead (`Plan Template`, `Final Plan Format`, `Audit Report Format`) as long as it defines that exact shape.
6. **Quality Gate**: a yes/no checklist that must fully pass before the work is considered done
## Style Rules
- Brief. Imperative. No filler.
- No redundant restrictions. Say a thing once, in the place it matters.
- Usable by any agent, not just one product's assistant.
- Prefer concrete checklists and steps over abstract advice.
- Use one governing idea in Core Principle, operational decisions in Workflow, and observable checks in Quality Gate.
- Distinguish facts, reversible defaults, assumptions and unknowns. Ask about material decisions; resolve routine choices from evidence and existing authorization.
- Check relevant requirements, likely failures and severe plausible failures; disclose inspection limits instead of promising exhaustive reliability.
- An unknown check never passes. A bounded assessment can finish with disclosed limits; implementation requires its acceptance evidence.
## Family Conventions
Skills install individually. Repeat the following contract locally in every skill that emits or consumes plans; do not require another skill's installation or template. Preserve valid existing plan structure.
### Shared Plan Contract
Preserve stable task IDs and existing user decisions. Every task carries: ID, What, How, Where, Depends on, Skills (list or `None`), Parallel, Effort, Priority, Done when (observable). Use `Priority: must | should | could`, `Effort: S | M | L`, and `Parallel: yes | no`. Only `should` and `could` sit below the cut line. Preserve phase checkpoints, change policy, replanning triggers, and objective-to-task traceability.
Plan status: `Draft` → `Ready for review` → `Reviewed and ready to execute`. Open blocking questions keep it `Draft`; a complete unreviewed plan is `Ready for review`; a passed review sets `Reviewed and ready to execute`. Material edits after review return it to `Ready for review` (or `Draft` if blocked). Execution progress belongs in the ledger, not the readiness status. An explicit user instruction to execute a supplied plan authorizes execution without a mandatory sibling review; record that basis without claiming a review occurred.
Every skill that produces a handover uses these six fields:
> Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
Each skill's **When to Use** ends with a single `- Not for ...` line naming its neighboring skills in backticks. Copy that line (without the bullet) into the README routing table's **Not for** column. The validator checks both destinations and exact synchronization.
### Optional Companion Convention
Every sibling skill, including future additions, starts its Workflow with this identical entry step. Token efficiency carries its own non-recursive consent handling.
### Entry: Optional Token Efficiency
Reuse explicit approval or refusal for `360-token-efficiency` from this session. If unknown and not already offered, ask once whether to enable it for this session; continue the main task with the overlay inactive while unanswered. Explicit user invocation counts as approval; merely appearing in a generated plan does not. On approval, discover and load it through the host's supported skill mechanism, reusing already-loaded instructions. If unavailable, explain briefly and continue; do not install automatically. Refusal disables the overlay, not ordinary efficient habits. Revocation takes effect immediately. Keep consent in-session only; it does not authorize cross-session memory writes. The overlay never invokes itself or restarts the parent skill.
Consent is a session decision, not a prerequisite for the main task. Keep unknown/offered, approved, refused, and revoked state in-session across skill handoffs. An unanswered offer is not approval and must not be repeated on each handoff. A new session starts unknown. Do not persist consent in faculty dossiers or other cross-session memory.
### Portable Delivery and Verification
Every sibling repeats these Delivery rules inside Workflow:
### Delivery
Prefer a recoverable file when supported, using the existing path or the default below. Honor explicit user output requests. If files are unavailable, deliver the same complete structure in-session and label it `in-session only; not persisted`; never claim a file was saved. With a saved file, chat normally carries a short briefing and its path. These delivery rules also apply to the templates and quality gate below.
Use observable preservation and acceptance checks. Structural validation cannot prove behavior or accuracy; evaluate realistic consent, capability, recovery, and handoff scenarios separately. Document unavailable telemetry and regressions; do not infer universal accuracy or token savings from finite tests.
### Interactive Completion Convention
Every skill, including future additions, ends its Workflow with the following completion step and local, result-dependent routing guidance. The token-efficiency companion shares its parent’s prompt; standalone use still offers a next action. Include a matching Quality Gate item. Preserve deliverable formats; this is a separate interaction.
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Use this exact Quality Gate item in each skill:
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
## Registering a New Skill
Adding a skill to `skills/` (not `.experimental/`) is not complete until both of these are updated:
1. **`README.md`**: add a row to the skills index table with name, description, and version, plus a row in the routing table.
2. **`llms.txt`**: add a bullet to the Skills list with the description copied verbatim from the frontmatter.
Optional, and local only: `.claude-plugin/marketplace.json`. That directory is gitignored and never ships with the repository, so keep it in sync only if you maintain a local copy for plugin testing. It does not gate "done".
Skills under `skills/.experimental/` are not registered anywhere until promoted.
## Before You're Done
Run `node scripts/build-llms-full.mjs`, `node --test scripts/*.test.mjs`, and `node scripts/check-consistency.mjs`. The checker validates public and experimental skill structure, local reference links, consent and plan contracts, routing/index registration, and compiled documentation. Experimental skills are excluded from indexes and the bundle. Fixture tests exercise both valid inputs and specific failures.
- [ ] Folder name is kebab-case and starts with `360-`
- [ ] Frontmatter `name` matches the folder name exactly
- [ ] `description` is two or three sentences, action-oriented, states when to use it
- [ ] `version` is valid semver
- [ ] Body follows the required section order
- [ ] Family conventions followed: priority vocabulary, handover fields, negative trigger
- [ ] `README.md` index, `README.md` routing table, and `llms.txt` are all updated (unless experimental)
- [ ] `node scripts/check-consistency.mjs` exits 0
Supporting Markdown references belong inside the skill directory and must be linked from its entrypoint or another reachable reference. Load them only when relevant. The compiled bundle includes these references for single-fetch use; routine skill loading does not.
---
## Contributing
# Contributing to 360-skills
Follow [AGENTS.md](AGENTS.md) for metadata, local contracts, optional session consent, routing, and independent installation requirements. Keep instructions and references scoped to the task; preserve user decisions and capability limits.
After changing skills or documentation, regenerate and verify:
```bash
node scripts/build-llms-full.mjs
node --test scripts/*.test.mjs
node scripts/check-consistency.mjs
```
Use Node 20 or later; no dependencies or Python packaging are required. Add validator fixtures for structural changes and realistic behavioral scenarios for instruction changes. Record actual outcomes, unavailable measurements, and regressions. See the [redesign evaluation](docs/validation/token-efficiency-redesign.md) for examples.
Tie each instruction change to a concrete failure scenario and the smallest correction. Preserve output schemas and existing user decisions; keep advisory judgment separate from evidence. Test structural conventions with positive and negative fixtures, then record bounded walkthrough observations separately from expectations. See the [skill logic review](docs/validation/skill-logic-review.md).
---
## Skills Catalog
### Skill: 360-backend-audit
Source: skills/360-backend-audit/SKILL.md
---
name: 360-backend-audit
description: Deep-audit backend code, write the full report to a file, and brief the user in chat with bugs, updates, and dead weight. Use before or after significant backend work, or when inheriting, refactoring, or handing off services.
version: 2.2.0
---
# 360 Backend Audit
## Purpose
Audit APIs, services, business logic, data access, integrations and jobs for correctness, structure, performance and observability risks. Produce evidence-backed recommendations within an explicit inspection boundary.
## When to Use
- Before or after merging significant backend work
- When inheriting, refactoring, or modernizing services or data layers
- When code feels heavy, fragile, slow, or hard to debug
- Before handing code to another engineer or AI agent
- Not for planning work that does not exist yet (`360-blueprint`) or executing a plan (`360-execute`)
## Core Principle
Judge observed behavior against the intended contract.
## Workflow
### Entry: Optional Token Efficiency
Reuse explicit approval or refusal for `360-token-efficiency` from this session. If unknown and not already offered, ask once whether to enable it for this session; continue the main task with the overlay inactive while unanswered. Explicit user invocation counts as approval; merely appearing in a generated plan does not. On approval, discover and load it through the host's supported skill mechanism, reusing already-loaded instructions. If unavailable, explain briefly and continue; do not install automatically. Refusal disables the overlay, not ordinary efficient habits. Revocation takes effect immediately. Keep consent in-session only; it does not authorize cross-session memory writes. The overlay never invokes itself or restarts the parent skill.
### Shared Plan Contract
Preserve stable task IDs and existing user decisions. Every task carries: ID, What, How, Where, Depends on, Skills (list or `None`), Parallel, Effort, Priority, Done when (observable). Use `Priority: must | should | could`, `Effort: S | M | L`, and `Parallel: yes | no`. Only `should` and `could` sit below the cut line. Preserve phase checkpoints, change policy, replanning triggers, and objective-to-task traceability.
Plan status: `Draft` → `Ready for review` → `Reviewed and ready to execute`. Open blocking questions keep it `Draft`; a complete unreviewed plan is `Ready for review`; a passed review sets `Reviewed and ready to execute`. Material edits after review return it to `Ready for review` (or `Draft` if blocked). Execution progress belongs in the ledger, not the readiness status. An explicit user instruction to execute a supplied plan authorizes execution without a mandatory sibling review; record that basis without claiming a review occurred.
### Delivery
Prefer a recoverable file when supported, using the existing path or the default below. Honor explicit user output requests. If files are unavailable, deliver the same complete structure in-session and label it `in-session only; not persisted`; never claim a file was saved. With a saved file, chat normally carries a short briefing and its path. These delivery rules also apply to the templates and quality gate below.
Use observable preservation and acceptance checks. Structural validation cannot prove behavior or accuracy; evaluate realistic consent, capability, recovery, and handoff scenarios separately. Document unavailable telemetry and regressions; do not infer universal accuracy or token savings from finite tests.
### 1. Map the Audit Scope
Read-only by default: do not change application code, dependencies, configuration, or observability unless implementation is explicitly requested. Writing the audit artifact is allowed. Use non-mutating inspection and authorized isolated tests; record checks that cannot safely run.
- Establish expected behavior from requirements, public contracts, tests and current user decisions; distinguish these sources from what the implementation actually does
- Identify callers, dependencies, state changes, and trust boundaries
- Record observed inputs, outputs and side effects as evidence, not proof of correctness. Preserve intended behavior when recommending changes; a defect need not be preserved
- If the intended contract is unclear, state competing interpretations and the next useful check
### 2. Hunt Dead Weight
Inspect suspected dead code, redundancy, incomplete logic, swallowed failures, stale flags, orphaned configuration and unused paths. Before recommending deletion, check indirect callers, registration, reflection or generated entrypoints, configuration and external consumers where relevant. State the search boundary. No direct callers or an inconclusive search means uncertain use, not proven dead code.
### 3. Verify Code Accuracy
Scrutinize business logic, data integrity, numeric accuracy, concurrency, trust boundaries, and failure semantics.
For each finding, record the triggering input/state, expected versus observed behavior, consequence, source or reproduction evidence, confidence, and inspection boundary. Use `likely` for strong indirect evidence, `possible` or `uncertain` for hypotheses, and name the next check. A local reproduction proves only that local case; production impact needs evidence of reachability and relevant conditions.
### 4. Recommend Structural Improvements
- Recommend unifying duplicated business logic only when it is truly the same
- Recommend replacing reinvented wheels with mature, maintained, lighter alternatives when justified
- Assess separation of transport, domain logic, and data access cleanly
- Avoid premature abstraction
### 5. Assess Performance Honestly
Find real computational and data-access risks: N+1 access, missing indexes, unbounded result sets, repeated work, blocking I/O, caching with explicit staleness trade-offs.
- No performance recommendation may weaken functionality, reliability, accuracy, or precision
- If you did not measure it, call it a risk, not a measured result
- Rank opportunities by impact versus effort
### 6. Assess Observability
Audit the debugging surface: structured logs, correct log levels, useful error context, correlation IDs, audit trails for critical writes, metrics for latency, errors, and throughput.
Proposed observability changes should be additive and non-breaking; do not implement them during the audit.
### 7. Record Verification and Gaps
- Cover relevant requirements, likely failures and severe plausible failures within scope. Report authorized tests and their results; state uninspected paths, unavailable access and the checks needed to close gaps
- Where coverage is thin, propose missing tests first: boundary, failure, and concurrency cases
- Classify every recommendation as safe now, needs tests first, or needs human decision
- Deliver recommendations as an audit; execute only separately requested implementation within its authorized scope
### 8. Write the Audit
- Write the full audit to a file using the work-file template
- Reuse the existing path if known; otherwise `plans/<short-slug>-audit.md`; create the folder if needed; ask once if ambiguous
- If the file cannot be written, use the in-session delivery fallback
- Keep new work and updates to existing work in separate lists
- With a saved file, print the terminal briefing and path
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Local routing: Recommend planning actionable correctness or reliability fixes with `360-blueprint`, or `360-execute` when a suitable authorized plan already exists. Recommend `360-optimize` when optimization is the remaining need; stop when there is no justified follow-up.
## Output Format
### Work file
1. Overview
2. Dead weight
3. Accuracy
4. Duplication and structure
5. Performance risks
6. Observability
7. Risk register
8. Implementation plan — grouped as New vs Updates to existing
9. Handover — Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
Keep the finding details from Workflow 3 inside the existing report sections. Call an area clean only within verified coverage; mark uninspected areas explicitly. A completed bounded audit does not verify proposed fixes or unavailable production behavior. Implementation tasks use the Shared Plan Contract. Label all proposed changes as recommendations, not completed fixes.
### Terminal briefing
Use this shape. Omit any section that would be empty. Follow the Delivery rules for file or in-session output.
```text
Backend audit — complete
Full audit: <path>
Bugs found
- <bug> — <proven|likely|possible|uncertain>
Updates to existing
- <fix or change to current behavior>
Features to add
- <new capability, if any>
Dead weight
- <remove or unused>
Need from you
- <decision required to proceed>
```
- Talk to the user, not the next agent
- A new artifact is Features to add. A change to an existing artifact, feature, or document is Updates to existing. Never mix them
- Confidence: `proven` evidence in hand; `likely` strong reason; `possible` suspected; `uncertain` hypothesis. Never numbers. Never say proven without evidence
## Quality Gate
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
The audit is complete only when every answer is yes:
- Every finding distinguishes expected and observed behavior, trigger, consequence, evidence or hypothesis, confidence and inspection boundary
- Clean claims are restricted to inspected coverage; unavailable checks remain limitations
- Deletion recommendations account for indirect, registered, configured and external use, or remain uncertain
- Local reproductions are not overstated as production impact
- No recommendation trades away functionality, reliability, accuracy, or precision without an explicit trade-off
- Unification is recommended only where logic is genuinely the same
- Library recommendations include maintenance, compatibility, and weight evidence or explicit verification gaps
- Performance claims are measured when measurement is possible, and labeled as risks when it is not
- The observability plan integrates without breaking behavior
- New work and updates to existing work are grouped separately
- The handover preserves evidence, limits and the next useful checks
- The deliverable is saved at the stated path, or honestly labeled in-session only
- The briefing omits empty sections and uses proven/likely/possible/uncertain, never numbers
Any "no" means the audit is not finished. Fix it and review again.
---
### Skill: 360-blueprint
Source: skills/360-blueprint/SKILL.md
---
name: 360-blueprint
description: Create an executable plan from a new objective with explicit tasks, constraints, and verification. Use when a goal exists but the path is unclear, or the request is "plan this".
version: 2.4.0
---
# 360 Blueprint
## Purpose
Turn a new objective into an executable plan with explicit assumptions, constraints, risks and acceptance checks.
## When to Use
- A new project, feature, migration, product, workflow, or initiative needs a plan
- A goal exists but the path is unclear
- Any request of the form "plan this"
- Not for hardening a plan that already exists (`360-expert-review`) or building one that is already final (`360-execute`)
## Core Principle
Resolve decisions that change success; make routine choices from evidence.
## Workflow
### Entry: Optional Token Efficiency
Reuse explicit approval or refusal for `360-token-efficiency` from this session. If unknown and not already offered, ask once whether to enable it for this session; continue the main task with the overlay inactive while unanswered. Explicit user invocation counts as approval; merely appearing in a generated plan does not. On approval, discover and load it through the host's supported skill mechanism, reusing already-loaded instructions. If unavailable, explain briefly and continue; do not install automatically. Refusal disables the overlay, not ordinary efficient habits. Revocation takes effect immediately. Keep consent in-session only; it does not authorize cross-session memory writes. The overlay never invokes itself or restarts the parent skill.
### Shared Plan Contract
Preserve stable task IDs and existing user decisions. Every task carries: ID, What, How, Where, Depends on, Skills (list or `None`), Parallel, Effort, Priority, Done when (observable). Use `Priority: must | should | could`, `Effort: S | M | L`, and `Parallel: yes | no`. Only `should` and `could` sit below the cut line. Preserve phase checkpoints, change policy, replanning triggers, and objective-to-task traceability.
Plan status: `Draft` → `Ready for review` → `Reviewed and ready to execute`. Open blocking questions keep it `Draft`; a complete unreviewed plan is `Ready for review`; a passed review sets `Reviewed and ready to execute`. Material edits after review return it to `Ready for review` (or `Draft` if blocked). Execution progress belongs in the ledger, not the readiness status. An explicit user instruction to execute a supplied plan authorizes execution without a mandatory sibling review; record that basis without claiming a review occurred.
### Delivery
Prefer a recoverable file when supported, using the existing path or the default below. Honor explicit user output requests. If files are unavailable, deliver the same complete structure in-session and label it `in-session only; not persisted`; never claim a file was saved. With a saved file, chat normally carries a short briefing and its path. These delivery rules also apply to the templates and quality gate below.
Use observable preservation and acceptance checks. Structural validation cannot prove behavior or accuracy; evaluate realistic consent, capability, recovery, and handoff scenarios separately. Document unavailable telemetry and regressions; do not infer universal accuracy or token savings from finite tests.
### 1. Ground the Objective in the Current Project
- Inspect available instructions, plans, decisions, relevant artifacts, recent changes, verification and any execution ledger. Establish the current stage and done/pending/blocked work; cite sources and distinguish observed facts from claims. Record unavailable evidence
- Reuse current evidence and user answers. Resolve discoverable facts by inspection; do not ask the user to rediscover them
- Separate remaining gaps: material user decisions change the outcome, scope, acceptance criteria, significant cost or irreversible effects; reversible implementation defaults stay within established intent and conventions. Choose supported defaults, record the reason and revisit condition, and proceed. Do not present a default or inference as a fact
- Ask only for unresolved material decisions, conflicting instructions, or inaccessible facts needed to judge success. Use the available interactive question tool in focused rounds of one to three questions, with the best supported recommendation, its trade-off, and a free-text/custom option. For missing facts, request the needed evidence without inventing it. If tools are unavailable, offer equivalent numbered choices and notes in chat
- Unanswered required questions block dependent planning and readiness, not independent inspection. Silence, a recommendation, a preselected option or elapsed time is not an answer. Honor explicit delegation or deferral within its limits
- After answers, update the existing context, assumptions or review sections and inspect newly relevant evidence. Stop questioning when material choices are settled; reopen only on new conflicting evidence
### 2. Establish Readiness
- Restate the intended outcome, observable success and what must not happen
- Keep outcome-changing unanswered choices in `Draft`, with the required decision and affected tasks; do not relabel a blocker as a default
- A complete plan with supported defaults may be `Ready for review`. An optional companion offer does not block it
- If a draft is requested despite blockers, deliver it honestly and actively prompt the needed decisions
### 3. Examine Relevant Risks
- First principles: what is known, what is assumed, what is unknown
- Inversion: what would guarantee failure
- Second-order effects: what each major step sets in motion
- Stakeholders: who is affected, who decides, who executes, who can block
- Constraints: time, people, skills, systems, dependencies, unknowns
### 4. Design the Strategy
- Choose the simplest path that fully satisfies the objective
- Sequence by dependency and by risk
- Put discovery work first when uncertainty could invalidate the plan
- Declare what is out of scope and how scope changes are handled
### 5. Craft the Plan
- Structure the plan as phases and tasks
- Separate new capabilities from changes to existing features, behavior, or documents
- Every task states what, how, where, and done when
- Mark effort, priority, and parallelization
- Define the cut line
- Use established stack choices and conventions. Record reversible defaults and falsifiable assumptions with supporting evidence or a pending validation method in `Validated by`; keep material unresolved choices as blockers
- Keep names and terms consistent from start to finish
- Make the plan self-contained for a fresh executor
### 6. Enforce Quality by Domain
Turn quality goals into task-specific acceptance checks.
- Code: identify required outputs, boundaries, failure behavior and compatibility; add performance or scaling checks only for relevant workloads and constraints
- Documentation: identify the intended reader action, facts and links to preserve, and how the result will be checked; update existing material when it serves that need
- Other domains: state observable success and likely or severe plausible failure checks
- Reuse existing structure where it fits. Require a concrete maintenance or correctness benefit before adding abstraction or consolidating superficially similar logic
### 7. Stress-Test Before Delivery
- Walk the plan end to end
- Run a premortem
- Falsify assumptions
- Check whether different readings change success or safety; permit equivalent reversible implementations
- Verify every objective maps to tasks and every task serves an objective
### 8. Deliver in Two Channels
- Write the full plan to a file using the work-file template
- Reuse the existing path if known; otherwise `plans/<short-slug>.md`; create the folder if needed; ask once if the location is ambiguous
- If the file cannot be written, use the in-session delivery fallback
- With a saved file, print the terminal briefing and path
- If the plan deserves adversarial review before build, say so in plain language — no skill names
- Ensure required answers were resolved before ready/final delivery; keep unanswered blockers in `Draft`; never conclude by dumping passive questions at the end of the conversation
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Local routing: Recommend `360-faculty` when tailoring or expertise would materially improve the prepared plan; otherwise recommend `360-expert-review`. Honor an explicit instruction to execute the supplied plan without imposing another review.
## Output Format
### Work file
```markdown
# Plan: <title>
| Field | Value |
|---|---|
| Objective | <one sentence> |
| Status | Draft / Ready for review |
| Version | <version or date> |
| Created | <date> |
## 1. Objective & Definition of Done
- Goal:
- Done when:
- Success measures:
- Must not happen:
## 2. Context & Constraints
- Background:
- Constraints:
- Stakeholders:
- Open questions and decision points:
## 3. Strategy
- Chosen path:
- Why it wins:
- Alternatives rejected:
## 4. Scope
- New:
- Updates to existing:
- Explicitly out of scope:
- Change policy:
- Cut line: <only should/could below it>
## 5. Assumptions
| # | Assumption | Validated by |
|---|---|---|
| A1 | | |
## 6. Phases & Tasks
### Phase 1: <name>
Checkpoint:
**Task 1.1 — <name>**
- What:
- How:
- Where:
- Depends on:
- Skills: <list or None>
- Parallel: yes | no
- Effort: S | M | L
- Priority: must | should | could
- Done when:
## 7. Risks & Countermeasures
| Risk | Impact | Countermeasure |
|---|---|---|
## 8. Verification & Replanning
- Per task:
- Overall:
- Replan when:
## 9. Traceability
| Objective | Covered by tasks |
|---|---|
## 10. Handover Summary
- Context:
- Decisions:
- State: done / pending / blocked
- Remaining tasks: what, how, where
- Verification:
- Risks and how to detect them early:
```
Handover fields: Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
Fill every field or write `N/A` with a one-line reason.
### Terminal briefing
Use this shape. Omit any section that would be empty. Follow the Delivery rules for file or in-session output.
```text
<what this plan is> — plan is ready
Full plan: <path>
Features to add
- <new capability as an outcome>
Updates to existing
- <change to something that already exists>
Not adding
- <left out on purpose>
Issues found
- <problem or uncertainty> — <proven|likely|possible|uncertain>
```
- First line is `plan is ready` or `draft — open questions remain`
- Talk to the user, not the next agent. Outcomes, not tasks
- A new artifact is Features to add. A change to an existing artifact, feature, or document is Updates to existing. Never mix them
- Confidence: `proven` evidence in hand; `likely` strong reason; `possible` suspected; `uncertain` hypothesis. Never numbers
- No phases, tasks, MoSCoW, status tables, handover, or skill names
- Never append open questions, design choices, or trailing bullet lists at the end of the conversation (no `Need from you` dump). Use the clarification loop for every required question before concluding
## Quality Gate
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
The plan is ready only when every answer is yes:
- The objective is clear and confirmed, or the output is explicitly marked `Draft`
- Current project stage and completed work were inspected, or unavailable evidence was explicitly recorded
- Existing answers were reused; follow-up rounds resolved the goal, scope, outline, and required details without treating silence or recommendations as consent
- No blocking question or decision was skipped
- Material unresolved decisions were actively prompted with evidence, recommendations and a free-text option; reversible defaults have a reason and revisit condition
- No invented technical details were presented as facts
- Every assumption is explicit and falsifiable
- Every task has what, how, where, done when, dependencies, skills, effort, priority, and parallel markings
- The cut line is defined
- Every objective maps to tasks and no orphan tasks remain
- New work and updates to existing work are grouped separately
- A fresh executor can distinguish requirements, defaults, assumptions and unresolved blockers
- Checkpoints and replanning triggers exist
- The premortem covered relevant requirements, likely failures and severe plausible failures; inspection limits are stated
- The deliverable follows the template and accurately states its location or in-session status
- The briefing omits empty sections and uses proven/likely/possible/uncertain, never numbers
Any "no" means the plan is not finished. Refine and review again.
---
### Skill: 360-execute
Source: skills/360-execute/SKILL.md
---
name: 360-execute
description: Execute a finalized plan task by task with a persisted coverage ledger, verify every item with evidence, and brief the user in chat. Use when a plan exists and work must begin, or when resuming a partial execution.
version: 2.2.0
---
# 360 Execute
## Purpose
Execute an authorized plan against its acceptance checks and keep a recoverable coverage ledger of results, deviations and unfinished work.
## When to Use
- A plan exists and work must begin — any domain, any scale
- Any request of the form "implement this plan", "build this", "execute this"
- Resuming a partially executed plan by reconciling its ledger with current evidence
- Not for creating plans (`360-blueprint`) or reviewing drafts (`360-expert-review`)
## Core Principle
Current evidence establishes progress; the ledger records it.
## Workflow
### Entry: Optional Token Efficiency
Reuse explicit approval or refusal for `360-token-efficiency` from this session. If unknown and not already offered, ask once whether to enable it for this session; continue the main task with the overlay inactive while unanswered. Explicit user invocation counts as approval; merely appearing in a generated plan does not. On approval, discover and load it through the host's supported skill mechanism, reusing already-loaded instructions. If unavailable, explain briefly and continue; do not install automatically. Refusal disables the overlay, not ordinary efficient habits. Revocation takes effect immediately. Keep consent in-session only; it does not authorize cross-session memory writes. The overlay never invokes itself or restarts the parent skill.
### Shared Plan Contract
Preserve stable task IDs and existing user decisions. Every task carries: ID, What, How, Where, Depends on, Skills (list or `None`), Parallel, Effort, Priority, Done when (observable). Use `Priority: must | should | could`, `Effort: S | M | L`, and `Parallel: yes | no`. Only `should` and `could` sit below the cut line. Preserve phase checkpoints, change policy, replanning triggers, and objective-to-task traceability.
Plan status: `Draft` → `Ready for review` → `Reviewed and ready to execute`. Open blocking questions keep it `Draft`; a complete unreviewed plan is `Ready for review`; a passed review sets `Reviewed and ready to execute`. Material edits after review return it to `Ready for review` (or `Draft` if blocked). Execution progress belongs in the ledger, not the readiness status. An explicit user instruction to execute a supplied plan authorizes execution without a mandatory sibling review; record that basis without claiming a review occurred.
### Delivery
Prefer a recoverable file when supported, using the existing path or the default below. Honor explicit user output requests. If files are unavailable, deliver the same complete structure in-session and label it `in-session only; not persisted`; never claim a file was saved. With a saved file, chat normally carries a short briefing and its path. These delivery rules also apply to the templates and quality gate below.
Use observable preservation and acceptance checks. Structural validation cannot prove behavior or accuracy; evaluate realistic consent, capability, recovery, and handoff scenarios separately. Document unavailable telemetry and regressions; do not infer universal accuracy or token savings from finite tests.
### 1. Load the Plan Completely
Never execute a plan you have not fully read.
- Read the entire plan before touching anything: objective, scope, assumptions, every phase, every task, every checkpoint
- Build the full task inventory: every task ID, its priority, its dependencies, its "done when" check
- If an acceptance check is missing, derive it from the approved objective; ask only when alternatives would change the intended outcome
- Resolve material ambiguity before dependent work; continue independent authorized work
Never automatically clear history. Use context management only when the host supports it. Before compaction or handoff, preserve and verify the six handover fields against the plan and ledger, including constraints, authorization, unresolved uncertainty, evidence locations, and exact values. Reopen original evidence if the summary cannot support the next decision. Without recoverable artifacts or context controls, keep state in-session and explain any continuity limit.
### 2. Persist the Coverage Ledger
Reconcile the plan, ledger, actual artifacts or external state, and relevant verification before resuming. A ledger verdict is a recorded claim until current evidence supports it.
- Write it to a file and keep it current after each task and significant partial effect
- Reopen stale or unsupported verified rows; record what changed and invalidate affected verification, including dependent results. Recheck only the affected acceptance conditions
- Reuse the existing path if known; otherwise `plans/<short-slug>-execution.md`; create the folder if needed; reuse repository conventions for routine location choices
- If the file cannot be written, use the in-session delivery fallback
- One row per task: ID, name, priority, acceptance check, status
- Statuses: `pending` / `in progress` / `done (verified)` / `blocked` / `dropped (approved)`
- Mark verified only when current acceptance evidence supports it; include source revision or state identifiers when relevant
- Record remaining uncertainty and the next check so a fresh agent can resume safely
### 3. Execute in Order
- Follow phase order and task dependencies exactly; honor parallel markers
- Load declared skills through the supported host mechanism when available. The optional `360-token-efficiency` always follows session consent; a plan listing is not approval. If another declared skill is unavailable, use the self-contained contract and available capabilities; block only tasks that actually require the missing capability
- Verify each phase checkpoint before advancing dependent work. A failure blocks affected descendants; continue independent authorized work where dependencies and parallel markers permit
- Keep work within the task and change policy
- Before a non-repeatable or external action, record intent and a recoverable operation identifier or reconciliation method. After interruption, inspect partial effects and external state before retrying. Reuse confirmed results; retry only if absence or safe repeatability is established. If state cannot be determined, block that action and identify the missing check
### 4. Verify Every Task
- Run the task's "done when" check and record the evidence in the ledger
- Done means the check passed with evidence — never "looks right", never "should work"
- If an implementation acceptance check cannot run, mark the task `blocked` with the missing evidence. An assessment task may finish with disclosed limits only when its own acceptance criteria permit them; an unknown check never becomes a pass
- Check affected behavior and the diff after each task; state the inspected boundary and regressions rather than claiming everything else is safe
### 5. Handle Deviations in the Open
When reality disagrees with the plan — a failed assumption, missing information, a visibly better path:
- Classify the departure under the plan's change policy and existing authorization
- Resolve routine reversible implementation choices within that authority; record the reason and affected checks without asking again
- Escalate changes to outcome, scope, acceptance criteria, significant cost, irreversible effects or authorization. Stop dependent work and present concrete options; continue independent authorized tasks
- Never silently absorb new scope, never silently skip a task
- A `must`-priority task is never dropped without an explicit user decision; `should`/`could` tasks follow the plan's cut line
### 6. Sweep and Write the Report
Before declaring completion, walk the ledger top to bottom:
- Every task has a final status — zero unaccounted items
- Cross-check the plan's traceability: every objective maps to verified work
- Run the plan's overall verification; confirm all checkpoints passed
- Sweep once more for regressions introduced across phases
- Write the execution report into the same ledger file
- With a saved file, print the terminal briefing and path
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Local routing: Recommend `360-backend-audit` after significant backend work, or `360-optimize` when working code has a relevant performance or weight concern. For other completed work, recommend a concrete domain-appropriate follow-up or stopping; do not invent backend work to fit the flow.
## Output Format
### Work file
The ledger file contains:
1. Coverage ledger — every task: ID, name, priority, final status, evidence for each `done (verified)`
2. Deviations — what diverged, how it was resolved, who approved it; or "None"
3. QC results — checks run, checkpoints verified, regression sweeps, outcomes
4. Unfinished items — pending, blocked, or dropped, with reason and approval; or "None"
5. Handover summary — Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
### Terminal briefing
Use this shape. Omit any section that would be empty. Follow the Delivery rules for file or in-session output.
```text
Execution — <n>/<m> tasks verified
Full report: <path>
Done
- <outcome delivered>
Updates to existing
- <change made to something that already existed>
Blocked
- <item and why>
Deviations
- <what changed and whether it was approved>
Issues found
- <bug or surprise> — <proven|likely|possible|uncertain>
```
- Talk to the user, not the next agent
- Done is new work shipped. Updates to existing is a change to something that already existed. Never mix them
- Confidence: `proven` evidence in hand; `likely` strong reason; `possible` suspected; `uncertain` hypothesis. Never numbers. Never say proven without evidence
- With file delivery, keep the briefing concise; name a missing skill when it explains a limitation
## Quality Gate
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
Execution is complete only when every answer is yes:
- Every task in the plan appears in the ledger file with a final status — zero unaccounted items
- Every `done (verified)` verdict is backed by evidence from the task's own acceptance check
- Every phase checkpoint passed before its dependent work advanced; independent work respected authorization and dependency/parallel markers
- Every deviation was recorded and resolved within existing authorization/change policy or an explicit user decision
- No `must`-priority task was dropped or skipped without explicit user approval
- Every objective in the plan's traceability maps to verified work
- Regressions and collateral damage were swept for, and the results are stated
- Declared skill availability and consent were respected; required unavailable capabilities are explicit blockers
- Unfinished items are stated honestly — pending, blocked, or dropped, with reasons
- The ledger matches current artifacts and verification; stale verdicts and partial actions were reconciled before resumption
- The next agent can recover evidence, uncertainty and the next check; affected verification was invalidated after relevant changes
- Delivery honors the requested format and accurately states persistence
- The briefing omits empty sections and uses proven/likely/possible/uncertain, never numbers
Any "no" means execution is not finished. Fix it and re-run the sweep.
---
### Skill: 360-expert-review
Source: skills/360-expert-review/SKILL.md
---
name: 360-expert-review
description: Stress-test a draft plan, write the finalized executable plan back to the same file, and brief the user in chat. Use before executing any plan where a missed case could cause real damage.
version: 3.4.0
---
# 360 Expert Review
## Purpose
Expose concrete failure cases and weak assumptions in a draft plan, then revise it in place until applicable readiness checks pass or blockers are explicit.
## When to Use
- A plan changes real systems, data, users, money, security, or operations
- A wrong assumption or missed case could create real damage
- A plan needs adversarial review before execution
- Not for drafting plans from scratch (`360-blueprint`) or executing a finalized plan (`360-execute`)
## Core Principle
Try to disprove readiness with concrete failure cases.
## Workflow
### Entry: Optional Token Efficiency
Reuse explicit approval or refusal for `360-token-efficiency` from this session. If unknown and not already offered, ask once whether to enable it for this session; continue the main task with the overlay inactive while unanswered. Explicit user invocation counts as approval; merely appearing in a generated plan does not. On approval, discover and load it through the host's supported skill mechanism, reusing already-loaded instructions. If unavailable, explain briefly and continue; do not install automatically. Refusal disables the overlay, not ordinary efficient habits. Revocation takes effect immediately. Keep consent in-session only; it does not authorize cross-session memory writes. The overlay never invokes itself or restarts the parent skill.
### Shared Plan Contract
Preserve stable task IDs and existing user decisions. Every task carries: ID, What, How, Where, Depends on, Skills (list or `None`), Parallel, Effort, Priority, Done when (observable). Use `Priority: must | should | could`, `Effort: S | M | L`, and `Parallel: yes | no`. Only `should` and `could` sit below the cut line. Preserve phase checkpoints, change policy, replanning triggers, and objective-to-task traceability.
Plan status: `Draft` → `Ready for review` → `Reviewed and ready to execute`. Open blocking questions keep it `Draft`; a complete unreviewed plan is `Ready for review`; a passed review sets `Reviewed and ready to execute`. Material edits after review return it to `Ready for review` (or `Draft` if blocked). Execution progress belongs in the ledger, not the readiness status. An explicit user instruction to execute a supplied plan authorizes execution without a mandatory sibling review; record that basis without claiming a review occurred.
### Delivery
Prefer a recoverable file when supported, using the existing path or the default below. Honor explicit user output requests. If files are unavailable, deliver the same complete structure in-session and label it `in-session only; not persisted`; never claim a file was saved. With a saved file, chat normally carries a short briefing and its path. These delivery rules also apply to the templates and quality gate below.
Use observable preservation and acceptance checks. Structural validation cannot prove behavior or accuracy; evaluate realistic consent, capability, recovery, and handoff scenarios separately. Document unavailable telemetry and regressions; do not infer universal accuracy or token savings from finite tests.
### 1. Understand the Project First
- Inspect available instructions, plans, decisions, relevant artifacts, recent changes, verification and any execution ledger. Establish the current stage and done/pending/blocked work; cite sources and distinguish observed facts from claims. Record unavailable evidence
- Reuse current evidence and user answers. Resolve discoverable facts by inspection; do not ask the user to rediscover them
- Separate remaining gaps: material user decisions change the outcome, scope, acceptance criteria, significant cost or irreversible effects; reversible implementation defaults stay within established intent and conventions. Choose supported defaults, record the reason and revisit condition, and proceed. Do not present a default or inference as a fact
- Ask only for unresolved material decisions, conflicting instructions, or inaccessible facts needed to judge success. Use the available interactive question tool in focused rounds of one to three questions, with the best supported recommendation, its trade-off, and a free-text/custom option. For missing facts, request the needed evidence without inventing it. If tools are unavailable, offer equivalent numbered choices and notes in chat
- Unanswered required questions block dependent planning and readiness, not independent inspection. Silence, a recommendation, a preselected option or elapsed time is not an answer. Honor explicit delegation or deferral within its limits
- After answers, update the existing context, assumptions or review sections and inspect newly relevant evidence. Stop questioning when material choices are settled; reopen only on new conflicting evidence
- Compare the supplied plan with current artifacts before judging it. Identify stale assumptions, duplicated completed work and conflicting decisions; preserve valid task IDs and progress records
### 2. Review Through Expert Lenses
Use only the lenses this project needs:
- Architecture: boundaries, simplicity, trade-offs
- Engineering: correctness, edge cases, integration, performance
- Product and UX: user value, friction, recovery, accessibility
- QA: scenario coverage, regressions, acceptance criteria
- Operations: deployment, observability, rollback
- Security and privacy: access, data exposure, abuse cases
- Domain: business rules, terminology, real-world accuracy
Do not role-play personas. Extract findings directly from each lens.
### 3. Check Relevant User Outcomes
- Trace the real need through normal use, likely failures and severe plausible failures
- Check empty, error, interrupted and recovery states where they affect this task
- Include affected user groups, accessibility and connectivity constraints when relevant; state inspection limits
### 4. Make Findings Actionable
- For each actionable finding, name its trigger, consequence, evidence or explicit hypothesis, and smallest correction or next check
- Rank consequence separately from confidence. A plausible severe failure can justify investigation without being a proven defect
- Require work only when it protects a requirement or addresses a concrete material failure; speculative improvements remain optional
- Map only affected components, interfaces and state. Check retries, duplicates, concurrency, dependencies and rollback where those mechanisms exist
### 5. Check Reliability and Traceability
- Trace relevant requirements to implementation tasks and observable verification
- Add validation, failure handling, security, compatibility, monitoring and ownership only where a concrete risk demands them
- Prefer the simplest correction that protects the intended outcome; do not introduce operational machinery into an unrelated documentation task
### 6. Define Proportionate Verification
- Match each acceptance check to the consequence it detects; “works correctly” is not observable
- Distinguish a reviewed plan from a verified implementation. A planned test has not run
- Mark irrelevant checks non-applicable with reasons. Missing evidence remains unknown; obtain it or disclose the resulting limitation/blocker
### 7. Attack the Plan
Switch to hostile critic. Ask:
- What is assumed but unproven?
- What is missing: a user, a state, a sequence, a permission, a failure mode?
- What is the most likely failure?
- What is the most damaging failure?
- Which step could be read two different ways?
- What cannot be detected, reproduced, or reversed?
- What can be simpler?
Fix actionable findings and recheck affected risks. Stop when applicable checks pass and material blockers are resolved. An unavailable check stays unknown; if it prevents judging readiness, keep the plan non-final and name the missing evidence or decision. Otherwise disclose the limit without manufacturing required work. Use the clarification loop only for material user decisions; record routine defaults. Do not loop indefinitely.
### 8. Write the Final Plan
- Update the draft in place using the local minimum contract below. Preserve valid existing headings and structure; no sibling installation is required
- Preserve task IDs; add, split, or drop a task only with a stated reason
- Keep new work and updates to existing work in separate scope lists
- Put review findings, key decisions, and remaining risks in a short appendix in that same file
- Include the executor handover in the chosen delivery format
- If the file cannot be written, use the in-session delivery fallback
- With a saved file, print the terminal briefing and path
- Ensure required answers were resolved before ready/final delivery; keep unanswered blockers in `Draft`; never conclude by dumping passive questions at the end of the conversation
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Local routing: Recommend `360-execute` for a finalized executable plan. For a blocked or incomplete review, offer resolution of the concrete blocker or `360-blueprint` for substantial replanning; do not present execution as ready.
## Output Format
### Work file
Preserve the existing plan and fill only missing contract elements. Minimum standalone shape:
1. Metadata: objective, status, version/date
2. Objective and observable definition of done
3. Context, constraints, scope, assumptions, open questions and decisions
4. Strategy, alternatives, change policy and cut line
5. Phases and tasks using the Shared Plan Contract, with checkpoints
6. Risks and countermeasures
7. Per-task and overall verification, replanning triggers and traceability
8. Handover: Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
9. Review appendix: findings, decisions, remaining risks, applicable checks and blockers
Set `Reviewed and ready to execute` only when the applicable Quality Gate passes. Otherwise use `Draft` for blockers and `Ready for review` for incomplete review without blockers. Preserve user-approved trade-offs; record unresolved high-severity risks as blockers. A typical appendix (retain an existing heading when present):
```markdown
## 11. Review appendix
- Findings:
- Decisions:
- Remaining risks:
- Applicable checks and evidence:
- Non-applicable checks and reasons:
- Blockers and resolution needed:
```
Do not replace the task list with a narrative plan.
### Terminal briefing
Use this shape. Omit any section that would be empty. Follow the Delivery rules for file or in-session output.
```text
<what this plan is> — plan is final
Full plan: <path>
What changed
- <material delta from the draft>
Features to add
- <new capability as an outcome>
Updates to existing
- <change to something that already exists>
Issues found
- <hole, bug, or weak assumption> — <proven|likely|possible|uncertain>
Remaining risks
- <accepted risk> — <likely|possible|uncertain>
```
- First line is `plan is final` or `not final — <specific blocker or remaining review>`
- Talk to the user, not the next agent. Outcomes, not tasks
- A new artifact is Features to add. A change to an existing artifact, feature, or document is Updates to existing. Never mix them
- Confidence: `proven` evidence in hand; `likely` strong reason; `possible` suspected; `uncertain` hypothesis. Never numbers. Never say proven without evidence
- No phases, tasks, skill names, or review-essay dump
- Never append open questions, design choices, or trailing bullet lists at the end of the conversation (no `Need from you` dump). Use the clarification loop for every required question before concluding
## Quality Gate
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
The plan is final only when every applicable answer is yes (record a reason for each non-applicable item):
- Real user need understood and served
- Current project stage and completed work were inspected, or unavailable evidence was explicitly recorded
- Existing answers were reused; follow-up rounds resolved the goal, scope, outline, and required details without treating silence or recommendations as consent
- The right expert lenses were applied
- Relevant requirements, likely failures and severe plausible failures were checked; inspection limits are explicit
- Every actionable finding has a trigger, consequence, evidence or hypothesis, and smallest correction; severity and confidence are separate
- Relevant requirements are traceable to tasks and observable verification; monitoring is included only where needed
- Material failure cases have proportionate detection, diagnosis and recovery checks
- The design is clear, consistent, and as simple as possible
- Testing matches the risk
- Release and rollback checks address the plan’s relevant failure cases
- The plan survived hostile review
- Material decisions were resolved through the clarification loop; defaults and unknown evidence are explicit, never silent passes
- Task IDs were preserved or changed with a stated reason
- New work and updates to existing work are grouped separately
- The finalized plan preserves its original location where supported, or follows the Delivery fallback
- The briefing omits empty sections and uses proven/likely/possible/uncertain, never numbers
Any unresolved "no" keeps the plan non-final. Fix actionable findings or report the blocker and the evidence or decision needed to resolve it.
---
### Skill: 360-faculty
Source: skills/360-faculty/SKILL.md
---
name: 360-faculty
description: Seat a living, tailored expert team on a plan or task. Use when work must fit this developer's goals, taste, mindset, and strategy, when a plan needs the right expertise chosen for its complexity, depth, and nature, or when a named faculty team must be created, called, or updated. Recommends a short list, asks only the questions that still change the work, polishes immediately, and keeps upgradable memory.
version: 1.3.0
---
# 360 Faculty
## Purpose
Use advisory expert lenses to tailor work to the developer’s goals, taste, mindset and strategy. Produce a fit assessment, roster, authorized memory and polished plan or concrete guidance, according to the selected mode.
## When to Use
- A plan or task needs experts fitted to this developer, not a generic panel
- The developer asks which expertise this work needs, and wants the best fit chosen from its complexity, depth, and nature
- A named faculty team must be created, called, or updated
- A project needs a house style before more planning
- The developer says "faculty this", "add experts", or "seat a team"
- A seated faculty must update its understanding and retouch the plan
- Not for writing the plan itself (`360-blueprint`), attacking a finished draft (`360-expert-review`), or building one (`360-execute`)
## Core Principle
Expert lenses improve decisions; evidence validates them.
## Workflow
### Entry: Optional Token Efficiency
Reuse explicit approval or refusal for `360-token-efficiency` from this session. If unknown and not already offered, ask once whether to enable it for this session; continue the main task with the overlay inactive while unanswered. Explicit user invocation counts as approval; merely appearing in a generated plan does not. On approval, discover and load it through the host's supported skill mechanism, reusing already-loaded instructions. If unavailable, explain briefly and continue; do not install automatically. Refusal disables the overlay, not ordinary efficient habits. Revocation takes effect immediately. Keep consent in-session only; it does not authorize cross-session memory writes. The overlay never invokes itself or restarts the parent skill.
### Shared Plan Contract
Preserve stable task IDs and existing user decisions. Every task carries: ID, What, How, Where, Depends on, Skills (list or `None`), Parallel, Effort, Priority, Done when (observable). Use `Priority: must | should | could`, `Effort: S | M | L`, and `Parallel: yes | no`. Only `should` and `could` sit below the cut line. Preserve phase checkpoints, change policy, replanning triggers, and objective-to-task traceability.
Plan status: `Draft` → `Ready for review` → `Reviewed and ready to execute`. Open blocking questions keep it `Draft`; a complete unreviewed plan is `Ready for review`; a passed review sets `Reviewed and ready to execute`. Material edits after review return it to `Ready for review` (or `Draft` if blocked). Execution progress belongs in the ledger, not the readiness status. An explicit user instruction to execute a supplied plan authorizes execution without a mandatory sibling review; record that basis without claiming a review occurred.
### Delivery
Prefer a recoverable file when supported, using the existing path or the default below. Honor explicit user output requests. If files are unavailable, deliver the same complete structure in-session and label it `in-session only; not persisted`; never claim a file was saved. With a saved file, chat normally carries a short briefing and its path. These delivery rules also apply to the templates and quality gate below.
Use observable preservation and acceptance checks. Structural validation cannot prove behavior or accuracy; evaluate realistic consent, capability, recovery, and handoff scenarios separately. Document unavailable telemetry and regressions; do not infer universal accuracy or token savings from finite tests.
### 0. Pick the Mode
Four entry points into one workflow. Choose from what the developer asked; clarify only if modes would materially differ in authorized work. In `suggest` mode, keep profile and fit notes in-session; do not edit the plan, dossiers or teams.
| Mode | Trigger | Does | Never does |
|---|---|---|---|
| `suggest` | "who should look at this", "which experts", "help me pick" | Steps 1–4, then stops with the fit assessment, short list, and rationale | Seat, ask past fit, or touch the plan |
| `seat` | "faculty this", "add experts", "seat a team" — the default | Steps 1–8: recommend, confirm, seat one at a time, polish after each | Batch the polish to the end |
| `team` | "save this as X", "call the X team", "add Y to X" | Step 7: create, call, or edit a named team | Skip the drift check |
| `quiet` | the developer replies `quiet` | Seats the confirmed set, asks blocking questions only, polishes from current evidence, records gaps as explicit assumptions | Block on equivalent reversible implementation choices |
### 1. Load What Exists
Read the current task or draft plan. Then read dossiers if present:
- `faculty/developer.md` — goals, taste, mindset, strategy, house style
- `faculty/roster.md` — who has been seated on this project, and when
- `faculty/teams.md` — named teams this developer created
- `faculty/<role>.md` — each faculty's claims about this developer and project; `<role>` is kebab-case
`faculty/` sits beside the plan file. If the plan is not on disk, put it at the project root. If files are unavailable, keep the same structure in-session and label it not persisted. Persist dossiers only within the user-authorized faculty scope; token-efficiency consent does not authorize memory writes.
If `faculty/developer.md` is missing or empty, do not seat yet. Measure first.
### 2. Measure the Developer
Ask at most 3 questions. Stop early if all four measures are already known. Each question must reveal at least one measure:
- Goals — what this project must become, and must never become
- Taste — what they find elegant, ugly, overbuilt, or cheap
- Mindset — how they decide, what they protect, how they take challenge
- Strategy — the live bet: speed, simplicity, control, craft, learning, or revenue
Prefer a forced choice over an essay. Never ask what the plan, repo, or dossiers already answer.
Record the developer profile in the authorized file or in-session structure before seating. Derive the house style from the answers:
- How to ask — choices vs open, blunt vs gentle
- Standing refusals — what this developer will not accept
- Push strength per area — where to spar, where to finish quietly
- Maturity map — which frames are already strong, which are still forming
House style is one shop. Every seated faculty wears it; none invents a private personality.
### 3. Assess the Fit
Read three dimensions off the plan, or off the stated task when no plan exists yet. Never skip this — it is what makes the short list defensible instead of habitual.
- **Nature** — what the work touches: user-facing, data, money, security and privacy, infrastructure, ML and agents, docs and content, org and process. Decides **which** lenses.
- **Depth** — blast radius and reversibility: reversible-local, reversible-shared, hard to reverse, irreversible or regulated. Decides **seat urgency**.
- **Complexity** — moving parts: task count, phase count, cross-system dependencies and unresolved assumptions. Helps size the distinct expertise needed.
Use these bands as guidance, not quotas. Each seat must add a distinct decision or failure check; combine overlapping lenses and use fewer seats when sufficient:
| Complexity | Depth | Seats |
|---|---|---|
| low | reversible | 1–2 |
| moderate | reversible | 2–3 |
| moderate | hard to reverse | 3–4 |
| high | any | 4–5 |
| any | irreversible or regulated | Consider 4–5; include the relevant security, privacy, compliance or reliability lens when its concrete risk demands it |
Five is the cap unless the developer asks for more. State the dimensions and chosen band, explaining distinct value and any reduction. A single seat may cover several related checks.
### 4. Recommend a Short List
Recommend only faculties that add distinct value to this task. Never dump the Seating Roster unless the developer asks for it.
| # | Faculty | Seat | Why now | If absent |
|---|---|---|---|---|
| 1 | <role> | must / should / could | <the task ID, assumption, or risk row that demands this lens> | <the concrete missed decision or failure check> |
- `must` — a concrete material failure requires this expertise
- `should` — silence creates likely waste or rework
- `could` — the developer may want this lens; their call
`Why now` cites a real element of the plan or task. A recommendation that cannot cite one is a habit, not a fit — drop it.
Choose expertise from actual failure modes and cite the relevant task or risk. Derive an unlisted role when needed. Read the optional [seating roster](references/seating-roster.md) only when expertise selection needs examples; do not load or scan the catalog by default.
Accept these replies: `accept`, `1-3`, `drop 4 add privacy`, `only architect`, `quiet`.
Do not seat until the developer confirms. If they name a faculty you did not recommend, seat it.
In `suggest` mode, stop here and deliver:
- **Asked before planning** — the fit assessment, the short list, and the house-style questions worth answering first. Then offer `360-blueprint`.
- **Asked before review** — a prioritized lens list for `360-expert-review`: which of its lenses this plan needs, in what order, plus any lens the plan needs that its seven do not name. Hand it over; do not run the review.
### 5. Seat One Faculty at a Time
For each confirmed faculty, in the listed order:
**Load.** Read `faculty/developer.md` and `faculty/<role>.md`.
**Ask.** Hard caps:
- New faculty on this project: at most 3 questions
- Returning faculty: at most 1, and only if this task contradicts or extends its memory
- Zero when memory plus the plan is enough
- `quiet` mode: zero unless the answer changes the intended outcome, scope, acceptance criteria, significant cost or irreversible effects
A question must do exactly one job: intention, taste, or a sharper frame offered as a choice. Never ask two questions that one answer would cover.
**Follow or raise.** Follow when the developer's cut matches their goals, their taste, and a professionally sound path. Raise only for a material decision supported by task evidence, such as:
- The frame is unfinished — they asked for a feature when they need a decision
- A better cut exists — same goal, cleaner seam, less future pain
- The thought can be shaped — they are close, but the frame is wrong
Raise in their language. Put two cuts side by side — theirs and the tailor's — and name what each protects. Then wait.
- They take the raise → polish to the tailor's cut; upgrade memory: this mind can be moved on this point
- They keep their cut → tailor that cut; record the accepted risk; do not fight it again without new evidence. `360-expert-review` may reopen it later with evidence — that is its job, not a contradiction of this one
**Polish now.** Edit only the steps, states, checks, and risks this faculty advises on. Regenerate the whole plan only if the faculty proved the objective itself wrong. If no plan exists, write concrete guidance within this faculty’s scope, then offer `360-blueprint`.
Hold the plan contract while polishing. `360-execute` reads these fields directly, and prose in their place breaks execution:
- Preserve every field of any task touched: ID, `Depends on`, `Skills`, `Parallel`, `Effort`, `Priority`, `Done when` — along with phase checkpoints, the change policy, the replanning triggers, and the traceability table
- A task a faculty adds gets a stable ID continuing that phase's numbering, `Effort: S | M | L`, `Priority: must | should | could`, `Parallel: yes | no`, a `Skills` list or `None`, and an observable `Done when`
- Never claim a review occurred. Preserve an unchanged reviewed plan's status; after material edits return it to `Ready for review`, or `Draft` if blocking questions remain
Write findings into the plan's existing semantic sections. The blueprint uses the numbers below; other valid plans may use different headings. Preserve an existing review appendix (normally section 11), without inventing sections 12–14:
- Accepted risks → section 7: record actual acceptance, countermeasure and operational owner (a real person/team, or explicitly unassigned). A faculty is only the advisory lens, never an operational owner
- Explicit assumptions → section 5: `Validated by` cites actual evidence and its boundary, or `Pending: <validation method>`. A seated role or its agreement cannot validate an inference
- Open questions → Context & Constraints; decisions → Handover Summary
- One header row so a fresh reader knows a faculty pass happened: `| Tailored by | faculty/roster.md @ <date> |`
**Resolve disagreements in the open.** Check conflicting advice against evidence and existing decisions. Resolve equivalent reversible details with a recorded reason; present material trade-offs needing user choice with what each protects.
**Remember.** Update dossiers (Step 6), show the faculty block (Output Format), then move to the next seat — or stop when the developer stops or the remaining faculties would not change the plan.
Block dependent planning only when a missing answer changes success, scope, acceptance criteria, significant cost, irreversible effects or authorization. Otherwise choose a supported reversible default, record its basis and revisit condition, and continue.
### 6. Upgrade Memory
Every claim keeps this shape:
| Claim | Source | Status | Touches |
|---|---|---|---|
| <one sentence> | user-said / inferred / observed | active / superseded / disputed | <plan sections this claim may change> |
Rules:
- New answer agrees → keep, refresh
- New answer sharpens → replace the claim; mark the old one superseded
- Keep provenance in Source: user statement or artifact reference, date/revision and observed boundary. Preserve the source kinds and status vocabulary above
- Explicit current user instructions override stale preferences; supersede the old claim with the current source without asking for redundant confirmation. Ask only when current instructions materially conflict or their intended scope is unclear
- Inferred never overwrites user-said or becomes validated through faculty agreement
- Stale or unused means freshness is uncertain, not disputed. Note the freshness limit in Source and recheck before reuse; reserve disputed for contradictory evidence, superseded for an actual replacement
- `faculty/developer.md` changes only when goals, taste, mindset, strategy, or a standing refusal changes
- `faculty/<role>.md` changes after every seating of that role
- `faculty/roster.md` records every seated role and its last seated date
Memory must stay short enough that a fresh agent can read it cold.
### 7. Create and Call Teams
A team is a named lineup this developer can reuse. Teams live in `faculty/teams.md` beside the dossiers:
```markdown
## Team: <name>
- For: <the kind of work this team is cut for>
- Members: <role>, <role>, <role>
- Created: <date> · Last called: <date>
- Notes: <what this team has learned about this kind of work>
```
- **Create** — `save this as <name>` turns the session's seated set into a team; `team create <name>: <roles>` builds one from scratch
- **Call** — `call <name>` loads the team and skips Step 4's recommendation, going straight to seating. Dossiers and every question cap still apply
- **Edit** — `team <name> add <role>` or `drop <role>`; record the change and its date
**Drift check, always.** After loading a team, still run Step 3 against the current plan. If the fit assessment surfaces a `must` seat the team lacks, name it in one line with its plan evidence and ask before seating. A stale team must never silently under-cover a new plan — the roster is a floor for teams too.
Keep every team short enough to read cold.
### 8. Close the Session
After the last seated faculty:
- Present the session report (Output Format)
- Use the requested handover format or the Delivery default. State whether dossiers and teams were persisted; ask about persistence only when required and not already authorized
- If the next agent would have to guess the house style, the roster, or any accepted risk, the handover is not finished
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Local routing: Recommend `360-blueprint` when guidance still needs a plan, `360-expert-review` when a prepared plan is ready for review, or resuming the active task after an expertise-only consultation. Apply this at the end of the selected mode, including suggestion-only mode, not after each faculty seat.
## Output Format
Canonical handover fields: Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
Before recommending, and always in `suggest` mode:
```markdown
### Fit assessment
Nature: <what the work touches>
Depth: <reversible-local / reversible-shared / hard to reverse / irreversible or regulated>
Complexity: <low / moderate / high> — <the counts that decided it>
Seats sized: <n>–<n>
```
After each faculty:
```markdown
### Faculty: <role>
Questions: <the questions asked, or "none — memory enough">
Follow or raise: <followed / raised — tailor taken / raised — developer kept cut>
Plan touch:
- <section or task>: <what changed and why>
Memory:
- <claim> (<source>, <status>)
Next: <next faculty> or stop
```
After the session:
```markdown
# Faculty Session: <title>
## Developer fit
- Goals:
- Taste:
- Mindset:
- Strategy:
- House style:
## Fit assessment
- Nature:
- Depth:
- Complexity:
- Seats sized: <band> — seated: <n>
## Seated
| Faculty | New or returning | Questions asked | Follow or raise |
|---|---|---|---|
## Plan changes
- <each surgical edit, grouped by faculty>
## Teams
- <team name>: <created / called / edited / none this session>
- Drift check: <missing must seats raised, or "none — team covered the plan">
## Memory written
- faculty/developer.md: <created / updated / unchanged>
- faculty/roster.md: <created / updated>
- faculty/teams.md: <created / updated / unchanged>
- faculty/<role>.md: <created / updated>
- Persist: <path / in-session only; not persisted>
## Accepted risks and explicit assumptions
- <risk or assumption> — advisory responsibility: <faculty>; operational owner: <person/team or unassigned>; evidence or pending validation: <source/check>; written to plan section <5 / 7 / 10>
## Handover
- Context:
- Decisions:
- State: done / pending / blocked
- Remaining tasks: what, how, where
- Verification:
- Risks and how to detect them early:
- Issued as: <prompt / document / declined by user>
## Next
- 360-expert-review / 360-blueprint / more faculties / stop
```
## Quality Gate
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
The session is finished when every gate applicable to the selected mode passes. In `suggest` mode, finish after the fit assessment, evidence-backed short list, and rationale; seating, plan edits, dossier writes, and team management gates do not apply. State why any other gate is non-applicable.
- The mode was clear, and `suggest` mode stopped before seating and left the plan untouched
- Goals, taste, mindset, and strategy were known from memory or asked in this session
- The fit assessment stated nature, depth and complexity; each seat adds distinct value, overlapping seats were reduced, and departures from guide bands are explained
- Every recommendation cited the plan element that demands it, chosen by failure mode rather than habit — including rare roles and roles the roster does not name
- A short list was shown and seating was confirmed, with no catalog dump unless the developer asked
- Every faculty stayed inside its question caps, and no question was asked whose answer the plan, repo, or dossiers already held
- Every faculty followed or raised; no silent override; no flattery
- Every raise placed two cuts side by side and waited for the developer
- The plan or guidance was polished after each seat, never batched to the end
- Every Plan Template field survived on every task touched: ID, Depends on, Skills, Parallel, Effort, Priority, Done when, checkpoints, change policy, replanning triggers, traceability
- Every task a faculty added carries a stable ID, declared skills, `Effort` / `Priority` / `Parallel` values from the family vocabularies, and an observable `Done when`
- Readiness reflects material edits and blockers; this skill never claimed a review occurred
- Accepted risks, assumptions, and open questions landed in the plan's existing sections; no new numbered section was invented
- Material faculty conflicts needing user choice were presented; routine evidence-backed resolutions were recorded
- Memory retains claim / source / status / touches with provenance; current instructions prevail, and stale evidence is not labeled disputed without contradiction
- Assumptions cite evidence or a pending validation method; advisory roles are distinct from operational owners
- Any called team was drift-checked against this plan, and a missing `must` seat was raised before seating
- House style is one shop, shared across all seated faculties
- The hostile critic was never seated — that role belongs to `360-expert-review`
- The handover carries context, decisions, state, remaining tasks, verification, and risks with how to detect them early
- Delivery follows the user's format and authorization; persistence or its absence is recorded
- A fresh agent can recover decisions, provenance, uncertainty and pending checks from the authorized record
An unresolved applicable gate needs a specific correction or an honest blocker. Do not seat faculties, edit plans, or write dossiers to satisfy gates excluded by the selected mode.
### Reference: skills/360-faculty/references/seating-roster.md
Source: skills/360-faculty/references/seating-roster.md
# Seating Roster
Optional examples. Read only when a task's risks leave a gap in expertise selection. Derive unlisted faculties when concrete failure modes demand them; combine overlapping lenses and require distinct value from every seat. These are advisory lenses, not validators of facts or operational risk owners. Never show it unless the developer asks.
**Human mind**
- UX psychologist — cognition, attention, memory, decision load
- Behavioral scientist — habits, defaults, incentives, dark-pattern refusal
- Emotional-design expert — trust, anxiety, delight, recovery after failure
- Inclusive-cognition expert — neurodiversity, literacy, aging, first-time vs expert
- Motivation specialist — why people start, stall, and abandon
- Trust psychologist — credibility, risk perception, permission to act
**Research**
- UX researcher — interviews, usability, evidence before opinion
- Ethnographer — real context of use, not lab tasks
- Jobs-to-be-done analyst — the job hired, not the feature requested
- Market researcher — alternatives, switching cost, category norms
- Accessibility researcher — who is excluded by the current path
- Support-insight analyst — tickets and complaints as product signal
**Experience design**
- Product designer — whole problem-to-interface path
- Interaction designer — flows, states, gestures, timing
- UI / visual designer — hierarchy, density, visual language
- Information architect — findability, navigation, mental model
- Service designer — cross-channel journey, handoffs, waiting
- Design-systems designer — tokens, consistency, reuse without sameness
- Content designer / UX writer — words as interface
- Conversational designer — chat, voice, agent tone, turn-taking
- Motion designer — feedback, orientation, reduced-motion respect
- Data-visualization designer — charts that tell truth, not decoration
- Onboarding designer — first-run, empty states, competence growth
- Error-experience designer — blame-free recovery, undo, next action
**Product and strategy**
- Product manager — outcome, priority, trade-off against goals
- Product strategist — positioning, bets, what not to become
- Product owner — backlog truth, acceptance, scope discipline
- Growth specialist — activation, retention, loops without coercion
- Monetization specialist — pricing, packaging, value exchange
- Roadmap / portfolio manager — sequencing across bets
- Opportunity discoverer — problem worth solving vs solution theater
- Competitive-intelligence analyst — copy nothing; steal only the job
**Delivery**
- Project manager — time, dependencies, cut line, status without theater
- Technical program manager — multi-team, multi-system sequencing
- Scrum master / delivery coach — flow, blockers, team health
- Release manager — ship window, freeze, comms, rollback clock
- Change manager — adoption inside the org that must live with it
- Risk officer — what can kill the plan, ranked by damage not drama
- Estimator / uncertainty specialist — ranges, not fake precision
**Architecture**
- Software architect — boundaries, simplicity, irreversible choices
- Systems architect — runtime, data, failure domains
- Solution architect — integration across existing systems
- Domain-driven design strategist — bounded contexts, language
- API designer — contracts, versioning, consumer empathy
- Data modeler — truth in storage, migrations, identity
- Integration architect — third parties, sync, eventual consistency
- Migration / legacy specialist — strangler paths, coexistence
- Complexity reductionist — delete before add
**Build**
- Backend engineer — logic, integrity, authorization, failure
- Frontend engineer — state, accessibility in pixels, perceived speed
- Mobile engineer — lifecycle, offline, store, device limits
- Desktop / native engineer — OS integration, install, updates
- CLI / developer-experience engineer — flags, scripts, composability
- Full-lifecycle feature engineer — one slice, all layers, no orphans
- Platform engineer — paved roads, internal products
- Build / tooling engineer — local loop, CI, reproducibility
- Performance engineer — latency, memory, budgets
- Concurrency / realtime specialist — races, ordering, backpressure
- Search / relevance engineer — find vs dump
- Offline / sync specialist — conflict, merge, user-visible truth
- Embedded / IoT specialist — when hardware is in the loop
- Gameplay / simulation specialist — loops, feedback, fairness
**Quality**
- QA strategist — what must be proven, at what cost
- Exploratory tester — the nasty path a script will never write
- SDET / automation engineer — durable checks, not brittle theater
- Test architect — pyramid, fixtures, environments
- Accessibility QA — WCAG as behavior, not a badge
- Localization tester — language, locale, cultural fit
- Chaos / resilience tester — kill dependencies on purpose
- UAT / acceptance specialist — "done" in the user's words
- Regression historian — what broke last time and why
**Operate and survive**
- SRE — SLOs, error budget, toil
- DevOps / delivery engineer — pipeline, environments, promotion
- Incident commander — detect, mitigate, communicate, learn
- Observability engineer — logs, traces, metrics, reproduction
- Capacity planner — load, cost, degradation
- FinOps / cost engineer — unit cost, waste, surprise bills
- Reliability engineer — graceful failure, idempotency, rollback
- Disaster-recovery specialist — backups that actually restore
- Environment / secrets steward — config, credentials, least privilege
**Security, privacy, abuse**
- Application-security engineer — threats in the actual design
- Threat modeler — assets, attackers, entry points
- Identity / auth specialist — sessions, tokens, account recovery
- Privacy engineer — collection, retention, consent, deletion
- Compliance officer — regulation that actually applies
- Cryptography specialist — only when crypto is the domain
- Abuse / fraud specialist — misuse, spam, automation, social attack
- Supply-chain security specialist — dependencies, provenance
- Security-UX specialist — safe defaults people will still use
**Data and intelligence**
- Product analyst — behavior vs intention
- Data analyst — questions, definitions, honest charts
- Data engineer — pipelines, quality, lineage
- Data scientist — prediction only when it beats a rule
- ML engineer — training, eval, drift, fallback
- MLOps specialist — reproducibility, promotion, rollback of models
- Evaluation / benchmarking specialist — claims vs numbers
- Information-retrieval / RAG specialist — grounding, citation, miss
- Agent architect — tools, memory, handoff, refusal
- Computer-vision specialist — data, labels, failure in pixels
- Human-in-the-loop designer — when the model must ask a person
**Words, brand, adoption**
- Brand strategist — promise, voice, what the product stands for
- Naming specialist — product, feature, and company language
- Technical writer — docs a stranger can finish a task from
- Developer advocate — when other builders are users
- Customer-success lead — activation after the sale
- Support engineer — first-line reality
- Sales engineer — what was promised vs what can ship
- Training designer — competence, not a tour
- Community / open-source maintainer — contribution, governance, tone
- Localization / i18n strategist — expansion without rewrite
**Business, legal, ethics**
- Business analyst — rules, processes, acceptance language
- Domain expert — the real-world craft the software sits inside
- Operations designer — the work around the software
- Stakeholder diplomat — who decides, who blocks, who is surprised
- Legal counsel — IP, contracts, liability, terms
- Licensing specialist — OSS and proprietary mix
- Procurement / vendor specialist — lock-in, SLA, exit
- Ethicist — who is harmed if this works as designed
- Digital-wellbeing specialist — attention, addiction, after-hours
- Sustainability specialist — energy, hardware waste, long life
- Accessibility policy lead — legal plus moral floor
- AI-policy / safety specialist — autonomy, consent, audit of agents
**Meta**
- Hostile critic — never seated here; owned by `360-expert-review`
- Handover specialist — the next agent needs zero guesses
- Simplifier — shortest design that still covers every required angle
- First-principles philosopher — separate fact from habit
- User-feeling advocate — thought, experience, and emotion as first-class
- Traceability clerk — requirement, task, test, production signal
- Professor-panel chair — convenes only the seated faculties this request needs
---
### Skill: 360-optimize
Source: skills/360-optimize/SKILL.md
---
name: 360-optimize
description: Audit working code for zero-cost speed, weight, and reliability gains — measure first, rank drop-in upgrades and restructures, and report conservative expected effects on this system. Use when existing code must run faster, lighter, or more robustly without changing what it does.
version: 2.2.0
---
# 360 Optimize
## Purpose
Assess working code for supported speed, weight and reliability gains that preserve its intended contract. Produce an audit with conservative expectations, unresolved candidates and rejected ideas; implementation requires separate authorization.
## When to Use
- Working code feels slow, heavy, or fragile under load, and behavior must stay the same
- A stack, library, or runtime may have a drop-in successor that is lighter or more robust
- A restructure might cut copies, I/O, or serialized work on a hot path
- Before or after `360-backend-audit` when the remaining question is performance and weight, not correctness
- Not for greenfield design (`360-blueprint`), not for executing changes (`360-execute`), not for correctness or dead-weight hunts (`360-backend-audit`)
## Core Principle
Recommend gains only as strongly as their evidence permits.
## Workflow
### Entry: Optional Token Efficiency
Reuse explicit approval or refusal for `360-token-efficiency` from this session. If unknown and not already offered, ask once whether to enable it for this session; continue the main task with the overlay inactive while unanswered. Explicit user invocation counts as approval; merely appearing in a generated plan does not. On approval, discover and load it through the host's supported skill mechanism, reusing already-loaded instructions. If unavailable, explain briefly and continue; do not install automatically. Refusal disables the overlay, not ordinary efficient habits. Revocation takes effect immediately. Keep consent in-session only; it does not authorize cross-session memory writes. The overlay never invokes itself or restarts the parent skill.
### Shared Plan Contract
Preserve stable task IDs and existing user decisions. Every task carries: ID, What, How, Where, Depends on, Skills (list or `None`), Parallel, Effort, Priority, Done when (observable). Use `Priority: must | should | could`, `Effort: S | M | L`, and `Parallel: yes | no`. Only `should` and `could` sit below the cut line. Preserve phase checkpoints, change policy, replanning triggers, and objective-to-task traceability.
Plan status: `Draft` → `Ready for review` → `Reviewed and ready to execute`. Open blocking questions keep it `Draft`; a complete unreviewed plan is `Ready for review`; a passed review sets `Reviewed and ready to execute`. Material edits after review return it to `Ready for review` (or `Draft` if blocked). Execution progress belongs in the ledger, not the readiness status. An explicit user instruction to execute a supplied plan authorizes execution without a mandatory sibling review; record that basis without claiming a review occurred.
### Delivery
Prefer a recoverable file when supported, using the existing path or the default below. Honor explicit user output requests. If files are unavailable, deliver the same complete structure in-session and label it `in-session only; not persisted`; never claim a file was saved. With a saved file, chat normally carries a short briefing and its path. These delivery rules also apply to the templates and quality gate below.
Use observable preservation and acceptance checks. Structural validation cannot prove behavior or accuracy; evaluate realistic consent, capability, recovery, and handoff scenarios separately. Document unavailable telemetry and regressions; do not infer universal accuracy or token savings from finite tests.
### 1. Lock the Contract
- Map purpose, inputs, outputs, ordering, error shapes, and caller-visible side effects
- Confirm that map against requirements and caller evidence; known defects are handed off, not preserved as intended behavior
- Honor already authorized contract changes; otherwise preserve the intended contract and ask only about material unresolved trade-offs
- If the contract is unclear, mark questions. Do not optimize guessing
### 2. Find Time and Weight
- Locate hot paths from profiles, traces, logs, complexity, or data volume
- Name the resource: latency, CPU, memory, I/O, binary size, dependency surface
- Record baseline and candidate under comparable representative workloads, input sizes, environment and run conditions; note variability and measurement limits
- Expose resource trade-offs: latency, CPU, memory, I/O, dependency/operational weight and money. A local speedup may move cost elsewhere
- Cold paths stay untouched
- If you did not measure, write `unmeasured` and keep the finding as a candidate
### 3. Hunt Waste First
Attack in this order. Stop at the first level that removes the cost.
1. Needless work — extra copies, repeated compute, unbounded scans, chatty I/O, work on the wrong side of a boundary
2. Algorithm and data structure on the hot path
3. Data movement — batching, pushdown, streaming, pagination that the contract already allows
4. Drop-in upgrade of what is already in the stack
5. Restructure
Do not micro-tune a cold path. Do not strip checks, reduce precision, or weaken failure handling to go faster.
### 4. Apply the Zero-Cost Gate
Admit a change only when every check is yes:
- Function: observable outputs and caller-visible side effects stay the same
- Reliability: failure semantics, integrity, bounds, and precision are not weaker
- Weight: net dependencies, operational surface, and code complexity do not grow
- Money: no new paid obligation unless it replaces a larger one
- Reversal: a rollback path exists
- Gain: a supported improvement on a named axis from a representative comparison on this system, with uncertainty and resource trade-offs stated
A failed check means Rejected, even if another axis improved. An unknown check leaves an unresolved candidate with the missing evidence and next check in Measurement/Verification; it is not admitted to Ranked gains, Upgrades or Restructures. Only all-pass items are supported recommendations. An audit can complete with no admitted gains.
### 5. Gate Upgrades and Libraries
Prefer, in order: delete the work; use the language and current stack; a compatible version of an existing dependency; a maintained drop-in with the same contract and less weight; a new library only when it removes more than it adds.
For every candidate tool or library, record:
- Contract match (API, errors, types, concurrency, numeric behavior)
- Maintenance and license compatibility
- What it removes, not only what it adds
- The named axis it wins on, with evidence from this system — not from the vendor's bench
Stdlib or an existing dependency wins when it meets the need. Do not add a dependency to look modern.
For an emergent or young tool, check an adapter, rollback and evidence of improvement on a named axis. Known absence of a required safeguard fails; an uninspected safeguard remains unknown under the same gate.
Never name a tool because it is popular. Name it because it passed this gate on this code.
### 6. Restructure Only When Shape Is the Cost
Restructure when the current shape forces extra I/O, extra copies, or serialized work on a hot path.
Forbidden as "optimization": layering fashion, package renaming, architecture theater, new abstractions that do not remove cost.
Callers and error shapes stay identical unless the user approved a contract change.
### 7. Prove No Cost
- Existing tests that cover the contract must still pass
- Where coverage is thin on a touched path, propose contract tests first: outputs, errors, ordering, bounds, concurrency
- Classify each item: `safe now` / `needs tests first` / `needs measurement first` / `needs human decision`. Items lacking gate evidence remain candidates in Measurement/Verification, regardless of a promising local result
- When in doubt, leave it in the audit and wait. Do not rewrite production code under this skill
Correctness bugs found while scanning are out of scope. Record them as `handoff to 360-backend-audit`, do not disguise them as optimizations.
### 8. Bound the Expectation
Every admitted gain gets a conservative expected effect. This step is not optional.
- Baseline is this codebase, this load shape, this environment
- Report the typical production path. Best-case is not the expectation
- A local win stays local until you can scale it by that path's share of end-to-end cost. If the share is unknown, say `end-to-end unknown`
- Do not add overlapping recommendations as if they stack
- Vendor benches, blog numbers and other projects do not establish this system’s gain. A local microbenchmark supports only its tested path and workload; do not generalize it without representative evidence
- Unmeasured candidates: no magnitude, rank, or `proven`; name the check needed before promotion
- Rank `high` only when evidence shows this path dominates the named resource
- Round the gain down and the cost of the change up (effort, overhead, risk)
- If the sentence still sells after the evidence is removed, rewrite it
- State `Holds when` and `Falsified when` for every expected effect
### 9. Write the Audit
- Write the full audit to a file using the work-file template
- Reuse the existing path if known; otherwise `plans/<short-slug>-optimize.md`; create the folder if needed; ask once if ambiguous
- If the file cannot be written, use the in-session delivery fallback
- Rank admitted gains by conservative impact versus effort versus risk
- Keep new work and updates to existing work in separate lists
- With a saved file, print the terminal briefing and path
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Local routing: Recommend `360-blueprint` to turn supported findings into scoped work, or `360-execute` when a suitable authorized plan already exists. Recommend stopping if no worthwhile candidates remain; findings alone do not authorize implementation.
## Output Format
### Work file
1. Contract — what must not change
2. Measurement — what was profiled, load, environment, what is `unmeasured`; unresolved candidates and missing gate evidence
3. Ranked gains — only items that passed the Zero-Cost Gate. Each row:
- Change and axis (speed / weight / reliability)
- Evidence (this system)
- Baseline
- Expected effect — conservative typical case, scoped `local` or `end-to-end`
- Holds when
- Falsified when
- Rank: `high` / `medium` / `low` (supported gains only)
- Effort and classification
4. Upgrades — drop-in runtime, library, or tool changes that passed
5. Restructures — shape changes that passed, with why the current shape is the cost
6. Rejected — attractive ideas that failed the gate, with the failed check
7. Handoff bugs — correctness or dead-weight items for `360-backend-audit`; or "None"
8. Verification — tests run/to run, rollback, how to detect a silent contract break; next checks for unresolved candidates
9. Implementation plan — grouped as New vs Updates to existing
10. Handover — Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
Implementation tasks use the Shared Plan Contract. Keep candidates separate from supported gains using the existing Measurement/Verification sections. State inspection limits; completion of the assessment does not verify a candidate or its implementation. No magnitude without a measurement on this system.
### Terminal briefing
Use this shape. Omit any section that would be empty. Follow the Delivery rules for file or in-session output.
```text
Optimize — complete
Full audit: <path>
Gains
- <change, axis, local|end-to-end> — <proven|likely|possible|uncertain>
Upgrades
- <drop-in replacement that passed the gate>
Restructures
- <shape change that passed the gate>
Rejected
- <idea and the check it failed>
Handoff
- <bug or dead weight for a correctness audit>
Need from you
- <decision required to proceed>
```
- Talk to the user, not the next agent
- A new artifact is Features to add only if a harness or tool must be introduced; a change to current code is Updates to existing. Never mix them
- Confidence: `proven` evidence in hand; `likely` strong reason; `possible` suspected; `uncertain` hypothesis. Never numbers. Never say proven without evidence
- Never hype. If you cannot state a conservative expected effect, the gain is not ready to brief
- Keep the briefing concise when the audit is saved to a file
## Quality Gate
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
The audit is complete only when every answer is yes:
- The contract is written and treated as the floor
- Hot paths were located from evidence, or marked `unmeasured`
- Waste was considered before upgrades, and upgrades before restructure
- Every admitted change has affirmative evidence for every Zero-Cost Gate check; unknowns remain candidates and failures are Rejected
- Every Rejected item names the check it failed
- No recommendation weakens function, reliability, precision, or failure semantics
- No new dependency was proposed that fails net-lighter
- Emergent tools meet the same pass/fail/unknown rules for adapter, rollback and improvement evidence
- Every admitted gain has baseline, conservative expected effect, scope, Holds when, and Falsified when
- No magnitude, rank, or `proven` on an unmeasured candidate
- Comparisons use representative workloads and expose resource trade-offs and variability
- No vendor, blog, or other-project number is presented as this system's gain
- Best-case is not reported as typical; overlapping gains are not stacked
- End-to-end claims are scaled by path share, or marked `end-to-end unknown`
- Correctness bugs are handed off, not sold as optimizations
- New work and updates to existing work are grouped separately
- The handover distinguishes supported gains, rejected ideas, unresolved candidates and their next checks
- The deliverable is saved at the stated path, or honestly labeled in-session only
- The briefing omits empty sections, uses proven/likely/possible/uncertain, never numbers, and does not hype
Any "no" means the audit is not finished. Fix it and review again.
---
### Skill: 360-token-efficiency
Source: skills/360-token-efficiency/SKILL.md
---
name: 360-token-efficiency
description: Reduce avoidable context and tool-output overhead with capability-aware retrieval, reuse, and verified handovers. Use as an optional, session-approved companion when token cost or context growth matters, with strict preservation of task requirements by default.
version: 2.1.0
---
# 360 Token Efficiency
## Purpose
Reduce avoidable overhead during the main task using capabilities the host actually exposes. Preserve task requirements and verification; strict mode forbids intentional accuracy sacrifice but cannot guarantee error-free outcomes.
## When to Use
- Multi-step or tool-heavy tasks with growing context or repeated retrieval
- A user explicitly invokes this companion, or approves a sibling's session offer
- Not for creating plans (`360-blueprint`), executing the task itself (`360-execute`), or optimizing application performance (`360-optimize`)
## Core Principle
Remove redundant work, not required evidence. Judge preservation against observable acceptance checks, and claim savings only from actual measurements.
## Workflow
### 1. Respect Session Consent
Explicit user invocation counts as approval for this session. Otherwise reuse the session's explicit approval or refusal; merely appearing in a generated plan is not approval. If unknown and not already offered, ask once and continue the main task with the overlay inactive while unanswered. Refusal or revocation disables the overlay immediately without disabling ordinary efficient habits. A new session starts unknown; a verified continuation in the same session preserves the decision. Do not infer consent from an old artifact or cross-session memory.
Discover and load through the host's supported skill mechanism only after approval; reuse already-loaded instructions. If unavailable, explain briefly and continue the main task without installing anything. Never invoke this skill recursively or restart the parent skill. Session consent does not authorize cross-session memory writes, configuration changes, or edits to installed skills.
### 2. Assess Available Capabilities
Make a brief internal, in-session assessment from exposed tools, instructions, and observed behavior. Unknown capabilities remain unavailable until established; do not infer them from the agent's brand.
| Capability | Use when established | Fallback |
|---|---|---|
| Search and selective reads | Locate evidence, then retrieve relevant spans with source locations | Read supplied inputs; request only missing material needed for a decision |
| Tool discovery | Discover needed tools on demand | Use the exposed tool set |
| Code execution | Filter and aggregate large results before returning them to model context | Request bounded results or process manageable chunks |
| Recoverable artifacts | Keep task state and evidence references in authorized files or artifacts | Keep the complete required state in-session; do not claim persistence |
| Context management | Use supported compaction with a verified handover | Keep a state summary; never automatically clear history |
| Caching controls | Use exposed controls when appropriate and authorized | Make no claim of cache control or savings |
| Delegation | Delegate only when permitted and the independent work justifies coordination | Work locally |
| Usage telemetry | Record comparable observed usage | Label savings `UNMEASURED` |
Do not modify the agent's configuration or installed skill to adapt it. Read [evidence and optional techniques](references/evidence-and-techniques.md) only when choosing caching, compression, delegation, or measurement techniques, or when the user asks for supporting evidence.
### 3. Retrieve and Reuse Carefully
- Fully read mandatory instructions and required task inputs. Progressive retrieval must not bypass them.
- Search progressively for additional evidence. Keep source paths, ranges, IDs, versions or timestamps needed to reopen it.
- Filter or aggregate large tool results before returning them to context when supported. Check pagination, truncation, counts, and omitted boundaries; a partial result cannot establish completeness.
- Reuse verified facts while checking whether their sources changed. Reopen stale, conflicting, or insufficient evidence before deciding.
- Avoid repeated explanations, whole-artifact regeneration for local edits, and redundant verification. Repeat checks after relevant changes, failures, or new uncertainty.
- Delegation must respect host permissions; account for duplicated instructions, worker context, tool use, retries, and coordination in its cost.
### 4. Preserve State Through Continuation
Before supported compaction or handoff, verify this summary against the task, source evidence, and ledger:
> Context · Decisions · State (done / pending / blocked) · Remaining tasks (what, how, where) · Verification · Risks and how to detect them early
Include constraints, exact values and units, citations and evidence references, user decisions and authorization boundaries, session consent or revocation and whether an unanswered offer was already made, unresolved uncertainty, acceptance criteria, and the next action. Preserve stable task IDs if present. Do not replace source evidence with a summary that cannot support the next decision; reopen the original when needed. Never automatically clear history. If a handover cannot retain or recover required detail, keep fuller context and disclose the continuity limit.
### 5. Enforce Strict Preservation
Never weaken requirements, exact values, citations, uncertainty, acceptance criteria, or required verification to save tokens. Keep the main task's quality bar and requested output format. Prefer files when supported; honor explicit user delivery requests, and provide complete in-session output labeled `in-session only; not persisted` when files are unavailable.
Potentially lossy techniques such as learned prompt compression require separate approval of the specific technique, task-specific quality metric, and tolerance before use. General session consent is insufficient. Even an approved experiment cannot silently weaken the main task's acceptance criteria; a failed comparison restores the fuller context or ordinary workflow. Missing information, conflicting evidence, or verification failure also triggers restoration and rechecking of affected decisions.
### 6. Measure Without Inventing Savings
No extra report by default. Never invent token counts or an unrun baseline. Distinguish measured input/output tokens, cached tokens, cost, latency, and estimates. Caching may reduce processing cost without reducing context presented to the model.
For a comparison, use the same task inputs and acceptance criteria, record host/model settings and run conditions, and include loading, summarization, tool discovery, retries, and delegation overhead. Report quality regressions and unavailable metrics. Reserve “validated savings” for actual comparable measurements with passing task checks; a finite test set does not establish universal accuracy. Unmeasured guidance is allowed when clearly labeled. Do not persist learned rules or memory unless separately authorized.
### Completion: Choose the Next Action
- Finish and verify the current deliverable first; make the result and its location available before asking about follow-on work. Keep artifact readiness separate from the next-action choice: an unanswered suggestion does not reopen completed work, and a blocked job is not complete
- Check the result, current project stage, remaining risks, and prior user instructions. Recommend the next useful action from the local routing guidance below; skip irrelevant stages and prefer stopping when no useful work remains
- Use an available interactive question tool to offer one concise next-action choice. Put the best recommendation first, explain why it fits this result, and include a stop/pause choice. Always allow a free-text note or custom direction, including work outside the 360 flow; never force the user into a sibling skill
- If interactive tools are unavailable, offer equivalent numbered choices with the recommendation, rationale, and explicit custom-note option in chat. This completion prompt is separate from the deliverable briefing; skill names are allowed here, and it is not a passive list of unresolved task questions
- Reuse an already explicit next-step instruction instead of asking again; continue work it authorizes. Otherwise wait for the user's choice before starting follow-on work. Silence, a preselected recommendation, and elapsed time are not authorization. An explicit stop or request for no suggestions suppresses the prompt
- When the user chooses, follow that direction and clarify only missing information needed for it. Discover and load a selected skill through the host's supported mechanism; do not assume it is installed or install it automatically. If unavailable, explain and offer an equivalent action. Carry forward artifact paths, decisions, verification, remaining risks, and session consent without restarting intake
Local routing: While accompanying another skill, let the parent own the single completion prompt; do not issue a duplicate or interrupt its work. For a standalone efficiency task, recommend resuming the main task or addressing an evidenced remaining issue, and allow stopping. Never recommend invoking this overlay recursively.
## Output Format
Default: complete the main task in its requested format, with no efficiency report.
When requested, report:
1. Capabilities used and techniques applied
2. Context reused, omitted, or compacted, with recoverable evidence locations
3. Preservation checks, acceptance results, and any observed regressions
4. Restorations or escalations and their reasons
5. Measurements: baseline and overlay, input/output tokens, cached tokens, cost, latency, overhead; use `UNMEASURED` for unavailable metrics and label estimates
6. Limitations and residual uncertainty
Use the six canonical handover fields above when a continuation is needed.
## Quality Gate
- Completion includes the interactive next-action offer with a recommendation, stop choice, and custom-note option, or the explicit-instruction/parent-owned exception; unanswered suggestions do not block the completed deliverable
Check every applicable item; record a concrete limitation if one cannot be checked:
- Session approval is explicit, current, and honored after refusal or revocation
- Techniques use only established, permitted capabilities
- Required inputs, constraints, exact values, citations, uncertainty, and acceptance criteria were preserved against source evidence
- Pagination and truncation were checked wherever completeness mattered
- Reused evidence is sufficiently current for the decision
- The handover retains required state and recoverable evidence; no automatic history clearing occurred
- Required task verification ran, or is explicitly unresolved; failed checks triggered restoration
- Any potentially lossy method had separate technique, metric, and tolerance approval
- Delivery matches the user's request and honestly states persistence
- Measurement claims have real evidence and include available overhead; absent telemetry is `UNMEASURED`
An unresolved check is a limitation or blocker, never proof of equivalent accuracy to a run that did not happen.
### Reference: skills/360-token-efficiency/references/evidence-and-techniques.md
Source: skills/360-token-efficiency/references/evidence-and-techniques.md
# Evidence, limitations, and optional techniques
Read only for a relevant technique or an evidence request. These sources support conditional techniques; they do not establish zero accuracy loss for every task or agent. Consult current host documentation before relying on a platform-specific control.
## Retrieval and compaction
Anthropic describes just-in-time retrieval using lightweight references and progressively loaded evidence. It also warns that aggressive compaction can lose subtle but critical context. Apply this as a reason to preserve exact constraints, evidence locations, uncertainty, and task state, then verify the handover against original inputs. Mandatory instructions and required inputs still need full reads. [Effective context engineering for AI agents](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents).
A summary is useful only when enough original evidence remains recoverable for the next decision. If retrieval is unavailable, retaining fuller context is safer than assuming a summary is sufficient. Check pagination and truncation before concluding that a search covered everything.
## Tool discovery and programmatic processing
Anthropic describes deferred tool discovery and programmatic tool calling, including filtering tool results before they reach model context. These techniques depend on host support; a skill cannot create those capabilities by instruction. Count discovery and execution overhead, and retain source identifiers and completeness checks when filtering. [Introducing advanced tool use](https://www.anthropic.com/engineering/advanced-tool-use).
## Caching
Google documents explicit and implicit context caching, model-dependent eligibility, and cache usage metadata. Cached content still forms part of the model's input context; pricing and reuse differ from removing input. Verify exposed controls and current usage metadata instead of assuming a cache hit. Report cached tokens separately from total input tokens, processing cost, and latency. [Gemini context caching](https://ai.google.dev/gemini-api/docs/caching).
Changing cache settings is optional and remains subject to host permissions and task authorization. Do not change agent configuration for this overlay.
## Learned compression
LLMLingua-2 learns token retention using distilled data and reports evaluations across selected tasks and models. Those benchmarks do not prove losslessness for arbitrary requirements or unseen tasks. Treat learned token removal as potentially lossy: obtain separate approval for the technique, a task-specific metric, and a tolerance; preserve original inputs, compare against the ordinary workflow, and restore on failure. [LLMLingua-2 research](https://arxiv.org/abs/2403.12968).
Do not use an aggregate score to excuse a lost exact value, citation, authorization boundary, or mandatory acceptance criterion.
## Delegation and comparisons
Delegate only within permissions and when independent work justifies the extra context and coordination. Small tasks can cost more with multiple agents. Include parent and worker usage, skill loading, tool calls, summaries, retries, and coordination when telemetry exists.
A useful comparison records:
- Same task inputs, acceptance checks, host/model settings, and tool access
- Baseline and overlay outputs, evidence, acceptance results, and observed regressions
- Observed input/output tokens, cache usage, cost and elapsed time, with unavailable metrics explicitly marked
- All available overhead; no assumed free summarization, discovery, retries, or delegation
- Run count and variation; no extrapolation to universal accuracy or savings
A manual or scripted walkthrough can expose instruction defects. It is not a measured model benchmark. Artifact byte or word counts measure document size, not billed tokens or task-level savings.
---
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
No one has posted yet. Be the first.

