agentleFS
Sign inSign up

skills

scenario-labs/skills/AGENTS.md

Guidance for AI agents (and humans) working in this repository. Public Agent Skills for Scenario: procedural knowledge that teaches AI coding agents how to create production-ready content (images, video, audio, textures, skyboxes, 3D, custom models) through the Scenario MCP server. The workflows serve games, entertainment, and any creative vertical. Skills follow the Agent Skills format, install with npx skills add scenario-labs/skills, and are listed on skills.sh. The expert tools are the one exception to the MCP focus: teams of skills…

AGENTS.md646 starsChanged 5 days ago
  • Installs packages
# AGENTS.md

Guidance for AI agents (and humans) working in this repository.

## What this repo is

Public Agent Skills for [Scenario](https://scenario.com): procedural knowledge that teaches AI coding agents how to create production-ready content (images, video, audio, textures, skyboxes, 3D, custom models) through the [Scenario MCP server](https://mcp.scenario.com). The workflows serve games, entertainment, and any creative vertical. Skills follow the [Agent Skills](https://agentskills.io) format, install with `npx skills add scenario-labs/skills`, and are listed on [skills.sh](https://www.skills.sh/scenario-labs/skills). The expert tools are the one exception to the MCP focus: teams of skills that drive DCC software and game engines installed on the user's machine (see Expert tools).

## Layout

```
skills/<name>/SKILL.md                         # name must equal the directory name
skills/<category>/<family>/README.md           # expert tools: the family's build notes
skills/<category>/<family>/<name>/SKILL.md     # expert tools: dcc/ or game-engines/
tests/<name>/                                  # test suite for any script shipped with skill <name>
```

Every check finds skills through `scripts/lib/skills.mjs`, which is where the layout is defined. The skills CLI installs every skill flat under its name, whatever folder it sits in, and silently keeps only the first of two skills sharing a name, so names are unique repo-wide (`pnpm skill-files` checks it).

Supporting files (heavy references, scripts) may sit next to a SKILL.md only when the content is too large to inline. Link supporting files directly from SKILL.md: agents resolve file references one level deep, so a reference chained through another supporting file may never be read.

A skill may also carry a `README.md` documenting how it was built and why it makes the choices it does. That file is for the agents and humans working on the skill, not for the agent running it, so it is the one supporting file exempt from the link rule (`pnpm skill-files` skips it): linking it would spend body words and invite a runtime agent to read maintainer notes as instructions, which has been observed to make an agent discount a fact it needed.

Every script shipped with a skill has a test suite in `tests/<name>/` at the repo root, never inside the skill directory: published skill directories carry only what an agent needs at runtime. Python scripts are tested with stdlib `unittest`; TypeScript scripts with vitest. See Validation and testing.

## Public content only

This repository is public. Everything in it, including commit messages, PR text, and issue text, must be limited to publicly shareable language:

- Reference only public surfaces: scenario.com, app.scenario.com, help.scenario.com (the Scenario Knowledge Base), mcp.scenario.com and its `/docs`, docs.scenario.com, and the public model catalog.
- Never reference internal repositories, source file paths, internal hostnames or environments, internal project or team names, customer names, real team/project/API identifiers, credentials, pricing internals, or unreleased features.
- Facts must be verifiable from public surfaces (the tool reference at mcp.scenario.com/docs/tools, live public catalog searches). If a fact is only knowable from internal sources, leave it out.
- When in doubt, treat it as internal: ask, or drop it. Example: no internal repository naming.

## Authoring contract

CI enforces the mechanical parts of this contract on every push and PR: [`skills-ref validate`](https://github.com/agentskills/agentskills/tree/main/skills-ref) (the Agent Skills reference validator) for the spec rules, plus house-style greps, formatting (prettier), spell checking (cspell), and supporting-file checks. Run the content checks locally with `pnpm run validate` (commit messages and PR titles are linted separately with commitlint), or spec-validate a single skill with:

```bash
uvx --from "git+https://github.com/agentskills/agentskills.git#subdirectory=skills-ref" skills-ref validate skills/<name>
```

- Published skill frontmatter: `name`, `description`, and `license: MIT`, nothing else. The spec caps `name` at 64 characters and `description` at 1024; the other spec-optional fields (`compatibility`, `metadata`, `allowed-tools`) are not used in this repo.
- `name`: lowercase letters, numbers, and hyphens; must equal the directory name.
- `description`: third person, starts with "Use when", describes triggering conditions only (never a summary of the skill's workflow), under 500 characters, rich in keywords an agent would search for.
- Body: 1000 words is the house target; the build fails past 2500 words (`pnpm style`). Structure: Overview, Quick reference, one excellent worked example, Common mistakes.
- Measure before trimming: `awk '/^---$/{c++; next} c>=2' skills/<name>/SKILL.md | wc -w` is the exact number `pnpm style` checks.
- Why the budget: agents load only `name` and `description` at startup; the body enters context only when the skill triggers, and then every word competes with the user's task. Spend words on facts an agent would otherwise guess wrong, not on prose.
- Where the numbers come from, so they are not re-litigated: nothing outside this repo caps a body by word count. The spec says the body has "no format restrictions" and `skills-ref` validates a 200,000-word body clean, and the widely repeated "5000 words" is a debugging remedy for a skill already behaving badly, not a budget. Anthropic states exactly one size rule, "keep SKILL.md body under 500 lines", and its own "Token budgets" section only restates it. That rule assumes list-heavy and code-heavy skills with short lines, so it does not transfer here: this repo writes dense paragraphs and its largest body spends about 40 lines, which is why no line cap is enforced. The one mechanism that actually bites is Claude Code's post-compaction re-attach, which keeps the first 5000 tokens of each invoked skill from a 25,000-token pool shared across them, filled from the most recently invoked so older skills are dropped entirely once it overflows. Hence a word cap: at this repo's density (about 7 characters per word, so roughly 1.75 tokens per word) 2500 words is near 4400 tokens, inside the per-skill ceiling with margin for identifier-dense prose, and past it a body trades content for truncation. The cap is the guardrail and the target is the craft. What protects the shared pool is not the cap, which binds a handful of skills, but the mean, which is what the fan-out notice watches.
- Body over budget? Trim before you split. The body loads on every trigger; a supporting file loads only when the agent follows its link, so facts every run needs stay in the body, and a linked reference holds only content that is genuinely situational (deep reference tables, per-mode detail, scripts). Splitting to dodge the budget adds a hop to the same context cost: a reference the agent must always read is a longer body in disguise.
- Context fan-out: a skill plus the siblings it names is the closest stand-in for what one session actually loads, and after compaction those share the 25,000-token re-attach pool. `pnpm style` prints a notice (never a failure, since the total moves when a skill nobody touched grows) once a fan-out approaches the pool. Answer it by trimming the body or by dropping a sibling reference the workflow does not need.
- Ground every tool and parameter claim in the [tool reference](https://mcp.scenario.com/docs/tools). Never present a generative model's id as a constant: availability differs per team, and the generation a skill names today is superseded within months.
- Name a model id only for a platform utility Scenario itself provides with no competing alternative, or for the one member a skill is dedicated to running. Where Scenario ships exactly one deterministic tool for an operation and nothing third-party does the same job, the id is a constant and discovering it only re-derives one: the compositors (`model_scenario-compose-video`, `model_scenario-compose-image`), the timeline tools (`model_scenario-video-cut`, `model_scenario-video-split`, `model_scenario-video-concat`, `model_scenario-video-to-image-seq`, `model_scenario-image-seq-to-video`), the exact-dimension resize tools (`model_scenario-resize-video`, `model_scenario-resize-image`), the audio structural tools (`model_scenario-audio-cut`, `model_scenario-audio-split`, `model_scenario-audio-extract`), the small image utilities (`model_scenario-grid-maker`, `model_scenario-image-slicer`, `model_scenario-padding-remover`) and the `model_scenario-postprocessing-*` effects. Say in the body why the id is fixed, so the next author does not read it as an oversight. A skill built around one first-party tool model also names its id outright (scenario-caption-studio runs `model_scenario-caption-studio`): running that member is the skill's whole purpose, so a discovery step would only re-derive the subject, and the contested-lane test below governs skills reaching the capability, never the skill dedicated to the member. A generative model is never named, whoever built it: availability differs per team and the generation named today is superseded within months.
- Being Scenario's own is necessary and not sufficient. `modelProvider: "Scenario"` on a `tool`-tagged model is the marker, and an upstream vendor there is the tell that the operation is a market: of the 64 tool-tagged Scenario-branded models, 61 report Scenario and the three that report BFL are the FLUX upscale variants, upscaling being the one lane among them with competitors. So check the provider first, then confirm nothing third-party does the job, because these lanes are Scenario-built and still contested: reframing and expanding (`model_scenario-smart-reframe`, `model_scenario-gemini-reframe` against Photoroom Expand and Uncrop), splitting an image into layers (`model_scenario-image-layers-extractor` against Ideogram and Seedream), transcription (`model_scenario-audio-to-text` against Gemini Transcribe), subtitles (`model_scenario-caption-studio` and `model_scenario-video-subtitles` are two Scenario tools for one job, so the lane fails the test internally) and segmentation (`model_scenario-detection` is unique for ControlNet preprocessor maps, but SAM covers object masks). Ids also split across `model_scenario-` and `model_sc-`, so confirm one against the catalog rather than reconstructing it.
- Discovery is `recommend` when the skill needs a capability, `search` when the member is already known by name or is private. `recommend` takes the capability plus the user's own words and ranks members on measured cost and latency, naming the purpose-built pick and warning when a generative model is the wrong instrument. `search` ranks by keyword and its `filters` hold no capability key, so a capability-worded query is the failure case: `query="image edit"` returned 318 hits with nothing to choose between them, and a skill that shipped it had its member picked by reading marketing blurbs. Lane tables carry capabilities (`img2video`), not query strings. `dry_run` still prices the exact payload, because `recommend`'s figures are modeled: one quoted 48.2 CU where 55 was billed. Leave the `next_step` discipline to the `scenario` skill and reference it.
- Model-family skills group one provider family per skill and are named without version numbers (`scenario-seedance`, never `scenario-seedance-2-5`). A skill name is a permanent identifier (installs, README rows, `skills.sh.json` groupings, cross-references from sibling skills) while model generations churn every few months: a versioned name goes stale the day the next generation ships and forces either a breaking rename or a pile of near-duplicate skills competing for the same trigger. It is the model-ID rule one level up: versions are data, not identity. Version and generation keywords belong in the `description` (the seedance description carries "Seedance 2.5 and 2.0") so an agent searching a specific version still triggers the skill, and the body teaches per-member differences by reading caps off `model_schema_get` with authoring-time numbers hedged as such. Split a family only when coexisting members target different output types, never by version.
- One skill, one output type: image, video, 3d, or audio. When a brand spans types, split by surface (`scenario-grok-imagine-image` and `scenario-grok-imagine-video`, `scenario-luma-image` and `scenario-luma-video`) and leave the other type's members to the sibling skill. An intermediate asset of another type inside one pipeline (a skybox still on the way to a splat) does not change the skill's type.
- MCP is the surface a skill teaches. Every step the body has an agent perform itself goes through an MCP tool, including the ones reachable another way: do not route the agent through the Scenario REST API, the official SDK, or a CLI where an MCP tool exists. A capability the MCP server does not expose is a gap worth reporting rather than quietly bridging with an API call. Fetching a URL an MCP tool already returned (`asset_download` then `curl -L`) is not an API call and stays fine.
- Scripts are the exception. A script under `skills/<name>/scripts/` may use the official public SDK wherever the SDK is the better tool, whether a maintainer runs it or a skill has an agent run it. What the SDK must not become is the route the body teaches: keep the agent's own steps on MCP, and have the script's link in SKILL.md say what it does and who runs it, so no agent reads it as the workflow. `pnpm style` prints a notice when skill markdown mentions the SDK; it prompts that judgment call and never fails the build.
- Cross-reference the `scenario` skill for connection setup instead of repeating it.
- Every skill that names a sibling skill also carries the sibling-install invitation, byte-identical across skills (copy it from any SKILL.md: "If a sibling skill named here is missing from your available skills, ask the user to install it..."). Neither the Agent Skills spec nor the skills CLI resolves dependencies (issue #49 tracks the upstream work), so this sentence is the mechanism.
- No disclosure marks: skills never instruct an agent to burn AI-disclosure marks, badges, or compliance overlays into generated content, by default or as an option. The skills serve professionals who own their distribution compliance; the creative pipeline does not make that decision for them.
- American spelling (color, behavior, center, modeling, optimize). cspell's en-US dictionary accepts some British forms, so those are listed as forbidden words, prefixed with `!`, at the top of `project-words.txt`; add a form there when one slips through.
- Style: no em dashes, ever (use a comma, a colon, parentheses, or two sentences). No marketing language. Agent-agnostic wording: do not assume a specific agent outside clearly labeled setup snippets.
- Say it once, then stop. This applies to skill bodies, scripts, commit messages, PR text, and replies alike. A comment earns its place only by recording what the code cannot say: a constraint, a trap that already bit someone, a why that is not visible from the lines below it. Do not narrate what the next line does, restate a rule that already lives in this file (link to it instead), or spend a paragraph where a clause serves. The same discipline governs prose: no preamble, no restating the question, no summary of what you just wrote. Noise is not free, it is paid on every future read, and a comment that drifts out of date costs more than the one that was never written. Before cutting a passage as a duplicate, compare the two item by item: topic overlap is not content overlap, and nothing here tests whether a fact that used to be present still is.

## Expert tools

`skills/dcc/` (ZBrush, Blender, Maya) and `skills/game-engines/` (Unreal Engine, Unity) hold teams of skills, ported from Emmanuel de Maistre's expert-skill repositories, that drive applications installed on the user's machine through their Python, bridges, and command lines. Each family is a lead `scenario-<app>-expert` plus specialists named `scenario-<app>-<topic>`; the specialists import the lead's `scripts/`, so a family installs as one set. The family folder's `README.md` records how the family was built and how far it was verified; it sits outside every skill folder, so no install ships it, and it carries no version (this repository's releases do).

The authoring contract applies with these differences, each because the tier drives an application rather than the MCP server:

- The MCP rules (MCP as the surface, `recommend` discovery, model ids) do not apply; the rest of the contract does, including the frontmatter, the 2500-word cap, house style, public content only, and the sibling-install sentence.
- Descriptions follow the spec's 1024-character cap rather than the 500-character target, and bodies run near the cap rather than the 1000-word target: both are the author's choice, kept until a live run shows what to trim.
- Shipped scripts need no suite of their own under `tests/<name>/`: the behavior suites live in the author's build project and run against the application. One suite per family, `tests/scenario-<app>-expert/` (the same file in each), imports every script with system Python and checks that each skill name a script uses to find the lead's `scripts/` exists; `pnpm skill-files` still parses every Python script.
- A link to a subfolder (`scripts/AgentKit/`, never `scripts/` itself) covers the files under it, because these skills ship code trees an agent copies into a project whole.
- cspell skips code, fenced and inline, and YouTube video ids; prose is checked as everywhere else (`cspell.json` overrides).
- Each family is one `skills.sh.json` grouping titled "Expert tools: <app>", and every expert grouping comes after every core grouping, so the expert tools stay at the bottom of skills.sh and the installer picker (`pnpm groupings` enforces both). A new core grouping goes above them, which renumbers their plugin ids.
- The application test protocol grades MCP plans; for an expert tool, grade the plan against the lead skill's execution channels and the application's own documentation instead.

## Authoring aids

Two of Anthropic's skills are vendored as dev skills in `.agents/skills/`, symlinked from `.claude/skills/`, so agents working in a clone of this repo pick them up automatically: [skill-creator](https://www.skills.sh/anthropics/skills/skill-creator) (Apache-2.0, from anthropics/skills) and [skill-development](https://github.com/anthropics/claude-code/tree/main/plugins/plugin-dev/skills/skill-development) (MIT per the plugin-dev README, from anthropics/claude-code). `skills-lock.json` records each source and hash; refresh with `npx skills update`. Vendored dev skills live in agent directories and are excluded from CLI discovery by their entries in `skills-lock.json`. Agent directories alone do not prevent discovery: the CLI also scans `.agents/skills/` and `.claude/skills/`. Where their generic guidance and this file disagree, this file wins.

Repository commands live in regular `.agents/skills/skills-*/SKILL.md` files. Codex exposes `$skills-pr-summary`, `$skills-squash-message`, `$skills-pr-handle`, and `$skills-validate`; the [contributor guide](CONTRIBUTING.md#shared-agent-commands) maps their Claude names. Claude command paths under `.claude/commands/` are relative symlinks to those files. Edit the canonical files; run `pnpm sync:agent-commands` to create or repair links (also run by `pnpm format`). Explicit-only commands carry `disable-model-invocation: true` and `argument-hint` in shared frontmatter for Claude, plus `allow_implicit_invocation: false` in `agents/openai.yaml` for Codex. Codex CLI 0.154.0 was verified to load both commands with those Claude fields present. These maintainer skills retain their command structure and stay outside the published `skills/` catalog. Set `metadata.internal: true` (a YAML boolean) in each command's frontmatter to exclude it from normal CLI discovery and bulk installs; explicit skill requests or `INSTALL_INTERNAL_SKILLS=1` can still include it. Existing skills.sh listings are not automatically removed by this flag: request removal from the maintainers (see [vercel-labs/skills#1578](https://github.com/vercel-labs/skills/issues/1578)). Add new command mappings in `scripts/sync-agent-commands.mjs`. `pnpm validate` checks command metadata, argument hints, invocation guards, and symlinks; strict spec validation covers only published `skills/*/` because the spec rejects Claude frontmatter extensions. `pnpm test` runs the command regression suite as well as shipped-script suites.

## Repo tooling

One-time setup after cloning: `pnpm install`. It installs commitlint, cspell, prettier, and the husky git hooks. A Claude Code SessionStart hook (`.claude/hooks/ensure-husky.sh`) runs it automatically when the hooks are missing. In Codex sessions, run `pnpm install --frozen-lockfile` if `node_modules` is missing or `git config --local core.hooksPath` is not `.husky/_`; the Claude Code hook does not run there.

- Commit messages and PR titles follow Conventional Commits, enforced by commitlint (`commitlint.config.js`) in three places: the husky `commit-msg` hook, a commitlint job on PR commits, and the `pr-name-linter` workflow on the PR title. Valid scopes are the skill directory names (derived automatically from `skills/`) plus `skills`, `agents`, `ci`, `deps`, `docs`, and `tooling`.
- The husky `pre-commit` hook runs `pnpm words:sort` (sorts `project-words.txt` in place, re-staging it only when it is part of the commit), `pnpm manifest` (regenerates `.claude-plugin/marketplace.json` from `skills.sh.json`, re-staging it only when it is part of the commit), then `pnpm validate`, the same checks CI runs as separate steps: `pnpm style` (house style, the body budget, and the context fan-out notice), `pnpm format:check` (prettier), `pnpm skill-files` (supporting files next to a SKILL.md are linked and runnable, a skill's own README.md excepted), `pnpm groupings` (every skill sits in a `skills.sh.json` grouping, every listed skill exists, and the file satisfies the published skills.sh schema), `pnpm manifest:check` (the plugin marketplace manifest matches `skills.sh.json`), `pnpm readme` (the README Skills section mirrors `skills.sh.json`: one subsection per grouping in the same order with the same title, description, and rows, and every skill has exactly one non-stale row, plus the install commands described below), `pnpm words:check` (`project-words.txt` is sorted and deduplicated), `pnpm spell` (cspell), and `pnpm spec` (spec validation). Each is defined once in `package.json`; most run a script in `scripts/`, while `format` and `format:check` invoke prettier directly.
- `pnpm groupings` proves `skills.sh.json` is valid, not that skills.sh has applied it. Per the [customize docs](https://www.skills.sh/docs/customize), the directory reads the file only after its telemetry has seen the repository (in practice, after a fresh `npx skills add scenario-labs/skills` run with telemetry enabled), and repo pages are cached, so the skill content on [the repo page](https://www.skills.sh/scenario-labs/skills) trails `main` by hours. Our groups were confirmed visible on September 22, 2026; the earlier failure is recorded at [vercel-labs/skills#2136](https://github.com/vercel-labs/skills/issues/2136). Ungrouped listings appear under "Other skills"; `notGrouped` controls their position, not their visibility.
- `.claude-plugin/marketplace.json` is generated from `skills.sh.json` by `pnpm manifest` (`scripts/generate-plugin-manifest.mjs`); the groupings file is the single source of truth, so never edit the manifest by hand. The manifest is what gives the `npx skills add` picker its group headers: the skills CLI groups the picker by the plugin entries in this file and never reads `skills.sh.json`, which only drives the repo page on skills.sh. Plugin names carry a two-digit position prefix (`01.-getting-started`, which the picker displays as "01. Getting Started") because the picker sorts group names alphabetically; the prefix makes it follow the `skills.sh.json` order, and inserting or merging a grouping renumbers the plugin identifiers on regeneration. It also makes the repo installable as a Claude Code plugin marketplace, one plugin per grouping. `pnpm manifest:check` fails validation and CI when the two files drift; the pre-commit hook regenerates the manifest automatically.
- `npx skills add scenario-labs/skills` opens a picker with nothing preselected, so a reader who runs the bare command and hits enter installs nothing: the README says so on the command. It no longer offers `--skill "*"`: that takes every skill, the expert tools included, and an agent loads every installed skill's description at startup, so a reader who never opens a DCC tool would carry 57 of them. The Install section leads instead with a default command that writes out every core skill, then one command per goal: the goal's lead skills plus the siblings they name in backticks (one level, not transitively, which would pull in nearly everything). `pnpm readme` checks that every install command names real skills, that the default lists exactly the core skills, and that each family README's command lists exactly its family; after adding a skill, add it to any goal it leads or serves. `INSTALL.md` holds the same kind of commands by role (a base of `scenario` plus the troubleshooting skills, the role's lead skills, and the siblings they name), and `pnpm readme` checks that its commands name real skills. A skills.sh pack URL (`npx skills add https://skills.sh/p/<id>`) is the only install that preselects a set. Packs are created manually on skills.sh behind a Vercel sign-in (no CLI or API); pack links belong in README.md's Install section.
- Known bug: committing from a git worktree fails the pre-commit hook at `pnpm spec`. Git exports `GIT_DIR` to hooks, uv's git subprocesses inherit it, and spec validation resolves the agentskills repo to this repo's own HEAD, erroring with "has no subdirectory `skills-ref`" (the upstream repo is fine). Until `GIT_DIR`, `GIT_WORK_TREE`, and `GIT_INDEX_FILE` are unset at the top of `.husky/pre-commit`, run `pnpm validate` from the worktree in a shell without those variables, then commit with `--no-verify`.
- Formatting is enforced by prettier with the config committed in `.prettierrc.json`, so the CLI and the VS Code extension agree. `pnpm format` rewrites the repo; `pnpm format:check` only verifies. Coverage is everything prettier can parse (markdown, JSON, YAML, JS/TS) plus shell scripts via `prettier-plugin-sh`, which formats with mvdan-sh, the engine behind shfmt. Generated files and the vendored dev skills are excluded in `.prettierignore`.
- After authoring markdown run `pnpm format` before `pnpm validate`: prettier reflows prose and rewrites tables, so `format:check` rejects correct hand-written markdown.
- Spelling: add legitimate project terms to `project-words.txt`; never disable cspell inline. The file stays sorted (byte order, deduplicated): the pre-commit hook sorts it automatically, and `pnpm words:sort` does it manually.
- PRs are squash-merged, so the PR title becomes the commit header on `main`. The `/squash-message` command drafts that message and `/pr-summary` refreshes the PR body (both in `.claude/commands/`). `/skills:validate <name>` (`.claude/commands/skills/`) runs the application test for one skill and reports it to the PR; see Validation and testing.
- Releases are cut by release-please (`.github/workflows/release-please.yml`, configured in `release-please-config.json`): feat, fix, docs, perf, and revert commits squash-merged to `main` land in the next release, which release-please stages as an open release PR. Every other Conventional Commits type is invisible to the changelog, chore and ci by design but also refactor, test, style, and build, and a release window holding only those cuts no release at all, so a change to shipped skill content must be typed feat, fix, or docs. release-please also never lets you pick the version; to force a release at an exact one, dispatch the workflow by hand (`gh workflow run release-please.yml --ref main -f release_as=X.Y.Z`, which must be higher than `.release-please-manifest.json`), never an empty `Release-As` commit pushed to `main`. The dispatch runs the release-please CLI (pinned in `package.json` to the version the action bundles), not the action: release-please-action v5 silently drops `release-as` in manifest mode, which this repo's `release-please-config.json` selects. The override is one-shot, the next push to `main` re-derives the version and rewrites the release PR, so merge it promptly; a pin that survives pushes is `"release-as"` in `release-please-config.json` through a PR, removed once released. Merging the release PR tags `skills-vX.Y.Z` and titles the GitHub release `skills: vX.Y.Z`, rewrites `CHANGELOG.md`, `version.txt`, and `.release-please-manifest.json`, and publishes the GitHub release; `.github/workflows/publish-changelog.yml` then posts the notes to the public changelog at [scenario.com/changelog](https://www.scenario.com/changelog). That publish hangs off the `release: published` event rather than the release-please run, because release-please reports a created release exactly once: a failed post is re-run from the Actions tab, or dispatched by tag. Those three files are release-please's alone, never hand-edited (same rule as the marketplace manifest), which is why `CHANGELOG.md` sits in `.prettierignore` (release-please emits `*` bullets that prettier would rewrite) and, as a guard should the spell-check globs ever widen, in cspell's ignorePaths, though `scripts/check-spelling.sh` does not currently reach it. In `.release-please-manifest.json`, exactly `0.0.0` is the sentinel for "no release yet" and is what makes `initial-version` and the root `bootstrap-sha` apply; editing it to match `initial-version` looks like a sync fix and silently changes the first release number. The release PR title (`chore: release skills X.Y.Z`) passes commitlint because an omitted scope only warns. The `skills-` tag prefix matches the convention the sibling repositories already follow, and is set by hand here as `component` plus `include-component-in-tag` in `release-please-config.json`: a `release-type: node` package gets its component free from the `package.json` name, while `release-type: simple` has none to derive one from, which is why releases through 0.37.3 tagged bare `vX.Y.Z`. Changing the tag format orphans release-please from its own history unless the existing tag and its GitHub release move with it, and the next changelog then replays every commit since `bootstrap-sha`.
- The `skill-files-reviewer` agent (`.claude/agents/`) reviews supporting files added or changed next to a SKILL.md: justified, linked from SKILL.md, runnable, MCP for the agent's own steps, public content only, house style. The `skill-tester` agent in the same directory is the clean-room runner `/skills:validate` spawns when a separate process cannot reach the MCP server; it is not for ordinary work.

## Validation and testing

- `skills-ref validate` must pass for every published skill before any commit (the pre-commit hook and CI both run it via `pnpm spec`).
- Every script shipped with a skill has a test suite in `tests/<name>/`. Python scripts use stdlib `unittest` (files named `test_*.py`); TypeScript scripts use vitest (files named `*.test.ts`). `pnpm test` runs every suite and fails when a shipped script has no suite; CI runs it on every push and PR. It is not part of the pre-commit hook (suites may need system tools such as ffmpeg), so run it manually when touching a script. A `tests/<name>/requirements.txt` declares extra Python dependencies for CI.
- Before merging a new or changed skill, run the application test below. Mechanical validation checks the format; the application test checks whether the skill actually teaches.
- Catalog-only MCP tools (`collection_*`, `asset_quality_gate_run`) take their arguments under `parameters`, not `arguments`, in `scenario_tool_execute_read/write`. The wrong key fails with a misleading "team_id and project_id are required".
- `usage` accepts inclusive date bounds and defaults to the last 31 days when both dates are omitted. Set explicit dates and scope for attribution. Overall consumed CU is `totals.totalCU` (sum of `consumption.value`); model CU is generation activity and can differ. Both are already after discounts. Never substitute job billing snapshots or a sum of your own calls for consumed CU.
- Verify a delivered video by sweeping it into contact sheets (`ffmpeg -vf "fps=2"`), not one frame per shot: a defect that fades in mid-shot passes a spot check. One such check read clean at 10.5s on a clip whose artifact appears at 11.0s.
- Pre-register what each outcome would mean before a paid run, so conclusions are not fitted to results afterwards.
- Stop hunting a mechanism after roughly two inconclusive controlled runs: ship the rule that holds regardless and record the open question.
- `/skills:validate <name>` (`.claude/commands/skills/`) drives that test: it writes a use case for the skill, installs the working-tree copy into a clean-room agent that has no repository context, runs it end to end against the real MCP server in a team and project you choose, grades the transcript, and posts the report to the PR for the current branch (or asks what to do with it when there is none). `--plan-only` runs the zero-cost planning variant below instead of live generation.

### Application test protocol

1. Spawn a fresh agent (no conversation history). Give it only: a framing line ("you are an agent connected to the Scenario MCP server; the skill document below is installed"), the SKILL.md under test (plus the `scenario` SKILL.md when testing any other skill, since real installs ship both), and one realistic task.
2. Ask for a numbered tool-call plan with exact tool names and argument shapes. Planning only: the agent must not execute tools, browse, or consult anything beyond the provided documents, and must flag uncertainty instead of guessing.
3. Pick a task that forces the skill's non-obvious facts (upload flow, job-wait re-calls, dry runs, launch semantics), not one answerable with generic MCP intuition.
4. Grade the plan against the [tool reference](https://mcp.scenario.com/docs/tools), fetched fresh rather than recalled:
   - Every tool and parameter named in the plan exists. One invented name is a fail.
   - Correct flow: the loop the skill under test teaches. For a generation task: discovery, `model_schema_get`, `model_run`, then `jobs_wait` re-called with `pending_job_ids` for any job still running (fast models return complete inline; `job_get` polling is never correct), then `asset_display` / `asset_download`.
   - Generative model ids come from a `recommend` or `search` step, never asserted as constants. A named first-party tool model is not a violation when it is a sanctioned singleton or the member the skill under test is dedicated to; a named generative model always is.
   - The agent's own steps stay on MCP. A plan that reaches for the REST API, the SDK, or a CLI where an MCP tool exists is a fail, and so is one that invents an API call to cover a capability MCP lacks. Running a script the skill ships is not, whatever that script uses internally, and neither is saving a URL an MCP tool returned.
   - The task's trap steps are handled the way the skill teaches.
   - Anything asserted that appears in neither the SKILL.md nor the tool reference counts as a guess, even when it happens to be right.
5. A failure is a defect in the skill text: fix the missing or ambiguous sentence, then re-run with a new fresh agent (a failed agent is contaminated by its own mistake).
6. Baseline probe, once per new skill (not per edit): run the same task with no skill installed to confirm the skill earns its context cost.

## Codex commit attribution

For Codex-assisted commits and prepared squash messages, include a co-author
trailer naming the model and, when verified, the reasoning effort and speed tier:
`Co-authored-by: Codex <model> <effort> <tier> <noreply@openai.com>`.
For example: `Co-authored-by: Codex gpt-6-astra low fast <noreply@openai.com>`.
Use the actual settings for the contributing session, not repository defaults
or the example above. Omit unknown fields rather than guessing; if the model
is unavailable, use `Co-authored-by: Codex <noreply@openai.com>`. Preserve the
human author and existing contributor trailers. Carry these trailers into the
final squash message so attribution survives the repository's squash workflow.

## Conventions

- Conventional Commits (`feat:`, `fix:`, `docs:`, `chore:`), enforced by commitlint (see Repo tooling).
- PRs target `main` and are squash-merged; the PR title is the future commit header.
- `CLAUDE.md` is a symlink to this file, so Claude Code and Codex read the same conventions. Shared Codex defaults live in `.codex/config.toml`; model selection stays in user configuration.

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.