Consult shared Claude Code probe results or test its harness behavior through a real interactive Claude session in tmux, driven by Codex, Grok CLI, or Copilot CLI. Excludes tests of those outer CLIs and application tests.
Generate AGENTS.md files for pertinent repository folders — only where missing. Provides coding agents with build commands, testing instructions, code style, project structure, and boundaries.
Nine conformance checks the vally-tests skill emits for .agent.md artifacts (including consolidated subagent structural template), with contract citations, stimulus shapes, and Vally grader recommendations
Comprehensive Rust coding guidelines with 179 rules across 14 categories. Use when writing, reviewing, or refactoring Rust code. Covers ownership, error handling, async patterns, API design, memory optimization, performance, testing, and common anti-patterns. Invoke with /rust-skills.
Comprehensive Rust coding guidelines with 179 rules across 14 categories. Use when writing, reviewing, or refactoring Rust code. Covers ownership, error handling, async patterns, API design, memory optimization, performance, testing, and common anti-patterns. Invoke with /rust-skills.
skill philosophy differs from Anthropic's published guidance on writing skills. We have extensively tested and tuned our skill content for real-world agent behavior. PRs that restructure, reword
CHILD_ENV_MARKER` valued as the fenced Kanban board root, not a bare flag.
## Tests
Loop/phase tests go in `tests/agent/`; patch the binding the phase actually reads (siblings often
`from
provider, and turn the competing PRs into plugins against it.
- **Behavior contracts over snapshots.** Tests assert how two pieces of data relate, never
freeze a current value (see Testing
root, so a child working against a scratch `HERMES_HOME` keeps a writable board.
## Tests
`tests/cron/`, `tests/hermes_cli/test_kanban*.py`, `tests/tools/test_kanban*.py`. Schedule parsing
and catch-up windows are pure functions — test
exists. A sealed native stream is a
regular message — `chat.update` works (live-verified).
Contract tests: `tests/gateway/test_stream_final_contract.py` (mutation-checked). Slack ground
truth: `chat.*Stream` speaks STANDARD markdown, not mrkdwn; `stopStream.markdown_text
ROOT_KEYS`). For a new
key: `rg -n '" "' hermes_cli/config_defaults.py` AND `rg -n ' ' --glob '!tests' .` both hit,
and one invariant test sets it in a temp `config.yaml` and asserts
process warning, documented replacement + migration note, ≥ 2 subsequent
minor releases before removal.
- Compat tests load **frozen plugins through the real discovery path** and assert outcomes — never
exact registry/catalog counts, source
every call — ship a helper script and reference it by skill-relative path.
7. **Tests at `tests/skills/test_ _skill.py`**, stdlib + pytest + `unittest.mock` only, no
live network. Run `scripts/run_tests.sh tests/skills/test_ _skill.py
spotify, terminal, todo,
tts, video, vision, web, yuanbao` (don't assert the list in tests). Per-platform enable/disable via
`hermes tools` (curses) or `tools. .enabled/disabled` in config.yaml. `browser_exec`
replaces
import; the dispatcher rejects unknown param keys
(`4000` + key path) and, under `HERMES_TEST_ISOLATION=1`, raises `ContractViolation` when a handler's
result or an emitted payload does not match
surface; a new surface is a new `web_routers/ .py`, not a growing
`web_server.py`.
- Tests: Python in `tests/hermes_cli/` (routers, pty bridge); JS in the `web/` vitest suite. Python
tests must