pytest-skill-engineering
sbroenne/pytest-skill-engineering/.github/copilot-instructions.md
NO LEGACY CODE. NO FALLBACKS. CLEAN CODE ONLY. JSON files are TEST OUTPUT. Never hand-edit them. Always say "system prompt" when referring to agent instructions. Never abbreviate to just "prompt". Custom agent ≠ subagent. Never use these interchangeably: - loadcustomagent() loads a custom agent definition — it does NOT create a subagent - custom_agents= on CopilotEval registers custom agents — Copilot may dispatch to them as subagents at runtime - Write "custom agent dispatch" (not "subagent dispatch") when describing what…
- Reads credentials
- Installs packages
# Copilot Instructions for pytest-skill-engineering
## CRITICAL: No Backward Compatibility
**NO LEGACY CODE. NO FALLBACKS. CLEAN CODE ONLY.**
- Never add backward compatibility code
- Never add fallback logic for "old" formats
- Never synthesize data that should be explicit
- If something is missing, it's an error - not a fallback opportunity
- Remove legacy code, don't maintain it
## CRITICAL: Never Edit JSON Test Data
**JSON files are TEST OUTPUT. Never hand-edit them.**
- Fixture JSONs (`tests/fixtures/reports/*.json`) are generated by running `tests/fixtures/scenario_*.py`
- Results JSONs (`aitest-reports/*.json`) are generated by running integration/showcase tests
- Hero report JSON is generated by running `tests/showcase/test_hero.py`
- If a JSON file has the wrong format, **re-run the test** that generates it
- If a field is missing or empty, the code that generates it needs fixing — not the JSON
- Manually patching JSON masks bugs instead of fixing them
## CRITICAL: Terminology
- **System Prompt** = Instructions given to the agent (what configures agent behavior)
- **Prompt** = What you tell the agent to execute (the test query / user message)
- **Custom Agent** = A `.agent.md` file that defines a specialist agent (name, description, instructions, tools). It is a **definition**, not a runtime concept.
- **Subagent Dispatch** = The runtime mechanism where Copilot's orchestrator routes a task to a custom agent. A custom agent becomes a subagent *when dispatched to* — but it is NOT inherently a "subagent."
- **Eval** = The test harness/configuration (`CopilotEval`). Not the thing being tested.
- **Coding Agent** = The real GitHub Copilot agent that runs user sessions.
Always say "system prompt" when referring to agent instructions. Never abbreviate to just "prompt".
**Custom agent ≠ subagent.** Never use these interchangeably:
- `load_custom_agent()` loads a custom agent definition — it does NOT create a subagent
- `custom_agents=` on `CopilotEval` registers custom agents — Copilot *may* dispatch to them as subagents at runtime
- Write "custom agent dispatch" (not "subagent dispatch") when describing what CopilotEval tests
- Only use "subagent" when specifically describing the runtime event (`result.subagent_invocations`)
## CRITICAL: Use uv, Not pip
**This project uses `uv` exclusively. Never use pip.**
- Install packages: `uv add <package>` (not `pip install`)
- Run commands: `uv run <command>` (not direct invocation)
- Install editable: `uv pip install -e .` (not `pip install -e`)
- Sync deps: `uv sync` (not `pip install -r requirements.txt`)
In documentation, always show `uv add` instead of `pip install`.
## CRITICAL: Never Force Commit
**Never use `git commit --no-verify`. Pre-commit hooks exist for a reason.**
- If a hook fails, **fix the issue** and commit normally
- `ruff format` reformats files → re-stage and retry
- `pyright` reports errors → fix the code
- If a hook is genuinely broken (e.g., false positive with 0 errors), fix the hook config — not bypass it
- Force-committing creates a habit of ignoring quality gates
## CRITICAL: What We Test
**We do NOT test agents. We USE agents to test:**
- **MCP Servers** — Can an LLM understand and use these tools?
- **CLI Tools** — Can an LLM operate this command-line interface?
- **Eval Skills** — Does this domain knowledge improve performance?
- **Custom Agents** — Do these `.agent.md` instructions produce the right behavior and custom agent dispatch?
**System prompts configure behavior.** A custom agent's body IS its system prompt. Use `CopilotEval.instructions` when comparing raw system prompt variants; there is no public `Eval` harness or `system_prompt=` parameter.
**The Eval is the test harness**, not the thing being tested. It bundles an LLM provider with the tools/prompts/skills/custom agents you want to evaluate.
## Why This Project Exists
Your MCP server passes all unit tests. Then an LLM tries to use it and:
- Picks the wrong tool
- Passes garbage parameters
- Can't recover from errors
- Ignores your system prompt instructions
**Why?** Because you tested the code, not the AI interface.
For LLMs, your API isn't functions and types — it's **tool descriptions, system prompts, skills, and schemas**. These are what the LLM actually sees. Traditional tests can't validate them.
**The key insight: your test is a prompt.** You write what a user would say ("What's my checking balance?"), and the LLM figures out how to use your tools. If it can't, your AI interface needs work.
## CRITICAL: HTML Report Development Workflow
**For report rendering changes, regenerate from existing JSON without live LLM calls:**
1. **REGENERATE ALL REPORTS** (non-negotiable)
```bash
uv run python scripts/generate_fixture_html.py
```
- Generates fixture reports in docs/reports/
- Do NOT skip this step
2. **RUN COPILOT INTEGRATION TESTS WHEN EXECUTION BEHAVIOR CHANGES**
```bash
uv run python -m pytest tests/integration/copilot/test_01_basic.py -q
```
- Run relevant test files sequentially when eval, tool, or engine behavior changes
- Do not re-run live tests just for templates, CSS, JS, or documentation edits
3. **VERIFY CHANGES IN ACTUAL HTML** (non-negotiable)
```bash
Select-String -Path "docs\reports\01_single_agent.html" -Pattern "YOUR_SEARCH_TERM"
```
- Confirm change is present in generated HTML
- Open in browser when feasible
**WHAT NOT TO DO:**
- ❌ Skip report regeneration
- ❌ Modify fixture JSONs directly - regenerate via pytest
- ❌ Commit without verifying in browser
## Technology Stack & Code Style
### Language & Runtime
- **Python 3.11+** - Use modern syntax (match statements, `|` union types, Self)
- **Fully async** - All eval execution is async (`async def`, `await`)
- **Type hints everywhere** - All public APIs have type annotations
- **`from __future__ import annotations`** - Always at top of every module
### Core Dependencies
| Package | Purpose | Pattern |
|---------|---------|---------|
| `github-copilot-sdk` | LLM abstraction | Eval execution, coding agent testing |
| `mcp` | MCP protocol | Server process management, tool discovery |
| `pydantic` | Validation | Config validation (used sparingly) |
| `pytest` | Test framework | Plugin system, fixtures, markers |
| `htpy` | Components | HTML report generation (Python-native) |
### Data Modeling
- **Use `@dataclass(slots=True)`** for all data objects
- **Use `frozen=True`** for immutable config
- **Use `TypedDict`** for template data contracts (see `components/types.py`)
- **No Pydantic models** for core types - just dataclasses
```python
# Good - immutable config with slots
@dataclass(slots=True, frozen=True)
class Wait:
strategy: WaitStrategy
timeout_ms: int = 30000
# Good - mutable data with slots
@dataclass(slots=True)
class EvalResult:
turns: list[Turn]
success: bool
```
### Async Patterns
- **pytest-asyncio auto mode** - Tests are async by default (no `@pytest.mark.asyncio`)
- **Use `asyncio.TaskGroup`** for parallel operations (Python 3.11+)
- **Context managers** - Use `async with` for server lifecycle
```python
# Tests are async by default (asyncio_mode = "auto" in pyproject.toml)
async def test_create_file(copilot_eval, tmp_path):
agent = CopilotEval(
name="coder",
model="gpt-5.6-luna",
instructions="You are a Python developer.",
working_directory=str(tmp_path),
)
result = await copilot_eval(agent, "Create hello.py")
assert result.success
```
### Build & Quality Tools
| Tool | Config | Purpose |
|------|--------|---------|
| `uv` | `pyproject.toml` | Package management, virtual envs |
| `hatch` | `pyproject.toml` | Build backend |
| `ruff` | `pyproject.toml` | Linting + formatting (replaces black, isort, flake8) |
| `pyright` | `pyproject.toml` | Type checking (basic mode) |
| `pre-commit` | `.pre-commit-config.yaml` | Git hooks |
### Commands
```bash
# Run lints
uv run ruff check src tests
# Fix lint issues
uv run ruff check --fix src tests
# Format code
uv run ruff format src tests
# Type check
uv run pyright src
# Run copilot integration tests (ALWAYS use these, not unit tests)
uv run python -m pytest tests/integration/copilot/ -v
# Re-run only failures
uv run python -m pytest --lf tests/integration/copilot/ -v
```
### Import Conventions
- Group imports: stdlib → third-party → local
- Use `TYPE_CHECKING` block for type-only imports
- Import from package root when possible (`from pytest_skill_engineering.copilot import CopilotEval`)
```python
from __future__ import annotations
import asyncio
from typing import TYPE_CHECKING, Any
from github_copilot_sdk import CopilotSDK
from mcp import ClientSession
from pytest_skill_engineering.copilot.result import CopilotResult
if TYPE_CHECKING:
from pytest_skill_engineering.copilot.eval import CopilotEval
```
## What We're Building
**pytest-skill-engineering** is a pytest plugin for testing MCP servers and CLIs with real coding agents. You write tests as natural language prompts, and GitHub Copilot executes them against your tools. Reports tell you **what to fix**, not just **what failed**.
### Core Features
1. **Base Testing**: Define evals, run tests against MCP/CLI tool servers
- Eval = Model + Instructions + MCP/CLI Servers + optional Skill + optional Custom Agents
- Use `copilot_eval` fixture to execute eval and verify tool usage
- Assert on `result.success`, `result.tool_was_called("tool_name")`, `result.final_response`
2. **Eval Leaderboard**: When you test multiple evals, the report shows a leaderboard
- 1 eval → Just results
- Multiple evals → Eval Leaderboard (always)
- AI detects what varies (Model, Skill, Custom Agent, Server) to focus its analysis
3. **Winning Criteria**: Highest pass rate → Lowest cost (tiebreaker)
- Use `--aitest-min-pass-rate=N` to fail the session if overall pass rate falls below N%
4. **Multi-Turn Context**: Test compound instructions that span multiple tool calls in one turn
- Each test is independent — context is provided in the prompt
- Inspect call order with `[call.name for call in result.all_tool_calls]`
5. **Skill Testing**: Validate eval domain knowledge
- Load skills from markdown files with `Skill.from_path()`
- Skills inject structured knowledge into eval context
- Reports analyze skill effectiveness and suggest improvements
6. **Custom Agent Testing**: Test `.agent.md` custom agent files (VS Code / Claude Code format)
- `load_custom_agent(path)` + `CopilotEval(custom_agents=[...])` — test real custom agent dispatch through Copilot
- Tests whether custom agents are invoked correctly and produce the expected behavior
7. **Semantic Assertions**: Check response behavior with `llm_assert`
- For example, assert that the response answers the request rather than asking an unnecessary question
- There is no automatic clarification-detection configuration on `CopilotEval`
### AI Analysis (KEY DIFFERENTIATOR)
Reports are **insights-first**, not metrics-first. AI analysis is **mandatory** when generating reports.
Reports include:
- **🎯 Recommendation**: Deploy recommendation with cost/performance analysis
- **❌ Failure Analysis**: Root cause + suggested fix for each failure
- **🔧 MCP Tool Feedback**: Improve tool descriptions, with copy button
- **📝 Prompt Feedback**: System prompt improvements
- **📚 Skill Feedback**: Skill restructuring suggestions
- **⚡ Optimizations**: Reduce turns/tokens
```bash
# Run tests with AI analysis (mandatory --aitest-summary-model)
uv run python -m pytest tests/ --aitest-html=report.html --aitest-summary-model=copilot/gpt-5.6-sol
# Regenerate report with new AI insights from existing JSON (no re-run)
uv run pytest-skill-engineering-report results.json --html report.html --summary --summary-model copilot/gpt-5.6-sol
```
### Key Types
```python
from pytest_skill_engineering.copilot import CopilotEval
from pytest_skill_engineering import load_custom_agent
# Define an eval
agent = CopilotEval(
name="financial-assistant",
model="gpt-5.6-luna",
instructions="You are a helpful financial assistant.",
skill_directories=["skills/financial-advisor"], # Optional domain knowledge
max_turns=10,
)
# Register a custom agent definition for Copilot to dispatch to at runtime
copilot_agent = CopilotEval(
name="coder",
model="gpt-5.6-luna",
instructions="You are a coding assistant.",
custom_agents=[load_custom_agent("skills/my-skill/agent.md")],
)
# Run test
result = await copilot_eval(agent, "Do something with tools")
assert result.success
assert result.tool_was_called("my_tool")
```
### Instructions
Instructions can be plain strings or loaded from `.md` files.
```python
# Use with pytest parametrize
INSTRUCTIONS = {
"concise": "Be brief and direct.",
"detailed": "Explain your reasoning step by step.",
}
@pytest.mark.parametrize("instruction_name,instructions", INSTRUCTIONS.items())
async def test_with_instructions(copilot_eval, instruction_name, instructions):
agent = CopilotEval(
name=f"assistant-{instruction_name}",
model="gpt-5.6-luna",
instructions=instructions,
)
result = await copilot_eval(agent, "What's my balance?")
assert result.success
```
## CRITICAL: Testing Philosophy
**Unit tests with mocks are WORTHLESS for this project. Do NOT run them.**
This is a testing framework that validates AI interfaces (MCP tool descriptions, system prompts, skills, custom agents). A mock that pretends to be an LLM proves nothing — you're testing the mock, not the AI interface.
**The only valid test is a real LLM call against real tools.**
### When you make a code change — run integration tests immediately:
```bash
# Run ALL copilot integration tests — do this after every change
uv run python -m pytest tests/integration/copilot/ -v
# Run a specific test file
uv run python -m pytest tests/integration/copilot/test_01_basic.py -v
# Re-run only failed tests after a full run
uv run python -m pytest --lf tests/integration/copilot/ -v
```
**Always use `uv run python -m pytest`** — bare `pytest` won't find the installed package.
**Fast test execution (< 1 second) is a RED FLAG.** Real LLM calls take time.
### What NOT to do:
- ❌ **NEVER run `tests/unit/` to validate a feature** — unit tests do not validate AI behavior
- ❌ **NEVER declare a feature "done" based on unit tests passing**
- ❌ **NEVER say "all tests pass" when you only ran unit tests**
- ❌ Do NOT write unit tests with mocked LLM responses
- ❌ Do NOT use `unittest.mock.patch` on agent execution
### What TO do:
- ✅ Run `tests/integration/copilot/` after EVERY code change — one file at a time, sequentially
- ✅ Start with `test_01_basic.py`, fix failures, then `test_02_models.py`, etc.
- ✅ Write integration tests that call real GitHub Copilot models
- ✅ Use the cheapest model (`gpt-5.6-luna`) via Copilot
- ✅ Test with Banking or Todo MCP server (built-in test harnesses)
- ✅ Accept that integration tests take 5–30+ seconds per test
- ✅ Run integration tests BEFORE declaring a feature complete
## CRITICAL: No Such Thing as a Pre-Existing Failure
**Every test failure is YOUR responsibility to fix. There are no "pre-existing" or "unrelated" failures.**
- NEVER skip a failing test because it was failing before your change
- NEVER label failures as "pre-existing" and move on
- NEVER say "this failure is unrelated to my change"
- If a test fails, fix it — full stop
- If a test is wrong (not the code), fix the test
- The baseline must be **zero failures** before and after your change
## CRITICAL: Efficient Test Execution
**Integration tests are EXPENSIVE. Never re-run passing tests unnecessarily.**
### pytest caching commands:
```bash
# Run ONLY tests that failed last time (MOST COMMON)
uv run python -m pytest --lf tests/integration/copilot/
# Run failed tests first, then the rest
uv run python -m pytest --ff tests/integration/copilot/
# Check what's in the cache (see last failures)
uv run python -m pytest --cache-show
# Clear the cache (fresh start)
uv run python -m pytest --cache-clear
# Run specific failing test(s) only
uv run python -m pytest tests/integration/copilot/test_01_basic.py -v
```
### Rules for the AI assistant:
1. **After making changes, ALWAYS run copilot integration tests** — `tests/integration/copilot/`, never `tests/unit/`
2. **Run test files ONE AT A TIME, sequentially** — start with `test_01_basic.py`, fix all failures, then move to `test_02_models.py`, etc.
3. **After fixing a test, run ONLY that specific test** to confirm
4. **Use `--lf` to re-run only failed tests** after a full run
5. **Fix ALL failures** — no exceptions, no "pre-existing" excuses
6. **Quote the specific test paths when running individual tests**
7. **Do NOT say "tests pass" without running `tests/integration/copilot/`**
8. **NEVER run `tests/unit/` as validation** — unit tests prove nothing in this project
## Copilot SDK
Uses GitHub Copilot SDK for all LLM calls. Authenticate with `gh auth login --hostname github.com` or an explicit token. Eval and judge sessions select `GITHUB_TOKEN` first, then `GH_TOKEN`; otherwise the SDK uses its signed-in user.
```bash
# Use Copilot for AI insights
uv run python -m pytest tests/ --aitest-summary-model=gpt-5.6-luna
# Use Copilot for llm_assert / llm_score
uv run python -m pytest tests/ --llm-model=gpt-5.6-luna
```
The Copilot SDK is REQUIRED (not optional). All eval execution goes through it.
Available models are dynamic (whatever Copilot exposes).
## Project Structure
```
src/pytest_skill_engineering/
├── core/ # Core types
│ ├── skill.py # Skill, load from markdown
│ └── errors.py # AITestError, ServerStartError, etc.
├── copilot/ # GitHub Copilot SDK integration
│ ├── eval.py # CopilotEval - main eval harness
│ ├── result.py # CopilotResult - test result data
│ ├── runner.py # Copilot SDK agent execution
│ ├── fixtures.py # copilot_eval fixture
│ └── judge.py # LLM judge (llm_assert, llm_score)
├── execution/ # MCP/CLI server process management
│ └── servers.py # MCPServer, CLIServer, MCPServerProcess, CLIServerProcess
├── fixtures/ # Additional pytest fixtures
│ ├── factories.py # skill_factory (Skills only - agents created inline)
│ ├── llm_assert.py # llm_assert fixture
│ └── llm_score.py # llm_score fixture
├── reporting/ # AI analysis & reports
│ ├── collector.py # TestReport, SuiteReport dataclasses + build_suite_report()
│ ├── generator.py # generate_html(), generate_json(), generate_mermaid_sequence()
│ ├── insights.py # AI analysis engine → InsightsResult
│ └── components/ # htpy report components
│ ├── types.py # TypedDicts for component data shapes
│ ├── report.py # Full report layout
│ ├── agent_leaderboard.py # Eval ranking table
│ ├── agent_selector.py # Eval comparison toggles
│ ├── test_grid.py # Test results grid
│ ├── test_comparison.py # Side-by-side agent comparison
│ └── overlay.py # Fullscreen expanded view
├── templates/ # Static assets for HTML reports
│ └── partials/ # Static assets
│ ├── report.css # Hand-written CSS (edit directly)
│ └── scripts.js # JS (Mermaid, copy buttons, filtering)
└── testing/ # Test harnesses
├── todo.py # TodoStore for CRUD tests
├── todo_mcp.py # Todo MCP server
├── banking.py # BankingService for financial tests
└── banking_mcp.py # Banking MCP server
tests/
├── integration/ # REAL LLM tests (the only tests that matter)
│ ├── conftest.py # Constants + server fixtures
│ ├── copilot/ # CopilotEval harness tests (Copilot SDK)
│ │ ├── conftest.py # MODELS, DEFAULT_MODEL, integration_judge_model
│ │ ├── test_01_basic.py # File create + refactor (parametrize models)
│ │ ├── test_02_models.py # Model comparison
│ │ ├── test_03_instructions.py # Instruction differentiation + excluded_tools
│ │ ├── test_05_skills.py # Skill A/B comparison
│ │ └── test_12_custom_agents.py # Custom agent definitions and dispatch
│ ├── agents/ # .agent.md test fixtures (banking-advisor, todo-manager, minimal)
│ └── skills/ # Test skills
└── unit/ # Pure logic only (no mocking LLMs)
```
## Test Configuration (conftest.py)
Integration tests use centralized constants from `tests/integration/copilot/conftest.py`:
```python
# Models
DEFAULT_MODEL = "gpt-5.6-sol" # Default integration model
MODELS = ["gpt-5.6-sol", "gpt-5.6-luna"] # For model comparison
# Turn limits
DEFAULT_MAX_TURNS = 5
# Server fixtures: todo_server, banking_server
# Evals are created INLINE in each test using these constants
```
**Pattern for writing copilot tests** (in `tests/integration/copilot/`):
```python
from pytest_skill_engineering.copilot.eval import CopilotEval
from .conftest import MODELS
@pytest.mark.parametrize("model", MODELS)
async def test_create_file(copilot_eval, tmp_path, model):
agent = CopilotEval(
name=f"coder-{model}",
model=model,
instructions="You are a Python developer.",
working_directory=str(tmp_path),
)
result = await copilot_eval(agent, "Create hello.py with print('hello')")
assert result.success
```
## Semantic Assertions with llm_assert
Use the built-in `llm_assert` fixture for AI-powered semantic assertions (powered by GitHub Copilot SDK):
```python
async def test_response_quality(copilot_eval, llm_assert):
agent = CopilotEval(...)
result = await copilot_eval(agent, "Show me my account balances and recent transactions")
# Semantic assertion - AI evaluates if condition is met
assert llm_assert(result.final_response, "includes account balances and transaction details")
```
## CRITICAL: Report Development
### Report Generation Architecture
The report pipeline flows:
**Test Execution → Plugin (list[TestReport]) → build_suite_report() → AI Insights → generate_html()/generate_json() → HTML/JSON**
```
src/pytest_skill_engineering/reporting/
├── collector.py # TestReport, SuiteReport dataclasses + build_suite_report()
├── insights.py # AI analysis engine (mandatory) → InsightsResult
├── generator.py # generate_html(), generate_json() module-level functions
└── components/ # htpy report components + types.py
```
### Data Contracts (IMPORTANT)
Every htpy component has a **typed data contract** in `components/types.py`. This is the explicit interface between Python and components:
| Contract | Used By | Purpose |
|----------|---------|---------|
| `AgentData` | `agent_leaderboard.py`, `agent_selector.py` | Eval metrics: pass_rate, cost, tokens, is_winner |
| `TestResultData` | `test_comparison.py`, `test_grid.py` | Per-agent test result: tool_calls, mermaid, outcome |
| `TestData` | `test_grid.py` | Test with `results_by_agent` dict |
| `TestGroupData` | `test_grid.py` | Session or standalone group containing tests |
| `ReportContext` | `report.py` | Full report context with all data + helper functions |
**Pattern**: When modifying components, always check `components/types.py` for the expected shape. When adding component fields, add them to the contract first.
### Template Structure
```
src/pytest_skill_engineering/templates/
└── partials/
├── report.css # Hand-written CSS (edit directly)
└── scripts.js # JS: Mermaid, copy buttons, expand/collapse, eval filtering
```
**Note:** HTML rendering is handled by htpy components in `reporting/components/`, not by template files.
### CSS Development
The project uses hand-written CSS in `partials/report.css`. No build step needed — edit and regenerate reports.
The CSS provides:
- Design tokens as CSS custom properties (colors, shadows, radii, fonts)
- Utility classes matching the patterns used by htpy components
- Semantic component classes (`.card`, `.leaderboard-table`, `.agent-chip`, etc.)
- Markdown content styling
To modify styles:
1. Edit `src/pytest_skill_engineering/templates/partials/report.css`
2. Regenerate report: `uv run pytest-skill-engineering-report aitest-reports/results.json --html aitest-reports/test.html`
3. Open in browser and verify
### Report Sources
1. **Integration tests demonstrate capabilities** - Each test file shows a feature:
- `test_01_basic.py` → Basic tool usage
- `test_02_models.py` → Model comparison leaderboard
- `test_03_prompts.py` → System prompt comparison
- `test_04_matrix.py` → Model × System Prompt grid
- `test_05_skills.py` → Skills integration
- `test_06_sessions.py` → Multi-turn sessions
- `test_07_clarification.py` → Clarification detection
- `test_08_scoring.py` → LLM scoring / rubrics
- `test_09_cli.py` → CLI server testing
- `test_10_ab_servers.py` → A/B server comparison
- `test_11_iterations.py` → Iteration reliability
- `test_12_custom_agents.py` → Custom agent files (`load_custom_agent`, `CopilotEval.custom_agents`)
2. **Showcase tests for hero report** - Located in `tests/showcase/`:
- `test_hero.py` → Curated tests for README showcase
- Demonstrates ALL capabilities in a single cohesive report
- Output: `docs/demo/hero-report.html` (committed to repo)
3. **Report config is in pyproject.toml**:
```toml
addopts = """
--aitest-summary-model=copilot/gpt-5.6-sol
--aitest-html=aitest-reports/report.html
"""
```
### Generate Reports
```bash
# Copilot integration tests (development)
uv run python -m pytest tests/integration/copilot/ -v
# Hero report for README showcase
uv run python -m pytest tests/showcase/ -v --aitest-html=docs/demo/hero-report.html
```
### Manual Testing Workflow (Pre-PR)
**Integration tests use real LLM calls and are expensive. Run the relevant integration tests after every change.**
```bash
# 1. Run the copilot integration test(s) that cover your change
uv run python -m pytest tests/integration/copilot/test_01_basic.py -v
# 2. Run all failed tests (if a broader run previously failed)
uv run python -m pytest --lf tests/integration/copilot/ -v
# 3. Generate hero report (before major releases)
uv run python -m pytest tests/showcase/ -v --aitest-html=docs/demo/hero-report.html
```
## CRITICAL: Template/Report Changes - NEVER Re-run Tests
**For template, CSS, JS, or report generation changes:**
```bash
# CORRECT - Regenerate HTML from existing JSON (instant, no LLM cost)
uv run pytest-skill-engineering-report aitest-reports/results.json --html aitest-reports/test.html
# WRONG - Never re-run tests just to see template changes!
# uv run pytest tests/showcase/ ... ← DON'T DO THIS
```
**When to re-run tests:**
- Test code itself changed
- Eval/tool/engine functionality changed
- Need fresh data with different models/prompts
**When to regenerate from JSON:**
- Template changes (HTML, CSS, JS)
- Report generator code changes
- Layout/styling experiments
- Data contract changes (if backward compatible)
**Workflow for template development:**
1. Edit CSS in `src/pytest_skill_engineering/templates/partials/report.css` or htpy components
2. Regenerate report from existing JSON:
```bash
uv run pytest-skill-engineering-report aitest-reports/results.json --html aitest-reports/test.html
```
3. Open the HTML file in browser
4. Repeat steps 1-3 until satisfied
This workflow is **instant and free** - no LLM calls, no API costs.
### Key Files for Report Development
| File | When to Modify |
|------|----------------|
| `reporting/components/types.py` | Adding new data fields for components |
| `reporting/generator.py` | Changing how context is built for components |
| `reporting/components/*.py` | Individual UI components (htpy) |
| `templates/partials/scripts.js` | Interactivity, Mermaid diagrams |
| `templates/partials/report.css` | Styles, colors, layout |
| `cli.py` | CLI commands for report regeneration |
### Report Design Principles
- **Material Design** - Match mkdocs-material indigo theme
- **Roboto fonts** - Via Google Fonts
- **Test details expanded by default** - Users want to see results immediately
- **Human-readable names everywhere** - All user-facing reports (HTML, Markdown) must use human-readable names, not internal IDs. Use docstrings or humanize function names for tests. Use `eval_name` (not `agent_id`) for evals. Use `system_prompt_name` (not raw keys). If a human-readable name isn't available, humanize the ID (e.g., `test_check_balance` → `Test check balance`).
- **Assertions visible** - Show tool_was_called, semantic assertions
- **Mermaid diagrams readable** - Use neutral theme with good contrast
- **AI insights prominent** - Verdict section at top with clear recommendation
- **Contract-first development** - Define TypedDict before touching templates
### Markdown Style Guidelines
- **No horizontal rules (`---`)** - Headings provide sufficient visual separation
- Use `##` headings for major sections, `###` for subsections
- Horizontal rules inside code blocks are fine (e.g., YAML frontmatter examples)
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
No one has posted yet. Be the first.

