How to verify AutoCSS work before it ships, in a zero-dependency architecture. Covers the air-gap test, driving the real app in a browser to observe behavior, verifying CSS-only state/visibility, and the accessibility passes. Test tooling is dev-only and never enters the shipped static shell. Use before committing any nontrivial change, or when asked to confirm something works.
Estratégia e escrita de testes production-grade: pirâmide vs troféu, unit/integration/e2e, o que testar (comportamento, não implementação), test doubles, contract e mutation testing, cobertura útil. Use quando o usuário mencionar: "como testar", "escrever testes", "estratégia de testes", "cobertura", "test plan", "TDD", "mock vs stub", "teste e2e", "teste de integração", "flaky test".
Choose, write, keep, or delete tests using the smallest real oracle. Use when adding, changing, deleting, or justifying tests; fixing brittle tests; choosing a proof oracle; or deciding no test is appropriate.
Testing strategy and coverage requirements — component-specific testing approaches, test types, coverage targets, and test organization. Use when planning test strategy, determining what types of tests to write, setting coverage goals, organizing test projects, or deciding between unit vs integration tests. Platform, language, and architecture agnostic — does not assume any specific architectural pattern or layering model.
Write tests that catch real regressions — choosing what to test, structuring tests, testing behavior over implementation, handling databases and async, and killing flakiness. Use when adding or improving tests, setting up a test for a bug fix or new feature, deciding what to cover, or dealing with flaky/slow tests. Stack: vitest, `pnpm test`.
Testing strategy, patterns, and evaluation for software and LLM/AI systems. Use when: writing tests, choosing test boundaries, designing test data, structuring test suites, evaluating LLM outputs, building evaluation pipelines, setting coverage thresholds, auditing test coverage gaps in existing projects, or improving test quality and structure.
Write, repair, debug, audit or delete automated tests. Use when adding coverage for new behavior, writing a regression test for a defect, fixing a failing or flaky test, choosing the right test level, or auditing a suite's quality. Not for explaining what an existing test does, and not for running a suite as a routine verification step.
Where a repo's test recipes live — TESTING.md at root, per-directory deltas in monorepos, nearest scope wins — and how implementing and shipping skills scope their test runs to a change. Consult before running tests as a spawned agent, filling a workflow's test-notes slot, or setting up a repo's TESTING.md.
Pragmatic testing guidance focused on confidence, behavior over implementation details, and integration-first coverage. Use when designing a test strategy, writing or reviewing tests, reducing brittle mocks, or deciding what is worth testing in an application or library.
Generate SOX sample selections, testing workpapers, and control assessments. Use when planning quarterly or annual SOX 404 testing, pulling a sample for a control (revenue, P2P, ITGC, close), building a testing workpaper template, or evaluating and classifying a control deficiency.
Covers the testing pattern for every layer — domain (package:test with a hand-written Fake data provider, no mocktail), data (an ephemeral real instance or a mocked client), ui Cubit/BLoC (blocTest + mocktail against the domain repository), ui widgets (the pumpApp helper), and router tests (a fresh buildAppRouter() per test); triggers whenever a test is being written for a new entity, provider, repository, Cubit/BLoC, widget, or route.
Guides risk-based selection and clear construction of useful automated tests. Use when deciding whether code needs tests or when creating, modifying, debugging, or reviewing automated tests.
Unit, bloc, widget, and integration testing conventions — bloc_test, mocktail, http_mock_adapter, test naming, and what to test. Use when writing or fixing any test, mocking a dependency, or deciding what tests a change requires.
Use when writing tests, deciding what to test, reviewing test coverage, or verifying that a spec's acceptance criteria are actually covered. Also when a test is failing and the fix is unclear.
Write comprehensive, maintainable tests following TDD and AAA pattern. Use when writing unit tests, integration tests, setting up fixtures, mocking dependencies, or improving test coverage. Covers Python (pytest), TypeScript (Jest), Go (testing), and Rust (built-in + proptest). Do NOT use for mutation testing specifics (use mutation-testing skill).
Plain text files in a repository that tell a coding agent how the project works: commands to run, conventions to follow and things to avoid. CLAUDE.md, AGENTS.md, cursor rules and skills are the common kinds.
CLAUDE.md or AGENTS.md?
CLAUDE.md is read by Claude Code. AGENTS.md is an open format that Codex, Cursor and other agents read. Many projects keep one and point the other at it.
What is a skill?
A folder with a SKILL.md that describes one capability, such as filling PDFs or reviewing code. The agent loads it only when the task calls for it.
Can I search my own team's files too?
Your agents already can, over MCP, limited to the files you're allowed to read. Searching them from this page is coming.