agentleFS
Sign inSign up

ratchet / rules

OmkarKamble1/ratchet/.cursor/rules/ratchet.mdc

Development loop for shipping AI-written code without rework. Front-loads the thinking, separates test-writing from code-writing, verifies by mutation. Scales rigor to blast radius.

Cursor rule1 starsChanged 34 days ago
  • Commits and pushes
---
description: Development loop for shipping AI-written code without rework. Front-loads the thinking, separates test-writing from code-writing, verifies by mutation. Scales rigor to blast radius.
globs:
alwaysApply: false
---

<!-- generated from skills/ratchet/SKILL.md — edit that file, then run: node scripts/sync.mjs -->

Agents write code fast. The slow part is finding out what they got wrong.

This loop moves that cost to the front, where a question is cheap, instead of the back,
where a wrong answer is a debugging session.

## Persistence

ACTIVE UNTIL THE WORK SHIPS OR THE USER SAYS STOP. Do not drift back to
write-then-check after a few turns. If unsure whether it still applies, it does.

## It never commits, pushes, or opens anything on your behalf

The loop reads, writes files in the working tree, runs tests, and reports. It does not run
`git commit`, `git push`, open pull requests, or file issues in a tracker. Those are the
human's call, every time, and an agent that takes them silently has removed the last review
step there was. Where a gate wants history — sealing tests, proof by deletion — it says so
and waits for you to do it, or degrades to a fallback that needs no repository at all.

## Prerequisites — none

Nothing needs installing before the loop runs. Every dependency below has a fallback, and
each fallback is weaker than the real thing. Use it, name it, and say so in the report.

**No test runner, or a repo that should not gain test files.** The test session writes the
suite to a temp file, runs it, deletes it — or runs the assertions inline. It still writes
them from the spec alone, and they still must all fail first. Nothing about gates 7 and 8
changes except where the file lives.

**No version control.** Sealing is a commit when there is one. Without it, the test session
states the full suite in the transcript and that text is frozen: the builder re-runs it
verbatim and may not restate it. Proof by deletion becomes comment out the implementation,
run, restore.

**No mutation tool.** Plant three bugs by hand in the new code — flip a comparison, drop a
term, remove a guard — and confirm three failures. Report `3/3 hand-planted`, never a
percentage you did not measure.

**No second session.** Ask the same builder to critique, after a hard reset: it is the
critic now, the spec is the only authority, and every finding must quote the spec line it
violates. This is the weakest substitution in the loop — a model defends what it just
wrote. Prefer clearing the context or switching tools. When neither is possible, run it
anyway and label the gate as same-session.

A degraded gate is still a gate. An unreported degraded gate is a lie about how verified
the code is.

## Gate 0 · Triage — decide how much loop to spend

A senior engineer does not apply the same rigor to a copy change and a payment path.
Decide once, out loud, before starting.

| Blast radius | Mode |
|---|---|
| Reversible, isolated, no user data | **lite** |
| Ships to users, touches shared state | **full** |
| Money, auth, medical, safety, migrations, anything one-way | **strict** |

Two questions decide it:
- **If this is wrong, who finds out and how?** A failing test, a user, or an auditor six
  months later. The further down that list, the more loop it needs.
- **Can we undo it?** A feature flag is reversible. A data migration is not. One-way doors
  get strict, always.

State the mode and the reason in one line. Then run the loop at that level.

## Gate 1 · Readiness — four answers before the first line

Not a feeling that you understand the task. Four written answers, stated in the reply
before any edit. A missing one is reported as missing, not discovered halfway through.

| | Question | How you get it |
|---|---|---|
| 1 | **How does this work in the real world?** What do the systems you integrate with actually send, and what do practitioners call it? | Gate 3, vendor docs, real samples |
| 2 | **What already exists here, and is it right?** | Read the code, per piece |
| 3 | **Which rules and invariants does this touch?** | The project's own conventions, the constraints in the spec |
| 4 | **How will we know it worked?** | Acceptance criteria and the command that proves them |

Answer 2 has exactly three shapes, and you must say which:

- **exists and correct** — do nothing, show where it lives.
- **exists and wrong** — that is the finding. Fix it or record it; do not build a second
  copy beside it.
- **missing** — now you may build.

Assuming "missing" without looking is how a codebase grows two implementations of one rule
that drift apart, and neither is wrong enough to be obviously the bug. Grep for the
*concept*, not the name you would have given it: the guard you are about to write often
already exists under a name you would never have picked.

## Gate 2 · Frame

Rewrite the task as an issue: background, goal, acceptance criteria, constraints, test
command, diff cap.

If acceptance criteria cannot be written, the task is not understood. Go to gate 4 and
stay there until it can.

## Gate 3 · Research, don't recall

Two moments, not one.

**Before building.** Anything about the outside world gets looked up: library versions, API
shapes, rate limits, pricing, config defaults, standards, formats, legal status. Cite the
source. Never assert an external fact from memory — training data has a date and the world
moved. Read the vendor's own words where they exist; a blog summarising them is second-hand.

**Before believing a bug.** Half of what looks like a defect is the domain behaving the way
it behaves everywhere, and much of the rest is an artifact of the harness that found it. A
finding is a claim about the world, so it has to survive the world. Four questions, in
order, and any "no" stops it:

1. **Is it reachable in production?** Trace the path real data takes to this code. A probe
   that calls a function directly can construct states no caller can.
2. **Is the behaviour actually wrong?** Look up how the domain does it before calling it a
   defect.
3. **Which direction does it fail?** Does it over-report or under-report, over-charge or
   under-charge, claim more or less? Both are wrong and only one is usually recoverable. A
   finding without a direction is not finished.
4. **Does the number survive being computed twice?** Once by the code, once by hand from
   the evidence. A mismatch you cannot explain is the finding.

Contradiction between two sources is itself a finding, not a tie to break silently. No
result at all: say so and name what is still unknown, rather than filling the gap with a
plausible number.

## Gate 4 · Grill — no code yet

List every ambiguity, unstated assumption, and place the spec reads two ways. Ask. Wait.

A guess here becomes a syntactically valid, contextually plausible, semantically wrong
implementation that passes every test. It is the most expensive failure available and it
arrives looking finished.

- One question at a time when they depend on each other. Batch them when they don't.
- Always propose the options you see and say which you'd pick, and why.
- "I'll assume X" is acceptable only when nobody is available to answer — and it goes in
  the spec, flagged, not buried in the code.

## Gate 5 · Walk the cases

The happy path is the case you already thought of. Before writing the branch, walk the
ladder out loud. It takes a minute and it is most of the job.

| Case | What it looks like |
|---|---|
| **Absent** | The field is not in the payload at all |
| **Zero, meaning absent** | A `0`, an empty string, an empty list standing in for "unknown". The single most common shape there is |
| **Negative** | A refund, a reversal, a correction that runs the other way |
| **Partial** | Two of three fields present. Not absent, broken, and provably so |
| **Repeated** | The same record restated on two rows and counted twice |
| **Split across sources** | One logical thing arriving in two files, two pages, two events |
| **Too many** | 50,000 rows: serial writes, memory, a timeout that only appears at scale |
| **Foreign format** | Another locale's dates, another convention's decimal separator, another vendor's casing |
| **Two at once** | Concurrent runs, a retry landing beside the original, a job that dies halfway |
| **The mirror** | Every rule has an opposite. If you handled "short", handle "over" |
| **The reader** | Who receives this output, and does it mean to them what it means to you? |

**Two questions catch most of it:**

*"Is this a gap or a value?"* A blank, a zero, a `false`, an empty list and an error all
arrive looking like data. Decide which is which **before** the arithmetic, not inside it.

*"Which direction is this wrong in?"* Errors are rarely symmetric. Name the direction that
costs more and make the code fail the other way when it is unsure.

**After any fix, check its mirror.** A fix is new code and gets no review of its own unless
you go looking. Ask what the opposite input is to the one you just fixed, and whether it
still works.

**Three failure modes specific to generated code**, all of which look correct in review:

- **A symbol that does not exist.** A plausible method or field name that compiles nowhere.
  Grep the symbol before calling it.
- **A test that agrees with the code instead of the requirement.** If it was never seen
  red, it proves nothing.
- **Plausible drift.** Code that reads correctly and implements something slightly
  different from what was asked. Re-read the request *after* writing, not before.

## Gate 6 · Attack the spec — fresh session

A session with no memory of the drafting tries to break it. What's underspecified? What
edge case is missing? What would a hostile reader build instead?

Cheapest fix in the loop. Do not skip it because the spec feels done — it always feels
done to the person who wrote it.

## Gate 7 · Tests — separate session, spec only

A test session sees the specification and nothing else. Not the implementation, not the
builder's plan, not how the codebase currently solves it.

**Do not read the existing tests for the expected answer.** They were written beside the
implementation and carry its assumptions; a test asserting today's behaviour is not
evidence that behaviour is right. Compute the expected value yourself, from the spec, by
hand, then compare.

It writes the full suite: happy path, plus every rung of gate 5's ladder, plus every
failure mode the grill surfaced.

**Run them. Every one must fail.** A test that passes before the implementation exists is
asserting on a mock, a constant, or nothing.

*Fanning this out across several agents?* One area per agent, and never two agents in the
same directory — each will keep hitting the other's half-written files and read the compile
error as its own. Give each a step budget and have it append findings to its own file as it
goes, so an agent that dies mid-run still leaves its evidence.

## Gate 8 · Seal

Freeze the tests. The implementer may not modify a test file, fixture, or CI config from
here. With version control that is a commit you make; without it, it is the frozen
transcript from gate 7.

If a test is genuinely wrong, that is an escalation: name the test, say why it contradicts
the spec, propose what it should assert. Someone else amends it. The builder never does.

This is the gate the whole loop rests on. Without it every gate below is advisory.

## Gate 9 · Build

Spec plus sealed tests, nothing else.

Stop at the first rung that holds:

1. Does this need to exist at all? Speculative need: skip it, say so in one line.
2. Does it already exist in this codebase? Reuse it. Re-implementing what lives a few files
   over is the most common waste there is.
3. Does the standard library do it?
4. Does the platform do it? A native control over a dependency, a database constraint over
   application code.
5. Does an already-installed dependency do it? Never add a new one for what a few lines can
   do.
6. Can it be one line?
7. Only then: the smallest code that works.

The ladder shortens the solution, never the reading. Trace the flow end to end first; the
smallest change in the wrong place is a second bug, not economy.

- **One vertical slice per session.** Deployable on its own, not a layer of a larger thing.
- **Read before write.** Editing code you have not read produces drift that compounds
  silently.
- **Don't rewrite what works.** The best code is the code you did not write.
- Stay under the diff cap. Set it before generation, never after.
- Mark a deliberate shortcut with a comment naming its ceiling and the upgrade path, so the
  next reader knows it was a choice.

## Gate 10 · Mutation gate

Plant bugs in the new code — flip a comparison, drop a term, invert a sign, return zero
instead of erroring, remove a guard — and confirm the suite catches each one.

**Report mutation score, not coverage.** Coverage counts lines executed. Mutation counts
bugs caught. A suite at 95% coverage that survives a flipped operator is decoration.

**A mutation that does not compile is not a mutation.** It prints no failure line, so a
grep for the word "FAIL" reads as a pass. Count failures; do not grep for a word. If your
planted bug produced a build error, plant a different one.

Stryker (JS/TS, C#, Scala) · mutmut, cosmic-ray (Python) · go-mutesting (Go) · PIT (JVM) ·
cargo-mutants (Rust) · infection (PHP). No tool for your stack: plant three by hand,
confirm three failures.

## Gate 11 · Proof by deletion

Remove the implementation. The specific test must fail. Restore it. It must pass.

If deleting the code leaves the suite green, the test was never testing the code.
Mandatory in strict mode, always worth it on anything money- or auth-adjacent.

## Gate 12 · Scope trace

Every hunk maps to a line of the spec. Anything that doesn't gets deleted: the extra
helper, the opportunistic refactor, the defensive branch nobody asked for, the dependency
added while passing through.

This is where over-engineering and training-data bleed-in enter. Both look reasonable in a
diff, and once merged they read as intent.

## Gate 13 · Attack the code — fresh session

New session. Spec and diff only. No history, no builder reasoning, adversarial framing:
find the spec violation.

**Read CI and test-file changes first.** A diff that weakens the referee defeats every gate
below it, and that hunk looks innocuous.

Then attack from angles rather than at random. Most of these will not apply — say so in a
line each, and test the ones that do:

| Lens | The question |
|---|---|
| The mirror | What is the opposite input? |
| The sibling caller | Who else calls this? A guard on one path leaves the others as they were |
| The second door | Is there another way into the same state? |
| The direction | If this is wrong, does it over- or under-report? Only one is usually recoverable |
| The gap | What does absent do here, as against zero, blank, negative, unreadable? |
| Downstream | Who reads this value next, and does the change alter what they conclude? |
| The reader | Does the artifact a human receives still say what we mean? |
| The trap | What would a naive version of this fix break? Test that it does not |
| The scale | 1 row, 50,000 rows, one row repeated, the same input twice |
| Provenance | Does the fix make a value look measured when it was inferred? |

Same-session self-review does not count. A model that just wrote code will affirm it, or
explain why the bug is a feature.

## Gate 14 · Run it cold

Every gate so far asked whether the code does what the diff intended. This one asks a
different question: **run the thing as a user would, and is the output what they expect?**

- No implementation context. Expectations come from the spec and from what a reasonable
  user would call correct.
- Drive the real entry points with realistic inputs, including the messy ones from gate 5.
- **Verify every finding yourself before it leaves the session.** A subagent's finding is a
  lead, not a result: re-run it, read the code, and put it through gate 3's four questions.
  Expect a third of them not to survive.
- Report what did *not* survive as well. A list of what is genuinely safe is worth as much
  as the failures.

A suite written beside the implementation agrees with the code about what the answer should
be. This gate is the only one with no such loyalty, and it is where the defects a green
suite was hiding tend to appear.

## Gate 15 · Report

The report is what the reader actually reads. Four parts, in order.

**1. The evidence it stands on.** What the research said and where it came from, what
already existed and whether it was right, which invariant was at stake. Gate 1's four
answers, restated as findings. A report that opens with a list of changed files is asking
the reader to take the reasoning on trust.

**2. Files touched, one line of why each.** Path, then the reason it was opened — not
"updated the parser", but what changed and what it is for. A file you cannot write a reason
for should not be in the diff.

**3. Before and after, in numbers.** The measured pair, not an adjective. Where there is no
number, quote both outputs verbatim: what the user saw, what they see now. "Improved" and
"more robust" cannot be checked.

**4. Proof, and what was left.** The commands you ran and their output — the suite, the
mutation that went red, the deletion that failed. Then what you did **not** do and why. A
skipped item reported is scope; a skipped item unreported is a bug somebody finds later.

Style: numbers rather than adjectives, no code in the report — the diff has it — and terse
per line, because the structure carries the weight.

**Documentation moves with the change.** A doc that lies is worse than no doc, because it
is believed. If the change alters documented behaviour, the doc changes in the same edit,
not the next one — including the comments that carry the reasoning.

## Gate 16 · Ship, then learn

Before it goes out, answer: **how would we know in production if this is wrong?** A log, a
metric, an alert, or an honest "we wouldn't", which is itself a finding.

Defects found along the way get recorded before they get fixed — in whatever tracker the
project uses, or a file in the repo if it has none. The record is what was wrong; the diff
is only what changed. One record per defect, closed with the evidence that it is fixed, and
a closed one that reproduces gets reopened rather than filed again: "the fix did not hold"
is the more useful fact.

After: what did the loop miss? Every escape becomes a guard, a lint, or a line in the
decisions record. Prefer executable — a test fails loudly, a document decays quietly.

## Escalation — stop and ask when

Pushing through these produces confident nonsense. Stop instead.

- The spec contradicts the code and you cannot tell which is right
- A test fails for a reason the spec does not cover
- The fix requires touching something outside the stated scope
- Two gates disagree
- You are three attempts into the same failure — the approach is wrong, not the code
- You are about to widen scope to make something pass

Say what you found, what you'd do, and what you need. Do not silently choose.

## Where to spend the loop

The loop costs real time. Senior judgment is knowing where it pays.

| Spend it on | Skip it for |
|---|---|
| Money, pricing, billing | Copy and content changes |
| Auth, permissions, PII | Prototypes with a delete date |
| Migrations and anything one-way | Internal scripts run once |
| Concurrency and ordering | Spikes answering "is this possible" |
| Anything a customer's auditor could ask about | Config with an obvious rollback |

Applying full rigor everywhere trains people to bypass it. Applying it nowhere is how
production breaks. Choose deliberately and say which you chose.

## Working in a repo with no tests

Common, and not a reason to abandon the loop.

1. Write a characterization test for current behaviour first. Not correct behaviour —
   *current*. It documents what you must not break.
2. Seal it.
3. Then run the loop on the change.

If the existing behaviour is wrong, that is a separate item with its own loop. Do not fix
it while passing through.

## Modes

| Mode | Gates | Use |
|---|---|---|
| **lite** | 0, 1, 2, 4, 5, 7, 8, 9, 15 | Reversible, isolated. Readiness, sealed tests and the report still apply. |
| **full** | All sixteen | Default for anything shipping. |
| **strict** | Full, plus two critics that must agree, a mutation threshold enforced in CI, and no gate skipped without a written reason | One-way doors. |

## Anti-patterns

**Not:** "Tests pass, 94% coverage, done."
**Yes:** "Mutation score 87%. Three survivors — here they are and why."

**Not:** "The test was asserting the wrong thing, so I updated it."
**Yes:** "Test `X` asserts `Y`, contradicting spec line 4. Escalating."

**Not:** "Let me review my implementation." *(same session)*
**Yes:** "Fresh session, spec and diff only, told to find the violation."

**Not:** "I assumed the timeout is in milliseconds."
**Yes:** "Milliseconds or seconds? Docs say one, the config example implies the other."

**Not:** 400 lines: the feature, a refactor, and a new util.
**Yes:** 60 lines: the feature. The refactor is its own item.

**Not:** "Added error handling everywhere to be safe."
**Yes:** "Handled the two failure modes in the spec. Others surface loudly by design."

**Not:** "The suite is green." *(not run)*
**Yes:** the command and its output, pasted.

**Not:** "The reviewer agent found six bugs."
**Yes:** "Six reported, four reproduced on the real path, two were the domain working as
intended."

**Not:** "Fixed — the blank case now returns zero."
**Yes:** "Fixed — the blank case now holds. Note it used to return zero, which is the
opposite direction of error; check callers that relied on it."

## Configure per repo

Three things are repo-specific and belong in your AGENTS.md, not here:

```
test command:      <command, or "none — write to a temp file, run, delete">
mutation command:  <command, or "none — plant three by hand">
diff cap:          <lines>
```

If your host cannot open a genuinely separate session, approximate: clear the context, or
run the critique in a different tool. The requirement is that the critic cannot see the
builder's reasoning — not which product it runs in.

---

MIT. Fork it, cut the gates you don't need, keep gate 8.

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.