ratchet / rules
OmkarKamble1/ratchet/.cursor/rules/ratchet.mdc
Development loop for shipping AI-written code without rework. Front-loads the thinking, separates test-writing from code-writing, verifies by mutation. Scales rigor to blast radius.
Cursor rule1 starsChanged 34 days ago
- Commits and pushes
--- description: Development loop for shipping AI-written code without rework. Front-loads the thinking, separates test-writing from code-writing, verifies by mutation. Scales rigor to blast radius. globs: alwaysApply: false --- <!-- generated from skills/ratchet/SKILL.md — edit that file, then run: node scripts/sync.mjs --> Agents write code fast. The slow part is finding out what they got wrong. This loop moves that cost to the front, where a question is cheap, instead of the back, where a wrong answer is a debugging session. ## Persistence ACTIVE UNTIL THE WORK SHIPS OR THE USER SAYS STOP. Do not drift back to write-then-check after a few turns. If unsure whether it still applies, it does. ## It never commits, pushes, or opens anything on your behalf The loop reads, writes files in the working tree, runs tests, and reports. It does not run `git commit`, `git push`, open pull requests, or file issues in a tracker. Those are the human's call, every time, and an agent that takes them silently has removed the last review step there was. Where a gate wants history — sealing tests, proof by deletion — it says so and waits for you to do it, or degrades to a fallback that needs no repository at all. ## Prerequisites — none Nothing needs installing before the loop runs. Every dependency below has a fallback, and each fallback is weaker than the real thing. Use it, name it, and say so in the report. **No test runner, or a repo that should not gain test files.** The test session writes the suite to a temp file, runs it, deletes it — or runs the assertions inline. It still writes them from the spec alone, and they still must all fail first. Nothing about gates 7 and 8 changes except where the file lives. **No version control.** Sealing is a commit when there is one. Without it, the test session states the full suite in the transcript and that text is frozen: the builder re-runs it verbatim and may not restate it. Proof by deletion becomes comment out the implementation, run, restore. **No mutation tool.** Plant three bugs by hand in the new code — flip a comparison, drop a term, remove a guard — and confirm three failures. Report `3/3 hand-planted`, never a percentage you did not measure. **No second session.** Ask the same builder to critique, after a hard reset: it is the critic now, the spec is the only authority, and every finding must quote the spec line it violates. This is the weakest substitution in the loop — a model defends what it just wrote. Prefer clearing the context or switching tools. When neither is possible, run it anyway and label the gate as same-session. A degraded gate is still a gate. An unreported degraded gate is a lie about how verified the code is. ## Gate 0 · Triage — decide how much loop to spend A senior engineer does not apply the same rigor to a copy change and a payment path. Decide once, out loud, before starting. | Blast radius | Mode | |---|---| | Reversible, isolated, no user data | **lite** | | Ships to users, touches shared state | **full** | | Money, auth, medical, safety, migrations, anything one-way | **strict** | Two questions decide it: - **If this is wrong, who finds out and how?** A failing test, a user, or an auditor six months later. The further down that list, the more loop it needs. - **Can we undo it?** A feature flag is reversible. A data migration is not. One-way doors get strict, always. State the mode and the reason in one line. Then run the loop at that level. ## Gate 1 · Readiness — four answers before the first line Not a feeling that you understand the task. Four written answers, stated in the reply before any edit. A missing one is reported as missing, not discovered halfway through. | | Question | How you get it | |---|---|---| | 1 | **How does this work in the real world?** What do the systems you integrate with actually send, and what do practitioners call it? | Gate 3, vendor docs, real samples | | 2 | **What already exists here, and is it right?** | Read the code, per piece | | 3 | **Which rules and invariants does this touch?** | The project's own conventions, the constraints in the spec | | 4 | **How will we know it worked?** | Acceptance criteria and the command that proves them | Answer 2 has exactly three shapes, and you must say which: - **exists and correct** — do nothing, show where it lives. - **exists and wrong** — that is the finding. Fix it or record it; do not build a second copy beside it. - **missing** — now you may build. Assuming "missing" without looking is how a codebase grows two implementations of one rule that drift apart, and neither is wrong enough to be obviously the bug. Grep for the *concept*, not the name you would have given it: the guard you are about to write often already exists under a name you would never have picked. ## Gate 2 · Frame Rewrite the task as an issue: background, goal, acceptance criteria, constraints, test command, diff cap. If acceptance criteria cannot be written, the task is not understood. Go to gate 4 and stay there until it can. ## Gate 3 · Research, don't recall Two moments, not one. **Before building.** Anything about the outside world gets looked up: library versions, API shapes, rate limits, pricing, config defaults, standards, formats, legal status. Cite the source. Never assert an external fact from memory — training data has a date and the world moved. Read the vendor's own words where they exist; a blog summarising them is second-hand. **Before believing a bug.** Half of what looks like a defect is the domain behaving the way it behaves everywhere, and much of the rest is an artifact of the harness that found it. A finding is a claim about the world, so it has to survive the world. Four questions, in order, and any "no" stops it: 1. **Is it reachable in production?** Trace the path real data takes to this code. A probe that calls a function directly can construct states no caller can. 2. **Is the behaviour actually wrong?** Look up how the domain does it before calling it a defect. 3. **Which direction does it fail?** Does it over-report or under-report, over-charge or under-charge, claim more or less? Both are wrong and only one is usually recoverable. A finding without a direction is not finished. 4. **Does the number survive being computed twice?** Once by the code, once by hand from the evidence. A mismatch you cannot explain is the finding. Contradiction between two sources is itself a finding, not a tie to break silently. No result at all: say so and name what is still unknown, rather than filling the gap with a plausible number. ## Gate 4 · Grill — no code yet List every ambiguity, unstated assumption, and place the spec reads two ways. Ask. Wait. A guess here becomes a syntactically valid, contextually plausible, semantically wrong implementation that passes every test. It is the most expensive failure available and it arrives looking finished. - One question at a time when they depend on each other. Batch them when they don't. - Always propose the options you see and say which you'd pick, and why. - "I'll assume X" is acceptable only when nobody is available to answer — and it goes in the spec, flagged, not buried in the code. ## Gate 5 · Walk the cases The happy path is the case you already thought of. Before writing the branch, walk the ladder out loud. It takes a minute and it is most of the job. | Case | What it looks like | |---|---| | **Absent** | The field is not in the payload at all | | **Zero, meaning absent** | A `0`, an empty string, an empty list standing in for "unknown". The single most common shape there is | | **Negative** | A refund, a reversal, a correction that runs the other way | | **Partial** | Two of three fields present. Not absent, broken, and provably so | | **Repeated** | The same record restated on two rows and counted twice | | **Split across sources** | One logical thing arriving in two files, two pages, two events | | **Too many** | 50,000 rows: serial writes, memory, a timeout that only appears at scale | | **Foreign format** | Another locale's dates, another convention's decimal separator, another vendor's casing | | **Two at once** | Concurrent runs, a retry landing beside the original, a job that dies halfway | | **The mirror** | Every rule has an opposite. If you handled "short", handle "over" | | **The reader** | Who receives this output, and does it mean to them what it means to you? | **Two questions catch most of it:** *"Is this a gap or a value?"* A blank, a zero, a `false`, an empty list and an error all arrive looking like data. Decide which is which **before** the arithmetic, not inside it. *"Which direction is this wrong in?"* Errors are rarely symmetric. Name the direction that costs more and make the code fail the other way when it is unsure. **After any fix, check its mirror.** A fix is new code and gets no review of its own unless you go looking. Ask what the opposite input is to the one you just fixed, and whether it still works. **Three failure modes specific to generated code**, all of which look correct in review: - **A symbol that does not exist.** A plausible method or field name that compiles nowhere. Grep the symbol before calling it. - **A test that agrees with the code instead of the requirement.** If it was never seen red, it proves nothing. - **Plausible drift.** Code that reads correctly and implements something slightly different from what was asked. Re-read the request *after* writing, not before. ## Gate 6 · Attack the spec — fresh session A session with no memory of the drafting tries to break it. What's underspecified? What edge case is missing? What would a hostile reader build instead? Cheapest fix in the loop. Do not skip it because the spec feels done — it always feels done to the person who wrote it. ## Gate 7 · Tests — separate session, spec only A test session sees the specification and nothing else. Not the implementation, not the builder's plan, not how the codebase currently solves it. **Do not read the existing tests for the expected answer.** They were written beside the implementation and carry its assumptions; a test asserting today's behaviour is not evidence that behaviour is right. Compute the expected value yourself, from the spec, by hand, then compare. It writes the full suite: happy path, plus every rung of gate 5's ladder, plus every failure mode the grill surfaced. **Run them. Every one must fail.** A test that passes before the implementation exists is asserting on a mock, a constant, or nothing. *Fanning this out across several agents?* One area per agent, and never two agents in the same directory — each will keep hitting the other's half-written files and read the compile error as its own. Give each a step budget and have it append findings to its own file as it goes, so an agent that dies mid-run still leaves its evidence. ## Gate 8 · Seal Freeze the tests. The implementer may not modify a test file, fixture, or CI config from here. With version control that is a commit you make; without it, it is the frozen transcript from gate 7. If a test is genuinely wrong, that is an escalation: name the test, say why it contradicts the spec, propose what it should assert. Someone else amends it. The builder never does. This is the gate the whole loop rests on. Without it every gate below is advisory. ## Gate 9 · Build Spec plus sealed tests, nothing else. Stop at the first rung that holds: 1. Does this need to exist at all? Speculative need: skip it, say so in one line. 2. Does it already exist in this codebase? Reuse it. Re-implementing what lives a few files over is the most common waste there is. 3. Does the standard library do it? 4. Does the platform do it? A native control over a dependency, a database constraint over application code. 5. Does an already-installed dependency do it? Never add a new one for what a few lines can do. 6. Can it be one line? 7. Only then: the smallest code that works. The ladder shortens the solution, never the reading. Trace the flow end to end first; the smallest change in the wrong place is a second bug, not economy. - **One vertical slice per session.** Deployable on its own, not a layer of a larger thing. - **Read before write.** Editing code you have not read produces drift that compounds silently. - **Don't rewrite what works.** The best code is the code you did not write. - Stay under the diff cap. Set it before generation, never after. - Mark a deliberate shortcut with a comment naming its ceiling and the upgrade path, so the next reader knows it was a choice. ## Gate 10 · Mutation gate Plant bugs in the new code — flip a comparison, drop a term, invert a sign, return zero instead of erroring, remove a guard — and confirm the suite catches each one. **Report mutation score, not coverage.** Coverage counts lines executed. Mutation counts bugs caught. A suite at 95% coverage that survives a flipped operator is decoration. **A mutation that does not compile is not a mutation.** It prints no failure line, so a grep for the word "FAIL" reads as a pass. Count failures; do not grep for a word. If your planted bug produced a build error, plant a different one. Stryker (JS/TS, C#, Scala) · mutmut, cosmic-ray (Python) · go-mutesting (Go) · PIT (JVM) · cargo-mutants (Rust) · infection (PHP). No tool for your stack: plant three by hand, confirm three failures. ## Gate 11 · Proof by deletion Remove the implementation. The specific test must fail. Restore it. It must pass. If deleting the code leaves the suite green, the test was never testing the code. Mandatory in strict mode, always worth it on anything money- or auth-adjacent. ## Gate 12 · Scope trace Every hunk maps to a line of the spec. Anything that doesn't gets deleted: the extra helper, the opportunistic refactor, the defensive branch nobody asked for, the dependency added while passing through. This is where over-engineering and training-data bleed-in enter. Both look reasonable in a diff, and once merged they read as intent. ## Gate 13 · Attack the code — fresh session New session. Spec and diff only. No history, no builder reasoning, adversarial framing: find the spec violation. **Read CI and test-file changes first.** A diff that weakens the referee defeats every gate below it, and that hunk looks innocuous. Then attack from angles rather than at random. Most of these will not apply — say so in a line each, and test the ones that do: | Lens | The question | |---|---| | The mirror | What is the opposite input? | | The sibling caller | Who else calls this? A guard on one path leaves the others as they were | | The second door | Is there another way into the same state? | | The direction | If this is wrong, does it over- or under-report? Only one is usually recoverable | | The gap | What does absent do here, as against zero, blank, negative, unreadable? | | Downstream | Who reads this value next, and does the change alter what they conclude? | | The reader | Does the artifact a human receives still say what we mean? | | The trap | What would a naive version of this fix break? Test that it does not | | The scale | 1 row, 50,000 rows, one row repeated, the same input twice | | Provenance | Does the fix make a value look measured when it was inferred? | Same-session self-review does not count. A model that just wrote code will affirm it, or explain why the bug is a feature. ## Gate 14 · Run it cold Every gate so far asked whether the code does what the diff intended. This one asks a different question: **run the thing as a user would, and is the output what they expect?** - No implementation context. Expectations come from the spec and from what a reasonable user would call correct. - Drive the real entry points with realistic inputs, including the messy ones from gate 5. - **Verify every finding yourself before it leaves the session.** A subagent's finding is a lead, not a result: re-run it, read the code, and put it through gate 3's four questions. Expect a third of them not to survive. - Report what did *not* survive as well. A list of what is genuinely safe is worth as much as the failures. A suite written beside the implementation agrees with the code about what the answer should be. This gate is the only one with no such loyalty, and it is where the defects a green suite was hiding tend to appear. ## Gate 15 · Report The report is what the reader actually reads. Four parts, in order. **1. The evidence it stands on.** What the research said and where it came from, what already existed and whether it was right, which invariant was at stake. Gate 1's four answers, restated as findings. A report that opens with a list of changed files is asking the reader to take the reasoning on trust. **2. Files touched, one line of why each.** Path, then the reason it was opened — not "updated the parser", but what changed and what it is for. A file you cannot write a reason for should not be in the diff. **3. Before and after, in numbers.** The measured pair, not an adjective. Where there is no number, quote both outputs verbatim: what the user saw, what they see now. "Improved" and "more robust" cannot be checked. **4. Proof, and what was left.** The commands you ran and their output — the suite, the mutation that went red, the deletion that failed. Then what you did **not** do and why. A skipped item reported is scope; a skipped item unreported is a bug somebody finds later. Style: numbers rather than adjectives, no code in the report — the diff has it — and terse per line, because the structure carries the weight. **Documentation moves with the change.** A doc that lies is worse than no doc, because it is believed. If the change alters documented behaviour, the doc changes in the same edit, not the next one — including the comments that carry the reasoning. ## Gate 16 · Ship, then learn Before it goes out, answer: **how would we know in production if this is wrong?** A log, a metric, an alert, or an honest "we wouldn't", which is itself a finding. Defects found along the way get recorded before they get fixed — in whatever tracker the project uses, or a file in the repo if it has none. The record is what was wrong; the diff is only what changed. One record per defect, closed with the evidence that it is fixed, and a closed one that reproduces gets reopened rather than filed again: "the fix did not hold" is the more useful fact. After: what did the loop miss? Every escape becomes a guard, a lint, or a line in the decisions record. Prefer executable — a test fails loudly, a document decays quietly. ## Escalation — stop and ask when Pushing through these produces confident nonsense. Stop instead. - The spec contradicts the code and you cannot tell which is right - A test fails for a reason the spec does not cover - The fix requires touching something outside the stated scope - Two gates disagree - You are three attempts into the same failure — the approach is wrong, not the code - You are about to widen scope to make something pass Say what you found, what you'd do, and what you need. Do not silently choose. ## Where to spend the loop The loop costs real time. Senior judgment is knowing where it pays. | Spend it on | Skip it for | |---|---| | Money, pricing, billing | Copy and content changes | | Auth, permissions, PII | Prototypes with a delete date | | Migrations and anything one-way | Internal scripts run once | | Concurrency and ordering | Spikes answering "is this possible" | | Anything a customer's auditor could ask about | Config with an obvious rollback | Applying full rigor everywhere trains people to bypass it. Applying it nowhere is how production breaks. Choose deliberately and say which you chose. ## Working in a repo with no tests Common, and not a reason to abandon the loop. 1. Write a characterization test for current behaviour first. Not correct behaviour — *current*. It documents what you must not break. 2. Seal it. 3. Then run the loop on the change. If the existing behaviour is wrong, that is a separate item with its own loop. Do not fix it while passing through. ## Modes | Mode | Gates | Use | |---|---|---| | **lite** | 0, 1, 2, 4, 5, 7, 8, 9, 15 | Reversible, isolated. Readiness, sealed tests and the report still apply. | | **full** | All sixteen | Default for anything shipping. | | **strict** | Full, plus two critics that must agree, a mutation threshold enforced in CI, and no gate skipped without a written reason | One-way doors. | ## Anti-patterns **Not:** "Tests pass, 94% coverage, done." **Yes:** "Mutation score 87%. Three survivors — here they are and why." **Not:** "The test was asserting the wrong thing, so I updated it." **Yes:** "Test `X` asserts `Y`, contradicting spec line 4. Escalating." **Not:** "Let me review my implementation." *(same session)* **Yes:** "Fresh session, spec and diff only, told to find the violation." **Not:** "I assumed the timeout is in milliseconds." **Yes:** "Milliseconds or seconds? Docs say one, the config example implies the other." **Not:** 400 lines: the feature, a refactor, and a new util. **Yes:** 60 lines: the feature. The refactor is its own item. **Not:** "Added error handling everywhere to be safe." **Yes:** "Handled the two failure modes in the spec. Others surface loudly by design." **Not:** "The suite is green." *(not run)* **Yes:** the command and its output, pasted. **Not:** "The reviewer agent found six bugs." **Yes:** "Six reported, four reproduced on the real path, two were the domain working as intended." **Not:** "Fixed — the blank case now returns zero." **Yes:** "Fixed — the blank case now holds. Note it used to return zero, which is the opposite direction of error; check callers that relied on it." ## Configure per repo Three things are repo-specific and belong in your AGENTS.md, not here: ``` test command: <command, or "none — write to a temp file, run, delete"> mutation command: <command, or "none — plant three by hand"> diff cap: <lines> ``` If your host cannot open a genuinely separate session, approximate: clear the context, or run the critique in a different tool. The requirement is that the critic cannot see the builder's reasoning — not which product it runs in. --- MIT. Fork it, cut the gates you don't need, keep gate 8.
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

