agentleFS
Sign inSign up

docx-cli

kklimuk/docx-cli/CLAUDE.md

CLI for AI agents to read, edit, and comment on .docx files. JSON-AST output, locator-based addressing, full format fidelity via in-place XML mutation. Bun, not Node. Use Bun.file, Bun.write, Bun.env, Bun.$. Bun loads .env automatically — no dotenv. Subsystem-specific guidance lives in nested CLAUDE.md files that load when you edit those folders. If you need to add a new CLAUDE.md to describe a new practice for a part of the system, do so. These conventions are NOT SUGGESTIONS. These are…

CLAUDE.md214 starsChanged 3 months ago
  • Reads credentials
# docx-cli

CLI for AI agents to read, edit, and comment on `.docx` files. JSON-AST output, locator-based addressing, full format fidelity via in-place XML mutation.

**Bun, not Node.** Use `Bun.file`, `Bun.write`, `Bun.env`, `Bun.$`. Bun loads `.env` automatically — no dotenv.

Subsystem-specific guidance lives in nested CLAUDE.md files that load when you edit those folders. If you need to add a new CLAUDE.md to describe a new practice for a part of the system, do so.

## Conventions

These conventions are NOT SUGGESTIONS. These are rules.

- **All stdout goes through `respond()` (JSON ack) or `writeStdout()` (text)** from `src/cli/respond.ts` — never `process.stdout.write`. The production sinks use lazy `Bun.stdout.writer()` / `Bun.stderr.writer()` FileSinks and await each write and flush before exit. Never switch these to `Bun.write(Bun.stdout, ...)` or `Bun.stdout.write(...)`: large OS-pipe writes can repeat bytes or hang after a dependency accesses `process.stdout` (issue #8). Stderr goes through `writeStderr()` for the same reason.
- **File naming**: kebab-case, named after the primary export (`xml-node.ts` → `XmlNode`).
- **Newspaper ordering.** The entry point (primary export) goes at the top; its dependencies follow in the order it uses them, then _their_ dependencies, and so on — a file reads top-to-bottom like a newspaper. Use hoisted `function` declarations for internal helpers so this works at runtime; arrow functions only for inline callbacks and short utilities. Types are usually not the primary exports and should go below the functions/classes that are.
- **Feature nesting** When a file accumulates too many dependencies to be read well with newspaper ordering (> 300 lines), split them into a separate folder/file named after the feature they're working on. It should be a folder if it is going to represent a logical feature of dependencies. This nesting can continue indefinitely if subfeatures have subfeatures of their own.
- **JSX is for emitters only.** Files that construct fresh XML can be `.tsx`; readers/locators/analysis stay `.ts`. Components are PascalCase, accept props, may return `NullableXmlNode` (null skipped by flatten). Attribute names with colons use the hyphen shortcut (`w-val="x"` → `w:val="x"`) or JSX spread.
- **Component vs view vs lens vs free function vs transient cursor**: five shapes, one decision tree.
  - A pure `props → XmlNode` builder is a PascalCase **component** — destructure its props in the signature (no `props.x` access), and don't take a `Document` (or any package state).
  - Stateful OOXML state lives in **tree-owning views**, embedded as fields on `Document`: `Body`, plus one view per OPC part — `StylesView`, `NumberingView`, `CommentsView`, `NotesView`, `RelationshipsView`, `ContentTypesView`, `SettingsView`, `CorePropertiesView`, `MarginalsView`. Each owns its part's `XmlNode` tree and any maps keyed to it, and exposes a `fromPackage`/`fromXml`/`writeTo` lifecycle (`register` too, for the lazily-provisioned ones). `MarginalsView` is the one that owns MANY parts (every `word/header{N}.xml` / `word/footer{N}.xml`) keyed by part name rather than one — so it has no single `register`; the `Marginals` lens mints each part's rel + content-type as it allocates it. Cross-view dependencies (e.g., `NotesView.ensureNoteStyles(stylesView)`) are passed as method arguments — no view reaches up to `Document`.
  - **Cross-cutting lenses** (`Images`, `Hyperlinks`, `Equations`, `TrackChanges`, `Comments`, `Fonts`, `Marginals`) are NOT fields on `Document` — they're stateless, constructed at the call site: `new Images(document).add(source)`, `new TrackChanges(document).accept(["tc0"])`, `new Marginals(document).set(sectPrs, "footer", "default", spec)`, `await new Fonts(document).setDefault("Times New Roman")`. They hold only a back-reference; the embedded views are the state they reach through. `Marginals` is the header/footer authoring lens — the noun pair `docx headers`/`docx footers` share it via `MarginalKind` the way `footnotes`/`endnotes` share `Note` — reaching through `MarginalsView` (part trees) + relationships/content-types (part registration) + settings (the even/odd toggle) + the live `<w:sectPr>` reference nodes (see [src/core/marginals](src/core/marginals/CLAUDE.md)). `Fonts` is the one that also touches an UNMODELED part — the document font lives in BOTH `word/styles.xml` `<w:docDefaults>` (owned by `StylesView`) and `word/theme/theme1.xml`'s `<a:fontScheme>` (not a view: read/mutated/staged through `Pkg` only when `set-default-font` runs, so unrelated saves never re-serialize the theme blob).
  - **Free functions** are reserved for: pure builders (the components above), the AST reader (`src/core/ast/read.ts` — Document's construction pass; populates the embedded views from XML and is the sole assigner of `tcN` ids), and emitter helpers in `src/core/blocks`/`table`/`sections` that thread a `Document` because they touch many slices in one call. If a free function's body operates on one slice of `document`, make it a method on that slice's view instead.
  - **Transient cursors** are the one stateful shape that is NEITHER a view nor a lens: a short-lived object holding position state over a SINGLE node's child list, valid only within one mutation pass (today `CellInsertionCursor` in [src/core/table](src/core/table/CLAUDE.md), which keeps a batch's inserts into one `<w:tc>` in entry order). It holds no `Document`, isn't embedded on one, and dies with the pass — so don't file it as a lens. It lives beside the primitives it sequences, and the CLI constructs one per target and calls it.
- **JSX.Element = XmlNode** (single, not nullable). `Fragment` returns a `#fragment` sentinel unwrapped in `flatten()` and `serialize()`. Components return `null` to render nothing; `jsx()` converts that to an empty fragment. The `jsx`/`jsxs`/`jsxDEV` runtime exports are distinct functions, not `= jsx` aliases (knip flags aliased re-exports as duplicates) — don't collapse them.
- **Path aliases**: `@core` → `src/core/index.ts`, `@core/*` → `src/core/*`. Use these in `src/cli/*`; `src/core` itself uses relative sibling imports. Import the body emitters from the `@core/blocks` and `@core/table` subpaths, not the `@core` barrel — `ast/types` already exports `Paragraph`/`Table`/`TableCell`/`TableRow` as _types_, and barrel-merging the same-named value emitters is confusing.
- **Variable names**: descriptive, no single/two-letter (`paragraph` not `p`). Exception: regex-match destructuring (`const [, prefix, idx] = match`).
- **Inline props in the signature.** When a component's props type is used only by that component, write it inline (`function HeadingStyle({ styleId }: { styleId: BaselineStyleId; … })`) rather than declaring a separate named `Props` type. Extract a named type only when it's shared.
- **knip runs strict** (`bun run check`, no rule overrides in `knip.json`). An unused export is dead code — delete it (this is a CLI app, not a library; there are no external `@core` consumers). The one exception: an export staged for a named upcoming tier with no caller yet gets a `@public` JSDoc tag whose comment names the future consumer (knip honors `@public`) — e.g. `HorizontalRule` (S8) and the `r`/`a`/`wp`/`pic` image namespaces (S5). Don't silence knip by re-adding rule suppressions.
- **Style**: tabs, double quotes (Biome enforced). Early returns over else-if chains.

## Invariants

These invariants are NOT SUGGESTIONS. These MUST be followed.

- **Changes must survive the write-read loop.** This tool is primarily designed for agent use, and specifically by weaker agents like Haiku. Thus, any change made should be retrievable for observation on the next turn in the AST. Whenever possible, it should be available in Markdown, which is the default read view.
- **We do not handle all possible OOXML elements: so make in-place XML mutations, not AST/Markdown round-trips.** The AST or Markdown are just views. Mutate `XmlNode` refs, only emit fresh XML for inserted nodes. This way, elements we cannot parse will not be broken. See [src/core](src/core/CLAUDE.md).
- **Most things, even those that are not easily representable, should be in Markdown.** For these, we (1) build comments into the markdown to give folks a better sense of what is going on, (2) make sure that these comment do not break the markdown and survive the roundtrip, and (3) provide intuitive verbs to deal with concepts. If some structure is truly not representable in Markdown, we nudge agents to render the pages and look. We should guarantee that the things that we're not doing the above for have no possible representation.
- **Authoring has two channels: parsed and literal.** `--markdown`/`--from` parse GFM (+ math + CriticMarkup + inline HTML); `--text-file` (on `insert` and `create`, `-` = stdin) is the PARSER-FREE channel — every character lands verbatim, each newline starts a new paragraph. It exists because GFM corruption of literal prose isn't always escapable: a bare URL autolinks with NO escape sequence at all, and CriticMarkup eats `{++…++}` regardless of backslashes (`*x*`, `3.`, `[t](u)` ARE backslash-escapable, but the agent — often Haiku — must get the whole dialect right, and the field evidence is they don't). So "put this exact text in, untouched" gets its own flag rather than an escaping burden on the weakest actor. `--text` stays inline single-paragraph; both it and the bulk-literal `--text-file` insert markdown-looking content VERBATIM (`--text "**bold**"` writes literal `**` — use `--markdown` to parse). The only value `--text` still refuses is a shell-gutted currency amount (`rejectShellMangledValue`: `"$300"` → `.00`), which is never intentional. **Every inline `--text`/`--markdown` argv value gets its `\n`/`\t`/`\r`(`\r\n`) decoded at the CLI ingress** (`decodeInlineEscapes` in `cli/parse-helpers.ts`) — a weak agent writes `--text "a\nb"` for a line break, but bash double-quotes hand the CLI the literal backslash-n, which used to land verbatim in the run. It's applied by EVERY inline authoring surface, not just the body: `edit`/`insert`/`create`, `footnotes`/`endnotes add`+`edit` (`--text` and inline `--markdown`), `comments add`+`reply` (`--text`), `headers`/`footers set` (`--text`), and the `find`/`replace` positionals (QUERY/PATTERN/REPLACEMENT) — where a decoded `\n` is ALSO the gate that routes the command to the cross-paragraph matcher (a `\n` in a find/replace pattern matches an in-paragraph line break OR the boundary between consecutive paragraphs — in the body or within one table cell, never across a cell wall; a `\n` in a replace REPLACEMENT means a paragraph mark, editor-style — single-line replacement across a boundary MERGES paragraphs, replacement `\n` SPLITS; untracked only, refuses under tracking). Decoded whitespace flows through the same real-character handling everything uses — the shared **`textToRunElements`** emitter in `core/blocks.tsx` (`textToRuns` + `RunElement`, nulls filtered) turns a real `\n`/`\t` into `<w:br/>`/`<w:tab/>`, so the note/comment/marginal emitters (which previously emitted a single raw `<w:t>` that swallowed a newline) now render breaks just like a body paragraph; the markdown block splitter handles the `--markdown` side. Plain single-line text still collapses to one `<w:t>` run, byte-identical to the old shape. ONLY those whitespace escapes decode: a `\`-before-punctuation (markdown's own `\*`/`\[`/`\.`, and a literal `\\`) passes through UNTOUCHED, so decoding can't corrupt markdown. The decode is INLINE-argv-only — `--text-file`/`--markdown-file`/`--from` read real characters from a file and `--batch` JSONL is already `JSON.parse`-decoded, so none are re-decoded; `--text-file` therefore stays the one channel where a bare `\n` really means backslash-n.
- **Comments are never anything but hints.** Every `<!-- … -->` comment `read` writes into Markdown — locators (`<!-- pN -->`) and the structural visibility annotations (`docx:section` — rendered at the section's START with an `applies-to="pX..pY (below)"` scope on deviating sections so the "which side of the boundary?" off-by-one is explicit, `docx:page`, `docx:table` (widths/borders + the `tables format` properties align/style/repeat-header/row-heights), `docx:cell` merge/shading/vAlign/halign/borders, `docx:track-changes on|off` ALWAYS at the head (the one hint that states its default too — weak agents can't reliably read "no hint" as "off," and a wrong tracking guess is high-cost, so it's stated outright), `docx:header`/`docx:footer` for page headers/footers (content in a `text` attr, fields as `{page}`/`{date}`/… tokens — authored via `docx headers`/`docx footers`; a UNIFORM marginal rides the head, a marginal that DIFFERS by section renders at that section's START, next to the content it governs), `docx:textbox tbxN anchor="pM"` … `docx:textbox-end tbxN` bracketing a text box's story, printed right after the paragraph that anchors it (the story's own paragraphs carry bare `tbxN:pK` locators), `docx:layout` flagging tab-aligned content that wraps in render — inside a multi-column section, OR a line whose trailing content rides a right-edge LEFT tab so a long value overflows the margin (the résumé `San`/`Francisco` split); the latter is also consolidated into ONE top-of-read `fix-all` summary naming the single-call cure `edit --at pN-pM --tabs right` (a render-only break Markdown can't show), and future image/style hints, all via `formatNote` in `cli/read/annotations.ts`) — is a READ-TIME VISIBILITY HINT. **Naming rule: a BARE comment is a locator (an address you pass to `--at`: `<!-- p0 -->`, `<!-- t0:r0c0:p0 -->`; a plain empty cell uses `<!-- t0:r0c0 -->` so it can be filled directly); a `docx:TYPE` comment is docx-cli METADATA** (structure/formatting). Metadata never rides a bare locator — it's its own `docx:` annotation, which may carry the relevant locator as a bare leading token (`<!-- docx:cell t0:r0c0 gridSpan="2" -->`). So `docx:` is greppable as "everything docx-cli added beyond addressing," and locators stay short. The Markdown importer DROPS them all (`walker.tsx`'s `case "html"` returns `[]`; the inline walker drops inline comments); they NEVER drive reconstruction or carry data through `read → create`. A comment can't silently corrupt the round-trip or smuggle structure back — what an agent sees is a hint, nothing more. The lossless view is `read --ast`; in-place `edit` preserves structure; the authoring verbs (`docx sections`, `docx tables …`) change it. Emit hints **deviation-only** (only what differs from the document default — even columns, default geometry, `single` borders, `align=left` emit nothing), with TWO deliberate exceptions that always emit even at their default: `docx:section` (every section gets a marker so the `sN` is always addressable) and `docx:track-changes` (states `off` too, so a weak agent never infers state from a missing hint). Run FORMATTING is the one thing that's Markdown-representable and DOES round-trip, but it rides HTML ELEMENTS (`<span>`/`<mark>`/`<sup>`/`<u>`, via `gatherHtmlSpans`), NOT comments. Even `<!-- docx:base -->` (the dominant font/size note) is a hint: `read` omits the baseline from runs and shows the note so an agent can match it, but the importer drops it — a full `read → create` rebuild falls back to the template docDefaults (Calibri 11pt) for the dominant font/size. No comment, anywhere, is parse-back.
- **Everything we parse MUST be represented in the AST**. Markdown is always either a subset or equal to what is in the AST, never less than it. `RUN_BEARING_WRAPPER_TAGS` in `src/core/parser/run-ops.ts` is the AST↔XML offset bridge; every offset-aware walker reads it. See [src/core](src/core/CLAUDE.md).
- **`<mc:AlternateContent>` is transparent at EVERY level, and text boxes are first-class.** ECMA-376 Part 3 lets a producer wrap a block, a paragraph's runs, or a run's children in a markup-compatibility wrapper and nest wrappers inside branches; Word writes every modern shape/text box that way (`wps` under `<mc:Choice>`, VML under `<mc:Fallback>`, SAME story in both). Every walker resolves it through `alternateContentBranch` in `src/core/mc.ts` (first Choice, else Fallback) — the reader at block/paragraph/run level, the offset walkers via `wrapperContent(node)` in `parser/run-ops.ts` (NEVER `node.children` for a wrapper — that is what keeps a multi-Choice wrapper counted exactly once), and the text-box collector while descending a shape. Dropping the wrapper was issue #4: `replace --all` reported 2 and exited 0 while a third "Acme" sat in a text box a human reads first. A `<w:txbxContent>` story is a block container like a cell (`TextBoxRun` on the anchor paragraph, `tbxN:pK` ids, `blockReferences` parents inside the Choice copy), so `find`/`replace`/`edit`/`insert`/`delete`/`comments` reach it by default; `Document.save` (and the raw gate, BEFORE validating) mirrors each Choice story onto its Fallback twin (`syncTextBoxFallbacks`, post-order so nested boxes sync inner-first) — verified against Word for Mac 365: Word regenerates a box's Fallback when IT edits the box and leaves an untouched box's Fallback verbatim (a stale one stays stale forever), so mirroring on our writes is exactly its behavior. `wc` counts what `read` shows — text boxes included, everywhere. **Word discards comments and footnotes/endnotes inside a text box on its next save** (verified: the markers vanish from the story and the comment from comments.xml, regardless of the Fallback), so `comments add` (`--at` AND `--anchor`) and `footnotes`/`endnotes add` REFUSE a `tbxN` target with `UNSUPPORTED` + a pointer to the anchor paragraph; tracked changes and hyperlinks inside a box survive Word and are allowed. The box rides its anchor RUN: a whole-paragraph `edit --text` lifts object-bearing runs (drawing/pict/object/AlternateContent) out of the token diff and puts them back (`extractObjectRuns`), a `replace` on the anchor run's text keeps a trailing zero-width shape (`sliceRun`), and a tracked delete of the anchor wraps the run in `<w:del>` WITHOUT rewriting the story's `<w:t>` (`TextBoxRun.trackedChange` hides it in the accepted view; `iterateBlocks({ view })`).
- **Every element must have stable positional ids for locators** (`p0`, `t0`, `c0`, `img0`, `link0`, `tc0`, `hdr0`/`ftr0` for page headers/footers, `tbx0` for a text box — its story chains like a cell: `tbx0:p1`, `tbx0:t0:r0c0:p0`). Block ids shift after structural edits. Re-read between non-trivial mutations. To not force agents to re-read on every modification, we create batch versions of any editing call such that the already made locators can be used.
- **Locators are the backbone of the system.** Agents use these to reliably identify elements, ranges (both bounded and one-side bounded). If you're introducing a new concept, it is CRITICAL that these work and work well.
- **The CLI surface should be thin.** The CLI surface is concerned with parsing user input. The core is where the actual domain logic happens.
- **Never delete a relationship or part something still references** — a dangling rId corrupts the file ("unreadable content"), but an unreferenced part is harmless. So pruning is gated on a reference check that scans **everything we don't model** (VML `<v:imagedata>` fallbacks, OLE objects, `<w:background>`, chart rels), not just the construct we authored: `isRelationshipReferenced(documentTree, rId)` before dropping a relationship, `hasRelationshipWithTarget` before deleting a shared media part (both in `core/relationships.ts`). When in doubt, leave the orphan.
- **Hyperlinks own a relationship, not their text.** `hyperlinks replace` updates the `<Relationship>` `Target` (mints a new rId if multiple `<w:hyperlink>` share one); `delete` unwraps and prunes the rId when unreferenced.
- **Track-changes is doc-level** — see [src/cli/track-changes](src/cli/track-changes/CLAUDE.md).
- **`docx render` is the only command that needs an external APPLICATION.** Every other verb works purely against the .docx zip + XML; `render` shells out to Word (macOS/Windows) or LibreOffice (cross-platform) to produce a PDF, then rasterizes in-process via the bundled `@hyzyla/pdfium` WASM package (no system tools needed for the rasterizer). The lens lives at [`@core/render`](src/core/render/CLAUDE.md) (`renderDocxPages` is the entry point); [`src/cli/render`](src/cli/render/CLAUDE.md) is a thin arg-parse shell. Agents that consume PNGs (and you, when verifying a fixture) use this command; the rest of the CLI never invokes it.
- **Raw OOXML enters only through the gate pipeline.** `docx raw insert`/`replace` is the one surface where the user hands us arbitrary XML, and every fragment runs `Raw.prepareFragment`'s gates before it may touch the tree (see [src/core/raw](src/core/raw/CLAUDE.md)): well-formedness via `XMLValidator` (our own parser NEVER throws — it silently mangles bad XML, so it can't be the gate), addressable roots only (`w:p`/`w:tbl`, or one `w:sectPr` replacing an `sN` — anything else has no locator and would vanish from `read`, breaking the write-read loop), inter-element whitespace strip (pretty-printed fragments corrupt Word), ECMA child-order checks against the same order tables the emitters use, namespace auto-declare/reject, reference integrity (dangling rIds/note ids reject; drawing/bookmark id collisions re-mint), and a final full-document baseline-diff schema validation (bundled ECMA-376 transitional XSDs via libxml2-wasm; a mutation may not ADD schema errors; `--no-validate` skips only this gate; `docx validate` exposes the same engine standalone). **Nothing is written when any gate fails** — all mutation is in-memory until the post-gate save. Inserted roots are stamped `dcx:raw="1"` (the `dcx` prefix rides `mc:Ignorable`, so Word and the validator skip it) and surface as a `raw` token on the `docx:p`/`docx:table` note + `rawXml` in the AST. Raw changes are never tracked (`--track` rejects); under doc-level tracking they apply untracked + audit comment, per the hyperlinks/images precedent. **Relationships are raw-addressable too** (`raw get --at rels|rIdN`, insert a `<Relationship>` fragment routed by shape, Id-preserving replace) with their own proportional gate set (Id/Type/Target/TargetMode is the whole rels schema; internal Targets must resolve to an existing part; colliding/renamed Ids reject). **Headers/footers are raw-GET-readable** by their `hdrN`/`ftrN` id — a read-only serialize of the `<w:hdr>`/`<w:ftr>` part (a read needs no gate); the id is a first-class locator (`headers`/`footers list` reports it, `read` leads its `docx:header`/`docx:footer` hint with it, `headers`/`footers set/clear --at ftrN` author by it), and authoring stays on the modeled `docx headers`/`footers` verbs. **OPC parts are addressable too, but NOT as locators** — locators address document content; parts are the container, so they get their own noun: `raw part list|get|add|replace|edit FILE --name NAME` (`add`: XML gated, binary verbatim, content type required or extension-Default; whole-part replace/edit with root tag preserved — unmodeled parts via `Pkg`, styles/numbering/settings/marginals via a VIEW reparse so the swap survives the save, and notes/comments/document.xml/rels rejecting toward their own surfaces because their ids pair with body markers). **`raw edit` / `raw part edit` is the patch loop as one gated call**: literal `--find`/`--with` over the target's exact serialization, all occurrences, zero matches errors, result runs the full replace gate set. Part names and native ids are the ONLY sub-document handles — no positional addressing inside a part, ever. The full embedded-object loop is `raw part add` → insert `<Relationship>` → insert the body fragment referencing the rId.
- **Installers fetch release ASSETS, never remote code.** Every install path resolves a release **tag** or the `latest` release pointer — never a branch — then downloads that release's prebuilt binary plus its published `SHA256SUMS` and verifies the digest before the binary lands. (`install.sh` is itself covered by `SHA256SUMS`, but nothing self-verifies it: checking the installer is the operator's step, which is why the README fetches it to a file instead of piping it.) None of them downloads a script and runs it — that was a Snyk CRITICAL on the skills.sh listing, because the skill's bootstrap used to `sh` a fetched `install.sh`, and neither a scanner nor a reviewer can audit a file the script goes off and gets. There is exactly ONE implementation of the install logic, `skills/docx-cli/scripts/install.sh`, delivered THREE ways, chosen by who is asking: as a release asset (the operator, by hand), inside the skill folder where `bootstrap.sh` runs it as a local sibling (an agent at session start), and EMBEDDED into the binary at build time via `import … with { type: "text" }` for `docx upgrade` (an installed human). Because of the third route install.sh must stay POSIX `sh` with no sibling or `$0` dependencies — `upgrade` pipes it to `sh -s`. **Don't collapse these into one route**: `bootstrap.sh` can't call `docx upgrade` (it must work when `docx` is absent, must not downgrade a local build, and must not trust the possibly-stale installer inside whatever binary is present), and `upgrade` can't fetch the script (that's the forbidden pattern). Note `upgrade` runs the installer from the release you're ON, not the one you're going TO, so an installer bug in vN can't be fixed by upgrading from vN. The repo-root `install.sh` is a byte-identical DEPRECATION SHIM, not a second implementation — bootstraps already published to npm/skills.sh (≤ v0.22.0) fetch `raw.githubusercontent.com/…/<tag>/install.sh`, so a tag whose tree lacks that path breaks every deployed skill's session-start check. `tests/cli/installer.test.ts` enforces the copy, the platform-table↔release-matrix agreement, and the absence of fetch-then-execute; delete the shim and its test once those bootstraps age out. It lives under `skills/` precisely so its sibling `bootstrap.sh` can delegate to it as a LOCAL file — executing a shipped sibling is not the flagged pattern; fetching a script and running it is. Keep the split by question: **install.sh answers "put `docx` here"** (detect platform, download, verify, place) and **bootstrap.sh answers "keep `docx` current"** (compare versions, update only if behind, never downgrade a local build, degrade to exit 0 offline, confirm the PATH `docx` is the one just installed). Don't reintroduce install logic into bootstrap. Posture differences are PARAMETERS, not forks: `bootstrap.sh` passes `REQUIRE_CHECKSUM=1` so an agent gets a refusal where a human running install.sh by hand gets a warn-and-install. Order the checks cheapest-first: platform, checksum policy, `PREFIX` writability, the few-hundred-byte manifest, and only then the ~100 MB binary.
- **No undo, no journal.** Mutating commands overwrite `FILE` in place; git is the history. `-o/--output PATH` writes a parallel file; `--dry-run` previews (wins over `--output`). Track changes is the closest thing to a journal, but not fully reliable because some operations are irreversible in OOXML. Our system denotes these with comments.

## Commands

`docx <verb>` and `docx <noun> <verb>`. Every command has `--help`. Mutating commands accept `--dry-run` and `-o/--output PATH` — **except `create`, whose positional FILE is already the output path** (it has no `-o`). Real-tracked-change mutators (`edit`/`insert`/`delete`/`replace`, the `tables` verbs, `images delete`, `headers`/`footers set`/`clear`) also accept `--track` to record one invocation as tracked even when the doc toggle is off. (Headers/footers track the `<w:sectPr>` *reference* change via `<w:sectPrChange>`; the part-body content is authored fresh — see [src/core/marginals](src/core/marginals/CLAUDE.md).)

**Locator flags are unified:** address an existing thing with **`--at LOCATOR`** everywhere (edit, insert, delete, tables, comments, footnotes/endnotes, images, hyperlinks, track-changes; `headers`/`footers` take `--at sN` to scope to one section, defaulting to the whole document). For `insert`, `--at` means AFTER an ordinary block and INTO a bare cell; `--before`/`--after` choose an explicit block side or a cell's start/end boundary, while `--at-start`/`--at-end` address document boundaries (no locator, single-shot only, not batchable). A bare-cell insert fills/reuses Word's mandatory empty paragraph and otherwise appends; bare-cell edit aliases exactly one direct paragraph. Both reject merged/grid-shifted cells, and edit rejects multi-block cells — use `CELL:pK` for precision. `read` uses `--from`/`--to` (a block slice), `comments add` also takes `--anchor PHRASE`, and `wc` takes a positional `[LOCATOR]`. No more `--id`/`--range`/`--to` for addressing.

**`--batch FILE.jsonl` (one JSONL object per line; `-` = stdin)** uses the shared `readJsonlObjects` reader (`cli/parse-helpers.ts`), with each command's path in a sibling `batch.ts`. `find --batch` is the read-only form: independent `{ query, … }` / formatting-filter entries all run against one immutable document read; text output flattens locators in entry order, while `--json` preserves request boundaries in a `batch` array. The mutation form lets `edit`, `insert`, `replace`, `delete`, and the `comments` verbs apply many changes from a single read — keys mirror the command's flags (`delete` entries are just `{ at }`). **All locators in a mutation batch address the document AS READ** — the load-bearing invariant — and each command keeps that true differently: `edit` resolves every entry to a live `XmlNode` ref _before any mutation_ (same-paragraph spans apply right-to-left; a paragraph takes one whole-paragraph edit OR several non-overlapping spans OR one line removal — empty `text` / `delete:true`, routed to the cell-safe `removeParagraphLine` so a form-fill fills cells AND drops leftover placeholder lines in one sweep); `replace` re-reads the live tree between entries (`document.reread()` re-walks `documentTree`, no re-parse, so node identity holds) for sed-like sequential semantics; `insert` pins block/cell anchors to live refs, builds all blocks with zero body mutation, reserves an initially empty cell's mandatory paragraph for its first bare-cell entry, rejects mixed bare-cell + explicit `CELL:pK` targeting in one batch, then splices with a per-anchor offset for block targets and one `CellInsertionCursor` (in `@core/table`) per cell, so stacked inserts keep entry order; `delete` resolves every locator to a live ref before any removal, then splices each by its live `indexOf` (so sibling shifts never misfire). Range/span `edit`s and `delete`s aren't batchable (do them one at a time); section (`sN`) and equation (`eqN`) targets aren't `edit`'s surface at all — use `docx sections` / `docx equations edit` (and `delete --at sN` for a section break).

**Output model — exit code is the success signal** (`0` ok, `1` general, `2` usage, `3` not-found, in `src/cli/respond.ts`), but every command also confirms in text — silence forced (weak) agents to re-read just to know a mutation took effect. The `ok` field appears ONLY on the `--verbose` ack. So:

- A mutator that mints a new addressable handle (`comments add`→`cN`, `comments reply`→`cN`, `footnotes/endnotes add`→`fnN`/`enN`, `hyperlinks add`→`linkN`, `insert`→the new `pN`) prints the bare locator(s) — one per line — by default (`respondMinted`); `--verbose` → full `{ok:true,…}` ack.
- A mutator with no new handle (`create`/`edit`/`delete`/`replace`/`comments resolve`/`tables *`/`track-changes *`/`toggle`) prints a one-line, text-first confirmation on success — `<operation> <target>` (e.g. `edit t1:r0c1:p0`, `edit 7 changes`, `replace 3 occurrences replaced`, `tables.insert-row t2 (1)`) via `respondAck`; `--verbose` → full `{ok,…}` JSON ack. The line is derived centrally in `respondAck`/`summarizeAck` from the ack payload's `operation` + most salient field (locator, count, id, table cell, …), so adding a mutator gets a confirmation for free as long as its ack carries `operation`.
- Query commands are text-first where there's a clean plain form: `find`→locator lines, `wc`→count, `outline`→indented `locator␉text` tree, `read`→markdown — each with `--json`/`--ast` for the structured payload. List verbs print a bare JSON array whose items' `id` is the `--at` handle.
- Errors print `{code, error, hint?}` (no `ok`) + a nonzero exit. Dry-run previews drop `ok` too.
- **A mutation that changes NOTHING is an error, not a silent success.** Weak agents key their done/retry decision off the exit code (they react to nonzero, ignore a cheerful `replaced: 0` line), so a zero-effect mutation that exits `0` bakes in a confidently-wrong document. Any mutator with a clean "matched N things" signal must exit nonzero (`MATCH_NOT_FOUND`) when N is 0 — `replace` and `comments add --anchor` both do. **`replace --batch` is the one wrinkle: it exits nonzero if ANY entry matched nothing, but (sed-like) still SAVES the entries that DID match — so a nonzero batch means "some entries missed," NOT "nothing changed," and the applied edits are already on disk** (the error names which entries missed and how many were saved). Locator-addressed mutators (`edit`/`delete`/`tables …`) fail `BLOCK_NOT_FOUND` when the target locator can't be RESOLVED — a narrower guard, not the same one: it catches a bad address, not a resolved target that changed nothing (an `edit` whose flags already match the current state still exits `0`). Keep new mutators consistent: never `respondAck` a no-op.

**Script fonts:** `--font` keeps its legacy ASCII/high-ANSI/complex-script scope.
`--font-east-asia` and `--font-complex-script` target independent slots, remove
only the corresponding theme reference, and apply after `--font`. Both travel
through edit/replace (including batch), styles set/create, AST, `--runs`, and
Markdown span data attributes; preserve them through the write-read loop.

## Testing

```bash
bun run test:unit         # core + cli (~3s)
bun run test:integration  # LibreOffice round-trip (auto-skips if no soffice)
bun test                  # everything
bun run check             # biome + knip + tsc
```

Fixture authoring (core-emitters-first, CLI-second) and rebuild instructions live in [tests/fixtures/setup/CLAUDE.md](tests/fixtures/setup/CLAUDE.md).

### Every new feature needs a fixture AND a weak-agent scenario

Unit tests are necessary but NOT sufficient — they prove the XML we emit is what we *think* it is, not that Word/LibreOffice accepts it or that a weak agent can *find and use* the feature. So a new verb/flag/surface is not "done" until BOTH of these exercise it (check each when you ship; this is how `set-default-font` slipped through with neither):

1. **A fixture** — a builder under `tests/fixtures/setup/` that **dogfoods the new verb** (`bun ${cliEntry} <new-verb> …`), added to `CORE_FIXTURES` so it round-trips through LibreOffice. First try to **fold it into the existing fixture for the same surface/noun** (e.g. `styles set`/`create`/`set-default-font` all live in [`style-defs.ts`](tests/fixtures/setup/style-defs.ts)); only add a new fixture if no existing one fits. This catches render-invalidity (bad child ordering, theme/part corruption) the unit XML assertions can't.
2. **A weak-agent-test scenario** — a task in the [`weak-agent-test`](.claude/skills/weak-agent-test) skill's `scenarios/` whose request would naturally lead a Haiku agent to reach for the feature (e.g. the `resume` scenario drives `styles set`). First try to **fold it into an existing scenario**; if none fits, **prompt the user before adding a new scenario** (a scenario is a heavier commitment — new fixture + render + judge). This is the only thing that proves the feature is *discoverable* and *usable* by the weak agents the tool is built for — the gap unit tests structurally can't see.

   **Scenario shape — two files, strict separation:**
   - `task.md` — AGENT-FACING. Written as **a human delegating real work**: the goal, the materials/data, the intent and quality bar. It contains **NO tool vocabulary** — no `docx` commands or flags, no locators (`pN`/`tcN`), no OOXML/internal terms (`twips`, `pPrChange`, `docx:base`, `--ast`). Use only words a Word *user* would say (tracked changes, heading style, margins, page numbers, points/inches). The agent must DISCOVER which features deliver the outcome — spelling out the commands defeats the entire test. (Was `task.md` + `brief.md`; merged — a human writes one request, not a task plus a brief.)
   - `criteria.md` — JUDGE-ONLY. The precise, tool-specific grading rubric (exact commands, `pPrChange`, hex values, …). The stage step **withholds it from the agent's run workspace** (`cp` then `rm`), and the judge reads it from the pristine source — so the agent structurally cannot read the answer key. A `# Grading rubric — <key> (JUDGE ONLY)` header + `## Pass conditions` + `## How to verify`.

When a feature's natural brute-force path already works for weak agents (verified: they apply doc-wide paragraph formatting via a range `edit`, never via the style default), a dedicated convenience verb may not be worth adding at all — but the feature that DOES exist still needs both kinds of coverage.

## Build

- `bun run build` → `dist/index.js` — bundled JS that npm publishes (runs under Bun). **Required**: path aliases (`@core/*`) and JSX runtime resolution don't work when consumed from `node_modules`; the bundle pre-resolves everything. Never ship raw `src/`.
- `bun run build:binary` → `dist/docx` — standalone executable for GitHub Releases.
- macOS CI/release builds run `sh scripts/sign-macos-binary.sh BINARY` after all binary modifications, then native version/help/read/validate smoke checks. Signing and strict verification must precede upload; release checksums cover the final signed bytes.

## Docs layout

Three docs, each for a different reader — don't cross the streams. When you change one, check the others:

- **README.md** — user-facing: install, examples, command reference, "How It Works."
- **CONTRIBUTING.md** — dev-facing: setup/test commands, LibreOffice install, project-structure tree, CI table.
- **CLAUDE.md** _(this file + nested)_ — agent-facing: conventions, invariants, per-subsystem playbooks.

New CLI command → README + `src/cli/CLAUDE.md`. New invariant → CLAUDE.md only. New build/test step → CONTRIBUTING.md + this Testing section. New runtime dep → README + CONTRIBUTING.md. Installer / supply-chain change → **SECURITY.md** (the canonical statement of install integrity — a fourth doc, user-facing) + the README install section + `skills/docx-cli/references/troubleshooting.md`; keep the two installers' posture identical rather than documenting a divergence.

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.