MatrixFounder/Universal-skills/skills/pdf/SKILL.md
Use when the user asks to create, combine, split, preview, or extract content from PDF files. Triggers include "markdown to pdf", "html to pdf", "webarchive/MHTML to pdf", "export this deck to pdf", "mermaid in pdf", "math/LaTeX in pdf", "merge PDFs", "split a PDF", "pdf to markdown", "extract text from pdf", "fill AcroForm", "preview pdf as image", and similar PDF generation, conversion, or manipulation tasks.
Skill1 starsChanged 5 months ago
- Installs packages
---
name: pdf
description: Use when the user asks to create, combine, split, preview, or extract content from PDF files. Triggers include "markdown to pdf", "html to pdf", "webarchive/MHTML to pdf", "export this deck to pdf", "mermaid in pdf", "math/LaTeX in pdf", "merge PDFs", "split a PDF", "pdf to markdown", "extract text from pdf", "fill AcroForm", "preview pdf as image", and similar PDF generation, conversion, or manipulation tasks.
tier: 2
version: 1.0
license: LicenseRef-Proprietary
---
# pdf skill
**Purpose**: Give the agent a small, deterministic set of CLIs for the
common PDF operations: render Markdown to a well-typeset PDF, merge
PDFs, split them by page range or into individual pages, and (via
references) extract text or fill forms. Picking the right library on
the fly is the single biggest source of PDF bugs; delegating to
scripts that embed those choices removes the variance.
## 1. Red Flags (Anti-Rationalization)
**STOP and READ THIS if you are thinking:**
- "I'll use `pypdf` to extract the text." → **WRONG** for layout-dependent content. `pypdf`'s text extraction is famously unreliable on anything with columns or complex layout; use `pdfplumber`. See [references/library-selection.md](references/library-selection.md).
- "I'll improvise a `pdfplumber` script to convert this PDF to Markdown." → **WRONG** to improvise from scratch. Run `pdf_extract.py` for the structured dump and follow [references/pdf-to-markdown.md](references/pdf-to-markdown.md) — and never skip the scan check: on a scanned PDF `pdfplumber` returns empty text *silently*; `pdf_extract.py` exits `10` instead.
- "I'll reach for `playwright` for Markdown → PDF because it handles everything." → **WRONG**. Playwright pulls a 200 MB Chromium install. `weasyprint` handles 95% of Markdown/HTML inputs with a fraction of the footprint.
- "I'll fill this XFA form with `pypdf`." → **WRONG**. `pypdf` doesn't fill XFA — only AcroForm. Detect the form type first (see [references/forms.md](references/forms.md)) and fail loudly if it's XFA rather than silently writing an unchanged file.
- "I'll skip checking exit codes from `pdf_merge.py`." → **WRONG**. Missing an input file produces exit 1 and the output is absent — silently assuming success ships a broken deliverable.
- "This HTML carries its own `@media print` and `@page` rules, so `html2pdf.py` will honour them." → **PARTLY WRONG, and the two engines fail differently.** `weasyprint` *does* apply `@media print` and `@page` (measured: a 41-slide deck rendered 41 pages at its declared 297×176 mm) but supports no **container queries** — a layout whose type scale is built on `cqw` collapses into overlapping text on most pages. `--engine chrome` has the layout engine but forces `media: screen` and scales to fit the paper (§2), so an authored print stylesheet never applies at all. For a document *authored for print* — its own `@page` size, its own break rules, modern layout — drive headless Chrome directly (`--headless=new --print-to-pdf`), which keeps `media: print`, then verify with `preview.py`.
- "The PDF has the right page count, so the render is fine." → **WRONG**. Page count catches split pages and nothing else. Render every page with `preview.py` and look: web fonts silently fall back to system faces, and a `@media (max-width:…)` block with no `screen` qualifier also applies to print, collapsing multi-column layouts.
- "I'll grep the PDF bytes for `/BaseFont` to confirm the fonts embedded." → **WRONG**. Font descriptors live in compressed object streams, so the grep reports only the few uncompressed entries and shows system fallbacks even when the real fonts *are* embedded. This misreads a working file as broken. Judge embedding from a rendered page, or with `pypdf`.
- "`pdf_profile.py` measured the word-gap ratio, so I'll use the number it printed." → **PARTLY RIGHT — read the `source:` it prints.** On a file that marks enough word boundaries with space glyphs the ratio is `source: labelled`: scored against those boundaries, reported with how many words it would glue and split, and safe to take. Otherwise it is `source: peaks` — the midpoint between the two gap modes, which **fails at the boundary** on justified text, where the inter-word population runs down into its own left tail. Either way, self-check on the file: count tokens of 22+ consecutive letters in the rebuilt text; any hit that is prose rather than a URL means go lower. Measured on a 13 pt Times body where the two welded lines' own inter-word gaps run 1.57–1.60 pt against a `0.125 × 13 = 1.625` pt threshold (the document's tightest labelled inter-word gap is 1.24 pt) — the peak midpoint said `0.125` and welded those two prose lines into single tokens; the labelled pass says `0.09` and welds none. A one-bin `safe_band` is the profiler telling you there is no margin either way.
## 2. Capabilities
- Render Markdown (+ optional custom CSS) to a typeset PDF via `weasyprint`. Fenced ```mermaid blocks pre-render to PNG via `mmdc`; bundled `scripts/mermaid-config.json` ships an office-friendly Cyrillic-capable font stack (override with `--mermaid-config PATH`, opt out with `--no-mermaid-config`). Inline `$…$` and display `$$…$$` **math** pre-render to MathML via the bundled `katex` (weasyprint typesets MathML natively; it runs no JS, so client-side KaTeX/MathJax can't be used) — currency (`$5`) and `$` inside code/fences are left untouched, and without node/KaTeX formulas degrade to literal text (opt out with `--no-math`, fail-hard with `--strict-math`). The PDF carries a navigable **outline (bookmarks)** auto-built from `h1`–`h6` headings — no flag needed.
- **Render HTML / web archives to PDF** via `html2pdf.py` — same weasyprint pipeline, natively handles `.html`/`.htm`, `.mhtml`/`.mht`, and `.webarchive`. Validated across **Fern (OpenRouter), Mintlify (Anthropic Claude Code, Discord, Berachain), GitBook (Hyperliquid), Confluence (Atlassian wikis), Хабр, vc.ru** and generic blogs — 34/34 fixtures pass in both modes. Bundled stylesheet on by default; `--no-default-css` for fully-styled inputs (BI dashboards, branded reports); `--css EXTRA.css` stacks on top. `--reader-mode` extracts the main article body (Safari Reader View parity). `--timeout 180` SIGALRM watchdog with `$HTML2PDF_TIMEOUT` override. Universal preprocessing handles draw.io/Confluence SVG diagrams, table-based code blocks (Fern/Mintlify shiki), Tailwind/FontAwesome icon strip, ARIA-role tables (GitBook), ad-network removal, and pathological-CSS protection (Хабр content-drop bug, vc.ru CPU-loop bug). Full pipeline + flag semantics + per-platform notes documented in [references/html-conversion.md](references/html-conversion.md). Output PDFs carry a navigable **outline (bookmarks)** from `h1`–`h6` headings — engine-agnostic (weasyprint and `--engine chrome`; the chrome engine emits a tagged PDF, the mechanism Chromium uses for the outline).
- Merge multiple PDFs into one preserving bookmarks (`pdf_merge.py`).
- Split a PDF by explicit page ranges, one-per-page, or fixed-size chunks (`pdf_split.py`).
- **Stamp a text or image watermark on every (or selected) page** via `pdf_watermark.py` (drafts, "CONFIDENTIAL", brand stamps). `--position center|top-left|top-right|bottom-left|bottom-right|diagonal`, `--opacity`, `--rotation`, `--pages "1-5,8"`. Builds one overlay per unique page mediabox, so heterogeneous decks (Letter+A4) keep correct proportions.
- **Detect, inspect, and fill AcroForm fields** via `pdf_fill_form.py` — three modes: `--check` (form-type triage with exit codes 0/11/12 = AcroForm/XFA/none; `13` = the stdout reader went away), `--extract-fields` (dump field schema as JSON for editing), and fill mode (`INPUT.pdf DATA.json -o OUT.pdf [--flatten]`). XFA forms are detected and refused with a clear message.
- Extract text, tables, and layout via `pdfplumber` (documented; inline usage from the agent is fine).
- **Dump a PDF's per-page text + tables to structured JSON** via `pdf_extract.py` — a structured *dump*, NOT a Markdown converter (it never emits Markdown). Its defining feature is **scan detection**: an image-only document exits `10` with a `DocumentScanned` signal instead of silently yielding empty text. **Robust word-splitting by default**: LaTeX/academic two-column PDFs encode inter-word spacing as positional gaps (no space glyphs), which pdfplumber's absolute tolerance glues into `ASurveyonBlockchain`; a font-relative `x_tolerance_ratio` (default `0.15`) splits them correctly without regressing real-space PDFs (tune/disable via `--x-tolerance-ratio R`, see [references/pdf-to-markdown.md §3.8](references/pdf-to-markdown.md)). **Two further silent-loss signals at exit `0`**: `figure_pages` — a page that is mostly diagram with too little text to be a text page, which the absolute char threshold misses as soon as the page carries a running header (§3.4); and `text_layer_lossy` — the document embeds no fonts and every encoding is single-byte Latin, so any non-Latin text was destroyed when the file was written and **OCR cannot recover it** (§3.9). **Two grouping knobs** for pdfplumber defaults that misread real documents: `--y-tolerance PT` (a list marker in a smaller point size otherwise becomes its own line, sorted *after* its item — §3.1) and `--table-strategy lines_strict` (background shading otherwise becomes a phantom table, or a bogus extra row on a real one — §3.2). **Hyperlinks are content, and they used to vanish**: every page now carries `links` (`{uri, text, bbox}` per `/URI` annotation, with the text the link covers) and the top level a `link_count` — measured at 578 links across a 20-document corpus that the dump dropped entirely, which for a web page printed to PDF is losing content, not formatting (§5). **Two further advisory counters** in `layout_hints`: `split_table_rows` (a table row cut in half by a page break arrives with an empty first cell while its label stays behind on the previous page — §3.3) and `multi_column_pages` (a full-height column gutter, whose x coordinate the hint reports because the repair is a crop per column, and no flag performs that crop yet — `--layout` does *not* separate columns, so the coordinate is the argument you crop with by hand until `--columns` lands; §3.1, backlog pdf-16). **Artwork extraction is opt-in** via `--extract-images DIR`: embedded rasters come out as pypdf re-encodes their *decoded pixels* into a container that can hold them — original pixels at native resolution, never resampled, but not the stored bytes (measured across both dogfood documents and the raster fixtures: no placement came out byte-identical to its stream) — and vector figures (diagrams with no image object) are cropped from the page through Poppler, each listed per page as `images` so Markdown references real paths — the second half of the `figure_pages` repair ([references/pdf-to-markdown.md §3.10](references/pdf-to-markdown.md)). Pairs with [references/pdf-to-markdown.md](references/pdf-to-markdown.md) for the PDF→Markdown decision tree and recipe; final Markdown composition stays agent judgement.
- **Profile a digital PDF's typographic template** via `pdf_profile.py` — the first step of any PDF→Markdown job. One pass reports the point-size histogram with a body-text hypothesis and a sample line per size, the font inventory, the **measured** word-gap threshold with the `source:` that produced it — scored against the boundaries the file itself marks with space glyphs where it marks enough of them, and the midpoint between the two gap modes only where it does not, the ligature-duplicate count, the running header/footer band located by *position and repetition* (never by point size — a measured document set its legal chapter in the same 6 pt as its footer, and a size filter deleted the chapter), fill-box colours (callouts / code samples), a ruled-table diagnostic that says when `extract_tables()` finds nothing on pages that *are* tables, list markers and the indent step, icon-font glyphs, and a figure-vs-icon image census — each ending in a `Recommendations` block. Read-only; exits `0` even when findings are alarming.
- **Line-level dump** via `pdf_extract.py --lines` — the shape composition actually needs. Per line: `text` (ligatures deduped, positional word gaps restored), `bbox`, modal `size`, modal `font`, the `styles` present, `marker_only`, and per-run `font`/`size`/`style`/`uri` wherever a line is not uniform; plus per page `rects` (filled boxes) and `rules` (horizontal rules grouped by y with their segment x boundaries — the column signature of a ruled table). Heading level, list nesting and inline emphasis come from point size, x position and font family, none of which survives the flat page `text`. Off by default: a dump taken without it is byte-for-byte unchanged.
- **Verify a conversion kept the PDF's words** via `pdf_verify_md.py` — token coverage of a `.md` against the PDF or its dump, normalising both sides (Markdown syntax and escapes stripped, both sides de-hyphenated, running furniture detected and excluded, contents pages skipped) and reporting overall loss, the worst pages, and the most-missing tokens. `--max-loss PCT` turns it into a gate — `2` is the smallest integer both measured editorial documents pass on their honest floors (1.03 % and 0.27 %), and `3` is that plus a point of headroom; `1` exits 1 on the 1.03 % document, i.e. on a good conversion. **Read the most-missing tokens before believing the number**: after furniture and hyphenation are normalised the remainder is dominated by de-hyphenation fragments the converter repaired *correctly* and by column leakage on the reference side, so the percentage is a pointer, not a verdict ([references/pdf-to-markdown.md §8](references/pdf-to-markdown.md)). This is the only step that catches silent structural loss; it found two real defects in a measured 1206-page conversion (a corpus not shipped with this skill).
- **OCR a scanned (image-only) PDF into a searchable PDF** via `pdf_ocr.py` — wraps `ocrmypdf` to overlay an invisible OCR text layer (default languages **`eng+rus`**), the remediation hop for `pdf_extract.py` exit `10`. The OCR engine is **soft-optional**: install with `bash scripts/install.sh --with-ocr` (+ system tesseract/eng/rus/ghostscript); a missing engine or language pack fails loud, never silent. See [references/ocr.md](references/ocr.md).
- Render any `.pdf` (or peer-skill `.docx`/`.xlsx`/`.pptx`) into a single PNG-grid preview via `preview.py` (uses Poppler directly for `.pdf`; LibreOffice + Poppler for OOXML).
- Emit failures as machine-readable JSON to stderr with `--json-errors` (uniform across all four office skills).
## 3. Execution Mode
- **Mode**: `script-first` for the bundled operations, `prompt-first` with library references for extraction and form filling.
- **Why this mode**: The bundled operations (render, merge, split) are stable recipes. Extraction and form filling depend heavily on the specific document and deserve inspection before running — the references guide the inline work.
- **`pdf_extract.py` — the bounded exception**: extracting per-page text + tables to a JSON *dump* IS a stable recipe, so it is bundled (it also closes the silent-scan failure with code). Markdown *composition* — heading levels, reading order, table stitching — stays `prompt-first` agent judgement: there is no Markdown-converter script, by design. See [references/pdf-to-markdown.md](references/pdf-to-markdown.md).
## 4. Script Contract
- **Commands**:
- `python3 scripts/md2pdf.py INPUT.md OUTPUT.pdf [--page-size letter|a4|legal] [--css EXTRA.css] [--base-url DIR] [--no-mermaid] [--strict-mermaid] [--mermaid-config PATH | --no-mermaid-config] [--no-math] [--strict-math]`
- `python3 scripts/html2pdf.py INPUT OUTPUT.pdf [--page-size letter|a4|legal] [--css EXTRA.css] [--base-url DIR] [--no-default-css] [--reader-mode] [--archive-frame N|main|all|auto] [--list-frames] [--timeout SECONDS] [--engine weasyprint|chrome] [--chrome-js]` — INPUT may be `.html`/`.htm`, `.mhtml`/`.mht`, or `.webarchive`; sub-resources in archives are extracted to a temp dir automatically. `--reader-mode` extracts the main article content (Confluence-priority candidate list with body-ratio guard for `<main>`; longest-match per selector handles archive pages with multiple `.entry` divs and Disqus comment threads + title-match LCS bonus for multi-article feed pages), stripping navigation, ads, sidebars, and SPA chrome (ARIA `role=navigation\|complementary\|banner\|contentinfo` + semantic `<aside>`/`<nav>`/`<footer>` + shallow `<header>`) — ideal for browser-saved news/blog/docs pages and hydrated SPAs. **`--archive-frame N|main|all|auto` (pdf-8, 2026-05-05)**: selects which inner frame in webarchive/MHTML to render — `main` (default) = main resource only; `N` (1-indexed) = specific inner frame; `all` = concat all "substantial" frames (≥ 1 KB + 0 `<script>` + ≥ 30 chars text + not single-`<img>`-only) with `<hr><h2>Frame N</h2>` separators + per-frame namespace + sha1 image-dedup + encoding parity; `auto` = deterministic (0 substantial → main, 1 → that frame with main-dominance guard, 2+ → all). Vendor-agnostic: validated on 9 real fixtures across Angular/Closure/Framer/bare-DOM SPA stacks without a single vendor name in the heuristic. **`--list-frames` (pdf-8)**: prints inner-frame inventory (index/kind/substantial/bytes/scripts/text-len/url) and exits without rendering — for picking `N` deterministically. `--timeout` (default 180s, `$HTML2PDF_TIMEOUT` env, 0 disables) caps weasyprint render via `signal.SIGALRM` for pathological inputs. Exit 1 with `RenderTimeout` envelope on watchdog fire; exit 2 with `NoSubstantialFrames` / `FrameIndexOutOfRange` envelopes for archive-frame errors. **`--engine weasyprint|chrome` (pdf-11, 2026-05-05)**: render engine selector. `weasyprint` (default) — pure-Python typeset PDF, no browser runtime. `chrome` — opt-in headless Chromium via Playwright (~150 MB, install with `bash install.sh --with-chrome`). Use `chrome` when weasyprint produces broken output: Material 3 calc/var bugs (Gmail-class), Framer infinite layout loops, ELMA365 inline.py assertion, JS-hydrated content, `<canvas>` charts. Chrome path skips weasyprint preprocess (calc-strip, font-face-strip, NORMALIZE_CSS) — those are weasyprint workarounds Chrome doesn't need; reader-mode and `--css EXTRA.css` remain engine-agnostic. **Chrome render hardenings (universal layout strategy, post-VDD-iter-8 — 8 adversarial iterations)**: `<script>` strip from HTML + JS-enabled at context level (page can't run own JS → no Gmail self-destruct, no Angular half-hydration; we keep `page.evaluate` for surgical DOM normalization; `--chrome-js` opts page-JS back in for canvas/hydration); `<base href>` stripped (webarchives carry `<base href="https://orig-site/">` which would route every relative URL to the offline-blocked origin); media forced to `screen` (default `print` triggers nav-hiding `@media print` rules in SPAs); 1280×1024 viewport (desktop CSS); layout-normalize CSS — high-specificity body release, icon-font ligature suppression with `:not(:has(*))` leaf-only guard (avoids font-size:0 cascading through CSS inheritance to children), `[class~="spinner"]` exact-word match (not substring — avoids hiding `class="spinner-class-banner"` and similar), image cap (200px), avatar-image cap (48×48 only on `... img`, not bare class so containers aren't shrunk); JS-based DOM normalize via `page.evaluate` — width-gate `offsetWidth ≥ 200` for overflow release (narrow icon-sidebars at 64px stay clipped, no label leak), substantial-modal release (`position:fixed → static` only when wide AND tall AND text-rich), modal-portal hide (when modal released, hide non-portal body children to remove underlying CRM page); **scale-to-fit `page.pdf(scale = pdf_usable / viewport_width)` ≈ 0.561** so 1280 px layout fits A4's ~718 px usable width without right-edge cutoff. Validated on 3 SPA archive shapes × 2 modes (Gmail Closure email, ELMA365 Angular dashboard, Yandex Cloud Console marketplace) — all 6 combinations produce full content with no overlap, no cutoff, no chrome-icon ligatures, no underlying-page noise. **Recommended composition**: `--engine chrome --reader-mode` for email/newsletter/article archives (cleaner article-only render); `--engine chrome` alone for dashboards/registries/structured UIs (preserves card layout). Exit 1 with `ChromeEngineUnavailable` envelope if Playwright not installed.
- `python3 scripts/pdf_merge.py OUTPUT.pdf INPUT1.pdf INPUT2.pdf [INPUT3.pdf ...]`
- `python3 scripts/pdf_split.py INPUT.pdf --ranges "1-5:part1.pdf,6-10:part2.pdf"`
- `python3 scripts/pdf_split.py INPUT.pdf --each-page OUTDIR/`
- `python3 scripts/pdf_split.py INPUT.pdf --every N OUTDIR/`
- `python3 scripts/pdf_watermark.py INPUT.pdf OUTPUT.pdf (--text "DRAFT" | --image STAMP.png) [--opacity 0.3] [--position center|top-left|top-right|bottom-left|bottom-right|diagonal] [--rotation 45] [--font-size 60] [--color "#888"] [--scale 0.5] [--pages "all"|"1-5,8,12-end"]`
- `python3 scripts/pdf_fill_form.py --check INPUT.pdf` — exit 0/11/12 = AcroForm/XFA/none, 13 = stdout went away. (Custom codes start at 10 to leave 0–9 for argparse / shell convention. Every JSON payload this script puts on stdout goes out as UTF-8 bytes, so it does not change with the caller's locale — `--extract-fields -o FILE` writes its JSON to the file instead, which was already UTF-8. A reader that goes away — `| head`, or a consumer that exits first — ends any of those writes as exit `13` + an `OutputWriteFailed` envelope, never a traceback and never the interpreter's substituted 120; that covers the one human-readable status line `--extract-fields -o FILE` prints too. Not covered on the locale axis: that status line is presentation, not JSON, so it is still encoded with the caller's codec — a character the codec cannot render is escaped, not fatal.)
- `python3 scripts/pdf_fill_form.py --extract-fields INPUT.pdf -o fields.json`
- `python3 scripts/pdf_fill_form.py INPUT.pdf DATA.json -o OUTPUT.pdf [--flatten]`
- `python3 scripts/preview.py INPUT OUTPUT.jpg [--cols 3] [--dpi 110] [--gap 12] [--padding 24] [--label-font-size 14] [--soffice-timeout 240] [--pdftoppm-timeout 60]`
- `python3 scripts/pdf_extract.py INPUT.pdf [-o OUT.json] [--layout] [--password PW] [--x-tolerance-ratio R] [--y-tolerance PT] [--table-strategy lines|lines_strict] [--extract-images DIR] [--image-dpi N] [--no-vector-images] [--json-errors]` — dumps per-page text + tables as structured JSON (NOT Markdown). `--x-tolerance-ratio` (default `0.15`) is the font-relative word-split threshold that un-glues LaTeX/academic PDFs; `0` disables it (legacy absolute tolerance). `--y-tolerance` (default: pdfplumber's 3 pt) is the absolute line-grouping threshold — raise it to `5` when a smaller-point-size list marker is split onto its own line and sorted after its item. `--table-strategy` (default `lines`) selects the table-edge strategy; `lines_strict` ignores fill-only background rectangles, which otherwise fabricate tables out of shading. The dump names its input in top-level `source` (resolved, so a relocated dump still says which PDF it describes), echoes all three effective values plus a document-level `fonts` list, and carries two further exit-`0` loss signals — `figure_pages` (page is mostly artwork: per-page `image_coverage` / `vector_coverage` ≥ `0.25` with `< 200` chars) and `text_layer_lossy` (no embedded font, no `/ToUnicode`, all-Latin encodings → non-Latin text was lost at export; OCR does **not** help). Both warn on stderr and leave the exit code alone. Every page also carries **`links`** — one `{uri, text, bbox}` record per `/URI` annotation, `text` being the characters the link covers (`null` over an image; internal `/GoTo` links are deliberately not reported) — and the top level a **`link_count`**. **`layout_hints`** (always present) counts four things: `orphan_list_markers` (a bullet in a smaller point size becomes its own line and sorts *after* its item), `single_column_tables` (background shading read as a table), `split_table_rows` + `split_table_pages` (a row severed by a page break: its continuation opens the next page's table with an empty first cell, and the stranded label — `2.11` — is named where the previous page's text still holds it; the dump does NOT stitch, that is composition) and `multi_column_pages` + `column_probe` (a full-height gutter, reported with its x coordinate and with what cropping there measured, because `--layout` preserves the interleaving rather than undoing it). While the matching knob is still at its default, the script **tries that knob on up to three affected pages and reports what it measured** — `--y-tolerance 5` reunited N of M, or changed nothing (an arXiv export interposes a line between marker and item; a Confluence export's gap is wider), `--table-strategy lines_strict` removed N of M, or kept them all (a real one-column table, a ruled box, or a fragment of a split wider table). Advisory only; the exit code never moves, the counts stay in the dump when the line is suppressed, and the probe costs ≤0.12 s and does not run when no hint would fire. **`--extract-images DIR`** writes the document's artwork out and lists it per page as `images` (`{file, kind, bbox, name, width, height, bytes, sha1}`) — **two classes**: embedded **rasters**, written as pypdf re-encodes their decoded pixels (`.png`/`.jpg`/`.tif`/`.jp2` names pypdf's *output* format, not the PDF's stored filter — the file is routinely several times larger than the stream, and a lossy source can be re-encoded lossily; what is preserved is the pixels, at native resolution, never resampled), and **vector** figures (diagrams drawn with path operators, for which no image object exists) cropped from the page via Poppler at `--image-dpi` (default `150`). Classification is by object model, never by looks: a block diagram that *looks* vector is often an RGBA PNG, and the raster branch covers it. Identical images are written once (sha1 dedup — one measured document placed a single backdrop 31 times), page-sized rasters are skipped (a background wash, or a scan whose repair is OCR), a raster declaring more than 80 MP is refused *before* decoding (pypdf allocates from the declared `/Width`x`/Height`), and three further rejections keep the directory usable: an **inline glyph** (an emoji from a colour font is a raster XObject — square, as tall as its text line and adjacent to it; 10 of 15 files on one measured document), a **flat-colour raster** (a blank page's only image, which also puts that page in top-level `blank_pages` so the scan warning can say there is nothing to OCR) and a **vector cluster enclosing the page's body text** (table ruling whose crop is a picture of the page — 345-385 KB of it on four measured pages; `preview.py` is the fallback). `DIR` is mandatory, and what did NOT come out is reported in `images_summary` + stderr — including the case where **nothing** came out, which is the normal outcome for an OCR'd scan (every page is one page-sized raster, the backdrop rule drops it): the run says so instead of leaving an empty directory and a silent stderr. Every `bbox` — raster and vector alike — is in page coordinates. **The one silent omission**: a vector figure drawn entirely with fills and no stroked path is not extracted and its page is not `figure_dominant` either (a flat pie measures 0.07 against the 0.25 threshold) — render the sheet with `preview.py` for those. Without the flag the dump is unchanged (no `images` key at all). Exit codes: `0` success; `1` failure (missing / not-a-PDF / corrupt / encrypted-without-password / `--extract-images` names a file); `2` usage error; `6` `SelfOverwriteRefused` (`-o` **or `--extract-images`** resolves to the input PDF); **`10` `DocumentScanned`** — the whole document is image-only, run OCR or read the pages as images. On exit 10 the dump is still emitted; exit 10 + stderr is the loud signal. Default output is stdout; `-o` writes a file (idempotent). The stdout dump is UTF-8 whatever the locale says, and a pipe closed early (`| head`) exits with the code its envelope declares. See [references/pdf-to-markdown.md](references/pdf-to-markdown.md).
- `python3 scripts/pdf_profile.py INPUT.pdf [--pages N] [--columns SPEC] [--json] [--password PW] [--json-errors]` — profile the typographic template before converting. `--columns` scopes the furniture probe to page regions, taking either bare increasing cuts (`"152.5,456.5"`) or named ranges tiling the page width (`"side:0-158,body:158-458,notes:458-612"`); a bad spec is exit `2`. Without it the probe runs page-wide and says `region: page` — on a multi-column page, lines welded across a gutter never repeat identically, so patterns are under-counted (measured on a 3-column magazine: 1 pattern page-wide, 4 scoped). Exit `0` profiled, `1` unreadable/encrypted, `2` usage.
- `python3 scripts/pdf_verify_md.py SOURCE OUT.md [--max-loss PCT] [--top N] [--password PW] [--json] [--json-errors]` — `SOURCE` is the PDF or a `pdf_extract.py` dump (`.json`, much faster). Exit `0` checked, `1` unreadable or loss above `--max-loss` (`CoverageBelowThreshold`), `2` usage.
- `python3 scripts/pdf_ocr.py INPUT.pdf OUTPUT.pdf [--lang eng+rus] [--skip-text|--redo-ocr|--force-ocr] [--sidecar OUT.txt] [--jobs N] [--password PW] [--deskew] [--rotate-pages] [--clean] [--json-errors]` — OCR a scanned PDF into a searchable PDF via `ocrmypdf` (default languages `eng+rus`). `--password` decrypts an encrypted input; `--rotate-pages` needs tesseract `osd` data; `--clean` needs `unpaper`. Exit codes: `0` success; `1` failure (`type` in the envelope: `OcrEngineUnavailable` / `LanguagePackMissing` / `EncryptedInput` / `InputUnreadable` / `PriorOcrFound` / `OutputWriteFailed` / `InputNotFound`); `2` usage; `6` `SelfOverwriteRefused`. **Soft-optional engine** — `bash scripts/install.sh --with-ocr` first. See [references/ocr.md](references/ocr.md).
- All scripts above accept `--json-errors` to emit failures as a single line of JSON on stderr (`{v, error, code, type?, details?}`). The schema version `v` is currently `1`; argparse usage errors are routed through the same envelope (`type:"UsageError"`).
- **Inputs**: positional paths; optional flags per command.
- **Outputs**: single PDF files (`md2pdf`, `pdf_merge`) or multiple PDFs under a directory (`pdf_split`). All stdout goes to the output path list.
- **Failure semantics**: non-zero exit on missing inputs, invalid range specs, or library errors. Error detail to stderr.
- **Idempotency**: all three scripts overwrite their outputs on re-run.
- **Dry-run support**: not applicable.
## 5. Safety Boundaries
- **Allowed scope**: only paths named on the command line.
- **Default exclusions**: do not fetch remote resources unless the user explicitly provides URLs; `md2pdf.py --base-url` defaults to the input's directory.
- **Destructive actions**: all three scripts overwrite their outputs without prompting.
- **Optional artifacts**: custom CSS via `md2pdf.py --css` is optional; defaults produce a reasonable layout.
## 6. Validation Evidence
- **Local verification**:
- `python3 -m venv .venv && source .venv/bin/activate && pip install -r scripts/requirements.txt` — installs pypdf, pdfplumber, weasyprint, markdown2, reportlab.
- `bash scripts/tests/test_e2e.sh` — runs the end-to-end smoke suite (md2pdf, merge, split, fill-form, mermaid, pdf_extract, pdf_ocr). Includes the html2pdf regression battery: ~37 unit tests for `html2pdf_lib/` helpers + data-driven fixture battery (6 synthetic micro-fixtures + 6 hand-stripped real-platform slices + N tmp/ originals when present on disk; per-fixture page-count / size / required+forbidden-needle assertions, see [tests/battery_signatures.json](scripts/tests/battery_signatures.json)).
- **Adding a new platform fixture** (e.g. you found a Notion/Stripe page that breaks): drop the `.webarchive`/`.html`/`.mhtml` file into `tmp/`, run `python3 scripts/tests/capture_signatures.py` (auto-captures page count + needles + size band; only ADDs new fixtures unless `--refresh` is passed), hand-add chrome strings to `forbidden_needles` in [battery_signatures.json](scripts/tests/battery_signatures.json), commit the JSON delta. Total ~5 min per new site. Detailed in [references/html-conversion.md](references/html-conversion.md) §Regression coverage.
- `python3 scripts/md2pdf.py examples/fixture.md /tmp/invoice.pdf --page-size letter` — produces a non-empty PDF.
- `python3 -c "from pypdf import PdfReader; r=PdfReader('/tmp/invoice.pdf'); print(len(r.pages))"` — returns at least 1.
- `python3 scripts/pdf_merge.py /tmp/merged.pdf /tmp/invoice.pdf /tmp/invoice.pdf && python3 -c "from pypdf import PdfReader; print(len(PdfReader('/tmp/merged.pdf').pages))"` — 2× the page count.
- `python3 scripts/pdf_split.py /tmp/invoice.pdf --each-page /tmp/pages/` — produces `/tmp/pages/invoice-001.pdf`.
- **Expected evidence**: `/tmp/invoice.pdf`, `/tmp/merged.pdf`, `/tmp/pages/invoice-001.pdf`.
- **CI signal**: `python3 ../../.claude/skills/skill-creator/scripts/validate_skill.py skills/pdf` — exit 0.
## 7. Instructions
### 7.1 Pick the library, not the script first
A full PDF→Markdown *converter* is deliberately not bundled — Markdown
composition (heading levels, reading order, stitching a table across pages) is
agent judgement. Form filling likewise depends on the document.
The corollary matters as much: everything *around* that judgement is
mechanical and IS scripted, so do not hand-roll it. Profile the template
(`pdf_profile.py`), take the line-level dump (`pdf_extract.py --lines`) so
size / position / font are available to judge with, and verify the result
(`pdf_verify_md.py`) — composition is the step that loses content silently.
1. Check [references/library-selection.md](references/library-selection.md) for which library matches the task.
2. For **PDF → Markdown**: follow [references/pdf-to-markdown.md](references/pdf-to-markdown.md) — its decision tree picks digital-vs-scanned, §1.1 profiles the template, `pdf_extract.py --lines` gives the structured line-level dump, and §8 verifies the result. You compose the Markdown from that dump; the script never emits Markdown.
3. For other extraction (a one-off text/table grab): write inline `pdfplumber` code, or run `pdf_extract.py` for a quick structured dump.
4. For form filling: follow [references/forms.md](references/forms.md) — detect AcroForm vs XFA first.
### 7.2 Creating PDFs from Markdown
1. `python3 scripts/md2pdf.py input.md output.pdf` covers the common case.
2. Pass `--css custom.css` when the user provides brand styling.
3. For images referenced with relative paths, either put them next to the Markdown file or pass `--base-url /absolute/image/root`.
4. For HTML-heavy inputs (embedded `<style>`, flexbox, columns), weasyprint handles those in the script — no extra work needed.
### 7.3 Merging PDFs
1. Order matters: `python3 scripts/pdf_merge.py out.pdf file1.pdf file2.pdf file3.pdf` appends in that order.
2. Bookmarks from each input are preserved and nested under a parent named after the source's stem.
### 7.4 Splitting PDFs
Three modes, exclusive:
- `--ranges "1-3:intro.pdf,4-8:body.pdf,9-12:appendix.pdf"`
- `--each-page OUTDIR/` — one PDF per input page, zero-padded filenames.
- `--every N OUTDIR/` — chunks of N pages each.
Page numbers are 1-indexed and inclusive. Invalid ranges exit 1.
### 7.5 Setup
1. **MUST** run `bash scripts/install.sh` once. It creates `scripts/.venv/` locally, installs `requirements.txt`, probes whether weasyprint can find its native libraries, and prints install hints if not. Idempotent.
2. **External system libraries** (checked by `install.sh`, installed manually per project plan §3.3 "внешние инструменты — не бандлятся"):
- **pango, cairo, gdk-pixbuf** — weasyprint native runtime; required by `md2pdf.py`. macOS: `brew install pango gdk-pixbuf libffi`. Debian: `sudo apt install libpango-1.0-0 libpangoft2-1.0-0 libharfbuzz0b libcairo2 libgdk-pixbuf2.0-0`. See [references/weasyprint-setup.md](references/weasyprint-setup.md) for fuller notes.
- **tesseract (+ eng/rus data) and ghostscript** — only for `pdf_ocr.py`; installed by `bash scripts/install.sh --with-ocr` (which installs `ocrmypdf` into the venv and probes these). macOS: `brew install tesseract tesseract-lang ghostscript`. Debian: `sudo apt install tesseract-ocr tesseract-ocr-eng tesseract-ocr-rus ghostscript`. See [references/ocr.md](references/ocr.md).
Commands that need them fail with a clear error until installed.
## 8. Workflows (Optional)
PDF → Markdown (the long-document workflow):
```markdown
- [ ] `python3 scripts/pdf_profile.py in.pdf` # template: sizes, furniture, tables
- [ ] `python3 scripts/pdf_extract.py in.pdf --lines --extract-images out/img -o /tmp/dump.json`
- [ ] Read the loss signals (scanned_pages / figure_pages / text_layer_lossy) and the profile's recommendations
- [ ] Measure the column map: the x range and the ROLE of each column (body / margin apparatus / notes). Crop each and run the lines builder per crop — layout_hints.multi_column_pages does not see an asymmetric page (§3.1)
- [ ] Measure the line pitch PER COLUMN: top-to-top gaps, header band excluded. It is bimodal; the second mode is the paragraph threshold. Bottom-to-top is wrong — a descender or a superscript moves the bottom
- [ ] Compose the Markdown from `lines` + `rules` + `rects` (judgement: heading levels, reading order, table stitching). Record both measurements in the run log so page 2 is not re-guessed
- [ ] `python3 scripts/pdf_verify_md.py /tmp/dump.json out.md` # then read worst_pages
- [ ] `python3 scripts/preview.py in.pdf preview.jpg` for any page the verifier flags
```
Markdown-driven PDF:
```markdown
- [ ] Draft the Markdown content
- [ ] `python3 scripts/md2pdf.py doc.md doc.pdf`
- [ ] Open the PDF, check layout (orphans/widows, table page breaks)
- [ ] Iterate on CSS if needed (`--css brand.css`)
```
Merge + split for distribution:
```markdown
- [ ] `python3 scripts/pdf_merge.py combined.pdf intro.pdf body.pdf appendix.pdf`
- [ ] `python3 scripts/pdf_split.py combined.pdf --each-page out/` (if per-page delivery is needed)
- [ ] Verify page count with pypdf or Preview
```
Extract text (inline, no bundled script):
```markdown
- [ ] Read references/library-selection.md, pick pdfplumber
- [ ] Inline: open the file, call page.extract_text(layout=True)
- [ ] For tables, page.extract_tables() with appropriate snap_tolerance
```
## 9. Best Practices & Anti-Patterns
| DO THIS | DO NOT DO THIS |
| :--- | :--- |
| Use `weasyprint` for Markdown/HTML → PDF. | Reach for `playwright` unless you actually need JS/modern CSS. |
| Use `pdfplumber` for text/table extraction. | Trust `pypdf.extract_text()` on column layouts — output is often garbled. |
| Detect AcroForm vs XFA before filling. | Try to fill XFA with `pypdf` and ship an unchanged file. |
| Pass `--base-url` so relative images resolve. | Assume weasyprint reads relative paths the same way your shell does. |
| Check exit codes of the bundled scripts. | Assume success because no exception was raised. |
| Render **every** page with `preview.py` before calling a PDF correct. | Judge a multi-page PDF from page 1 — `sips` and Quick Look convert only the first page. |
| Drive headless Chrome directly for HTML authored *for print* (own `@page`, own break rules, container queries). | Reach for `html2pdf.py` reflexively: weasyprint drops container queries, and `--engine chrome` overrides the document's print media. |
### Rationalization Table
| Agent Excuse | Reality / Counter-Argument |
| :--- | :--- |
| "All PDFs can be read with the same library." | Reading vs creation vs editing vs rendering are four different problem spaces; pick per task. |
| "The Markdown renderer doesn't matter, they're all similar." | `weasyprint` supports `@page` and page-break-inside; `markdown-pdf` and `mdpdf` don't. |
| "My script worked on one PDF, it'll work on all of them." | PDFs are wildly heterogeneous — scanned, image-only, XFA, flattened. Always test on the actual file. |
## 10. Quick Reference
| Task | Command |
|---|---|
| Markdown → PDF | `python3 scripts/md2pdf.py doc.md doc.pdf --page-size letter` |
| Markdown → PDF with custom mermaid theme | `python3 scripts/md2pdf.py doc.md doc.pdf --mermaid-config theme.json` |
| HTML → PDF | `python3 scripts/html2pdf.py report.html report.pdf` |
| HTML → PDF (skip bundled CSS, only embedded styles) | `python3 scripts/html2pdf.py dashboard.html out.pdf --no-default-css` |
| Web page / archive → PDF (reader mode, strips nav/ads) | `python3 scripts/html2pdf.py page.webarchive article.pdf --reader-mode` |
| List inner frames in webarchive/MHTML (pdf-8) | `python3 scripts/html2pdf.py --list-frames email-thread.webarchive` |
| Render only inner frame N (1-indexed, e.g. one email from a thread) | `python3 scripts/html2pdf.py --archive-frame 1 thread.webarchive single.pdf` |
| Render all "substantial" inner frames concatenated (e.g. full email thread) | `python3 scripts/html2pdf.py --archive-frame all thread.webarchive thread.pdf` |
| Auto-pick frame strategy (0 substantial → main, 1 → that, 2+ → all) | `python3 scripts/html2pdf.py --archive-frame auto archive.webarchive out.pdf` |
| HTML → PDF with custom render deadline | `python3 scripts/html2pdf.py page.html out.pdf --timeout 300` (or `HTML2PDF_TIMEOUT=300 python3 …`) |
| HTML → PDF, watchdog disabled (large book webarchive) | `HTML2PDF_TIMEOUT=0 python3 scripts/html2pdf.py book.webarchive book.pdf` |
| HTML → PDF via headless Chrome — for dashboards/registries (pdf-11; install with `bash install.sh --with-chrome`) | `python3 scripts/html2pdf.py registry.webarchive out.pdf --engine chrome` |
| Chrome + reader-mode — recommended for email/newsletter/article archives | `python3 scripts/html2pdf.py email.webarchive out.pdf --engine chrome --reader-mode` |
| Chrome engine with JavaScript on (rare; for canvas charts or pre-hydration HTML) | `python3 scripts/html2pdf.py page.webarchive out.pdf --engine chrome --chrome-js` |
| Merge PDFs | `python3 scripts/pdf_merge.py out.pdf a.pdf b.pdf c.pdf` |
| Split by ranges | `python3 scripts/pdf_split.py in.pdf --ranges "1-5:intro.pdf,6-10:body.pdf"` |
| Split one-per-page | `python3 scripts/pdf_split.py in.pdf --each-page pages/` |
| Split in chunks of N | `python3 scripts/pdf_split.py in.pdf --every N out/` |
| Text watermark on every page | `python3 scripts/pdf_watermark.py in.pdf out.pdf --text "DRAFT"` |
| Image watermark, bottom-right corner | `python3 scripts/pdf_watermark.py in.pdf out.pdf --image stamp.png --position bottom-right --scale 0.2` |
| Watermark only specific pages | `python3 scripts/pdf_watermark.py in.pdf out.pdf --text CONFIDENTIAL --pages "1-5,8"` |
| Inspect AcroForm fields | `python3 scripts/pdf_fill_form.py --check form.pdf` |
| Extract field schema as JSON | `python3 scripts/pdf_fill_form.py --extract-fields form.pdf -o fields.json` |
| Fill AcroForm from JSON | `python3 scripts/pdf_fill_form.py form.pdf data.json -o filled.pdf [--flatten]` |
| Preview as PNG-grid | `python3 scripts/preview.py file.pdf preview.jpg [--cols 3] [--dpi 110]` |
| **Profile a PDF's template (do this first)** | `python3 scripts/pdf_profile.py in.pdf` |
| Dump PDF text + tables to JSON | `python3 scripts/pdf_extract.py in.pdf -o dump.json` |
| Dump with the **line-level** view (sizes, fonts, runs, rules, boxes) | `python3 scripts/pdf_extract.py in.pdf --lines -o dump.json` |
| **Verify a conversion lost nothing** | `python3 scripts/pdf_verify_md.py dump.json out.md` |
| Same, as a gate | `python3 scripts/pdf_verify_md.py in.pdf out.md --max-loss 3` |
| PDF → Markdown **with images** | `python3 scripts/pdf_extract.py in.pdf --extract-images out/img > /tmp/dump.json` — the dump is an intermediate, keep it out of the delivered folder; the image dir must sit next to the `.md` (relative links) |
| OCR a scanned PDF (eng+rus) | `python3 scripts/pdf_ocr.py scan.pdf scan.ocr.pdf` (needs `install.sh --with-ocr`) |
| PDF → Markdown (approach + recipe) | follow [references/pdf-to-markdown.md](references/pdf-to-markdown.md) |
| Machine-readable failures | append `--json-errors` to any of the above |
## 11. Examples (Few-Shot)
Fixture: [examples/fixture.md](examples/fixture.md).
**Input** — user request:
> Turn this invoice Markdown into a letter-sized PDF.
**Output** — agent action:
```bash
python3 scripts/md2pdf.py invoice.md invoice.pdf --page-size letter
```
**Input** — user request:
> Join these three quarterly reports into one annual PDF.
**Output** — agent action:
```bash
python3 scripts/pdf_merge.py annual.pdf q1.pdf q2.pdf q3.pdf q4.pdf
```
**Input** — user request:
> Split the 120-page handbook into chapters of roughly 10 pages each.
**Output** — agent action:
```bash
python3 scripts/pdf_split.py handbook.pdf --every 10 chapters/
```
**Input** — user request:
> Convert this Confluence-export HTML report to PDF.
**Output** — agent action:
```bash
python3 scripts/html2pdf.py q1-report.html q1-report.pdf --base-url ./q1-report_files/
```
**Input** — user request:
> Stamp every page of this draft contract with "DRAFT" diagonally.
**Output** — agent action:
```bash
python3 scripts/pdf_watermark.py contract.pdf contract-draft.pdf --text "DRAFT"
```
## 12. Resources
- [references/library-selection.md](references/library-selection.md) — which PDF library for which task, installation shortcuts.
- [references/pdf-to-markdown.md](references/pdf-to-markdown.md) — PDF → Markdown: decision tree (digital vs scanned), extraction recipe, pitfalls (multi-column, borderless tables, cross-page tables, headings), and why Markdown composition stays agent judgement.
- [references/forms.md](references/forms.md) — AcroForm vs XFA, filling with pypdf, flattening, visual overlay fallback.
- [references/weasyprint-setup.md](references/weasyprint-setup.md) — install platform notes, `@page` recipes, font embedding, page breaks.
- [references/html-conversion.md](references/html-conversion.md) — `html2pdf.py` deep dive: 10-step preprocessing pipeline, `_NORMALIZE_CSS` rules, reader-mode candidate list with body-ratio guard, render-time hardening (offline URL fetcher + SIGALRM watchdog), PDF outline (bookmarks) from headings, per-platform notes (Fern / Mintlify / GitBook / Confluence / Хабр / vc.ru), honest-scope limitations.
- [scripts/md2pdf.py](scripts/md2pdf.py) — Markdown → PDF via weasyprint + markdown2; mermaid blocks pre-rendered to PNG via `mmdc`.
- [scripts/html2pdf.py](scripts/html2pdf.py) — HTML → PDF via the same weasyprint pipeline; reuses md2pdf's default stylesheet (opt-out via `--no-default-css`).
- [scripts/pdf_merge.py](scripts/pdf_merge.py) — bookmark-preserving merger via pypdf.
- [scripts/pdf_split.py](scripts/pdf_split.py) — range, per-page, or fixed-chunk splitter.
- [scripts/pdf_watermark.py](scripts/pdf_watermark.py) — text/image watermark overlay via reportlab + pypdf; per-mediabox overlay caching for heterogeneous decks; cross-7 same-path guard.
- [scripts/pdf_fill_form.py](scripts/pdf_fill_form.py) — AcroForm inspect/extract/fill/flatten via pypdf; XFA forms detected and refused.
- [scripts/preview.py](scripts/preview.py) — universal `INPUT → PNG-grid` renderer for `.pdf` (via Poppler) and `.docx`/`.xlsx`/`.pptx` (via LibreOffice + Poppler). Byte-identical across all four office skills.
- [scripts/pdf_profile.py](scripts/pdf_profile.py) — typographic-template profiler: sizes, fonts, measured word-gap threshold, ligature duplicates, running furniture (by position + repetition), fill boxes, ruled-table diagnostic, list indents, icon glyphs, figure-vs-icon images, plus recommendations. Read-only.
- [scripts/pdf_verify_md.py](scripts/pdf_verify_md.py) — token-coverage check of a Markdown conversion against the PDF or its dump; normalises Markdown syntax, escapes, hyphenation and running furniture before comparing; `--max-loss` gates.
- [scripts/_textlines.py](scripts/_textlines.py) — shared character-level line reconstruction (ligature dedupe, measured positional-space insertion, style runs, rule signatures, fill boxes). Imported by `pdf_extract.py --lines` and `pdf_profile.py`.
- [scripts/pdf_extract.py](scripts/pdf_extract.py) — dumps a PDF's per-page text + tables to structured JSON via `pdfplumber`, with scan detection (image-only document → exit `10`) and opt-in artwork extraction (`--extract-images DIR`: rasters re-encoded by pypdf from their decoded pixels + vector figures cropped via Poppler). A dump, not a Markdown converter.
- [scripts/pdf_ocr.py](scripts/pdf_ocr.py) — OCR a scanned PDF into a searchable PDF via `ocrmypdf` (default `eng+rus`); soft-optional engine (`install.sh --with-ocr`); imports `_errors.py` read-only (no cross-skill replication). See [references/ocr.md](references/ocr.md).
- [scripts/mermaid-config.json](scripts/mermaid-config.json) — bundled office-friendly mermaid config (Cyrillic-capable font stack, auto-applied unless overridden via `--mermaid-config`).
- [scripts/_errors.py](scripts/_errors.py) — `--json-errors` envelope helper (schema `v=1`).
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

