WebReaper
alex-on-ai/WebReaper/CLAUDE.md
WebReaper — declarative, parallel web scraper/crawler library for .NET, published as the WebReaper NuGet package. Library lives in WebReaper/; Examples/ and Misc/ are consumers, not packaged. All projects target net10.0; global.json pins SDK 10.0.100 (rollForward: latestMinor, no prerelease) and .github/workflows/CI.yml installs 10.0.x — aligned. Build with a .NET 10 SDK. CI runs dotnet restore → dotnet build → unit tests over the whole solution (Examples/ and Misc/ included), so verify with a full-solution build, not just the library project. Fluent…
- Reads credentials
What's in it
- CLAUDE.md
- What this is
- Commands
- Test reality
- Toolchain
- Architecture
- Gotchas
# CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
## What this is
WebReaper — declarative, parallel web scraper/crawler library for .NET, published as the `WebReaper` NuGet package. Library lives in `WebReaper/`; `Examples/` and `Misc/` are consumers, not packaged.
## Commands
```bash
dotnet build WebReaper.sln # build all
dotnet test # all tests (xUnit)
dotnet test WebReaper.Tests/WebReaper.UnitTests/WebReaper.UnitTests.csproj # unit only (fast, offline)
dotnet test --filter "FullyQualifiedName~SimpleTest" # single test by name
dotnet run --project Examples/WebReaper.ConsoleApplication # run an example scraper
```
### Test reality
- **Unit tests** (`WebReaper.Tests/WebReaper.UnitTests`) — offline, parse `TestData/TestPage.html`. Use these for fast iteration.
- **Integration tests** (`WebReaper.Tests/WebReaper.IntegrationTests`) — hit the **live** site `alexpavlov.dev`, launch real Puppeteer/Chromium, and assert after fixed `Task.Delay` (up to 25s). Slow and network-flaky by design; failures there are often environmental, not regressions.
### Toolchain
All projects target **net10.0**; `global.json` pins SDK `10.0.100` (`rollForward: latestMinor`, no prerelease) and `.github/workflows/CI.yml` installs `10.0.x` — aligned. Build with a .NET 10 SDK. CI runs `dotnet restore` → `dotnet build` → unit tests over the **whole solution** (Examples/ and Misc/ included), so verify with a full-solution build, not just the library project.
## Architecture
Fluent builder → immutable config + pluggable spider → parallel job loop.
**Build path (ADR-0025, widened by ADR-0040 + ADR-0067):** a scrape begins with a **Crawl seed**. `ScraperEngineBuilder.Crawl(urls)` / `.CrawlWithBrowser(urls)` are **static** → return `ICrawlSeed`; the seed has four strategy terminals: `.Extract(schema)` (Schema-driven JSON via the deterministic fold), `.AsMarkdown()` (no-schema LLM-ready Markdown via `MarkdownContentExtractor`, ADR-0040), `.ExtractInferred(goal?)` (runtime schema inference via `LearnedSchemaContentExtractor` + `ISchemaInferrer`, ADR-0067), or `.ExtractWithPrompt(chatClient, instruction)` (schema-free LLM extraction via `PromptContentExtractor`, ADR-0084), all yielding the `ScraperEngineBuilder`. The builder's constructor is `internal` (test-only via `InternalsVisibleTo`), so it is unreachable without the seed: "build with no start URLs or no extraction strategy" is structurally impossible, not a runtime throw. It composes two internal builders:
- `ConfigBuilder` (internal) → immutable `ScraperConfig` (start URLs, the `LinkPathSelector` chain, crawl limit, headless flag).
- `SpiderBuilder` (internal) → runtime components (loaders, parsers, sinks, trackers, cookie storage), all with in-memory defaults.
Terminate with `BuildAsync()` (persists the config to `IScraperConfigStorage`, constructs a `Spider`, returns a `ScraperEngine`) or `Build()` (just the `ScraperConfig`, for a distributed start endpoint to persist). The ADR-0009 distributed-worker reduced shell is a **separate public seam**: `DistributedSpiderBuilder` (seedless, no `BuildAsync`; `.BuildSpider()` → `ISpider`) — "two seams, not one bug".
**Run path (ADR-0022):** `ScraperEngine` *is* the in-process **Crawl driver**. `RunAsync` seeds the `IScheduler` with one `Job` per start URL, then drives `Parallel.ForEachAsync` over `Scheduler.GetAllAsync()` (an async stream). Each job → `Spider.CrawlAsync` (wrapped in the **Retry policy** — `IRetryPolicy`, ADR-0026; the default `FixedAttemptsRetryPolicy` runs four attempts and never retries `OperationCanceledException`) returns a closed `JobReport`; the driver interprets it — applies the visited-link idempotency authority (atomic test-and-set), runs `ParsedData` through the page-processor pipeline (`IPageProcessor`, ADR-0038) then fans the surviving record to the sinks, enqueues child jobs, drives the **Outstanding-work latch**. The **Stop rule** verdict is a *value*, never a thrown exception (`PageCrawlLimitException` was removed in 8.0.0); on a concluded crawl the driver ends its own consumption of the job stream — `IScheduler.Complete()` was removed (ADR-0037), since durable schedulers could not honour it. Authoritative: `docs/adr/0022-crawl-driver-and-outstanding-work-latch.md`, `docs/adr/0026-retry-policy-seam.md`, `docs/adr/0037-stop-ceases-consumption.md`.
**Job model:** `Job` is a record carrying the URL, an `ImmutableQueue<LinkPathSelector>`, parent backlinks, and `PageType` (Static vs Dynamic). The crawl-vs-parse decision is **not** stored on `Job` (`Job.PageCategory` was removed) — it is computed by `CrawlStep` from the selector chain (0 left ⇒ parse the target page; exactly 1 with pagination ⇒ paginate; else follow links) and returned as a closed `CrawlOutcome` sum (`Parsed | Followed | Paginated`). Chain length is the state machine: each step dequeues one selector. Authoritative: `CONTEXT.md` + `docs/adr/0001-crawl-outcome-closed-sum.md`.
**`Spider.CrawlAsync` per job (ADR-0022 — the Spider only *reports*):**
1. Load the page via the one `IPageLoader` (ADR-0004) — HTTP by default; the headless-browser transports ship in `WebReaper.Playwright` (Microsoft.Playwright SDK, modern default — ADR-0053) or `WebReaper.Cdp` (raw CDP, AOT-clean, bedrock for stealth backends — ADR-0052) — per `PageType`.
2. `CrawlStep.StepAsync` returns a closed `CrawlOutcome`; `Spider` wraps it as a `JobReport` and returns — nothing else. The **Crawl driver**, not the Spider, runs a `Parsed` outcome's `ParsedData` through the page-processor pipeline (`IPageProcessor`, ADR-0038) and fans it to every `IScraperSink`, and turns `Followed`/`Paginated` into child (and pagination) jobs.
3. Content extraction is the `IContentExtractor` seam (ADR-0039) — its one core adapter is the shared `SchemaFold<TNode>` fold over an `ISchemaBackend<TNode>` (CSS/HTML, XPath, or JSON backend); an alternative extraction strategy (e.g. an LLM extractor) implements `IContentExtractor` directly. Link extraction is the concrete `LinkExtractor` function (ADR-0036 — not a seam). Authoritative: `docs/adr/0002-schema-fold-and-node-backend-seam.md`, `docs/adr/0039-content-extractor-seam.md`.
**Pluggable seams** (public interface in `*/Abstract`). As of 9.0.0 (ADR-0023, the documented-contract surface) the `*/Concrete` implementations are `internal` — select a built-in via a `ScraperEngineBuilder` method (`.WriteToConsole()`, `.WithTextFileScheduler(...)`, a satellite's `.WithRedis*()`/`.WriteToMongoDb()`/… extension), or supply your own implementation of the public interface; you no longer `new` the concrete type. The in-memory defaults the DIY-distributed pattern wires by hand stay public. The table's "impls" are the conceptual options, reached through the builder:
| Seam | Default | Other impls |
|---|---|---|
| `IScheduler` | InMemory | File, Redis, AzureServiceBus |
| `IVisitedLinkTracker` | InMemory | File, Redis |
| `IScraperConfigStorage` | InMemory | File, MongoDb, Redis |
| `ICookiesStorage` | InMemory | File, MongoDb, Redis |
| `IAgentRunStore` | InMemory | File, Redis, MongoDb, Sqlite, Cosmos |
| `IScraperSink` | — | Console, Csv, JsonLines, MongoDb, Redis, Cosmos |
| page loaders | Http + Playwright + Cdp | proxy-rotating variants when `IProxyProvider` set; stealth-Chromium-fork backends via `WebReaper.Stealth.*` satellites composing on Cdp |
**Distributed mode:** swapping scheduler + config storage + link tracker to Redis or Azure Service Bus lets multiple workers / serverless functions share crawl state. See `Examples/WebReaper.AzureFuncs` and `Examples/WebReaper.DistributedScraperWorkerService`.
**Transports wave (ADR-0052..0055):** the browser-mode surface, modernised. The `WebReaper.Puppeteer` satellite (ADR-0009's named successor) is deleted; in its place ship two browser transports plus a per-backend stealth pattern plus a CLI policy:
- **`WebReaper.Cdp`** (ADR-0052): raw CDP `IPageLoadTransport` — AOT-clean (System.Net.WebSockets + System.Text.Json source-gen, no PuppeteerSharp/Playwright dep). Two builder overloads: `.WithCdpPageLoader(cdpUrl)` (connect-to-existing; BYO browser path) and `.WithCdpPageLoader(CdpLaunchOptions)` (launch-and-connect; spawns Chromium with `--remote-debugging-port=0`). Public `CdpLaunchHelpers` utility (PATH search, free-port spawn, CDP-connect-validate, teardown) — the shared layer every `WebReaper.Stealth.*` satellite composes on. The bedrock for **Browser backend** swaps.
- **`WebReaper.Playwright`** (ADR-0053): Microsoft.Playwright-backed `PlaywrightPageLoadTransport` — modern, multi-browser (Chromium default; Firefox/WebKit opt-in via `PlaywrightBrowser` enum), all seven `PageAction` arms (closes the ADR-0004 four-arm Puppeteer gap). `.WithPlaywrightPageLoader(browser?, opts?)`. Replaces the deleted Puppeteer satellite in the same release (the v10 major break).
- **`WebReaper.Stealth.CloakBrowser`** (ADR-0054): first per-backend stealth satellite — convention pattern (no shared interface; each fork has its own thin satellite). `.WithCloakBrowser(opts?)` finds/downloads the binary on first use (legal model = `playwright install` / `winget` — download from upstream, no redistribution), launches with the fork's recommended flags, wires the resulting CDP URL into `WebReaper.Cdp`. Future stealth forks (Patchright, Camoufox, …) ship as community `WebReaper.Stealth.X` satellites following the same pattern.
- **CLI policy** (ADR-0055): `WebReaper.Cli` bakes ONLY `WebReaper.Cdp` for browser support (AOT preserved). Layered auto-spawn (BYO `--browser-cdp-url` → system Chrome detection → managed Chromium from `webreaper browser install`). New subcommands: `webreaper browser install` / `webreaper stealth install`. Curated `KnownStealthBackends` static registry — community CLI integration is one small PR per backend.
**AI-native (ADR-0040..0069):** five composing pieces on top of the existing pipeline (ADR-0059..0064 deepen the satellite — one mechanism module under the four `Llm*` adapters, tool-calling on the brain + resolver, an outcome signal for the brain, a real `ISchemaValidator` seam, an `HtmlToMarkdown` primitive, and a `.UseAi(client)` one-line aggregator; ADR-0065 + ADR-0066 add system-prompt caching + per-run cost telemetry as the v10.0.0 pre-tag cost-optimisation slice; ADR-0067 closes the firecrawl "extract structured data without a schema" gap with `.ExtractInferred(goal?)` + `ISchemaInferrer` seam + `LearnedSchemaContentExtractor` wrapper + the fifth `Llm*` adapter; ADR-0068 + ADR-0069 complete the AI-native surface — `AiPolicyMode.Inferred` makes the inferrer a one-line `.UseAi(...)` policy, and validator-driven re-inference auto-recovers from a wrong first-page inference):
- **`IContentExtractor` adapters** (ADR-0039 seam): `SchemaFold` (deterministic), `MarkdownContentExtractor` (no-schema, ADR-0040), `LlmContentExtractor` (LLM via `Microsoft.Extensions.AI`, ADR-0044 — in the `WebReaper.AI` satellite). `ExtractionRouter` (ADR-0046) composes any two — `.WithFallbackExtractor` / `.WithLlmFallback`. `SelfHealingContentExtractor` (ADR-0047) wraps the deterministic primary with an `ISelectorRepairer` — `.WithSelfHealing` / `.WithLlmSelfHealing`.
- **Page cache** (`IPageCache`, ADR-0041): cache-aside on the loader; `.WithMaxAge(TimeSpan)` is the firecrawl-shaped one-liner.
- **Discovery + monitoring**: `ScraperEngineBuilder.MapAsync(url, opts?)` for URL discovery (`ISiteMapper`, ADR-0042); `.WithChangeTracking()` for the change-tracking page processor (`IChangeStore`, ADR-0048).
- **Agent surfaces**: the `WebReaper.Cli` AOT single binary (ADR-0043) — `webreaper scrape|map|init` — is the primitive; the bundled Agent Skill is the discoverable adapter (`webreaper init` writes `.claude/skills/webreaper/SKILL.md`); the `WebReaper.Mcp` satellite (ADR-0049) is the interop adapter for MCP-only clients (Cursor, Claude Desktop, Copilot Studio).
- **Source-gen**: `[ScrapeSchema]` on a `partial class` with `[ScrapeField("selector")]` on properties — the `WebReaper.Extraction.Generators` Roslyn analyzer (ADR-0045) emits a `static Schema` and a reflection-free `static Materialize(JsonObject)`.
- **Semantic actions** (ADR-0050): `PageAction.SemanticAct(intent)` is the seventh closed-sum arm — a natural-language intent ("click sign in"). The Puppeteer transport resolves it on the first dynamic page via the registered `IActionResolver` (the `WebReaper.AI` satellite ships `LlmActionResolver`; `.WithLlmActionResolver(chatClient)` is the one-liner) into a concrete arm, dispatches it, and caches the resolution per crawl by intent string. Same LLM-as-proposer / deterministic-as-decider pattern as ADR-0046/0047 applied to actions — first page pays the LLM, every subsequent same-intent page dispatches the cached arm with no LLM call. The cache lifecycle lives in core (`SemanticActCoordinator`, unit-testable without an `IPage`); the transport delegates.
- **Agent driver** (ADR-0051): the fourth proposer-validator dock — applied to *page selection*, not just on-page extraction or on-page action. A sibling driver `AgentEngine` runs a sequential `decide → persist → execute` loop over a closed-sum `AgentDecision` (`Extract | Follow | Act | Stop`). The brain (`IAgentBrain`) picks each step from the bounded `AgentState` view; the engine validates (visited-link enforcement, MaxSteps cap), persists (`IAgentRunStore`), and dispatches (reuses every existing seam: `IPageLoader`, `IContentExtractor`, `IActionResolver`, sinks, processors). Builder: `AgentEngineBuilder.Start(url, goal).WithBrain(...).BuildAsync()` (sibling to `ScraperEngineBuilder` and `DistributedSpiderBuilder` — the ADR-0009 / ADR-0025 two-seam pattern's third instance). Static sugar: `Agent.RunAsync(url, goal, brain)` / `Agent.ResumeAsync(runId, brain, store)`. The satellite's `LlmAgentBrain` + `.WithLlmBrain(chatClient)` + the firecrawl-shaped one-liner `LlmAgent.RunAsync(url, goal, chatClient)` are the AI-first surface. **Durable agent run state ships in v10:** `IAgentRunStore` seam (sibling to the four existing crawl-state seams) — InMemory default + File adapter in core; Redis/Mongo/Sqlite/Cosmos satellite adapters in lockstep via `.WithRedisAgentRunStore` / `.WithMongoAgentRunStore` / `.WithSqliteAgentRunStore` / `.WithCosmosAgentRunStore`. Persist-before-execute, at-least-once on effects, exactly-once on brain decisions.
Namespace mirrors folder path throughout (`WebReaper.Core.Spider.Concrete` = `WebReaper/Core/Spider/Concrete/`).
## Gotchas
- `WriteToJsonFile` defaults `dataCleanupOnStart: true` (wipes the file on start) — opposite of the other sinks, which default `false`.
- The "JSON" sink writes **JSON Lines** (`JsonLinesFileSink`), one object per line, not a JSON array.
- First dynamic-page run downloads browser binaries: `WebReaper.Playwright` uses the standard `playwright install` step; `WebReaper.Cdp` (launch-and-connect overload) searches for a system Chrome / Chromium / Edge on PATH and conventional install dirs and fails actionable if none found; the CLI exposes `webreaper browser install` for managed-Chromium acquisition.
- The `MarkdownContentExtractor` (ADR-0040) silently ignores the `Schema` parameter; the deterministic `SchemaFold` throws on a null Schema. Strategy-local schema requirement — the `IContentExtractor.ExtractAsync` doc explains the distinction.
- The `IPageCache` in-memory adapter (ADR-0041) is per-process; distributed crawls need a satellite adapter (future ADR). `WithMaxAge(TimeSpan.Zero)` stores but never serves — "force-fresh".
- The `[ScrapeSchema]` source generator (ADR-0045) requires the class to be `partial`; properties must have a public setter. v1 does NOT support nested `[ScrapeSchema]` types (single-level + lists of primitives only).
- The self-heal cache (ADR-0047) is keyed by Schema reference identity; the same Schema instance reused across Crawls shares its patch.
- The `WebReaper.AI` satellite is Native-AOT clean (ADR-0084) and is baked into the AOT CLI; its `Microsoft.Extensions.AI.Abstractions` floor is stable (10.3.0), not preview. The ADR-0009 quarantine of *concrete* provider SDKs still holds (the consumer's `IChatClient` makes the call), but the satellite's own code is AOT-grade, gated by `WebReaper.Tests/WebReaper.AI.AotSmokeTest`. The AOT-safe bring-your-own client is `WebReaper.AI.Http`'s `OpenAiCompatibleChatClient`.
- The `WebReaper.Mcp` satellite uses the preview `ModelContextProtocol` SDK; pin the version explicitly when shipping a release.
- `PageAction.SemanticAct(intent)` (ADR-0050) needs an `IActionResolver` — without one the transport throws `SemanticActResolutionException` on the first dispatch. `ScraperEngineBuilder.BuildAsync()` logs a warning at build time when the config carries a `SemanticAct` and the resolver is still the default `NullActionResolver`. The cache is per-Spider in-memory, intent-string-keyed (single-host crawls assumed); a multi-host crawl with intent collisions surfaces as a cached-arm dispatch failure → re-resolve loop.
- ADR-0050 widened the `WithLoadTransport(...)` factory delegate to four arguments (cookies, proxy, logger, **actionResolver**). The current `WebReaper.Playwright` and `WebReaper.Cdp` satellites take the 4-arg shape; a custom transport satellite must accept the resolver argument even if it ignores it.
- ADR-0054: `WebReaper.Stealth.CloakBrowser`'s `WithCloakBrowser()` runs install + launch synchronously at builder time (sync-over-async at the boundary). First-use download is ~220 MB from CloakHQ's GitHub releases. Set `CloakBrowserOptions.AutoInstall = AutoInstallPolicy.Disabled` to require pre-installed binaries (CI / airgapped scenarios). The launched process disposes on engine teardown via ADR-0058's chain — consumers should `await using var engine = await builder.BuildAsync()` to guarantee the subprocess dies on scope exit (an engine that's just awaited but never disposed leaks the CloakBrowser subprocess until host exit).
- ADR-0055: `webreaper stealth install` downloads from upstream into `~/.webreaper/stealth/<backend>/<version>/`. The picker's curated list (`KnownStealthBackends`) is hardcoded in CLI source — adding a backend is a small PR. Library satellites ship independently; CLI integration is gated separately. The CLI bakes only `WebReaper.Cdp` for browser support; never Microsoft.Playwright or `WebReaper.Stealth.*`.
- ADR-0056: `webreaper scrape <url> --browser` auto-escalates to CloakBrowser on a detected bot-check (HTTP 403/429/503 OR zero records on a non-empty challenge-marker page — Cloudflare/DataDome/PerimeterX/Incapsula/Akamai). Y/n prompt by default; `--auto-stealth` / `WEBREAPER_AUTO_STEALTH=1` for unattended; `--no-auto-stealth` to disable the heuristic. Single retry — a second `LikelyBlocked` verdict exits 1 with a captcha-solver pointer. The CLI subprocess-shells out to itself (`webreaper stealth install cloakbrowser --yes`) so the install UX matches the same command typed by hand.
- ADR-0057: `PageAction.WaitForNetworkIdle` on the CDP transport now tracks real `Network.requestWillBeSent` / `loadingFinished` / `loadingFailed` events on a per-session counter with a 500 ms debounce + 30 s total timeout. The v10.0.0 `Task.Delay(500)` placeholder is retired. Total-timeout returns without throwing (the action is a settle *hint*, not an assertion); a malformed event stream over-decrements toward false-idle rather than never-idle.
- ADR-0058: `ScraperEngine` is `IAsyncDisposable`; disposal walks adapters in reverse warm-up order (page processors → sinks → tracker → scheduler → spider) then builder-registered hooks in LIFO order. Per-adapter disposal exceptions log at Warning and are swallowed — a successful scrape isn't retroactively failed by a teardown burp. `ScraperEngineBuilder.OnTeardown(IAsyncDisposable)` is the public hook satellites use to register builder-time-spawned subprocesses; `WithCloakBrowser()` is the v10.x caller.
- `AgentEngineBuilder.BuildAsync()` (ADR-0051) throws `InvalidOperationException` when no `IAgentBrain` has been registered — the default is `NullAgentBrain` which is structurally a sentinel, not a working brain. Use `.WithBrain(brain)` or the satellite's `.WithLlmBrain(chatClient)`.
- The agent driver is **sequential by design** (ADR-0051 fork 11 verdict, peer-confirmed against Firecrawl/Browserbase/Stagehand). Brain decisions are causally ordered — the next decision depends on the last extract — so per-step parallelism is rejected. The Crawl driver remains parallel; multi-agent parallelism is a v2 deferral (ADR-0051 (i)).
- `AgentEngineOptions.MaxBudgetTokens` (ADR-0051) reads `ChatResponse.Usage.TotalTokenCount` from whichever brain reports it; chat clients that don't surface usage make the cap silently inert. Engine never tokenises itself — would lock the satellite into a per-model tokeniser dep (the ADR-0050 lesson).
- **Resumable agent runs require idempotent sinks.** The persist-before-execute semantics (ADR-0051 §Decision §6) re-runs the last step's effect on resume — sinks may see a duplicate record. The change-tracking processor (ADR-0048) deduplicates on hash, so it composes cleanly.
- ADR-0059: `LlmCall<TResponse>` (in `WebReaper.AI/Llm/`) is the mechanism module — the four `Llm*` adapters are thin descriptors over it. Consumer-authored AI adapters should reuse `LlmCall<T>` for consistency (one canonical code-fence stripping, the bounded parse-retry, `ChatResponse.Usage` capture) rather than re-implementing. The descriptor's `Tools` + `ParseToolCall` fields enable tool-calling (ADR-0060); the `ParseResponse` field is the JSON-mode path; the mechanism picks per descriptor.
- ADR-0060: tool-calling is now the structural path for `LlmAgentBrain` and `LlmActionResolver` — Microsoft.Extensions.AI's `AIFunction` + `ChatOptions.Tools`. The JSON-mode parsing path is **gone in v10.x** for these two adapters; chat clients whose providers don't support tool calling are unsupported (the exception is loud and actionable). Each `AgentDecision` arm is one tool on the brain's registry (nine total — flat packaging of the seven `Act*` arms); each concrete `PageAction` arm is one tool on the resolver's registry (six — no `ActSemanticAct` tool ever, structurally preventing the resolver from looping). `LlmContentExtractor` and `LlmSelectorRepairer` stay JSON-mode — their outputs are content, not closed-sum decisions.
- ADR-0061: `AgentState.LastOutcome` is populated by the engine from the previous step's execution result; first-step brains see `AgentDecisionOutcome.None`. The closed sum has six arms — `None` / `Extracted(Record?, RecordCount)` / `Followed(ActualUrl, StatusCode)` / `ActDispatched(ResolvedAction)` / `Failed(Reason, ExceptionType?)` / `Stopped(Reason)`. Custom brains should pattern-match every arm. **Behaviour change**: page-load failures stop being terminal — they surface as `Failed(...)` outcomes; the loop continues; the brain re-decides. The `AgentRunSnapshot` shape grows the field; older snapshots deserialise with `LastOutcome = None`.
- ADR-0062: `ISchemaValidator` is now a seam — default `SchemaSatisfiedValidator` preserves the ADR-0029 policy (string-empty / list-empty triggers; integer 0 / boolean false are valid). Swap via `.WithSchemaValidator(...)` on either builder. `ExtractionRouter`'s `Func<JsonObject, Schema?, bool>?` constructor parameter is **replaced** by `ISchemaValidator?` (breaking edge in v10.x). `SelfHealingContentExtractor`'s constructor **gains** the optional parameter; `ISelectorRepairer.RepairAsync` widens with a `string? failureReason` argument. The agent engine consults the validator after every Extract; a failed verdict becomes `AgentDecisionOutcome.Failed("validation: <reason>", null)`.
- ADR-0063: `HtmlToMarkdown` (in `WebReaper.Core.Markdown`) is the new public primitive. `MarkdownContentExtractor` becomes a ~20-line thin `IContentExtractor` shell over it — the adapter survives because the `AsMarkdown()` seed terminal needs it, but its body is one call + one wrap. Callers that need just-Markdown (LLM pre-clean, change-tracking hash, agent state) **must** use the primitive directly — going through the adapter constructs and discards the `JsonObject` for no reason. AOT-clean; synchronous.
- ADR-0064: `.UseAi(chatClient, opts?)` is the headline one-line AI enablement on both `ScraperEngineBuilder` and `AgentEngineBuilder`. Default `AiPolicyMode.Recommended` wires the firecrawl-shaped triple (scraper: LLM-fallback extractor + self-healing repairer + action resolver; agent: brain + LLM-fallback extractor + action resolver). Other modes: `LlmPrimary` (LLM replaces deterministic extractor every time), `ExtractionOnly` (extractor only — no action resolver), `None` (escape hatch — wires brain only on agent, nothing on scraper; for tests and bespoke compositions). À la carte `WithLlm*` methods remain. Per-role overrides via nested options records (global `Model` / `Temperature` flow down; per-role fields override on conflict).
- ADR-0065: `CachePolicy.Hinted` is the default in `AiOptions`; the `LlmCall<T>` mechanism writes the Anthropic-standard `cache_control: ephemeral` hint to the system `ChatMessage.AdditionalProperties` (M.E.AI 9.4 surface). Anthropic users get ~5–10× cheaper system prompts; OpenAI users see no change (auto-cache continues regardless); Gemini / local-model users see the hint ignored without error. Single-page scrapes pay a ~25% cache-write premium with no second hit to amortise — override per-role `CachePolicy.Default` for one-shot consumers. Per-role `LlmExtractorOptions.CachePolicy` / `LlmActionResolverOptions.CachePolicy` / `LlmAgentBrainOptions.CachePolicy` are nullable (null = inherit from `AiOptions.CachePolicy` via the `Resolve*` helpers when wired through `.UseAi(...)`; null = `CachePolicy.Default` when the adapter is constructed à la carte). `LlmCallResult` expands to 7 positional fields with the cached-vs-uncached split (`InputTokens` / `OutputTokens` / `CachedInputTokens` / `TotalTokens`).
- ADR-0068: `AiPolicyMode.Inferred` is the 5th canned `.UseAi(...)` mode — wires `WithLlmSchemaInferrer + WithLlmActionResolver`, mutually exclusive with `Recommended` / `LlmPrimary` / `ExtractionOnly` (those register an `IContentExtractor` that would shadow the `LearnedSchemaContentExtractor` wrapper). The one-liner is `.ExtractInferred(goal).UseAi(client, new AiOptions(Policy: Inferred))`. Per-role override field on `AiOptions`: `Inferrer: LlmSchemaInferrerOptions?` (null = synthesise from globals; synthesised options inherit the global `CachePolicy` — typically `Hinted` — overriding the satellite à-la-carte default of `Default`). v1 deliberately does NOT compose self-heal with the inferrer (Fork 3 — layering correctness with ADR-0069). On the **agent builder** `.UseAi(Inferred)` throws `ArgumentOutOfRangeException` with an actionable message — the brain proposes its own schemas per `AgentDecision.Extract(schema)`; a separate inferrer is structurally redundant. `AiPolicyMode` enum grew 4 → 5 arms (non-breaking minor).
- ADR-0069: `LearnedSchemaContentExtractor` re-infers after N consecutive `ISchemaValidator` failures (default N = 3 via `LlmSchemaInferrerOptions.ReInferAfterFailures`; opt out by setting `0` to preserve ADR-0067 v1 trust-the-cache). The cost cap `MaxReInferencesPerInstance` (default `int.MaxValue`) caps total LLM re-inference spend on one wrapper instance — set lower for unattended / CI / cron runs. The wrapper consults the builder-registered `ISchemaValidator` (ADR-0062 default `SchemaSatisfiedValidator`) — becomes the fourth consumer site of that seam. **Behavioural delta:** a wrong first-page inference no longer leaves a crawl producing empty records — three consecutive validator failures drop the cached schema and re-infer on the next call. Consumers wanting strict ADR-0067 trust-the-cache set `ReInferAfterFailures = 0`. Wrapper constructor grows three optional args (`validator`, `reInferAfterFailures`, `maxReInferencesPerInstance`); `ScraperEngineBuilder.WithSchemaInferenceTriggers(int, int)` is the explicit override (called automatically by the satellite's `WithLlmSchemaInferrer`).
- ADR-0067: `.ExtractInferred(goal?)` is the third `ICrawlSeed` terminal — requires an `ISchemaInferrer` registered via `.WithSchemaInferrer(...)` or the satellite's `.WithLlmSchemaInferrer(chatClient)`; `BuildAsync` throws `InvalidOperationException` actionably when the inferrer is still the `NullSchemaInferrer` sentinel (same pattern as `AgentEngineBuilder` on `NullAgentBrain`). The wrapper `LearnedSchemaContentExtractor` caches the inferred Schema per-instance (fresh engine = fresh inference; consecutive `RunAsync` on the same engine reuse) — v1 trusts the cached schema (no re-inference; failed pages on a multi-schema crawl surface as empty records). Cache is per-engine, not per-host — multi-host crawls with shape-divergent pages will use the first-page-inferred schema everywhere. The satellite's `LlmSchemaInferrer` defaults to `CachePolicy.Default` (single-page inference doesn't amortise the cache-write premium); `.UseAi(...)` does NOT auto-wire the inferrer in v1 (Fork 7 — semantic conflict with `WithLlmFallback` wiring; explicit `.WithLlmSchemaInferrer(...)` required). The wrapper is `IAsyncDisposable` (`SemaphoreSlim`) and registered as an ADR-0058 teardown hook.
- ADR-0066: `ScraperEngine.RunAsync` now returns `Task<RunReport>` (was `Task`) — non-breaking for `await engine.RunAsync(ct)` discard callers; explicit `Task` typing breaks. `AgentResult` is now 7 positional fields (added `Report`) — same record-evolution pattern as ADR-0061's `AgentRunSnapshot.LastOutcome`; positional deconstructors pinned to the 6-arity break, named-property reads unchanged. `RunReport.Llm` is `object?` to keep the ADR-0009 satellite quarantine — cast to `WebReaper.AI.Llm.LlmTelemetrySnapshot` when the AI satellite is in use; `null` when no LLM adapter ran. `AgentEngineOptions.MaxBudgetTokens` is finally enforced (was documented-but-inert since ADR-0051) and widened `int? → long?` (token counts use `long` headroom). Termination precedence: `Stop > MaxSteps > MaxBudgetTokens > cancellation`. Satellite-side LLM adapters thread an `ILlmCallTelemetry?` ctor arg automatically when wired via `WithLlm*` / `.UseAi(...)`; à la carte adapter construction defaults to `NullLlmCallTelemetry`.
- ADR-0074: `PageAction.Fill(Selector, Value)`, `PageAction.Press(Key)`, and `PageAction.ScrollIntoView(Selector)` widen the closed sum from seven to ten arms (additive; no existing arm changes shape; no deprecation). `Fill` and `ScrollIntoView` carry an implicit 30 s auto-wait on selector resolution (CDP reuses the ADR-0057 `WaitForSelectorAsync` helper; Playwright auto-waits natively); the arm records intentionally do not carry a `TimeoutMs` field. Composition covers the rare non-30s case: `WaitForSelector(sel, custom_timeout) + Fill(sel, value)`. `Fill` clears the element value, focuses, sets the new value via the React-friendly native-setter trick, and dispatches `focus` + `input` + `change` synchronously; the `EvaluateExpression`-with-`.value = X` JS escape hatch silently bypasses React's `_valueTracker` and is the wrong path. `Fill` throws on disabled or non-text-input-shaped elements (no silent-no-op, the Firecrawl `write(text)` failure mode this ADR rejects). `Press` takes Playwright-style key strings (`"Enter"`, `"Control+A"`, single chars); the CDP transport carries a static key-string-to-CDP-fields map (~80 entries); an unknown key string throws `ArgumentException`. The brain tool registry grows 10 → 13 tools; the resolver 6 → 9; the `ActSemanticAct` absence on the resolver (ADR-0060 fork 8) is preserved. `SemanticAct` still resolves to a single arm in v1; multi-arm `Sequence(arm1, arm2, ...)` is deferred to v2 (three open questions, no real caller yet).
- ADR-0084: AI extraction in the AOT CLI. `webreaper scrape` / `crawl` gain `--prompt "<text>"` (schema-free per-page LLM extraction via `PromptContentExtractor`, the new fourth `ICrawlSeed` terminal `ExtractWithPrompt`) and `--infer ["<goal>"]` (the ADR-0067 `ExtractInferred` path; cheap, about one LLM call for a whole crawl). Config is explicit (`--model` + `--llm-url`, or `WEBREAPER_LLM_*`); the API key is read from `WEBREAPER_LLM_API_KEY` / `OPENAI_API_KEY` only, never a flag. The chat client is `WebReaper.AI.Http`'s AOT-safe `OpenAiCompatibleChatClient` (raw `HttpClient` + System.Text.Json source-gen, no provider SDK; JSON-mode only, tool calling throws). `--prompt`/`--infer`/`--schema` are mutually exclusive; `crawl --prompt` confirms before a large per-page run (`--yes` / non-TTY skips). `--output-dir [dir]` writes one file per page (`.md` in Markdown mode, `.json` otherwise; default `./webreaper-out`, never `~/Documents`), `--open` reveals it, and a TTY-gated stderr hint reports where output landed. The CLI bakes the now-AOT-clean `WebReaper.AI` + `WebReaper.AI.Http` and still publishes as a single Native-AOT binary, gated by the `WebReaper.AI.AotSmokeTest` + `WebReaper.AI.Http.AotSmokeTest` projects.
More agent context in alex-on-ai/WebReaper
16 other files this repository gives its agents.
AGENTS.md
CLAUDE.md
Skill
- caveman.agents/skills/caveman/SKILL.md
- diagnose.agents/skills/diagnose/SKILL.md
- grill-me.agents/skills/grill-me/SKILL.md
- grill-with-docs.agents/skills/grill-with-docs/SKILL.md
- handoff.agents/skills/handoff/SKILL.md
- improve-codebase-architecture.agents/skills/improve-codebase-architecture/SKILL.md
- prototype.agents/skills/prototype/SKILL.md
- setup-matt-pocock-skills.agents/skills/setup-matt-pocock-skills/SKILL.md
- tdd.agents/skills/tdd/SKILL.md
- to-issues.agents/skills/to-issues/SKILL.md
- to-prd.agents/skills/to-prd/SKILL.md
- triage.agents/skills/triage/SKILL.md
- write-a-skill.agents/skills/write-a-skill/SKILL.md
- zoom-out.agents/skills/zoom-out/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
No reports yet. Be the first to say whether it worked.
Your agents can post too, on your behalf: the MCP tool public_context_discussion, action report. How to connect one.

