trawl
bbulb/trawl/CLAUDE.md
Selective web content extraction. URL + natural-language query → the few chunks most relevant to the query, ranked by dense embedding similarity. Exposed as a Python library and a stdio MCP server so agents (Claude Code, Claude Desktop, anything MCP-aware) can read only the parts of a page they actually need. This file is loaded automatically by Claude Code when working in the trawl directory. Humans should read README.md first, then CONTRIBUTING.md for dev setup. - Version: 0.4.6 (2026-07-06). Highlights:…
CLAUDE.md3 starsChanged 6 months ago
- Reads credentials
- Installs packages
# trawl
Selective web content extraction. URL + natural-language query → the few
chunks most relevant to the query, ranked by dense embedding similarity.
Exposed as a Python library and a stdio MCP server so agents
(Claude Code, Claude Desktop, anything MCP-aware) can read only the
parts of a page they actually need.
This file is loaded automatically by Claude Code when working in the
trawl directory. Humans should read `README.md` first, then
`CONTRIBUTING.md` for dev setup.
## Current status
- **Version**: 0.4.6 (2026-07-06). Highlights: rs-trafilatura
extraction candidate default on (`TRAWL_RS_TRAF=0` to opt out;
optional dependency), extraction-selector fixes (heading-density
sqrt smoothing + records-sentinel hard gate → +30 score bonus;
WCXB combined F1 0.777 → 0.818), profile mapper DIV→MAIN/ARTICLE
LCA promotion (profile eval IDEAL 16/36 → 26/37), per-call cache
freshness `max_cache_age_s` on `fetch_relevant()`/MCP `fetch_page`.
0.4.5 highlights: injection defense, hybrid retrieval, embed cache
all default on. Full list in `CHANGELOG.md`.
- **Parity matrix**: 15/15 cases pass (see `tests/test_cases.yaml`).
`kbo_schedule` pinned to a historical game day to survive KBO
off-days. `wanted_jobs` asserts the structural marker `합격보상금`
only (PR #72) — live job-board titles rotate spelling hourly.
- **Agent patterns**: 110 patterns across 8 shards; full run
101 PASS + 9 SKIP (`live: optional` anti-bot/DDG/VLM draws),
required FAIL 0. Coding shard 24/24.
- **Profile eval**: 37-site evaluation — 92% success rate, 26/37 IDEAL
selectors (2026-07-06, post DIV→MAIN/ARTICLE promotion; the 3 fails
are live bot-block/error pages).
- **Benchmark vs Jina Reader**: ~23x fewer tokens on average across 12
cases; profile-cached mode ~30x.
- **WCXB external benchmark**: trawl `html_to_markdown` F1 = 0.818 vs
Trafilatura baseline 0.750 on the 1,497-page dev split (0.777 at
v0.4.5; lifted by the rs-trafilatura candidate + selector fixes in
0.4.6).
- **Longform retrieval cost (default on, 2026-04-22)**: `TRAWL_CHUNK_BUDGET`
default flipped from `0` to `100` after re-validating on curl.se
manpage (275 KB / 760 chunks, p95 25149 ms → 3065 ms). Parity 15/15
+ agent_patterns coding 23/24 at that time preserved (the unrelated
`arxiv_pdf_lora` fetcher flake has since recovered — coding is 24/24
as of v0.4.5). Opt out via `TRAWL_CHUNK_BUDGET=0`.
### What a new session should do first
1. Read `README.md` (5 min) to understand what trawl does.
2. Read `ARCHITECTURE.md` if you need to modify the pipeline — it has
the "why each library was chosen" reasoning you'll need.
3. Activate the dev env: `mamba activate trawl` (create with
`mamba env create -f environment.yml` if missing).
4. Run `python tests/test_pipeline.py` as a smoke test. Requires a
running bge-m3 embedding server at `TRAWL_EMBED_URL` (default
`localhost:8081`).
### What NOT to do on a fresh session
- Don't re-tune `chunking.py` / `retrieval.py` parameters
"to see what happens". The "Things NOT to change" table below
exists for a reason.
- Don't add crawling, search, or page-rewriting features. See the
"In / out of scope" section below.
## Critical Rules
> **MUST follow these rules. No exceptions.**
- **Use the `trawl` mamba env.** Every command: run inside
`mamba activate trawl` or prefix with `mamba run -n trawl`. `trawl`
is `pip install -e .`-installed into this env; other envs won't
have the editable install.
- **Run the parity matrix before committing any change to `src/trawl/`.**
`python tests/test_pipeline.py` must stay 15/15. If a tuning change
breaks one case, it's almost certainly breaking something else too —
diagnose, don't just tighten ground truth.
- **Run the MCP smoke test before touching `src/trawl_mcp/`.**
`python tests/test_mcp_server.py` proves the stdio protocol still
works end-to-end.
- **Do not commit `tests/results/`.** Already gitignored, but watch
for timestamp directories sneaking in.
- **Do not change `tests/test_cases.yaml` ground truth to make a
failing test pass** without re-running the matrix to confirm the
change is principled, not a fudge.
- **llama-server endpoint map** (reference setup; override with env
vars, see `.env.example`):
- `:8081` — bge-m3 embeddings, **mandatory** (without it retrieval
fails).
- `:8082` — utility LLM (e.g. Gemma 4 E4B). HyDE only. Small context
(4K), auxiliary tasks.
- `:8083` — bge-reranker-v2-m3 cross-encoder. Default on; falls back
gracefully to cosine-only if absent. Run with
`--reranking --pooling rank`.
- `:8080` — vision-enabled main LLM. Used **only** for explicit
`profile_page` invocations (bounded, manual-trigger workload).
If slot contention shows up, set `TRAWL_VLM_URL` to a dedicated
vision server; no code changes needed.
- **Slot pinning** — `TRAWL_VLM_SLOT=<N>` / `TRAWL_HYDE_SLOT=<N>`
pin requests to a specific llama-server slot (via `id_slot`) to
avoid evicting other consumers' KV cache on shared servers with
prompt caching.
- **Raw passthrough** — JSON/XML/RSS/Atom responses are returned as-is
without extraction. URL suffixes (`.json`, `.xml`, `.rss`, `.atom`)
take an httpx fast path; suffix-less API endpoints are detected by
response `Content-Type`. Byte cap via `TRAWL_PASSTHROUGH_MAX_BYTES`
(default 256 KB).
- **Telemetry** (opt-in) — `TRAWL_TELEMETRY=1` appends one JSON line
per `fetch_relevant()` call to `~/.cache/trawl/telemetry.jsonl`
(override via `TRAWL_TELEMETRY_PATH`). Single-generation rotation
at 64 MB. Purpose: feed the C4 decision in `notes/RESEARCH.md`.
Schema: `src/trawl/telemetry.py` + the C4 spec doc.
- **Per-fetch cache** (C8, default on) — successful HTML/PDF fetches
are cached in `~/.cache/trawl/fetches/<sha256>.json` for 300 s.
Re-fetch of the same URL within TTL skips Playwright +
Trafilatura; chunking / embedding / retrieval still run fresh.
Disable via `TRAWL_FETCH_CACHE_TTL=0`; relocate via
`TRAWL_FETCH_CACHE_PATH`; size cap via `TRAWL_FETCH_CACHE_MAX_MB`
(default 100). `PipelineResult.cache_hit` flags the reuse.
- **Document embedding cache** (**default on since 2026-05-19**,
opt-out via `TRAWL_EMBED_CACHE_TTL=0`) — successful chunk
embeddings are cached in `~/.cache/trawl/embeddings/<sha256>.json`
for `TRAWL_EMBED_CACHE_TTL` seconds (default 3600). Cache key
partitions on model, base_url, text_sha256, contextual_mode,
`prefix_max_chars`, `prefix_version`, and `SCHEMA_VERSION` so
text edits / contextual on/off / model swaps all miss naturally.
Disk cap via `TRAWL_EMBED_CACHE_MAX_MB` (default 512 MB) with LRU
trim at 20 % headroom; relocate via `TRAWL_EMBED_CACHE_PATH`.
`PipelineResult.embed_cache_hits` / `embed_cache_misses` and the
opt-in telemetry JSONL surface per-call counters. Warm-repeat
measurement on 6 reader-comparison URLs: cold avg retrieval 1503
ms → warm 57 ms (−96.2 %). See
`docs/superpowers/specs/2026-05-18-embed-cache-default-on-design.md`.
- **Per-host adaptive ceiling** (C9, default on) — Playwright's
content-ready wait ceiling becomes `p95(host) × 1.5` once 5
observations accumulate, clamped to `[1500, 15000] ms`. New hosts
use the static 5000 ms default. Stats in
`~/.cache/trawl/host_stats.json`. Disable via `TRAWL_HOST_STATS=0`.
- **Hybrid dense + BM25 retrieval** (C6, **default on since
2026-05-19**, opt-out via `TRAWL_HYBRID_RETRIEVAL=0`) —
BM25 lexical ranking runs alongside dense cosine, fused via
Reciprocal Rank Fusion (`k=60`). Tokenizer is rule-based
multilingual (Latin word / Hangul bigram / CJK char) in
`src/trawl/bm25.py`. Reranker window unchanged (2x candidates).
Tune via `TRAWL_HYBRID_RRF_K` (default 60).
Initial C6 A/B (PR #29) on `code_heavy_query` was conservative
(no wins, no regression) — that conclusion holds for that
surface. The default flip is justified by a `reader_comparison`
A/B at 2026-05-19 (`hybrid_default_on_spike.py`) that PASSed
all 6 pre-registered gates after a companion wiki-fetcher
heading-preservation fix: reader-comp hybrid 6/6 vs dense 5/6
(`wiki_large_language_model` saved by BM25 fusion), parity 15/15
+ coding 24/24 in both modes, Korean 3/3, code_heavy_query
regression 0, retrieval p95 +1.3%. See
`notes/c6-hybrid-measurement.md` (original C6) and
`notes/hybrid-default-on-spike-outcome.md` (default-flip
measurement).
- **Chunk budget prefilter** (longform follow-up, **default on since
2026-04-22**, opt-out via `TRAWL_CHUNK_BUDGET=0`) —
`TRAWL_CHUNK_BUDGET=100` is the default; any positive int caps the
pool sent to bge-m3. When a page's chunk count exceeds the budget,
the BM25 scorer from C6 ranks the chunks and only the top-N reach
the embedding stage; reranker input window unchanged. Reuses the
C6 tokenizer, so the flag stacks with `TRAWL_HYBRID_RETRIEVAL=1`.
Original measurement at budget=100 on 4 longform cases cut
retrieval_ms.p95 from 6,002 ms to 1,890 ms (69%) with 4/4 rank-1
identity preserved. Re-validated 2026-04-22 on curl.se manpage
(p95 25149 ms → 3065 ms) and full parity + coding agent_patterns
(15/15, 23/24 baseline preserved). `PipelineResult.n_chunks_embedded`
reports the post-prefilter count. See
`docs/superpowers/specs/2026-04-20-longform-retrieval-cost-design.md`
and `docs/superpowers/specs/2026-04-22-chunk-budget-default-on-design.md`.
- **Indirect prompt-injection defense** (default on, opt out via
`TRAWL_INJECTION_SCAN=0`) — `src/trawl/sanitize.py` runs a
model-free scan per fetch (`pipeline._scan_injection`, both the
full-retrieval and profile return paths): strips Unicode tag chars
(U+E0000–E007F), flags instruction-like chunks
(`suspicious_injection`) and instruction-like CSS-hidden /
off-screen / aria-hidden segments (`suspicious_hidden`) —
annotate, never delete. Signals land in `PipelineResult.warnings`;
flag keys appear on a chunk only when true. MCP `fetch_page` /
`profile_page` carry `openWorldHint` + `readOnlyHint` annotations
plus the existing `content_boundary` response field. The
CSS-hidden branch is gated on the instruction-pattern matcher so
benign hidden content (sr-only text, collapsed menus) does not warn
— load-bearing for the benign-0-flag gate; do not loosen without a
companion false-positive check. Cost <1% on a 198 KB / 600-chunk
page. See
`docs/superpowers/specs/2026-06-05-injection-defense-design.md`.
- **Shadow-DOM unwrap for code-block custom elements** (default on)
— `fetchers/playwright.py` inlines each matching element's
`shadowRoot`'s `pre > code` textContent (wrapped in a fresh
`<pre><code>`) into the light DOM before `page.content()`.
Initial allow-list: `mdn-code-example`. Playwright's default
content capture skips shadow roots, so pages that render code in
Shadow DOM (notably MDN post-2024 redesign) previously fed the
extractor empty `<mdn-code-example></mdn-code-example>` tags;
the MDN fetch pattern's assertion keywords (`JSON.stringify`,
`method:`, `application/json`) were simply not in the extracted
markdown. Measurement on the 16 `code_heavy_query` patterns:
baseline 15/16 → on 16/16 (`claude_code_mdn_fetch_api` flipped
to PASS), top1 changed on 1/16 (MDN only, `n_chunks_total`
22 → 24), parity 15/15 in both modes. Disable via
`TRAWL_SHADOW_DOM_UNWRAP=0`. Grow `SHADOW_DOM_UNWRAP_TAGS` only
with a companion measurement: each addition must fix a specific
pattern and not regress the other 15. See
`docs/superpowers/specs/2026-04-20-playwright-shadow-dom-design.md`.
- **rs-trafilatura extraction candidate** (default on since
2026-07-04, opt-out via `TRAWL_RS_TRAF=0`) — a Rust extractor
(PyO3 bindings, `pip install rs-trafilatura`, optional dependency:
silently skipped when missing) joins `extract_html()`'s score-based
candidate selection. Output route is
`html_to_markdown(extract(html).content_html)` to preserve
headings/tables for the chunker. Measured 2026-07-03 on the WCXB
dev split: combined F1 0.809 vs prior 0.777 (+0.032); rs variant
alone 0.848 over-ok vs same-session Trafilatura 2.0.0 baseline
0.750, winning all 7 page types. Parity 15/15 in both modes,
coding 24/24. NOTE for A/B measurements: the C8 fetch cache stores
post-extraction results — disable it (`TRAWL_FETCH_CACHE_TTL=0`)
when comparing extraction modes or the candidate never runs. See
PR #83 and `notes/rs-trafilatura-verification-outcome.md`.
- **Reranker chunk-window cap** (default on) —
`src/trawl/reranking.py` clamps outbound documents to
`TRAWL_RERANK_MAX_DOCS` (default `30`), each individual document
to `TRAWL_RERANK_MAX_PER_DOC_CHARS` (default `1500`), and the
query+docs total to `TRAWL_RERANK_MAX_CHARS` (default `40000`)
before the POST. Lower-cosine-rank tail chunks are dropped
first; oversize docs are clamped to the per-doc cap; remaining
docs are proportionally truncated (floor `200` chars each) only
if the total still exceeds. A single `WARNING` is logged per
call when any cap fires. `<= 0` on any env var disables that
knob. Defaults were bracketed empirically: 40 k total passes
and 50 k fast-rejects on the 8 192-token total context limit;
the per-doc 1800 default mirrors `MAX_EMBED_INPUT_CHARS` and
sits safely under the per-document 512-token batch limit
(PR #41 D2 outcome — observed 2 056 chars → 514 tokens). See
`docs/superpowers/specs/2026-04-20-reranking-chunk-window-cap-design.md`
and `docs/superpowers/specs/2026-04-21-rerank-per-doc-char-cap-design.md`.
## Quick Reference
All commands assume you're inside the `trawl` mamba env
(`mamba activate trawl`) or prefixed with `mamba run -n trawl`.
```bash
# First time (creates the env + installs deps)
mamba env create -f environment.yml
mamba run -n trawl playwright install chromium
mamba activate trawl
# Parity matrix: 15 cases, non-zero exit on regression
python tests/test_pipeline.py
# Single case, verbose
python tests/test_pipeline.py --only kbo_schedule --verbose
# With HyDE enabled (adds ~15-20s, rarely useful)
python tests/test_pipeline.py --hyde
# Agent usage pattern matrix (workflow-shape regressions for openclaw/hermes/Claude Code)
python tests/test_agent_patterns.py --dry-run # schema only
python tests/test_agent_patterns.py --shard coding # one shard
python tests/test_agent_patterns.py --only <pattern_id> --verbose
# MCP server smoke test
python tests/test_mcp_server.py
# Benchmark vs Jina Reader (requires .env with JINA_API_KEY)
python benchmarks/run_benchmark.py
python benchmarks/run_benchmark.py --no-profile
# Profile eval: 36-site VLM prompt quality check (requires :8080 VLM)
python benchmarks/profile_eval.py
python benchmarks/profile_eval.py --category docs
# Start MCP server (stdio)
python -m trawl_mcp
# Library usage check
python -c "
from trawl import fetch_relevant
r = fetch_relevant('https://example.com/', 'what is this')
print(r.chunks)
"
# WCXB external extraction benchmark (one-shot)
python benchmarks/wcxb/fetch.py && python benchmarks/wcxb/run.py
```
## Architecture pointer
See `ARCHITECTURE.md` for:
- Full pipeline diagram
- Why each component was chosen
- Tuning decisions (adaptive k, max_chars, waitFor, stealth) and their
measured effect
- Known limitations and workarounds
The `README.md` is for users. `ARCHITECTURE.md` is the file to read
when you need to understand *why* something is the way it is.
## Code layout
```
src/trawl/ library — the pipeline
pipeline.py fetch_relevant() entry point
chunking.py heading + table preservation + sentence
fallback + markdown markup stripping
retrieval.py bge-m3 cosine top-k with adaptive k
reranking.py bge-reranker-v2-m3 cross-encoder rerank
extraction.py Trafilatura (precise+recall) + BS fallback
hyde.py optional query expansion (off by default)
records.py repeating-sibling record detection + sentinels
reranking.py bge-reranker-v2-m3 cross-encoder (title-injection)
telemetry.py opt-in JSONL telemetry collector
profiles/ VLM-based page profiling
prompts.py VLM prompt (v2: anti-sidebar anchor guidance)
mapper.py anchor→DOM→LCA→CSS selector (noise filter)
vlm.py llama-server VLM client
profile.py profile load/save/cache
cache.py per-host profile lookup for host-transfer
fetchers/
playwright.py sync_playwright + stealth, content-ready wait
pdf.py httpx + pymupdf
passthrough.py raw JSON/XML/RSS/Atom pass-through (httpx)
youtube.py youtube_transcript_api + playwright fallback
github.py GitHub REST API + playwright fallback
stackexchange.py Stack Exchange API v2.3 + playwright fallback
wikipedia.py MediaWiki parse API + playwright fallback
src/trawl_mcp/ MCP server wrapper (stdio default, --http opt-in)
server.py list_tools / call_tool handlers
http.py streamable-HTTP transport (--http)
__main__.py `python -m trawl_mcp [--http [HOST:PORT]]`
tests/
test_cases.yaml 12 golden cases (extraction-quality parity)
test_cases.yaml 15 golden cases (extraction-quality parity)
test_pipeline.py parity runner — compares against ground truth
test_agent_patterns.py agent workflow harness (single/repeat/host-transfer/compositional)
agent_patterns/ pattern catalog (one yaml per shard)
schema.py dataclass + YAML validator
loader.py shard loader + ID dedupe
coding.yaml coding-assistant patterns (~25)
README.md catalog rules + assertion DSL ref
test_mcp_server.py stdio protocol smoke test
results/ gitignored test outputs
benchmarks/
benchmark_cases.yaml 12 cases for trawl vs Jina comparison
run_benchmark.py trawl (base/profile/cached) vs Jina runner
profile_eval_cases.yaml 36 cases for VLM profile eval
profile_eval.py profile generation quality evaluator
wcxb/ external WCXB extraction benchmark (Phase 1)
fetch.py snapshot download + hash verify
run.py runner (trawl + Trafilatura baseline)
aggregate.py summary + report rendering
evaluate.py vendored WCXB word-F1 evaluator
manifest.json pinned SHA-256 manifest of dev split
results/ gitignored benchmark outputs
examples/
claude_code_config.json MCP server entry for Claude Code
mcp_gateway_config.yaml mcp-gateway style snippet
```
## Conventions
- Python 3.10+, typed where it helps readability. Not a mypy-strict
codebase yet.
- **No emoji in source or test files.** CLAUDE.md and README.md may
have them sparingly when the user explicitly asks, but the default
is no emoji.
- Docstrings on public functions; a one-line comment only when the
*why* is non-obvious.
- Commits: conventional commit prefixes (`feat`, `fix`, `docs`,
`test`, `refactor`, `chore`). Short subject, longer body if
the change is non-trivial or has tuning rationale.
- Test artefacts land in `tests/results/<timestamp>/` and are
gitignored. Don't bypass the gitignore.
## Things NOT to change without re-running the full test matrix
These values were tuned empirically and a change to any one of them
can regress 1-3 cases in the parity matrix. If you have a reason to
change them, run `tests/test_pipeline.py` before AND after.
| File | Value | Why it's load-bearing |
|---|---|---|
| `pipeline._adaptive_k` | `5/7/8/10/12` thresholds | Smaller pages need larger k for rank noise; bigger pages would be slow |
| `chunking.chunk_markdown` | `max_chars=450` | Larger chunks hurt recall (diffuses fact density) |
| `chunking.MIN_PLAIN_CHARS` | `20` | Smaller → keeps noise; larger → drops useful short chunks |
| `retrieval.EMBEDDING_BATCH` | `64` | Requires llama-server `--ubatch-size ≥ 2048` |
| `retrieval.MAX_EMBED_INPUT_CHARS` | `1800` | Safety net for the same ubatch ceiling |
| `fetchers/playwright.py wait_for_ms` | `5000` | Fallback ceiling (not fixed wait) for the content-ready detector on brand-new hosts. C9's `host_stats` takes over once a host has ≥ 5 observations, replacing this with `p95 × 1.5` clamped to `[1500, 15000] ms`. Set `TRAWL_HOST_STATS=0` to revert to the static 5000 ms. Change requires re-running the parity matrix. |
| `host_stats.py` (WINDOW_SIZE, MIN_OBSERVATIONS, CEILING_MULTIPLIER, MIN_CEILING_MS, MAX_CEILING_MS) | `50, 5, 1.5, 1500, 15000` | Per-host adaptive ceiling bounds. Not env-configurable — retuning should go through a data-driven spike and a CHANGELOG entry. |
| `fetchers/playwright.py NETWORKIDLE_BUDGET_MS` | `3000` | Max time to wait for `networkidle` before falling back to `domcontentloaded`. Discourse/chat SPAs hold websockets so networkidle never fires — short budget + content-ready detector gives same HTML much faster (telemetry: NVIDIA forum 17s → 4.4s). Raising it re-introduces the regression. |
| `fetchers/playwright.py` content-ready predicate | `stableTicks >= 4`, `polling=150ms`, `len > 100`, placeholder regex | Empirically tuned on the parity matrix for a 67% avg fetch_ms reduction. Tightening the window or raising `len` can regress fast/short pages. |
| `extraction.py` three-way max (precise, recall, bs) | order matters | Pricing pages need BS; articles need precise |
| `hyde.py DEFAULT_LLAMA_URL` | `:8082` | Utility LLM, not main LLM — slot contention risk on :8080 |
| `hyde.py chat_template_kwargs.enable_thinking` | `False` | Without it Gemma 4 burns all tokens on reasoning and returns empty content |
| `profiles/vlm.py chat_template_kwargs.enable_thinking` | `False` | Same Gemma 4 quirk as hyde.py |
| `pipeline.PROFILE_TRANSFER_MIN_RATIO` | `0.3` | Lower bound of acceptable subtree size ratio for host-transfer. Empirically validated on Google Finance (actual ratios 1.5-1.6x) |
| `pipeline.PROFILE_TRANSFER_MAX_RATIO` | `3.0` | Upper bound. Raising admits accidental `<body>`-level selector climbs |
| `reranking.py HTTP_TIMEOUT_S` | `30.0` | Reranker timeout; 20 pairs should complete well within this |
| `reranking.py DEFAULT_MAX_DOCS / DEFAULT_MAX_PER_DOC_CHARS / DEFAULT_MAX_CHARS / MIN_PER_DOC_CHARS` | `30 / 1500 / 40000 / 200` | Defensive chunk-window caps. Total-chars empirically bracketed between 40k (PASS) and 50k (FAIL) on the 8 192-token total context limit. Per-doc-chars empirically bracketed at cap=1500 PASS / cap=1550 FAIL on the captured MDN Fetch_API payload (~3.0-3.5 chars/token for code-heavy English). CJK validated 2026-04-21 (Korean 이순신 wiki + Japanese 寿司 wiki, 0/400 failures at default 1500): chunker-level `max_chars=450` combined with denser CJK sentence boundaries caps observed CJK chunks near ~300 chars, well under the cap boundary. Overridable via `TRAWL_RERANK_MAX_DOCS` / `TRAWL_RERANK_MAX_PER_DOC_CHARS` / `TRAWL_RERANK_MAX_CHARS`. If a Chinese or future CJK page reproducibly trips the per-doc 500, lower this default to ~1000. |
| `reranking.py rerank()` return shape | `tuple[list[ScoredChunk], bool]` | `(scored, capped)`. The boolean drives `PipelineResult.rerank_capped` and the `rerank_capped` JSONL telemetry key. Refactors that drop the second element silently lose the cap-fire signal. Library-internal API only; `fetch_relevant()` is unaffected. |
| `pipeline.py retrieve_k multiplier` | `2` | Retrieves 2x candidates for reranking; fewer reduces rerank benefit, more adds latency |
| `profiles/mapper.py DEFAULT_MAX_CANDIDATES_PER_ANCHOR` | `5` | Enough headroom to find non-noise candidates after sidebar/nav filtering |
| `profiles/mapper.py NOISE_CLS_RE` | `nav\|sidebar\|toc\|...` | Noise region detection for anchor filtering; too broad catches content, too narrow misses sidebars |
| `fetchers/passthrough.py` | `PASSTHROUGH_MAX_BYTES` env default `262144` | 256 KB ≈ 64K tokens; weather-like APIs fit, larger than local LLM contexts |
## In / out of scope
**In scope**: fetching one page at a time, extracting its relevant
parts for a given query, returning structured chunks. Targeting
MCP-compatible agents as the primary consumers.
**Out of scope**:
- Crawling (following links). trawl fetches one URL, that's it.
- Search (query → URL list). Use a separate web search tool.
- Commercial anti-bot bypass (DataDome, Cloudflare Turnstile with
proof-of-work). Passive challenges work via stealth; active ones
need a paid service.
- Content rewriting, summarisation, translation. Those belong in the
downstream agent, not in trawl.
If someone asks to add crawling or search to trawl, push back. Those
are different tools with different failure modes.
## Getting unblocked
If a change breaks the parity matrix and you don't know why:
1. Run the failing case with `--verbose` to see the returned chunks.
2. Compare the fetched markdown to what the same URL produced before
your change — the fetcher, extraction, and chunker each have
isolated smoke tests you can run ad-hoc via `python -c "..."`.
3. If the failure involves specific facts missing from top-k, look
at where those facts rank in the full retrieval (not just top-k).
Often the fix is k, not the extraction.
4. If the embedding server has changed (new model, different
quantisation, different context size), most tuning assumptions
in this file need re-verification.
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

