ketch
1broseidon/ketch/AGENTS.md
Fast, stateless CLI for agentic web search and scrape.
AGENTS.md644 starsChanged 3 months ago
# Ketch — Architecture Fast, stateless CLI for agentic web search and scrape. ## Module Layout ``` main.go Thin entry point → cmd.Execute() cmd/ root.go Cobra root, global flag (--json) search.go Search command: query → results, optional --scrape scrape.go Scrape command: URLs → markdown, concurrent batch extract.go Extract command: piped HTML → markdown (no fetch/cache/browser) crawl.go Crawl command: BFS/sitemap crawl with streaming output crawl_bg.go Background crawl: status, stop subcommands, worker mode code.go Code search command: query → snippet results, --lang/--repo filters, literal-qualifier warnings docs.go Docs search command: query → docs/snippet results, --library, --resolve config.go Config command: discovery, init, set, path cache.go Cache command: stats (page cache and tag index), clear tag.go Tag command: add, show, list, remove over the durable tag index browser.go Browser command: install, status doctor.go Doctor command: report formatting + exit-code gating over doctor.Run mcp.go MCP command: `mcp serve` runs the MCP server over stdio proc_unix.go Unix process management (detach, signals) proc_windows.go Windows process management stub search/ Searcher interface + Brave/DDG/SearXNG/EXA/Firecrawl/Keenable/Tavily/Parallel/SerpBase/Degoog/Serply/Youcom backends; NewFromConfig resolves the ordered provider registry for cmd/ and mcp/. auto.go is the default `auto` backend (keyless fallback chain, AutoRank-ordered), multi.go adds federated --multi search (RRF fusion, NewMultiFromConfig), random.go shuffled fallback, canonical.go the URL dedup keys code/ code.Searcher interface + GrepApp/Sourcegraph/GitHub backends; NewFromConfig resolves the ordered provider registry. Query.Repo (owner/name, filter.go) is exact on every backend: each translates it and narrows a looser native filter itself docs/ docs.Searcher interface + Context7 backend (FTS5 local is an unimplemented stub); NewFromConfig resolves the ordered provider registry mcp/ MCP server (search/code/docs/scrape/crawl/tag tools; the mcp_tools config key is an allowlist over the published set) over the go-sdk mcp package; Server struct holds the shared scraper + cache, tools call the same NewFromConfig constructors as the CLI scrape/ HTTP fetch + Page type, JS detection fallback, Rod browser; pipeline.go has the cache-aware scrape pipeline (CachedScrape*, ScrapeSelector, FetchLLMSTxt) shared by cmd/ and mcp/ extract/ structural extraction (landmark → uniform sections → prose root, chrome pruned by what it is; config extract_mode clean|complete) with readability fallback and html-to-markdown, charset decoding, JS shell detection (Detector: built-in + config spa_markers, modern hydration/streaming frameworks) crawl/ BFS crawler, work queue + worker pool, background status cookies/ Netscape cookies.txt jar loader + RFC 6265 domain/path matching (Jar.For); nil-safe, values never logged config/ JSON config loading/saving (~/.config/ketch/) doctor/ Health checks: concurrent read-only probes per backend + browser + cache + tag index (informational, never gates the exit code), status classification (ok/no_key/unreachable/misconfigured/skipped) health/ Shared bounded provider health checks (Status classification, HTTP probe helpers) used by doctor and the search/code/docs provider probes cache/ TTL page cache (Store interface, BBoltStore backend) and durable bookmarks in a separate tags.db (tags.go/tag_store.go). Both open their bbolt file per transaction through boltfile.go and never hold it between operations (ADR-0006); writes sweep expired pages at most hourly, and cache clear deletes the file without touching bookmarks httpx/ Shared tuned *http.Transport for all HTTP backends updatecheck/ "new release available" probe + throttled stderr hint site/ Builds ketch.run from MANUAL.md with inkcap (GitHub Pages, docs.yml) ``` Reusable packages live at the module root so external programs can `import "github.com/1broseidon/ketch/<pkg>"`. Shared implementation helpers, including the config model used by provider packages, live under `internal/`. ## Design Principles - **Stateless**: no daemon, no queue. Call → result → done. Background crawls are detached processes, not a server. - **Fast path first**: plain HTTP fetch; browser rendering kicks in automatically when JS shell is detected. - **Interface-driven backends**: `Searcher` for search engines, `Store` for cache backends, `BrowserConn` for browser rendering. - **Concurrent by default**: multiple URLs scraped in parallel via goroutines. - **Operator configures, agent consumes**: config sets defaults (backend, browser, cache TTL) so agents don't need to know infrastructure. - **Works before it is configured**: `ketch search` answers on a fresh install. The default `auto` backend is a fallback chain over the keyless providers, and configuring a key or an instance promotes that provider rather than requiring `backend` to be set too. Configuration raises limits and picks favourites; it is never the price of a first result. - **Three search surfaces**: `ketch search` finds web pages, `ketch code` greps real OSS code, `ketch docs` fetches library documentation. Each has its own backend interface and Result type — they never share backends. - **Smart input detection on scrape**: single URL, multiple positional args, JSON array string, file path, or stdin pipe all work — ketch routes automatically. No --batch flag needed. - **Context-aware interfaces**: all three Searcher interfaces (`search`, `code`, `docs`) take `context.Context` as first param for cancellation and timeout propagation. The reasoning behind each principle — and what ketch deliberately does *not* do — is in [design/DESIGN.md](design/DESIGN.md) ([Non-Goals & Scope](design/DESIGN.md#non-goals--scope)). Possible directions: [design/ROADMAP.md](design/ROADMAP.md). Decisions already made: [design/adr/](design/adr/). ## MCP Server `ketch mcp serve` runs an MCP (Model Context Protocol) server over stdio, exposing six tools: `search`, `code`, `docs`, `scrape`, `crawl`, and `tag`. The `mcp_tools` config key is an allowlist over that set (JSON array or comma-separated; unset or `[]` publishes all six) — unlisted tools are never registered, and `serverInstructions` is generated from the enabled set so a pruned server never advertises what it won't answer. Values are validated fail-loud at `config set`, on the `KETCH_MCP_TOOLS` env override, and at server startup. Tool handlers call the same packages as the Cobra commands, through the same config-driven constructors (`search.NewFromConfig` etc.), and resolve backends/API keys from the same `~/.config/ketch/` config — an agent talking MCP sees exactly what a human using the CLI sees. - **Lifecycle**: the go-sdk dispatches tool calls concurrently, so process-lifetime resources — the headless-browser scraper and the compiled URL rewriter — are constructed once in `mcp.NewServer`, shared by all calls, and released by `Server.Close` when `serve` exits. Never construct these per call. The page cache and the tag index are the opposite: the server keeps their paths and opens each bbolt file only for one transaction, so an idle server holds no lock and the CLI, crawls and other servers share the cache (ADR-0006). Never hold a bbolt handle across calls, and release the tag index before checking page warmth. - **Option parity**: each tool exposes the per-invocation options of its CLI command (`scrape` gets `selector`/`raw`/`force_browser`/`no_llms_txt`/`trim`/`max_chars`/`no_cache` plus a `urls` batch input; `search` gets `searxng_url`, `scrape`, and `multi` (federated RRF search, with an additive `errors` map for per-backend failures); `crawl` gets `depth`/`sitemap`/`allow`/`deny`/`max_pages`; `code` gets `lang`/`repo`/`regexp`). Config-level settings (API keys, cache TTL, browser binary) stay operator-configured and are never tool params. - **Error taxonomy**: every tool error starts with a stable machine-readable prefix mirroring the CLI exit codes — `[validation]` (exit 2), `[not_found]` (3), `[upstream]` (4), `[precondition]` (5), `[cancelled]` (6) — so agents can tell "fix your input" from "retry later". MCP has no structured tool-error field; the prefix is the contract. - **Bounded crawl**: the `crawl` tool is synchronous and capped (`max_pages` default 30, hard cap 100, 3-minute wall clock); partial results return with `stopped: "max_pages" | "timeout"`. Detached background crawls (`ketch crawl --background`, status/stop) remain CLI-only. - **CLI-only operator commands**: `config`, `cache`, and `doctor` are deliberately not MCP tools. They are operator actions (change credentials, clear state, diagnose the installation), not research surfaces — an agent that needs to know whether a backend is ready reads `ketch config`'s `*_set` booleans or the operator runs `ketch doctor`. Don't add them to the server. - **Annotations**: the five fetching tools are read-only network fetchers and declare `readOnlyHint: true` and `openWorldHint: true`. `tag` is the exception in both directions — it mutates the local index and never touches the network — so it declares `readOnlyHint: false` and `openWorldHint: false`. - **Tagging**: `tag` supports add/show/list/remove, plus a `tag` option on all five research tools. Show defaults to 50 newest entries (`limit: 0` for all), with `entries`/`cached` totals, `shown`, and `cache_status`. The independent index uses `os.UserConfigDir()/ketch/tags.db` or `KETCH_TAGS_PATH`; labs must override the tag path as well as the page cache. Membership follows the source URL, never cookies/UA/rewrite keys. Expired bodies leave durable bookmarks; an unavailable page cache returns `cached: false` with `cache_status: unavailable`. Optional write failures retain research output with MCP `warnings` or CLI stderr warnings. See [ADR-0004](design/adr/0004-tagged-cache-corpus.md). - **Security note**: the server performs no URL filtering — `scrape` and `crawl` fetch whatever URL the client supplies, including private or internal addresses reachable from wherever the server runs (their descriptions say so). Run it with the network posture you'd give the agent itself; don't point an untrusted agent at a server inside a sensitive network. - **Smoke test**: `go test -tags mcpsmoke ./mcp/... -v` exercises the real binary over stdio (live network; not part of `go test ./...`). ## Quality Standards - `golangci-lint run` must pass (gocyclo max 15) - `go test ./...` must pass - Extraction changes also run the [ketch-bench](https://github.com/1broseidon/ketch-bench) regression check in both modes (`go run . check` and `go run . check -extract-mode clean` from a sibling checkout); it is a separate repository so its corpus never weighs on this one - Pre-commit hook enforces both - CGO_ENABLED=0 — pure Go, cross-compile everywhere ## CLI Usage ``` ketch search "query" # search, return results (auto backend, no key needed) ketch search "query" --scrape # search + fetch full content ketch search "query" -b auto # keyless fallback chain (the default) ketch search "query" -b searxng # use SearXNG backend ketch search "query" -b exa # use Exa hosted MCP backend ketch search "query" -b firecrawl # use Firecrawl v2 search API (keyless by default) ketch search "query" -b keenable # use Keenable backend (keyless by default) ketch search "query" -b tavily # use Tavily search API (keyed; basic depth) ketch search "query" -b parallel # use Parallel Search MCP (keyless) ketch search "query" -b serpbase # use SerpBase Google Search API (keyed) ketch search "query" -b degoog # use a self-hosted Degoog instance (degoog_url) ketch search "query" -b serply # use Serply's Google Search API (keyed) ketch search "query" -b youcom # use You.com web search (keyless by default) ketch search "query" --multi # federate across every usable backend, RRF-fused ketch search "query" --multi=brave,ddg,exa # federate across a specific set (use the = form) ketch scrape <url> # single URL → markdown ketch scrape <url1> <url2> <url3> # concurrent batch scrape ketch scrape urls.txt # file with one URL per line ketch scrape '["url1","url2"]' # JSON array of URLs echo "url1\nurl2" | ketch scrape # stdin pipe curl -L https://example.com | ketch extract # piped HTML → markdown (no fetch) ketch crawl <url> # BFS crawl ketch crawl <url> --sitemap # sitemap-based crawl ketch crawl <url> --background # run in background ketch crawl status [id] # check crawl progress ketch crawl stop <id> # stop a background crawl ketch browser status # check browser config ketch browser install # download Chromium ketch code "query" # code search (grepapp) ketch code "query" --lang go # with language filter ketch code "query" --repo owner/name # one repository, exact on every backend ketch docs "query" # docs search (context7) ketch docs "query" --library /org/repo # skip resolve, fetch directly ketch docs --resolve "library name" # resolve library name → Context7 IDs ketch config # show effective config + backends (incl. *_key_set presence booleans) ketch code "query" --tag remote-access # record results under a tag (any surface) ketch scrape <url> --tag docs # fetch and record the page under a tag ketch tag add docs <url>... # tag URLs directly, cached or not (no network) ketch tag show docs # newest 50 bookmarks, with total counts ketch tag show docs --limit 0 # all bookmarks under a tag ketch tag list # every tag, with entry counts ketch tag remove docs [url...] # drop a tag, or just those pages from it ketch cache # show page-cache and tag-index stats ketch doctor # live health check of every backend + browser + cache + tag index (exit 5 if a configured surface is broken) ketch mcp serve # run as an MCP server over stdio (search/code/docs/scrape/crawl/tag tools) ``` ## Flags | Flag | Scope | Default | Description | |------|-------|---------|-------------| | --json | global | false | JSON output | | --backend, -b | search | auto | Search backend (auto/brave/ddg/searxng/exa/firecrawl/keenable/tavily/parallel/serpbase/degoog/serply/youcom); `auto` is a fallback chain, not a provider | | --multi | search | — | Federated search: comma list or bare/`=all` for every usable backend; RRF-fused, dedup'd, mutually exclusive with --backend (use the `=` form for a list) | | --limit, -l | search | 5 | Max results | | --scrape | search | false | Fetch full content | | --searxng-url | search | http://localhost:8081 | SearXNG URL | | --raw | scrape | false | Raw HTML output | | --no-cache | scrape, crawl | false | Bypass page cache | | --depth | crawl | 3 | Max BFS depth | | --concurrency | crawl | 8 | Worker pool size | | --sitemap | crawl | false | Treat seed URL as sitemap | | --background | crawl | false | Run in background | | --allow | crawl | — | Path substring filters | | --deny | crawl | — | Regex deny patterns | | --backend, -b | code | grepapp | Code backend (grepapp/sourcegraph/github) | | --backend, -b | docs | context7 | Docs backend (context7; local is planned, not implemented) | | --lang | code | — | Language qualifier (appended to query) | | --repo | code | — | One repository, owner/name or a GitHub URL; exact on every backend. A backend that lacks it exits 3 (sourcegraph, github) or warns on an empty result (grepapp, partial index) | | --library | docs | — | Context7 library ID, skips resolve | | --tokens | docs | 4000 | Context7 token budget | | --resolve | docs | false | Resolve library name instead of searching | | --max-chars N | scrape, search --scrape | 0 (off) | Truncate markdown output to N chars, appends `[truncated]` | | --trim | scrape, search --scrape | false | Strip markdown formatting syntax, keep content text only | | --minimal | search, code, docs | false | One result per line, tab-separated, no frontmatter (a 4th backends column is appended under `search --multi`, plain search only) | | --select \<css\> | scrape | — | Extract only elements matching CSS selector (skips content selection) | | --url | extract | — | Source URL for metadata and relative-link resolution (no fetch) | | --select \<css\> | extract | — | CSS selector to extract (skips content selection) | | --trim | extract | false | Strip markdown formatting, keep content text only | | --max-chars N | extract | 0 (off) | Truncate markdown output to N chars, appends `[truncated]` | | --no-llms-txt | scrape | false | Disable automatic /llms.txt detection for bare domains | | --concurrency | scrape | 5 | Max concurrent requests for multi-URL scraping | | --force-browser | scrape | false | Always render via the configured browser, skipping JS-shell auto-detection (composes with --raw/--select; errors without a browser) | | --tag <name> | search, code, docs, scrape, crawl | — | Bookmark each source URL with a bounded title and description derived from the result or fetched page. Composes with --no-cache; bodies remain in the separate page cache | | --limit | tag show | 50 | Newest entries to return; 0 returns all. Reports returned and total counts | | --minimal | tag show | false | One page per line, tab-separated (url/title/description) | | --cookie-file <path> | scrape, search --scrape, crawl | config `cookie_file` or off | Netscape cookies.txt jar; flag overrides config and an explicit empty value disables cookies | | --user-agent <ua> | scrape, search --scrape, crawl | config `user_agent` or built-in default | User-Agent override applied to HTTP and browser fetches; flag overrides config and an explicit empty value restores each fetch path's default. A configured UA is folded into the page-cache key, so pages cached under one UA are not reused under another | ## Adding a provider Use the provider registry for new search, code, and docs backends. Read [CONTRIBUTING.md](CONTRIBUTING.md#proposing-a-provider) for admission criteria. Place the implementation, descriptor, and health probe in one Go file under `search/`, `code/`, or `docs/`, with tests alongside it. Append its descriptor call to the package's ordered `providers` slice in `registry.go`; register explicitly, without `init()`. Config discovery, CLI/MCP descriptions, doctor, and search auto/multi/random eligibility follow the descriptor. Keep provider-specific branches and config fields out of shared consumers. - Define `ID`, `Name`, `Settings`, `Usable`, `New`, and `Probe`. - Search providers set `AutoRank` to join the default `auto` chain: zero keeps the provider out, and lower non-zero ranks are tried first. Rank is only the tiebreak — a provider with configured credentials is promoted ahead of the rest. Rank a keyless provider by how well it serves an unconfigured install; leave a provider out if it cannot answer without setup. Use `AutoEligible` only when `Usable` is deliberately laxer than chain membership should be (SearXNG's localhost default is the one case). - `Usable` checks configuration without network I/O. `Build` checks it before calling `New`. - `New` only constructs a client and must accept empty credentials. `Probe` performs the health check; construction must not contact the service. - Import `internal/configbase` as `config` to avoid an import cycle with the public config facade. Read values with `String`/`Strings`; use `SetProvider` for overrides so per-call changes do not mutate shared MCP config. - Declare settings on the descriptor: `config.KeyPool` for rotating credentials, `config.Scalar` for URLs, or `config.Setting` for custom secret/token behavior. Keep existing order values; new settings use the helpers' default ordering. Existing provider-specific accessors are compatibility helpers, not a pattern to extend. - Doctor checks are required when the provider is selected or a setting marked `GateDoctor` is configured. Missing credentials or an instance URL must fail a selected provider's check. Set `MinProbeTimeout` for slower search probes. - Code providers declare regex support in the descriptor, set `Qualifiers` when they apply `repo:`/`lang:` written into the query (without it, callers warn that such tokens were searched as text), and set `PartialIndex` when they index only some repositories. `Query.Repo` names exactly one repository: translate it into the provider's filter, narrow a looser match (substring, unanchored pattern) yourself, and return `code.ErrRepoNotFound` when the service says the repository is missing. Docs providers can implement `docs.LibraryResolver` for library operations. Keep the unfinished local docs provider hidden. Test requests, result mapping, authentication, cancellation, and relevant retry and error behavior without live services. Run `make lint` and `make test`. Regenerate config/doctor fixtures with `UPDATE_REGISTRY_GOLDENS=1 go test ./cmd/ ./doctor/`; preserve existing entries and add only the new provider's entries. Refactors must leave fixtures unchanged. See [the registration test](search/registry_test.go) for a provider flowing through shared consumers. Include documentation and changelog updates as described in [CONTRIBUTING.md](CONTRIBUTING.md#implementing-a-provider).
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

