agentleFS
Sign inSign up

ketch

1broseidon/ketch/AGENTS.md

Fast, stateless CLI for agentic web search and scrape.

AGENTS.md644 starsChanged 3 months ago
# Ketch — Architecture

Fast, stateless CLI for agentic web search and scrape.

## Module Layout

```
main.go                      Thin entry point → cmd.Execute()
cmd/
  root.go                    Cobra root, global flag (--json)
  search.go                  Search command: query → results, optional --scrape
  scrape.go                  Scrape command: URLs → markdown, concurrent batch
  extract.go                 Extract command: piped HTML → markdown (no fetch/cache/browser)
  crawl.go                   Crawl command: BFS/sitemap crawl with streaming output
  crawl_bg.go                Background crawl: status, stop subcommands, worker mode
  code.go                    Code search command: query → snippet results, --lang/--repo filters, literal-qualifier warnings
  docs.go                    Docs search command: query → docs/snippet results, --library, --resolve
  config.go                  Config command: discovery, init, set, path
  cache.go                   Cache command: stats (page cache and tag index), clear
  tag.go                     Tag command: add, show, list, remove over the durable tag index
  browser.go                 Browser command: install, status
  doctor.go                  Doctor command: report formatting + exit-code gating over doctor.Run
  mcp.go                     MCP command: `mcp serve` runs the MCP server over stdio
  proc_unix.go               Unix process management (detach, signals)
  proc_windows.go            Windows process management stub
search/                      Searcher interface + Brave/DDG/SearXNG/EXA/Firecrawl/Keenable/Tavily/Parallel/SerpBase/Degoog/Serply/Youcom backends; NewFromConfig resolves the ordered provider registry for cmd/ and mcp/. auto.go is the default `auto` backend (keyless fallback chain, AutoRank-ordered), multi.go adds federated --multi search (RRF fusion, NewMultiFromConfig), random.go shuffled fallback, canonical.go the URL dedup keys
code/                        code.Searcher interface + GrepApp/Sourcegraph/GitHub backends; NewFromConfig resolves the ordered provider registry. Query.Repo (owner/name, filter.go) is exact on every backend: each translates it and narrows a looser native filter itself
docs/                        docs.Searcher interface + Context7 backend (FTS5 local is an unimplemented stub); NewFromConfig resolves the ordered provider registry
mcp/                         MCP server (search/code/docs/scrape/crawl/tag tools; the mcp_tools config key is an allowlist over the published set) over the go-sdk mcp package; Server struct holds the shared scraper + cache, tools call the same NewFromConfig constructors as the CLI
scrape/                      HTTP fetch + Page type, JS detection fallback, Rod browser; pipeline.go has the cache-aware scrape pipeline (CachedScrape*, ScrapeSelector, FetchLLMSTxt) shared by cmd/ and mcp/
extract/                     structural extraction (landmark → uniform sections → prose root, chrome pruned by what it is; config extract_mode clean|complete) with readability fallback and html-to-markdown, charset decoding, JS shell detection (Detector: built-in + config spa_markers, modern hydration/streaming frameworks)
crawl/                       BFS crawler, work queue + worker pool, background status
cookies/                     Netscape cookies.txt jar loader + RFC 6265 domain/path matching (Jar.For); nil-safe, values never logged
config/                      JSON config loading/saving (~/.config/ketch/)
doctor/                      Health checks: concurrent read-only probes per backend + browser + cache + tag index (informational, never gates the exit code), status classification (ok/no_key/unreachable/misconfigured/skipped)
health/                      Shared bounded provider health checks (Status classification, HTTP probe helpers) used by doctor and the search/code/docs provider probes
cache/                       TTL page cache (Store interface, BBoltStore backend) and durable bookmarks in a separate tags.db (tags.go/tag_store.go). Both open their bbolt file per transaction through boltfile.go and never hold it between operations (ADR-0006); writes sweep expired pages at most hourly, and cache clear deletes the file without touching bookmarks
httpx/                       Shared tuned *http.Transport for all HTTP backends
updatecheck/                 "new release available" probe + throttled stderr hint
site/                        Builds ketch.run from MANUAL.md with inkcap (GitHub Pages, docs.yml)
```

Reusable packages live at the module root so external programs can `import "github.com/1broseidon/ketch/<pkg>"`. Shared implementation helpers, including the config model used by provider packages, live under `internal/`.

## Design Principles

- **Stateless**: no daemon, no queue. Call → result → done. Background crawls are detached processes, not a server.
- **Fast path first**: plain HTTP fetch; browser rendering kicks in automatically when JS shell is detected.
- **Interface-driven backends**: `Searcher` for search engines, `Store` for cache backends, `BrowserConn` for browser rendering.
- **Concurrent by default**: multiple URLs scraped in parallel via goroutines.
- **Operator configures, agent consumes**: config sets defaults (backend, browser, cache TTL) so agents don't need to know infrastructure.
- **Works before it is configured**: `ketch search` answers on a fresh install. The default `auto` backend is a fallback chain over the keyless providers, and configuring a key or an instance promotes that provider rather than requiring `backend` to be set too. Configuration raises limits and picks favourites; it is never the price of a first result.
- **Three search surfaces**: `ketch search` finds web pages, `ketch code` greps real OSS code, `ketch docs` fetches library documentation. Each has its own backend interface and Result type — they never share backends.
- **Smart input detection on scrape**: single URL, multiple positional args, JSON array string, file path, or stdin pipe all work — ketch routes automatically. No --batch flag needed.
- **Context-aware interfaces**: all three Searcher interfaces (`search`, `code`, `docs`) take `context.Context` as first param for cancellation and timeout propagation.

The reasoning behind each principle — and what ketch deliberately does *not* do — is in [design/DESIGN.md](design/DESIGN.md) ([Non-Goals & Scope](design/DESIGN.md#non-goals--scope)). Possible directions: [design/ROADMAP.md](design/ROADMAP.md). Decisions already made: [design/adr/](design/adr/).

## MCP Server

`ketch mcp serve` runs an MCP (Model Context Protocol) server over stdio, exposing six tools: `search`, `code`, `docs`, `scrape`, `crawl`, and `tag`. The `mcp_tools` config key is an allowlist over that set (JSON array or comma-separated; unset or `[]` publishes all six) — unlisted tools are never registered, and `serverInstructions` is generated from the enabled set so a pruned server never advertises what it won't answer. Values are validated fail-loud at `config set`, on the `KETCH_MCP_TOOLS` env override, and at server startup. Tool handlers call the same packages as the Cobra commands, through the same config-driven constructors (`search.NewFromConfig` etc.), and resolve backends/API keys from the same `~/.config/ketch/` config — an agent talking MCP sees exactly what a human using the CLI sees.

- **Lifecycle**: the go-sdk dispatches tool calls concurrently, so process-lifetime resources — the headless-browser scraper and the compiled URL rewriter — are constructed once in `mcp.NewServer`, shared by all calls, and released by `Server.Close` when `serve` exits. Never construct these per call. The page cache and the tag index are the opposite: the server keeps their paths and opens each bbolt file only for one transaction, so an idle server holds no lock and the CLI, crawls and other servers share the cache (ADR-0006). Never hold a bbolt handle across calls, and release the tag index before checking page warmth.
- **Option parity**: each tool exposes the per-invocation options of its CLI command (`scrape` gets `selector`/`raw`/`force_browser`/`no_llms_txt`/`trim`/`max_chars`/`no_cache` plus a `urls` batch input; `search` gets `searxng_url`, `scrape`, and `multi` (federated RRF search, with an additive `errors` map for per-backend failures); `crawl` gets `depth`/`sitemap`/`allow`/`deny`/`max_pages`; `code` gets `lang`/`repo`/`regexp`). Config-level settings (API keys, cache TTL, browser binary) stay operator-configured and are never tool params.
- **Error taxonomy**: every tool error starts with a stable machine-readable prefix mirroring the CLI exit codes — `[validation]` (exit 2), `[not_found]` (3), `[upstream]` (4), `[precondition]` (5), `[cancelled]` (6) — so agents can tell "fix your input" from "retry later". MCP has no structured tool-error field; the prefix is the contract.
- **Bounded crawl**: the `crawl` tool is synchronous and capped (`max_pages` default 30, hard cap 100, 3-minute wall clock); partial results return with `stopped: "max_pages" | "timeout"`. Detached background crawls (`ketch crawl --background`, status/stop) remain CLI-only.
- **CLI-only operator commands**: `config`, `cache`, and `doctor` are deliberately not MCP tools. They are operator actions (change credentials, clear state, diagnose the installation), not research surfaces — an agent that needs to know whether a backend is ready reads `ketch config`'s `*_set` booleans or the operator runs `ketch doctor`. Don't add them to the server.
- **Annotations**: the five fetching tools are read-only network fetchers and declare `readOnlyHint: true` and `openWorldHint: true`. `tag` is the exception in both directions — it mutates the local index and never touches the network — so it declares `readOnlyHint: false` and `openWorldHint: false`.
- **Tagging**: `tag` supports add/show/list/remove, plus a `tag` option on all five research tools. Show defaults to 50 newest entries (`limit: 0` for all), with `entries`/`cached` totals, `shown`, and `cache_status`. The independent index uses `os.UserConfigDir()/ketch/tags.db` or `KETCH_TAGS_PATH`; labs must override the tag path as well as the page cache. Membership follows the source URL, never cookies/UA/rewrite keys. Expired bodies leave durable bookmarks; an unavailable page cache returns `cached: false` with `cache_status: unavailable`. Optional write failures retain research output with MCP `warnings` or CLI stderr warnings. See [ADR-0004](design/adr/0004-tagged-cache-corpus.md).
- **Security note**: the server performs no URL filtering — `scrape` and `crawl` fetch whatever URL the client supplies, including private or internal addresses reachable from wherever the server runs (their descriptions say so). Run it with the network posture you'd give the agent itself; don't point an untrusted agent at a server inside a sensitive network.
- **Smoke test**: `go test -tags mcpsmoke ./mcp/... -v` exercises the real binary over stdio (live network; not part of `go test ./...`).

## Quality Standards

- `golangci-lint run` must pass (gocyclo max 15)
- `go test ./...` must pass
- Extraction changes also run the [ketch-bench](https://github.com/1broseidon/ketch-bench) regression check in both modes (`go run . check` and `go run . check -extract-mode clean` from a sibling checkout); it is a separate repository so its corpus never weighs on this one
- Pre-commit hook enforces both
- CGO_ENABLED=0 — pure Go, cross-compile everywhere

## CLI Usage

```
ketch search "query"                        # search, return results (auto backend, no key needed)
ketch search "query" --scrape               # search + fetch full content
ketch search "query" -b auto                # keyless fallback chain (the default)
ketch search "query" -b searxng             # use SearXNG backend
ketch search "query" -b exa                 # use Exa hosted MCP backend
ketch search "query" -b firecrawl           # use Firecrawl v2 search API (keyless by default)
ketch search "query" -b keenable            # use Keenable backend (keyless by default)
ketch search "query" -b tavily              # use Tavily search API (keyed; basic depth)
ketch search "query" -b parallel            # use Parallel Search MCP (keyless)
ketch search "query" -b serpbase            # use SerpBase Google Search API (keyed)
ketch search "query" -b degoog              # use a self-hosted Degoog instance (degoog_url)
ketch search "query" -b serply              # use Serply's Google Search API (keyed)
ketch search "query" -b youcom              # use You.com web search (keyless by default)
ketch search "query" --multi                # federate across every usable backend, RRF-fused
ketch search "query" --multi=brave,ddg,exa  # federate across a specific set (use the = form)
ketch scrape <url>                          # single URL → markdown
ketch scrape <url1> <url2> <url3>           # concurrent batch scrape
ketch scrape urls.txt                       # file with one URL per line
ketch scrape '["url1","url2"]'              # JSON array of URLs
echo "url1\nurl2" | ketch scrape            # stdin pipe
curl -L https://example.com | ketch extract # piped HTML → markdown (no fetch)
ketch crawl <url>                           # BFS crawl
ketch crawl <url> --sitemap                 # sitemap-based crawl
ketch crawl <url> --background              # run in background
ketch crawl status [id]                     # check crawl progress
ketch crawl stop <id>                       # stop a background crawl
ketch browser status                        # check browser config
ketch browser install                       # download Chromium
ketch code "query"                          # code search (grepapp)
ketch code "query" --lang go               # with language filter
ketch code "query" --repo owner/name       # one repository, exact on every backend
ketch docs "query"                          # docs search (context7)
ketch docs "query" --library /org/repo     # skip resolve, fetch directly
ketch docs --resolve "library name"        # resolve library name → Context7 IDs
ketch config                                # show effective config + backends (incl. *_key_set presence booleans)
ketch code "query" --tag remote-access      # record results under a tag (any surface)
ketch scrape <url> --tag docs               # fetch and record the page under a tag
ketch tag add docs <url>...                 # tag URLs directly, cached or not (no network)
ketch tag show docs                         # newest 50 bookmarks, with total counts
ketch tag show docs --limit 0               # all bookmarks under a tag
ketch tag list                              # every tag, with entry counts
ketch tag remove docs [url...]              # drop a tag, or just those pages from it
ketch cache                                 # show page-cache and tag-index stats
ketch doctor                                # live health check of every backend + browser + cache + tag index (exit 5 if a configured surface is broken)
ketch mcp serve                             # run as an MCP server over stdio (search/code/docs/scrape/crawl/tag tools)
```

## Flags

| Flag | Scope | Default | Description |
|------|-------|---------|-------------|
| --json | global | false | JSON output |
| --backend, -b | search | auto | Search backend (auto/brave/ddg/searxng/exa/firecrawl/keenable/tavily/parallel/serpbase/degoog/serply/youcom); `auto` is a fallback chain, not a provider |
| --multi | search | — | Federated search: comma list or bare/`=all` for every usable backend; RRF-fused, dedup'd, mutually exclusive with --backend (use the `=` form for a list) |
| --limit, -l | search | 5 | Max results |
| --scrape | search | false | Fetch full content |
| --searxng-url | search | http://localhost:8081 | SearXNG URL |
| --raw | scrape | false | Raw HTML output |
| --no-cache | scrape, crawl | false | Bypass page cache |
| --depth | crawl | 3 | Max BFS depth |
| --concurrency | crawl | 8 | Worker pool size |
| --sitemap | crawl | false | Treat seed URL as sitemap |
| --background | crawl | false | Run in background |
| --allow | crawl | — | Path substring filters |
| --deny | crawl | — | Regex deny patterns |
| --backend, -b | code | grepapp | Code backend (grepapp/sourcegraph/github) |
| --backend, -b | docs | context7 | Docs backend (context7; local is planned, not implemented) |
| --lang | code | — | Language qualifier (appended to query) |
| --repo | code | — | One repository, owner/name or a GitHub URL; exact on every backend. A backend that lacks it exits 3 (sourcegraph, github) or warns on an empty result (grepapp, partial index) |
| --library | docs | — | Context7 library ID, skips resolve |
| --tokens | docs | 4000 | Context7 token budget |
| --resolve | docs | false | Resolve library name instead of searching |
| --max-chars N | scrape, search --scrape | 0 (off) | Truncate markdown output to N chars, appends `[truncated]` |
| --trim | scrape, search --scrape | false | Strip markdown formatting syntax, keep content text only |
| --minimal | search, code, docs | false | One result per line, tab-separated, no frontmatter (a 4th backends column is appended under `search --multi`, plain search only) |
| --select \<css\> | scrape | — | Extract only elements matching CSS selector (skips content selection) |
| --url | extract | — | Source URL for metadata and relative-link resolution (no fetch) |
| --select \<css\> | extract | — | CSS selector to extract (skips content selection) |
| --trim | extract | false | Strip markdown formatting, keep content text only |
| --max-chars N | extract | 0 (off) | Truncate markdown output to N chars, appends `[truncated]` |
| --no-llms-txt | scrape | false | Disable automatic /llms.txt detection for bare domains |
| --concurrency | scrape | 5 | Max concurrent requests for multi-URL scraping |
| --force-browser | scrape | false | Always render via the configured browser, skipping JS-shell auto-detection (composes with --raw/--select; errors without a browser) |
| --tag <name> | search, code, docs, scrape, crawl | — | Bookmark each source URL with a bounded title and description derived from the result or fetched page. Composes with --no-cache; bodies remain in the separate page cache |
| --limit | tag show | 50 | Newest entries to return; 0 returns all. Reports returned and total counts |
| --minimal | tag show | false | One page per line, tab-separated (url/title/description) |
| --cookie-file <path> | scrape, search --scrape, crawl | config `cookie_file` or off | Netscape cookies.txt jar; flag overrides config and an explicit empty value disables cookies |
| --user-agent <ua> | scrape, search --scrape, crawl | config `user_agent` or built-in default | User-Agent override applied to HTTP and browser fetches; flag overrides config and an explicit empty value restores each fetch path's default. A configured UA is folded into the page-cache key, so pages cached under one UA are not reused under another |


## Adding a provider

Use the provider registry for new search, code, and docs backends. Read
[CONTRIBUTING.md](CONTRIBUTING.md#proposing-a-provider) for admission criteria.

Place the implementation, descriptor, and health probe in one Go file under
`search/`, `code/`, or `docs/`, with tests alongside it. Append its descriptor
call to the package's ordered
`providers` slice in `registry.go`; register explicitly, without `init()`.
Config discovery, CLI/MCP descriptions, doctor, and search auto/multi/random
eligibility follow the descriptor. Keep provider-specific branches and config
fields out of shared consumers.

- Define `ID`, `Name`, `Settings`, `Usable`, `New`, and `Probe`.
- Search providers set `AutoRank` to join the default `auto` chain: zero keeps
  the provider out, and lower non-zero ranks are tried first. Rank is only the
  tiebreak — a provider with configured credentials is promoted ahead of the
  rest. Rank a keyless provider by how well it serves an unconfigured install;
  leave a provider out if it cannot answer without setup. Use `AutoEligible`
  only when `Usable` is deliberately laxer than chain membership should be
  (SearXNG's localhost default is the one case).
- `Usable` checks configuration without network I/O. `Build` checks it before
  calling `New`.
- `New` only constructs a client and must accept empty credentials. `Probe`
  performs the health check; construction must not contact the service.
- Import `internal/configbase` as `config` to avoid an import cycle with the
  public config facade. Read values with `String`/`Strings`; use `SetProvider`
  for overrides so per-call changes do not mutate shared MCP config.
- Declare settings on the descriptor: `config.KeyPool` for rotating credentials,
  `config.Scalar` for URLs, or `config.Setting` for custom secret/token behavior.
  Keep existing order values; new settings use the helpers' default ordering.
  Existing provider-specific accessors are compatibility helpers, not a pattern
  to extend.
- Doctor checks are required when the provider is selected or a setting marked
  `GateDoctor` is configured. Missing credentials or an instance URL must fail
  a selected provider's check. Set `MinProbeTimeout` for slower search probes.
- Code providers declare regex support in the descriptor, set `Qualifiers`
  when they apply `repo:`/`lang:` written into the query (without it, callers
  warn that such tokens were searched as text), and set `PartialIndex` when
  they index only some repositories. `Query.Repo` names exactly one
  repository: translate it into the provider's filter, narrow a looser match
  (substring, unanchored pattern) yourself, and return `code.ErrRepoNotFound`
  when the service says the repository is missing. Docs providers can
  implement `docs.LibraryResolver` for library operations. Keep the unfinished
  local docs provider hidden.

Test requests, result mapping, authentication, cancellation, and relevant retry
and error behavior without live services. Run `make lint` and `make test`.
Regenerate config/doctor fixtures with
`UPDATE_REGISTRY_GOLDENS=1 go test ./cmd/ ./doctor/`; preserve existing entries
and add only the new provider's entries. Refactors must leave fixtures unchanged.
See [the registration test](search/registry_test.go) for a provider flowing
through shared consumers. Include documentation and changelog updates as
described in [CONTRIBUTING.md](CONTRIBUTING.md#implementing-a-provider).

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.