hydra
ankit373/hydra/docs/llms.txt
Hydra is the Trust Control Plane for AI development: an open source CLI written in Go that routes developer tasks across every AI model on your machine to a target confidence of correctness, not just the cheapest one, sampling models adaptively (SPRT) and stopping as soon as the evidence is strong enough, spending more only where accuracy demands it. It also reduces LLM API costs by 70-85% along the way, by using cheap or free models for simple tasks and…
llms.txt10 starsChanged 26 days ago
- Pipes a download into a shell
- Reads credentials
- Installs packages
# Hydra: The Trust Control Plane for AI Development > Hydra is the Trust Control Plane for AI development: an open source CLI written in Go that routes developer tasks across every AI model on your machine to a target *confidence of correctness*, not just the cheapest one, sampling models adaptively (SPRT) and stopping as soon as the evidence is strong enough, spending more only where accuracy demands it. It also reduces LLM API costs by 70-85% along the way, by using cheap or free models for simple tasks and reserving frontier models for work that needs them. ## What It Does - **10-tier routing**: Claude Opus (tier 1, complex reasoning) down to Qwen3 via Ollama (tier 10, free local) - **Interactive TUI cockpit**: `hyctl tui` opens an interactive cockpit whose chat **executes**: type a task and it runs, with a route line naming the tier, model and strategy the router chose and why. Modes: Auto (plan → edit → run your tests → fix → repeat, verified through `internal/oracle`), Plan, Edit, Ask, plus Architect/Careful/Unattended for more control. Six views: Chat · Agents · Models · Activity · Usage · Audit; `?` opens a shortcut glossary rendered from the keymap itself, and `ctrl+o` overrides where the next task runs. `hyctl tui --snapshot [--view 0..5]` renders one static frame for docs/previews - **Desktop app**: a Wails v2 GUI in `desktop/` that opens on **Chat**: ask for work, and each reply says which model answered, at which tier, and what it cost, with the run's timeline narrated live as it happens. The composer's model picker is grouped by **token pool**, so it shows when a choice spends a quota another model shares. Picking a model **pins** it: it is used or refused with a reason, never silently swapped for a different one, and when the router does fall back the reply names every model that could not answer and why. Beside the thread, a companion pane carries the active head, this run's confidence, per-model measured accuracy from the calibration record, and the files the run changed. Four more views sit behind an icon rail: **Models** (every head this machine can route to, grouped by the token pool it spends, each marked with whether anything can actually drive it right now and why not, plus the measured accuracy record, never a score without its outcome count), **Activity** (every request, grouped by who has to act: waiting on you, running now, something failed, done), **Usage** (spend, remaining context budget with the headroom in updates rather than a band name, and how many models a consensus check actually asked), **Audit** (what the agents did and whether the record can be trusted, OWASP LLM Top-10 coverage over the MCP audit log, the guardrail rules in force, and a trust verdict for every MCP server installed on this machine). Opening a request from Activity drills into **Session**: that run's timeline, a layered graph when it fanned out, and a Code tab whose diff marks what changed *within* a modified line and lets you accept or undo the change on disk. A ledger policy that answers `ask` parks the task and the question appears inline in the chat transcript, answerable there. Every release ships a build for macOS (universal), Windows x86-64/ARM64 and Linux x86-64/ARM64 on the releases page. The builds are not code-signed yet, so macOS requires right-click → Open on first launch. Can also be built from source with `cd desktop && wails build` - **Per-run event log**: every dispatch writes `~/.hydra/logs/runs/<run_id>.jsonl`, head selection, swarm attempts, SPRT samples with their running confidence, A2A handoffs, and file edits, which is what the cockpit's Activity view and the desktop app reconstruct a run from. `HYDRA_RUN_ID` groups invocations an external orchestrator spawns into one run - **~1µs routing overhead**: the Go routing engine (policy eval + head selection + budget check) adds ~1,130 ns per dispatch, benchmarked on Apple M1 via `go test -bench=BenchmarkRoutingPath` - **Swarm dispatch**: fan a single prompt to multiple models simultaneously, race (fastest), best (LLM-judged quality), or all (aggregate) - **Confidence routing**: `hyctl dispatch --confidence 0.95` runs an SPRT (sequential probability ratio test) ensemble that samples models adaptively and stops as soon as a target confidence of correctness is reached, using per-source calibration (`hyctl trust calibration`) you build up from real outcomes, fewer model calls on easy tasks, more only where accuracy demands it. Calibration is a prerequisite: an uncalibrated source contributes no evidence, so a domain with none is refused with the command to record one rather than sampled at full cost for a coin-flip answer - **Blast-radius aware**: `hyctl dispatch --file X` and `hyctl graph blast X` read a code dependency graph (`graph.json`) so an edit to a widely-depended-on file demands higher confidence than a leaf, the defect-cost model holds expected leaked-defect cost roughly constant - **Token budget enforcement**: pressure modes (Normal → Emergency) plus a rate-aware first-passage governor auto-downgrade tiers as the orchestrator's context window fills - **Cost analytics**: `hyctl stats` shows spend broken down by model, tier, and day - **Traces that answer counterfactuals**: every dispatch row records the probability the router chose that head (`act_prob`) and the probability the row was kept (`keep_prob`). Without them a sampled log cannot be corrected back to the population. In simulation, averaging a non-uniformly sampled log inverted the true ranking of two heads, which would then change routing. Storage is bounded by keeping statistics rather than text: mergeable quantile sketches for latency, oracle verdicts as a permanent labelled corpus, and old run logs sealed into compressed monthly segments - **Off-policy evaluation**: `hyctl trace evaluate --policy cheaper` estimates what a routing policy you never ran would have cost, with a confidence interval, and refuses where the log cannot answer. `hyctl trace export --otlp` ships the same dispatches to any OpenTelemetry collector (Langfuse, Grafana, Jaeger) without giving up Hydra's own schema - **Dynamic pricing**: live rates from OpenRouter API with 24h cache and static YAML offline fallback - **Any API key works**: set `ANTHROPIC_API_KEY`, `OPENAI_API_KEY`, `OPENROUTER_API_KEY`, `GEMINI_API_KEY`, `GROQ_API_KEY`, `MISTRAL_API_KEY`, `DEEPSEEK_API_KEY`, `XAI_API_KEY` (or Bedrock/Azure/Cohere/Replicate/Together/Fireworks/Perplexity credentials) and the provider becomes a routable head automatically, no config file to touch - **A2A handoff**: structured JSON context (with vector clocks for causal ordering) passed between agents across tasks ## CLI Commands - `hyctl dispatch`: route a task to the cheapest head that clears the bar (`--enum`/`--tier`, `--swarm`, `--confidence`, `--file`, `--local`, `--dry-run`, `--a2a`) - `hyctl probe`: scan the machine for every available head (CLI agents, API keys, local servers) - `hyctl status`: session state, the rate-aware `claude_pct` budget governor and discovered heads - `hyctl tui`: interactive cockpit whose chat executes (Auto/Plan/Edit/Ask, plus Architect/Careful/Unattended), plus five more views over the same data: Agents, Models, Activity, Usage, Audit - `hyctl cost`: spend report (estimated vs actual) from the dispatch cost log - `hyctl stats`: session cost rollup by model, tier, and day. `--latency` reports p50/p90/p99 per model from mergeable relative-error sketches held in per-day rollups, so a percentile is answered without rescanning the whole cost log and two machines' stats can be combined without shipping their traces - `hyctl pricing`: live pricing DB (`list`, `refresh`) - `hyctl trust`: confidence layer (`calibration`, `record`, `defect`, `stats`, `explain`) - `hyctl graph`: code dependency graph. `blast` gives blast radius + percolation-κ, `parallel` the optimal agent count - `hyctl context`: context signal density / entropy governor (`entropy`) - `hyctl mcp`: local MCP accountability ledger (`check`, `record`, `verify`, `log`, `report`, `verify-chain`). Access checks can be classification-aware (PII/prompt-injection-marker auto-detected from content) and bind a SHA256 hash of the invocation parameters, so `verify` can later prove the executed parameters match the ones that were approved. Every event is hash-chained and anchored: `verify-chain` detects post-hoc edits, deletions from the middle, **and deletion from the end** (a truncated tail leaves every surviving link valid, so it is caught by comparing the log against the chain anchor rather than by walking it) and can carry an optional MITRE ATLAS/OWASP LLM Top-10 `Framework` tag - `hyctl mcp registry`: trust score for the MCP servers actually installed on your machine. `sync` (pull the official MCP registry), `scan` (find installed servers across Claude Code/Desktop, Cursor, Windsurf, VS Code, identity only, never reads env/secret values), `audit` (resolve + score + advance lifecycle state), `export` (static `index.html`/`index.json` of audited servers, no hosting), `backtest` (validate the scoring pipeline against known real incidents, `postmark-mcp`'s rug-pull, CVE-2025-6514), `clear` (recover a server quarantined in error). Scoring follows the CSA MCP Selection Scorecard's four categories. Every version bump drops a server's trust state back to provisional until it re-earns it, the direct fix for a server that ships clean for months, then turns malicious in a later version. Auto-classifies ledger events (`mcp-unverified`/`mcp-flagged`/`mcp-quarantined`) so `hyctl mcp check` can gate on it automatically. Also flags `mcp-behavior-change`: a server whose local ledger history has only ever shown one action type (e.g. read) performing another (e.g. network) for the first time, catchable from this machine's own history alone, no registry data or CVE required. A server's first-ever call is never flagged - `hyctl security`: answers one question by default: what did the agents on this machine actually do, and can the record be trusted. Prints a **verdict** (act now / attention / ok) with the single condition that produced it, a state machine over facts, never a blended score, the **correlated incident** behind it (risky ledger events grouped per actor into an attack sequence: injection → recon → escalation → audit-tampering, rated by OWASP Risk Rating likelihood × impact) and the **evidence state** (how many events, how many hash-chained, and whether the chain is intact, broken, truncated or unanchored). Roughly a dozen lines, no tables. `--why` opens the full programme underneath it: a **risk register** where every finding is one governed object (stable ID, severity, SLA clock with breach detection, per-defect cost from the router's own defect model, and a curated crosswalk to OWASP LLM / NIST AI RMF / ISO 42001 / MITRE ATLAS / SOC 2), an OWASP LLM Top-10 coverage score (percentage of applicable categories with a live, evidence-backed mechanism, hard-overridden to "INTEGRITY COMPROMISED" if the ledger chain is tampered with), a **policy audit** (per-rule hit counts, rules that never matched, rules provably unreachable, and whether the default is fail-open), **sensitive-data exposure** (every PII detection named by type and resolved against the discovered heads to say whether it stayed local or reached a head that leaves the machine, with unidentified heads reported separately rather than counted as confirmed leaks), a **threat breakdown**, a **control-effectiveness audit** (a control that is configured but never applied is reported as inert rather than reading as protection), **confidence evidence quality**, **configuration stability**, **head-binary integrity** (change detection, not provenance), **edit blast radius** (a file the graph does not index is reported as unknown, never as low-risk), least-privilege review per agent, an AI-BOM of the model estate, and a prioritized action queue. `--exec` prints the executive summary; `--attest` emits a checkable attestation (posture + evidence state + rules in force + digest) that says plainly when the underlying log is not tamper-evident. `--json` and `--csv` for machine-readable output. Every ledger-derived string is stripped of control characters before it reaches a terminal, so a hostile tool name cannot rewrite the line it is printed on. The TUI cockpit's Audit view (`hyctl tui`) and the desktop app surface the same report visually - `hyctl ask`: the tasks Hydra parked waiting on a human decision (`list`, `answer`, `decline`). A ledger policy rule can answer `ask` rather than allow or deny; dispatch then stops **before invoking any executor** and stores the task under `logs/pending/<task-id>.json` so it survives the process. An `ask` is not a denial and does not fall through to the next head in the fallback chain, skipping the head that needs permission and running a cheaper one would mean the question is never asked. Approval is per head, so answering for one head never authorizes another. `decline` records a denial in the ledger and runs nothing; an empty answer is not an answer and leaves the task parked. The desktop app surfaces the same queue - `hyctl oracle`: verification oracles (tests/compile/lint) as calibrated evidence: `verify` - `hyctl eval`: the corpus of oracle-verified examples. `list` takes `--failed`, `--limit`, `--json`; `stats` reports pass rate by domain. An oracle verdict on a real candidate is ground truth, so unlike traces these are kept verbatim and never pruned, they live at `~/.hydra/evalset/`, deliberately outside `logs/` so nothing that prunes logs can reach them. Deduplicated on (task, candidate), and PII is marked rather than dropped, since dropping it would bias the corpus - `hyctl trace`: trace storage and analysis (`seal` (`--older-than`, `--dry-run`, `--json`), `payloads` (`--json`), `evaluate` (`--policy`, `--confidence`, `--clip`, `--days`, `--json`), `export` (`--otlp`, `--out`, `--limit`, `--days`). `seal` folds run logs older than a cutoff into one compressed segment per month. This is lossless relocation, not retention: every event still reads back through the same commands and the desktop app afterwards. What it recovers is disk, one file per run is charged a whole filesystem block, so 65 short runs holding 27 KB of events were measured occupying 240 KB. `payloads` reports the **opt-in** store of prompt and response text: off unless chosen at `hyctl init`, sampled (every blob records the probability it was admitted, so the set stays correctable to the population), and anything matching a secret detector is replaced before it is written, redaction happens before hashing, so the content address does not leak a fingerprint of the secret. Blobs are packed rather than stored one file each, measured 22.2x smaller on disk for 200 small blobs (allocated space, not logical bytes), and compressed with a dictionary trained from your own corpus. `evaluate` answers what a routing policy you did **not** run would have cost, from the dispatches you did, self-normalized inverse propensity scoring with a percentile bootstrap interval, never a bare point estimate. Where the router had no chance of doing what the candidate policy would do it **refuses** rather than answering: that question is unidentifiable, not merely uncertain, and a wide interval still reads as a number. It reports effective sample size and whether weights were clipped, and distinguishes a log written before propensity logging existed from one with no overlap, because only the second is fixable by raising `explore_rate`. `export` renders dispatches as OpenTelemetry spans over OTLP/HTTP; nothing leaves the machine unless `--otlp` names an endpoint, and Hydra's own fields (tier, enum, cost, propensity) are carried under `hydra.*` rather than dropped into a `gen_ai.*` shape that has no place for them) - `hyctl models`: model capability registry (`list`, `add`, `remove`, `sync`) - `hyctl parallel`: fan N independent tasks out to N heads at once, from a task JSON file - `hyctl edit`: scoped, validated, rollback-safe single-file edit - `hyctl review`: code review - `hyctl init`: first-run wizard: discover heads, choose your cortex - `hyctl upgrade`: re-runs the same curl installer a fresh install uses: downloads the latest release, verifies its checksum, and renames the new binary over the old one, so the running process finishes on the old code and the next invocation picks up the new. Deliberately skipped for a Homebrew install, where overwriting brew's symlink would desync its bookkeeping, `brew upgrade hyctl` there instead - `hyctl version`: build and version info ## Who It's For Developers who use Claude Code, Cursor, Windsurf, or any AI coding assistant and want to reduce API costs without changing their workflow. Also useful for anyone building multi-model AI pipelines in Go. ## Installation The CLI binary is named `hyctl` (the product is Hydra). Available via: ```bash brew install ankit373/hydra/hyctl # Homebrew npm install -g hyctl # npm npx hyctl # npx (no install) pip install hyctl # pip curl -fsSL https://raw.githubusercontent.com/ankit373/hydra/main/install.sh | sh # standalone (macOS/Linux) irm https://raw.githubusercontent.com/ankit373/hydra/main/install.ps1 | iex # standalone (Windows) ``` The desktop app installs separately, with its own script (macOS and Linux): ```bash curl -fsSL https://raw.githubusercontent.com/ankit373/hydra/main/install-app.sh | sh ``` It resolves the newest release tag (desktop asset names embed their version, so GitHub's `/latest/download/` shortcut cannot address them), verifies the download against the `.sha256` published beside it, installs the `.app` to `/Applications` on macOS or the binary to `~/.local/share/hydra` with a `~/.local/bin` symlink on Linux, and clears the macOS quarantine flag. `HYDRA_VERSION` pins a release; `HYDRA_APP_DIR` changes the install directory. Windows has no script, download `hydra-desktop_<version>_windows_amd64.zip` (or `_windows_arm64.zip`) from the releases page and unzip it. Platform support: `hyctl` is built for macOS, Linux and Windows on both x86-64 and ARM64, six targets, every release. Homebrew covers macOS and Linux; npm/npx and pip cover all three OSes; `install.sh` covers macOS and Linux and `install.ps1` covers Windows (per-user, no admin rights; both verify the release checksum). The desktop app ships five targets every release: macOS (universal), Windows x86-64, Windows ARM64, Linux x86-64 and Linux ARM64. `install-app.sh` picks the right one automatically. Or build from source (requires Go 1.22+): ```bash git clone https://github.com/ankit373/hydra && cd hydra && go build ./cmd/hydra ``` ## Pricing **Free and open source, MIT License.** No subscription. Bring your own API keys. You pay the underlying model providers directly: - Claude Opus: ~$15/M output tokens (Anthropic) - Gemini 3.8 Flash: $0.075-0.30/M input depending on tier (Google, via Antigravity) - Qwen local via Ollama: $0 Typical savings: 75-85% vs routing everything through Claude. Run `hyctl pricing list` for live rates. ## Key Links - Homepage: https://hydra.uvansa.com/ - Desktop app, the views (Chat, Models, Activity, Usage, Audit, plus Session on drill-in), what each answers, and the platform downloads: https://hydra.uvansa.com/app.html - Documentation, every command (what it does, how it works, the logic/calculation, how to use): https://hydra.uvansa.com/docs.html - Pricing, free and open source (MIT); bring your own keys, pay providers directly. Custom builds on request: https://hydra.uvansa.com/pricing.html - Source code: https://github.com/ankit373/hydra - Installation: https://hydra.uvansa.com/#start - First Principles, the math behind the router (Wald SPRT, calibration diagnostic power D via KL divergence, Bayes defect-cost, Molloy-Reed percolation-κ, Amdahl optimal parallelism, context entropy), with interactive derivations: https://hydra.uvansa.com/first-principles.html - The routing math teaser (Pareto, SPRT, percolation-κ charts): https://hydra.uvansa.com/#math - Where Hydra fits (local-first + trust quadrant): https://hydra.uvansa.com/#universe - Cost comparison vs gateways: https://hydra.uvansa.com/#compare - Changelog: https://github.com/ankit373/hydra/blob/main/CHANGELOG.md - Funding, machine-readable funding manifest (fundingjson.org): https://hydra.uvansa.com/funding.json ## License MIT, https://github.com/ankit373/hydra/blob/main/LICENSE ## Author Ankit Jha, https://github.com/ankit373
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

