agentleFS
Sign inSign up

slotstream

carloslfu/slotstream/llms.txt

Run a 105 GB AI model on a Mac that can't hold it. Local inference for Qwen3.8-Flash-Next (a 125B-parameter mixture-of-experts, 105 GB on disk at 4-bit) on Apple Silicon Macs with far less RAM than the model: measured at 15.86 tok/s on the development Mac, a 48 GB M5 Pro, at a 22 GB target with the 0.2.19 corrected forecast (1.10x faster decode than the 0.2.18 forecast), and at 13.47 tok/s at a 20 GB target in the pre-release benchmark…

llms.txt402 starsChanged 28 days ago
  • Pipes a download into a shell
# slotstream

> Run a 105 GB AI model on a Mac that can't hold it. Local inference for
> Qwen3.8-Flash-Next (a 125B-parameter mixture-of-experts,
> 105 GB on disk at 4-bit) on Apple Silicon Macs with far less RAM than the
> model: measured at 15.86 tok/s on the development Mac, a 48 GB M5 Pro, at a
> 22 GB target with the 0.2.19 corrected forecast (1.10x faster decode than the
> 0.2.18 forecast), and at 13.47 tok/s at a 20 GB target in the pre-release
> benchmark of the configuration shipped in 0.2.16 with decode lookahead enabled. The older 0.2.3 result was ~12 tok/s.
> The ~3.5 tok/s estimate for 16 GB overstates the measured 1.41 tok/s on a
> base-storage Mac mini M2.
> Experts stream from SSD through a fixed slot cache; one Swift binary serves
> Ollama-, OpenAI- and Anthropic-compatible chat APIs on 127.0.0.1:11434.

The installer follows the [latest published release](https://github.com/carloslfu/slotstream/releases/latest); a candidate benchmark or version bump on main does not establish publication.

The README, user guides, engineering notes, and changelog are combined in
[llms-full.txt](https://raw.githubusercontent.com/carloslfu/slotstream/main/llms-full.txt).

## Facts

- The only model is `qwen3.8-flash-next:4bit`; the engine is built around its geometry.
- Target range: Macs that cannot hold the model, 16 to 64 GB. From 96 GB the model fits in memory; Slotstream runs there but is not optimized for it, and engines that keep the model resident report faster decode. Decision: `db/records/decisions/target-range-macs-that-cannot-hold-the-model.md`.
- The current engine requires Apple Silicon, MLX, and Metal. Windows and Linux support for AMD and NVIDIA is planned for Sevra. Other models, including the dense Qwen3.8-27B, are unsupported.
- Native end to end: the CLI, server, planner, expert cache, governor and model are Swift on mlx-swift and Metal, with Slotstream's own run-time-compiled Metal kernels and a C decoder for the download format. No Python runtime, interpreter or bridge is on the request path; the Python under `Tools/` is the parity reference, benchmark and release tooling and never runs in the product. Speed is attributed to measured mechanisms, never to the stack itself. Details: `docs/ENGINEERING.md#native-stack`.
- Requirements: Apple Silicon, macOS 14+, ~110 GB of free disk. Weights: 105.3 GB in 25 files (the 1.5 GB draft head is optional), sha256-pinned, at `~/.slotstream/models/qwen38-flash-next-mlx-4bit`.
- Install: `curl -fsSL https://raw.githubusercontent.com/carloslfu/slotstream/main/install.sh | sh` puts the latest release in `~/.slotstream/bin` (binary plus `mlx.metallib`, no Python); rerun to upgrade.
- Images work on every dialect: Ollama `images` (base64) on `/api/chat` and `/api/generate`, OpenAI `image_url` parts, AI-SDK `file` parts on the fx gateway, `run --image`. Inline bytes only — a `data:` URL or bare base64; `http://` and `file://` URLs are refused and never fetched. A picture is one token per 32x32 pixels of its resized self, at most 2,304 tokens, out of the same context as the text; the vision tower is 0.9 GB and loads on the first image. Auto and `--memory-gb` reserve it inside the process target; explicit pool sizes retain their pool and add the tower to the expected footprint. Image workspace and real headroom are checked separately.
- Cache-size and live-resize gates check byte-identical greedy output with other generation settings fixed. Changing the total memory target can also change prefill grouping or speculative decoding, outside that equality claim.
- One model process per user (`/tmp/slotstream-model-<uid>.lock`). The server binds 127.0.0.1 only, accepts browser origins from loopback only, and has no auth.
- Hardware reports include a 48 GB M5 Pro, a 16 GB M2, a 24 GB M4 Pro, a 32 GB M5 Air, a 36 GB M4 Max, a 64 GB M3 Max, a 64 GB M4 Max measured from its internal SSD and from a 10 Gb/s external drive, and a 128 GB M5 Max. All but the M5 Pro are community reports with different settings. The README gives rough planning ranges; the hardware guide lists the tested configurations with installed RAM, release and process memory target. The upper High/Ultra estimates assume M5 Max-class hardware, a fast internal SSD and larger manual targets. They are extrapolations across machines and releases, not measured limits or confidence intervals. Larger manual targets improved the reported M5 Max speed; installed RAM alone does not define a performance tier. The hardware guide explains each range and keeps simulated allocation plans separate. The 18 GB size still needs reports; 8 GB Macs are refused because the smallest plan doesn't fit. Compatibility-tier support is coming soon; it is not available in the current release. Runtime testing on macOS 14/15 is still needed. Ollama CLI and symlinked-weight-directory bugs in 0.2.0 were fixed in 0.2.1.

Sevra for Mac is under development in `apps/macos`, using SwiftUI/AppKit and a
Swift-owned runtime with in-process Slotstream. Product behavior is shared as
text specifications; platform implementations are independent. This is a local
development app, not a published desktop release.

## Commands

- `slotstream doctor [--memory-gb G] [--sim-ram N] [--sim-available N] [--json]`: device report and the plan any flags would produce. Loads nothing, takes no lock.
- `slotstream pull [--dir D] [--connections N] [--transport automatic|compressed|raw] [--verify]`: compressed Hugging Face download by default, with **16.12% fewer bytes**, overlapping decode/writes, adaptive connections, exact resume and pinned raw fallback. No Hugging Face account is required. Explicit connections fix the count; automatic mode preserves existing raw resumes. `--verify` re-hashes without downloading.
- `slotstream run --prompt "..." [--max-tokens N] [--greedy] [--raw] [--think] [--image PATH]`: one generation, no server.
- Since 0.2.21, `slotstream launch [--port P] [--memory-gb X] [--idle-exit M] [--no-start] [--dry-run] [claude|codex|pi|opencode|hermes] [agent arguments...]` starts that coding agent connected to the server on the port, reading the served model, window and reply limit from `/v1/models`; without a name, a terminal is asked which installed agent to start. When nothing answers, it first checks the agent's settings, then starts `slotstream serve --port P [--max-context W] [--memory-gb X] --prefix-cache-dir ~/.slotstream/prefix-cache --idle-exit 30` in its own session with its log in `~/.slotstream/logs/serve.log`, shows the start until the model answers (Control-C stops that server), and registers the agent's pid with `POST /slotstream/clients`. The window is the automatic one unless the agent needs more. Such a server keeps running while a registered agent runs and stops 30 minutes after the last request and agent end (`--idle-exit 0` keeps it). A server launch started (recorded by pid and process start time in `~/.slotstream/launch/server-P.json`) that nothing uses is restarted when its window is too small; a server started by hand is used as it is and never restarted. Launches that may start a server hold `~/.slotstream/launch/start-P.lock`, so two at once start one server. `--no-start` only uses a running server. The model is never downloaded without asking. `slotstream stop [--port P]` stops the Slotstream server on the port, found through `GET /slotstream/status`, or the one launch recorded while it is still starting. `serve --idle-exit M` is that stop for any server: it exits after M minutes with no request and no registered client still running, answering 503 to requests once it has decided. Claude Code gets variables plus `--settings` (every cloud provider switch off, API key and key helper cleared, the hosted WebSearch tool denied); Codex gets `-c` settings placed after its deepest subcommand and a catalog with its version's base instructions under `~/.slotstream/launch/codex/`, and `codex cloud` and `--output-schema` are refused; Pi gets a `slotstream` provider set in `~/.pi/agent/models.json`, keeping every other entry as written; opencode gets `OPENCODE_CONFIG_CONTENT` with only the `slotstream` provider enabled and the build and plan agents on the served model; Hermes gets `HERMES_HOME=~/.hermes-slotstream`, `--profile default` and a configuration that pins every side task to the served model. Thinking starts off. `--dry-run` hides keys and tokens. Minimum windows: 32,768 tokens for Claude Code and Codex, 16,384 for Pi and opencode, 65,536 for Hermes.
- `slotstream serve [--port 11434] [--max-context auto|N] [--vision auto|on|off] [--no-elastic] [--no-prefix-cache] [--prefix-cache-dir D]`: the API server. Since 0.2.17 auto picks the prompt-plus-reply window per Mac: the largest of 32,768, 65,536, 131,072 and 262,144 tokens that keeps speculative decoding, retains one complete conversation and adds at most 10% to the estimated request time, without removing cache above the measured decode range (32,768 through 32 GB, 65,536 at 36 GB, 32,768 at 48 GB, 131,072 at 64 GB, 262,144 from 96 GB). `--max-context 65536` fixes a 65,536-token window. Any N up to 262,144 is accepted, image requests stay within 65,536, and the extra state and transient reserve are charged before allocating the cache.
- Since 0.2.18, `serve --prefix-cache-dir D` also keeps conversation states on disk, so a restarted server, or a conversation longer than memory retains, resumes from its last committed state instead of re-reading its prompt. Off unless set; `--prefix-cache-disk-gb`, `--prefix-cache-min-tokens` and `--prefix-cache-max-age-days` bound what is kept, and requests with images are not written. `slotstream prefix-cache --dir D [--clear] [--json]` lists or clears the directory without loading the model. Since 0.2.21 the prefix conversations share is kept too. While a prompt is processed, its system message (when it ends 512 tokens or more in) and the longest head it shares with a kept prompt before parting from it are stored as shared prefixes at the last prefill pass end at or before that boundary (256-token grid by default): forked into the memory cache and, at `--prefix-cache-min-tokens` or more, written to disk. The next conversation with the same system prompt resumes from it, in the same process or after a restart; outputs are unchanged since no pass is reshaped. Shared prefixes are kept once, never replaced by the conversations extending them, and listed by `prefix-cache`; one that two or more conversations start from is evicted after them. Library callers name the shared head with `request.sharedPrefixTokens`. In memory, a request with tools keeps its shared prefix like a conversation (`request.sharedPrefixRetention = .conversation`), and a shared prefix stays as recent as the conversation continuing from it; other shared prefixes remain optional snapshots that never displace a conversation. A plan that retained whole conversations keeps the largest retention that fits beside the vision tower when the first image loads, instead of falling back to the budget share.
- Since 0.2.14, `run`, `serve` and `doctor` share `--max-context` and `--max-prefill-wait`. The wait defaults to 30 minutes from accepted request to first sampled token, including queueing and preparation. `0` disables only time; memory and cancellation guards remain. Memory feasibility is separate from the request deadline.
- `slotstream context-check --tokens N [--ladder]`: measures a synthetic prompt plus its required reply, reserving both before load; preserves exact counts, memory, swap and abort observations. Incomplete work fails qualification. `slotstream prefill-schedule --chunk C --tokens N` prints the bounded pass ladder; unmeasured durations remain unknown.
- Memory options on `run`, `serve`, `doctor`: `--memory-gb G` (total for the process, minimum 8.1 for the 32,768-token window, higher for larger windows; the usual manual option), `--experts-per-layer N` (1–512; pool = N × 0.133 GB), `--pool-gb G`, `--max-ram-percent P` (auto only, default 70). Precedence: experts-per-layer beats pool-gb beats memory-gb. Explicit sizes disable automatic resizing; physical checks before loading and request growth still apply. Auto takes the lowest of 33 GB, 70% of RAM, and the Metal working set minus 2 GB (34.6 GB with the draft head at the 32,768-token window, plus the context charges of any larger window auto picks), and resizes between requests every 15 s.
- That ceiling is an intentional model-specific default based on measured tradeoffs, not a universal optimum or a hard limit on explicit targets. Keep it until comparable real measurements justify revision; unused RAM and flat estimator output alone do not establish a defect or a performance plateau. See [measured operating policies](db/records/design/measured-operating-policies.md) for evidence, scope, override and revision requirements for important tuning values.
- `--mtp auto|on|off` on `run`, `serve`, and `doctor`: speculative decode with the model's draft head, `mtp.safetensors`, which `pull` fetches with the weights (optional, mirror-hosted). `auto` (default) turns it on when the cache reaches 28 experts per layer after the head's charge and context charges, before the lookahead reservation. A 12 GB target qualifies at the 32,768-token window; larger windows or unavailable memory can keep it off even on a larger Mac; the floor was 120 before 0.2.16 and then 76. On a cache of 76 experts per layer or more after the full 1.6 GB charge the head keeps its 512 experts resident; below that they stream from the SSD through a 64-expert cache of their own, charged 0.4 GB, with the same output, and at a 12 GB target the head then decoded 1.23x faster than plain decode with the lookahead. `SLOTSTREAM_MTP_EXPERTS=resident|streamed` forces a placement; the automatic context window never trades a resident head for a streamed one. With the head on and its floor met, the decode lookahead is on too: router-reuse expert prefetch straight into cache slots, FP32 router weights and a GPU drain every four layers, 373 MiB charged, 1.11x decode on held-out prompts at a 20 GB target with identical output. Without the head it runs in plain decode from 20 experts per layer before its charge, where it made plain decode 1.11x faster at a 10 GB target. Since 0.2.19 the forecast reads the previous layer's attention output and applies a rank-128 learned correction (`lookahead/tap-correction-attention-rank128-v1.safetensors`, an optional 37.5 MB sidecar that `pull` fetches; 409 MiB charged with it): 1.10x faster decode than the 0.2.18 forecast on eight held-out prompts at a 22 GB target with identical output, 14.38 to 15.86 tok/s. `SLOTSTREAM_OPT_EXPERT_PREFETCH=0` turns it off; `SLOTSTREAM_EXPERT_PREFETCH_TAP=boundary` keeps the 0.2.16 forecast with the file present. `SLOTSTREAM_DRAFT_DEPTH` selects 1–16 drafts (default 2), an adopted mixed-workload choice rather than a universal optimum. Historical one-draft measurements at the former 28 GB memory target were ×1.24 decode greedy and ×1.18 with server sampling; they are not new two-draft measurements. The three-row verify pass runs two rows at a time through the vector attention kernel from 6,144 tokens of context instead of the dense kernel, whose cost grows with the context: the gain grows with the context, and on a quiet machine at a 22 GB target speculative decode measured 11.82 against 11.67 tok/s (x1.013) with a 16,356-token prompt, x1.085 on a second 16,356-token prompt and x1.29 with a 32,740-token prompt, with a fetch-free pass 32% cheaper at 32,740 tokens and 45% at 65,508; the split changes which drafts are accepted, so the 16k gain follows the prompt, 1% and 8% in the two measured, while at 32k both arms accept alike; `SLOTSTREAM_OPT_VERIFY_SPLIT=0` restores the dense pass. `SLOTSTREAM_OPT_ROW_INVARIANT=1` with `SLOTSTREAM_OPT_VERIFY_SPLIT_CONTEXT=0` is an exact mode, off by default, in which a speculative run's output equals a plain run's for draft depths up to 4; `mtp-rowcheck` gates it. See db/records/decisions/verify-pass-split-attention-default.md, db/records/decisions/draft-head-auto-floor-76-per-layer.md, db/records/decisions/decode-lookahead-default-with-the-draft-head.md, db/records/decisions/draft-depth-defaults-to-two.md and historical MEASUREMENTS.md M9.
- `--gpu-keepalive auto|on|off` on `run` and `serve`: a one-thread kernel on its own command queue keeps the GPU from clocking down between the short bursts of streamed decode, and cache misses are read straight into their cache slots instead of through staging arrays and a GPU scatter (`SLOTSTREAM_OPT_DIRECT_DEMAND=0` restores the old path). On the development Mac, with identical output, the two made decode 1.28x faster at a 10 GB target without the draft head and 1.22x faster at 22 GB with the draft head and lookahead, counting only pairs with no swap activity. The keepalive costs power: energy per generated token rose 7% at 16 GB, so `auto`, the default, keeps it on only with AC power outside Low Power Mode; `SLOTSTREAM_GPU_KEEPALIVE` sets the default. `decode-overlap-check` gates exactness. See db/records/decisions/gpu-keepalive-on-ac-power.md and db/records/decisions/direct-demand-reads-default.md.
- Checks that need no weights: `runtime-check`, `governor-check`, `sampler-golden`, `pull-check`. Checks that load the model (give them `--memory-gb 8.1` to `10`): `elastic-check`, `elastic-drill`, `prefix-check`, `sweep-check`, `decode-overlap-check`, `draft-stream-check`, `parity`, `template-check`, `ngram-golden`, `dequant-golden`, `mtp-parity`, `mtp-accept`, `mtp-check`, `mtp-rowcheck`. The battery is `Tools/verify.sh`; release acceptance is `Tools/e2e_release.sh`.
- Environment: `SLOTSTREAM_WEIGHTS_SOURCES`, `SLOTSTREAM_PULL_CONNECTIONS`, `SLOTSTREAM_PREFIX_CACHE=0`, `SLOTSTREAM_PREFILL_CHUNK`, `SLOTSTREAM_IO_QUEUE_DEPTH`, `SLOTSTREAM_EXPERT_LOAD_BATCH`, `SLOTSTREAM_SWEEP=0`, `SLOTSTREAM_SWEEP_ADMIT=0`, `SLOTSTREAM_SWEEP_TRACE=1`, `SLOTSTREAM_PREFILL_CACHE_MB`, `SLOTSTREAM_OPT_EXPERT_PREFETCH=0`, `SLOTSTREAM_OPT_ROUTER_WEIGHTS`, `SLOTSTREAM_DECODE_BARRIER_LAYERS`, `SLOTSTREAM_DRAFT_DEPTH`, `SLOTSTREAM_OPT_VERIFY_SPLIT`, `SLOTSTREAM_OPT_VERIFY_SPLIT_CONTEXT`, `SLOTSTREAM_OPT_ROW_INVARIANT`, `SLOTSTREAM_OPT_DIRECT_DEMAND=0`, `SLOTSTREAM_GPU_KEEPALIVE`, `SLOTSTREAM_MTP_EXPERTS`; installer: `SLOTSTREAM_ROOT_DIR`, `SLOTSTREAM_RELEASE_BASE`.

## API

- Ollama dialect: `POST /api/chat` and `POST /api/generate` (stream by default), `GET /api/tags`, `GET /api/ps`, `POST /api/show`, `GET /api/version`. OpenAI dialect: `POST /v1/chat/completions` (`stream: false` by default) and `GET /v1/models`; base URL `http://localhost:11434/v1`, any `api_key`.
- The OpenAI Responses API, which Codex requires: `POST /v1/responses` with function, namespace, and freeform tools, images, and streamed item events. See [Codex setup](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/CODEX.md).
- Since 0.2.21, the Anthropic Messages API, which Claude Code uses, is at `POST /v1/messages` (base URL `http://127.0.0.1:11434`, any token) and `POST /v1/messages/count_tokens`, with tools, `tool_choice`, base64 images, plain-text documents, thinking whose `signature` carries the reasoning, stop sequences, Anthropic's SSE event order with a `ping` every 10 seconds during a prompt read, and usage reporting reused prompt tokens as `cache_read_input_tokens`. Unknown top-level fields are ignored and listed in `X-Slotstream-Ignored-Fields`; errors use Anthropic's shape; an oversized prompt fails with `prompt is too long: N tokens > M maximum`. See [Claude Code setup](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/CLAUDE-CODE.md).
- The AI SDK gateway supports tool calling and images: `POST /v3/ai/language-model` (also `/v1/ai/language-model`), `GET /coding-agent/v1/models`, and `GET /coding-agent/v1/credits`. See [fx setup](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/FX.md).
- Ordinary-chat sampling defaults for the Ollama/OpenAI endpoints: `temperature` 0.7, `top_p` 0.8, `top_k` 20, `min_p` 0, `presence_penalty` 1.5, `num_predict` 512 on the Ollama endpoints and, without an output limit, a quarter of the served window up to 8,192 tokens on `/v1/chat/completions` and `/v1/responses`, `seed` random, `stop` none. Unknown fields return a 400 that names them, except top-level fields on `/v1/messages`; tools on the Ollama endpoints, JSON-schema output, strict tool schemas, logprobs, and embeddings return 400; a wrong model is a 400 (404 on `/api/show`); `/api/pull` and `/api/create` are 501. Prompt plus completion defaults to 32,768 tokens. OpenAI function tools, call/result history, streamed calls and reasoning are supported. Model discovery reports the served context; required/named tool choices and incomplete calls are checked before exposing executable output.

## Docs

- [docs/GETTING-STARTED.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/GETTING-STARTED.md): installation, first reply, model download, pictures, updates, and common questions
- [docs/ENGINEERING.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/ENGINEERING.md): developer index, performance, context and memory, architecture, project history, and credits
- [docs/EXPERT-LOOKAHEAD.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/EXPERT-LOOKAHEAD.md): a short explanation of expert lookahead, the experiments, measured results and limits
- [docs/HERMES-NOTES.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/HERMES-NOTES.md): Hermes protocol, configuration rationale, discovery, images, summary fallback, and integration checks

- [docs/CLIENTS.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/CLIENTS.md): connection settings for OpenAI and Ollama clients, Hermes and fx setup links, integration checks, troubleshooting, and issue evidence
- [docs/HERMES.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/HERMES.md): first-time Hermes setup, complete tested configuration, file-reading example, and troubleshooting
- [docs/CODING-AGENTS.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/CODING-AGENTS.md): `slotstream launch` for Claude Code, Codex, Pi, opencode and Hermes, what each agent receives, Pi and opencode setup, and troubleshooting
- [docs/CLAUDE-CODE.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/CLAUDE-CODE.md): Claude Code on Slotstream through `slotstream launch claude`, thinking, differences from Claude, shared prompts on disk, manual variables, and troubleshooting
- [docs/CODEX.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/CODEX.md): Codex through `slotstream launch codex`, the manual catalog and profile setup, a file-editing example, and troubleshooting
- [README.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/README.md): project overview, requirements, measured speed, quick start, guides, a plain explanation, limitations, FAQs, community, and star history
- [docs/SEVRA-MAC.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/SEVRA-MAC.md): native philosophy, development build, local workflow, tests, internal CLI and unpassed release gates
- [docs/CLI.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/CLI.md): everyday commands, common diagnostics, memory-option precedence, environment variables, file locations, the diagnostic checks
- [docs/API.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/API.md): every endpoint, accepted field, sampling default, streaming and error shapes
- [docs/LIBRARY.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/LIBRARY.md): using slotstream as a Swift package — the two products, the Metal library requirement, weights, planning, serving
- [docs/TESTING.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/TESTING.md): the check catalogue and its tiers, what runs in CI against what needs the real weights, coverage and its ratchet
- [docs/TROUBLESHOOTING.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/TROUBLESHOOTING.md): port clashes, the one-process rule, paging, slow decode, moving or verifying weights, and fixes for old versions
- [docs/HARDWARE.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/docs/HARDWARE.md): requirements, speed explained, credited results from real Macs, and separate estimates; measurement steps are in docs/TESTING.md
- [CHANGELOG.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/CHANGELOG.md): what each release changed

## Design and evidence

- [PLAN.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/PLAN.md): living design doc: architecture (§4), correctness invariants (§6), milestones, risk register
- [MEASUREMENTS.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/MEASUREMENTS.md): every measured number with its method, including the experiments that failed
- [AGENTS.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/AGENTS.md): contributor guide: serving invariants, memory-safety rules, the release process
- [db/DB.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/db/DB.md): the public brain, a db.md store: every measurement, design and plan section, public claim, decision, machine, and raw run as a record with frontmatter and wiki-links; PLAN.md and MEASUREMENTS.md are generated from it by Tools/projections.py, and a number on any surface has a claim record naming the measurement behind it
- [CONTRIBUTING.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/CONTRIBUTING.md): how to build, test, measure, and send a change; [SECURITY.md](https://raw.githubusercontent.com/carloslfu/slotstream/main/SECURITY.md): what runs where, supply chain, how to report

## Source map

- [Sources/Slotstream/](https://github.com/carloslfu/slotstream/tree/main/Sources/Slotstream): the engine: memory planner (Plan.swift), expert streaming (ExpertStore.swift), HTTP server (Server.swift), sampling (Generate.swift), elastic governor (Governor.swift), conversation reuse (PrefixCache.swift), speculative decode (MTP.swift), the one-process lock (ProcessMemory.swift)
- [Sources/slotstream-cli/](https://github.com/carloslfu/slotstream/tree/main/Sources/slotstream-cli): the CLI: subcommands in main.swift and the command files; the download engine and pinned manifest live in the Slotstream library
- [Tools/verify.sh](https://raw.githubusercontent.com/carloslfu/slotstream/main/Tools/verify.sh): the acceptance battery; [Tools/e2e_release.sh](https://raw.githubusercontent.com/carloslfu/slotstream/main/Tools/e2e_release.sh) re-tests the installed binary; [Tools/llms_full.sh](https://raw.githubusercontent.com/carloslfu/slotstream/main/Tools/llms_full.sh) regenerates llms-full.txt

Discussion

Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.

Posts are public.Sign in to post

No one has posted yet. Be the first.