horizon
peters/horizon/AGENTS.md
Source of truth for all contributors and AI agents working on this project. An AI agent given these instructions should be able to go from zero to a running Horizon binary. Repository: https://github.com/peters/horizon.git Pre-built binaries are attached to GitHub releases. Prefer the latest non-prerelease tag. Use gh release download --repo peters/horizon or download from https://github.com/peters/horizon/releases/latest. No Rust toolchain or system headers are needed for this path. If any step fails, read the error, fix the prerequisite, and retry. On Linux,…
AGENTS.md710 starsChanged 5 days ago
- Installs packages
- Commits and pushes
# Horizon — Agent Guidelines
> **Source of truth** for all contributors and AI agents working on this project.
## Quick Start
An AI agent given these instructions should be able to go from zero to a running Horizon binary.
**Repository:** `https://github.com/peters/horizon.git`
### Option A — Download a release binary (fastest)
Pre-built binaries are attached to GitHub releases. Prefer the latest **non-prerelease** tag. Use `gh release download --repo peters/horizon` or download from `https://github.com/peters/horizon/releases/latest`.
| Platform | Asset | Contents |
|----------|-------|----------|
| Linux x64 | `horizon-linux-x64.tar.gz` | Single `horizon` binary — extract and make executable |
| macOS x64 | `horizon-osx-x64.tar.gz` | Single `horizon` binary — extract and make executable |
| Windows x64 | `horizon-windows-x64.exe` | Ready-to-run executable |
No Rust toolchain or system headers are needed for this path.
### Option B — Build from source
#### Prerequisites
- **Rust stable ≥ 1.95** (edition 2024). Install via [rustup](https://rustup.rs) if not present.
- **Linux only:** the eframe/wgpu rendering stack needs system headers. Install them before `cargo build`:
- Debian/Ubuntu: `sudo apt install -y build-essential pkg-config libxkbcommon-dev libwayland-dev libxcb-render0-dev libxcb-shape0-dev libxcb-xfixes0-dev libvulkan-dev libgl-dev cmake`
- Fedora: `sudo dnf install -y gcc pkg-config wayland-devel libxkbcommon-devel vulkan-loader-devel mesa-libGL-devel cmake`
- Arch: `sudo pacman -S --needed base-devel wayland libxkbcommon vulkan-icd-loader cmake`
- **macOS:** Xcode Command Line Tools (`xcode-select --install`). Metal ships with the OS.
- **Windows:** MSVC build tools (installed automatically by `rustup` on the `msvc` target). DX12/Vulkan drivers ship with the GPU driver.
- **Speech input (`--features speech`, opt-in):** additionally needs **CMake** and a **C++ compiler** (to build the vendored transcribe.cpp), plus on Linux the ALSA headers (`libasound2-dev`/`alsa-lib-devel`) for microphone capture. The default build does not require these.
#### Build & Run
```bash
git clone https://github.com/peters/horizon.git
cd horizon
cargo run --release
```
#### Verify
```bash
cargo fmt --all -- --check
cargo test --workspace
cargo test --workspace --features speech # speech tier (needs CMake + ALSA headers)
cargo clippy --all-targets --features speech,trace-profiling -- -D warnings
```
If any step fails, read the error, fix the prerequisite, and retry. On Linux, missing system headers are the most common issue — look for `pkg-config` or linker errors and install the corresponding `-dev` package.
## Project Overview
**Horizon** is a GPU-accelerated terminal board — a visual workspace for managing
multiple terminal sessions as freely positioned, resizable panels on a canvas.
**Stack:** Rust (edition 2024) · eframe/egui (wgpu backend) · alacritty_terminal (VT parsing, PTY, event loop)
## Workspace Layout
```
crates/
horizon-core/ Core: terminal emulation, PTY, board & panel management
horizon-ui/ Binary: eframe application, UI rendering, input handling
```
### horizon-core
- `error.rs` — Typed error enum via thiserror
- `terminal.rs` + `terminal/` — alacritty_terminal wrapper split by lifecycle, event handling, resize policy, selection, replay, and support helpers
- `panel.rs` — Panel = terminal + PTY session + identity
- `board.rs` + `board/tests/` — Board orchestration and colocated behavior-focused tests
### horizon-ui
- `main.rs` — Entry point, tracing init, eframe launch
- `app/` — `eframe::App` orchestration split by canvas, panels, sidebar, settings, session, persistence, workspace helpers, and focused action submodules
- `remote_hosts_overlay.rs` + `remote_hosts_overlay/` — remote SSH chooser state/input shell with dedicated query, layout, and paint helpers
- `terminal_widget/` — Terminal widget split by layout, input, render, scrollbar logic
- `input/` — Keyboard translation, mouse reporting, escape-sequence building
- `theme.rs` — Color palette (Catppuccin Mocha), styling constants
## Development Workflow
### Pre-push validation (all must pass)
```bash
cargo fmt --all -- --check
./scripts/check-maintainability.sh
# `--features speech` is the widest runner-buildable set (GPU features need
# machine toolchains and add no Rust surface; needs libasound2-dev on Linux).
# horizon-ui declares the feature, so every other crate builds identically
# with and without it and the feature pass only needs that crate.
RUSTFLAGS="-D warnings" cargo test --workspace
RUSTFLAGS="-D warnings" cargo test -p horizon-ui --features speech
cargo clippy --all-targets --features speech,trace-profiling -- -D warnings
cargo clippy --workspace --lib --bins --examples --features speech -- -D warnings -D clippy::unwrap_used -D clippy::expect_used
cargo clippy --workspace --all-targets --features speech -- -D warnings -W clippy::pedantic
```
- Run the validation commands in the exact checkout you will push. If you split work across branches or `git worktree`s, rerun the blocking and strict clippy tiers in each final branch/worktree after applying the split, not only in the original combined checkout.
### Cloud Development Profiles
The repository's `.horizon/cloud.yml` defines CPU and GPU development profiles.
Use [the cloud development guide](.horizon/README.md) for image prerequisites,
isolated build caches, validation commands and native smoke requirements. The
GPU profile requires GPU capacity; a CPU result cannot qualify that lane.
### Browser Interface Parity
- Treat browser features as capabilities shared by the UI, CLI, and MCP by default.
Include equivalent operations and status through the public MCP contract and
CLI plan runner where applicable; a settings-only implementation is not complete.
- Keep provider, credential, capacity, and lifecycle logic in shared modules.
UI, CLI, and MCP must use the same provider policy and credential bindings.
- Cover multiple configured providers and separate credential sets on the same
provider. Reusing a credential reference, OS-store slot, or environment binding
intentionally shares it; a profile name alone does not create credential isolation.
- Validate each supported interface and document actual runtime requirements and
unsupported cases explicitly. Do not silently defer CLI/MCP support or require
the settings UI to be open for an agent-facing capability to work.
### Configuration Changes
- When changing default presets, CLI flags, or any config-related code in `horizon-core/src/config.rs`, always sync the user's local config file (`~/.horizon/config.yaml`) to match
### Code Quality Bar
- Self-documenting code preferred over comments; add comments only for invariants, non-obvious tradeoffs, or safety contracts
- Use idiomatic Rust naming: `snake_case` (functions/modules), `CamelCase` (types), `SCREAMING_SNAKE_CASE` (consts)
- Typed error enums (thiserror) — no `Box<dyn Error>` or `.unwrap()` in library code
- Keep APIs explicit: prefer `Result<T, E>` and typed structs/enums over ad-hoc tuples
- Keep default values and default-selection rules centralized. Prefer `Default`, associated consts, or focused helper functions over repeating literals or enum variants across call sites. Use `static` items only when stable global storage is actually required
- Prefer `tracing` for structured logging
- `#![forbid(unsafe_code)]` on all crates. Two exceptions carry `#![deny(unsafe_code)]` with narrowly scoped `#[allow]`s instead, because they call platform FFI that has no safe equivalent: `horizon-cursor` (Win32 `GetCursorPos`) and the `horizon-wayland` pinch bridge together with its single call site in `horizon-ui` (libwayland `wl_display` adoption, whose lifetime winit's `run_app(self)` makes impossible to encode in a type). Do not widen either exception; new crates get `forbid`
- Consolidate repeated helpers into shared modules in horizon-core
- Keep new or edited modules single-purpose; avoid mixing rendering, persistence, session bootstrap, and filesystem logic in one file
- If UI code needs shared layout math, state conversion, or template-sync logic, move it into `horizon-core` instead of duplicating it in `horizon-ui`
- Treat roughly 600 lines as the point to split a Rust source file; the CI guardrail fails non-test files above 1000 lines under `crates/horizon-core/src` and `crates/horizon-ui/src`
- Do not use `#[allow(clippy::too_many_lines)]` in core or UI source files; decompose the code instead
- Keep inline `#[cfg(test)]` modules at the end of the file so maintainability checks can measure production code cleanly
- Minimize allocations in the render hot path (per-frame code)
- Every `unsafe` block (if ever needed) must have a `// SAFETY:` rationale
- Avoid unnecessary `crate::` path prefixes in module-local code/tests when imports already provide the item
- Prefer `unwrap_or_else(std::sync::PoisonError::into_inner)` over manual `match` for poisoned mutex recovery
- `expect_used` is treated the same as `unwrap_used`: both can hide panic paths in runtime code
- Fix clippy/rustc warnings instead of suppressing them
### Maintainability Rules
- Prefer small module trees over large flat files: `mod.rs` should orchestrate, leaf modules should do one job
- UI modules render or collect UI actions; domain state mutation belongs in `horizon-core` unless it is purely presentational state
- When editing a file that is already large, land any purely mechanical split in a focused prerequisite PR before adding another responsibility
- When an inline test block starts dominating a source file, move it into a colocated `module/tests/` tree split by behavior instead of letting the parent file keep growing
- Keep architecture notes current in [`docs/architecture/maintainability.md`](docs/architecture/maintainability.md) when module boundaries or guardrails change
### CI Tiers (`.github/workflows/ci.yml`)
| Tier | Command | Status |
|------|---------|--------|
| Blocking | `cargo clippy --all-targets --features speech,trace-profiling -- -D warnings` | Must pass |
| Strict | `cargo clippy --workspace --lib --bins --examples --features speech -- -D warnings -D clippy::unwrap_used -D clippy::expect_used` | Must pass |
| Pedantic | `cargo clippy ... -W clippy::pedantic` | Advisory (will promote) |
### Commit Guidelines
- Concise imperative messages, optionally scoped: `feat(board):`, `fix(render):`, `ci:`
- One logical change per commit
- No assistant attribution anywhere in repository history: never add a `Co-authored-by` trailer, a session link or a similar tool or vendor identifier to a commit message, PR body or squash-merge message, and strip an inherited one when amending, rebasing or squashing someone else's work
- Always squash-merge pull requests; do not use merge commits or rebase merges for PRs
- PRs include: purpose, behavior impact, test evidence
- Fix Clippy warnings introduced or worsened by the PR and any warnings that block required tiers before committing; a commit must leave the blocking and strict CI tiers green in the exact branch/worktree that will be pushed for review
### Pull Request Scope
- Deliver one independently testable outcome per PR. Split multi-part issues, refactoring, migrations, and cleanup into serial PRs.
- Stop and request explicit user approval before a PR changes more than 10 source or test files, changes more than 1,500 non-generated source or test lines (additions plus deletions), or spans multiple independent subsystems. Temporary smoke-test plans do not count toward these limits.
- Migrate only the call sites required by the acceptance criteria. Treat similar pre-existing code as follow-up work.
- Fix only problems that the PR introduces or worsens, acceptance-criteria violations, security or data-loss risks, and merge blockers in the same PR.
- Keep local and agent self-review findings local and deduplicated. Do not publish automated self-review findings unless the user explicitly requests them; this does not replace the repository-mandated independent review below.
- Reassess scope after material fix rounds, but do not split solely because of the number of fix rounds or percentage diff growth. Split when the absolute source/test file or changed-line limits above are crossed, the work spans multiple independent subsystems, or review uncovers a separately testable outcome.
- Put purely mechanical module moves in a separate prerequisite PR.
- Request cross-machine smoke testing only after local review is complete and CI has stabilized on the candidate head.
### Pull Request Review and Merge Strategy
- For implementation work, create a focused branch in a separate worktree from fresh `origin/main` unless the user explicitly asks to use the current checkout. Keep unrelated files in the primary checkout untouched.
- Run the full Horizon validation matrix in the exact worktree and commit that will be pushed. Complete applicable local UI smoke before opening the PR. Any required cross-machine smoke must finish on the current head before reporting the PR ready to merge.
- Before opening the PR, review the full diff and run an independent local code review. Fix actionable in-scope findings and record valid out-of-scope findings as follow-up candidates.
- Open PRs ready for review by default, not as drafts, unless the user explicitly requests a draft. Include reproduction details for bug fixes, runtime or platform assumptions when relevant, and screenshots, logs, or completed smoke evidence for behavior-affecting changes.
- Every PR gets an independent Copilot review. Request it after the PR exists through the REST API, using the login `copilot-pull-request-reviewer[bot]`. The POST returns 200 whether or not it registered, so the only proof is that the PR gained a `review_requested` event — count them either side of the request:
```bash
copilot_requests() {
gh api --paginate --slurp repos/<owner>/<repo>/issues/<n>/timeline \
| jq '[.[][] | select(.event == "review_requested" and .requested_reviewer.login == "Copilot")] | length'
}
copilot_reviewed_head() {
head=$(gh api repos/<owner>/<repo>/pulls/<n> --jq .head.sha)
gh api --paginate --slurp repos/<owner>/<repo>/pulls/<n>/reviews \
| jq --arg h "$head" '[.[][] | select(.user.login == "copilot-pull-request-reviewer[bot]" and .commit_id == $h)] | length > 0'
}
before=$(copilot_requests)
gh api repos/<owner>/<repo>/pulls/<n>/requested_reviewers \
--method POST --input - <<< '{"reviewers":["copilot-pull-request-reviewer[bot]"]}'
if [ "$(copilot_requests)" -gt "$before" ]; then
echo "requested"
else
for _ in $(seq 1 30); do
[ "$(copilot_reviewed_head)" = "true" ] && break
sleep 20
done
[ "$(copilot_reviewed_head)" = "true" ] || {
echo "no new request registered and no review arrived within 10 minutes" >&2
exit 1
}
fi
```
An unchanged count is inconclusive, not a failure, which is why the else branch waits instead of exiting. Several repositories request Copilot automatically when the PR opens, so a second POST adds no event even though a review is already coming. Do not try to detect that state directly: between Copilot picking the request up and submitting, it is gone from `requested_reviewers` and not yet in `/reviews`, so every instantaneous check reads false. Waiting for the review itself is the only signal that covers the whole transition, and it is what you were going to do anyway.
Count the events rather than comparing timestamps: the PR almost always carries earlier Copilot requests, because the next rule re-requests after every push, and `created_at` has only second resolution, so a timestamp comparison can match a request made in the same second. The login filter keeps an unrelated human review request from counting.
The pagination handling is fussy and worth copying exactly. `gh api --paginate` runs `--jq` once per page, so `--jq '… | length'` prints one count per page and `$before` becomes a multi-line string that `-gt` cannot compare. `--slurp` wraps the pages into one array, but `gh` rejects `--slurp` together with `--jq` (`the --slurp option is not supported with --jq or --template`), so the pages have to be piped to a standalone `jq` — and because each page is itself an array, the filter opens with `.[][]`.
Do not substitute `requested_reviewers` for any of this — the reviewer disappears from that list as soon as Copilot picks the request up, so an empty list means nothing either way.
Every other spelling of the login fails, and most of them fail *silently*: `gh pr create --reviewer @copilot` and `gh pr edit --add-reviewer Copilot` error with `Could not resolve user with login 'copilot'`; `Copilot` over REST or GraphQL returns HTTP 200 and requests nothing; `copilot-pull-request-reviewer` without `[bot]` is rejected as not a collaborator; and GraphQL `requestReviews` with the `copilot-swe-agent` bot id reports success while recording nothing, because that bot is the coding agent rather than the reviewer.
Mind the asymmetry: the request must name `copilot-pull-request-reviewer[bot]`, but the timeline reports the reviewer as `Copilot`.
- A Copilot review is pinned to the commit it ran against, so re-request it after every push. Compare the review's `commit_id` with the current head (`gh api --paginate repos/<owner>/<repo>/pulls/<n>/reviews --jq '.[] | select(.user.login == "copilot-pull-request-reviewer[bot]") | .commit_id'`) before treating the review gate as met.
- Wait for the requested Copilot review and all repository-mandated checks on the current head. Triage every actionable comment against the PR scope: fix in-scope findings on the same PR, explicitly disposition valid out-of-scope findings as follow-up candidates, and leave no actionable thread unresolved. Before every push, rerun the repository-mandated local validation for that exact head as defined by the pre-push section above. After the push, refresh the review and checks for the new head and rerun affected smoke lanes. A behavior-affecting push invalidates smoke evidence from an older head.
- Apply repository-standard metadata only when the convention is unambiguous: assignee `@me`, `Awaiting Review` label, current milestone, and project. Otherwise report and skip the ambiguous item rather than guessing.
- Do not merge unless the user explicitly requests that specific merge. Immediately before merging, establish a stable exact head and inspect thread-aware `reviewThreads`; a flat comment list is not enough.
- The positive merge gate is: the PR is open and non-draft as intended, `mergeable` is `MERGEABLE`, readiness and merge state are neither blocked nor unknown, the required review decision is satisfied, zero actionable review threads remain unresolved, and every repository-mandated lane on the exact head has settled successfully. GitHub-required checks are only a minimum; the pedantic Clippy lane remains advisory as documented above. Abort on head drift or any new blocker. A skipped check counts only when its workflow explicitly marks the job non-applicable.
- For an authorized merge, use squash with an expected-head guard, for example `gh pr merge <PR#> --squash --match-head-commit <sha>`. Afterward, verify GitHub's merged state and resulting base commit and monitor all relevant post-merge workflows on the squash commit to a successful terminal state. Before cleanup, prove the task worktree is clean, the task branch still equals the guarded PR head with no later commits, and the squash commit contains the merged patch; then remove only task-created worktrees and branches, never a dirty or shared checkout.
- Merge approval is not release approval. Do not create a tag, publish a GitHub Release, trigger a release workflow, or claim deployment without a separate explicit request and verification.
- Treat pull-request evidence as public by default. Anonymize unrelated user, host, customer, or operational identifiers, and never publish credentials, tokens, private keys, signed URLs, or secret-bearing configuration.
### Versioning
- The authoritative base version lives in `Cargo.toml` under `[workspace.package].version`
- The `horizon-core` workspace dependency version in `Cargo.toml` must match the workspace package version
- CI validates this with `scripts/check-version-sync.sh`
### Release Flow
Releases are tag-driven and documented in [`docs/release-flow.md`](docs/release-flow.md).
- Use `./scripts/next-version.sh alpha`, `./scripts/next-version.sh beta`, or `./scripts/next-version.sh stable` to suggest the next tag for the current release line
- Save a **draft** GitHub Release using tags like `vX.Y.Z-alpha.N`, `vX.Y.Z-beta.N`, or `vX.Y.Z`; do not publish the GitHub Release by hand
- Mark alpha and beta tags as prereleases in GitHub; leave stable tags as normal releases
- Saving the draft (or dispatching the Release workflow with an existing tag) triggers CI to build and upload binaries, then publish the GitHub Release
- After a stable release, bump `Cargo.toml` to the next release line in a normal PR before cutting more prereleases
### Release Notes
When cutting a new release, generate concise release notes from the commits since the last tag (`git log <prev-tag>..HEAD --oneline --no-merges`). Group into **What's new** (features) and **Fixes** (bug fixes). Keep it scannable -- one line per item, no commit hashes. Pass the notes to `gh release create --notes --draft`.
### Dependencies
- Always check crates.io for the latest stable version before adding
- Prefer workspace-level dependencies (root `Cargo.toml`)
- New dependencies require justification
### Testing
- Unit tests close to code (`#[cfg(test)]`)
- Integration tests under `crates/*/tests/`
- Test panel creation, PTY lifecycle, resize, input routing
- CI runs the whole test suite on Linux, macOS and Windows; do not skip failing tests in the Windows job. When a test genuinely needs a Unix shell, PTY semantics or Unix paths, mark it `#[cfg_attr(windows, ignore = "<reason>")]`, or `#[cfg(unix)]` with a comment stating the reason when it cannot compile on Windows. Fix a real Windows bug instead, or file an issue for it.
- For UI/layout changes, verify with a live screenshot after launch and after resize/fit interactions; build success alone is not sufficient
- Unless release-specific behavior is the thing under test, prefer `target/debug/horizon` for smoke testing so iteration stays fast while validating UI and interaction correctness
- For any UI-related change, always create an extensive temporary smoke-test plan under `docs/testing/` that another agent or machine can execute without extra context. Cover baseline behavior, primary flows, edge cases, persistence/migration, and visual regressions.
- Temporary smoke-test plans are validation artifacts, not permanent docs. Delete them after the UI validation pass is complete unless the user explicitly asks to keep them.
### Isolated UI Testing Through Horizon Native VNC
- All interactive Horizon UI testing must run on a task-owned isolated desktop viewed live through a **Horizon native VNC Device panel** in the user's current workspace. Always use this panel; do not use noVNC or a browser viewer. Screenshots, recordings, or a viewer on another isolated desktop alone do not satisfy the live-view requirement. Unit tests, headless integration tests and static checks do not need a viewer.
- Run parallel application scenarios on distinct task-owned displays, control targets, ports, private state and evidence directories. Freeze candidate executables before launching and verify the actual application child PID and executable hash, not just a sandbox wrapper.
- Keep internal application names, workflows, screenshots, videos and operational details out of public issues, PRs, logs and attachments. Public demonstrations use generic synthetic fixtures only.
- On Linux, use the [local Horizon smoke fixture](scripts/device-smoke/README.md#horizon-inside-a-native-vnc-device-panel) with `--native-view`: a separate Xvfb display and window manager, private application state and a loopback VNC server. The separate desktop and private state provide isolation; mirroring the developer's desktop does not satisfy this rule.
- Use the public `device_panel` tool to create the viewer in the calling agent's current workspace using the fixture's `vnc_address`. Inspect `connection`, `image_received`, `image_displayed` and advancing `frame_sequence` during changing output before claiming a live view; creation or `visible: true` alone is insufficient. If the tool, supporting host or visible native panel is unavailable, report the blocked lane. Do not substitute noVNC, alter private runtime files, restart active sessions, or automate the developer's desktop.
- Diagnose viewer liveness automatically; do not ask a human to confirm visibility or whether the image is updating. Follow the [bounded viewer health procedure](scripts/device-smoke/README.md#automatic-viewer-health-check) and retain timestamped observations. Distinguish a static target from a disconnected or unpresented viewer; recover only the task-owned viewer through public lifecycle operations. Recheck actual presentation after recovery: setting visibility or reconnecting does not guarantee that an off-screen panel is brought into view. If the public contract cannot establish live presentation, record the precise blocked lane and a deduplicated follow-up issue, then continue independent checks. Once live presentation was established for a connection, a viewer that later reports `not_rendered` with `outside_canvas` was navigated away from by the person: keep the interactive test and its recording running, do not pause the lane, and do not reveal again just to advance counters. Never restart the user's Horizon instance to repair test visibility. Track the missing presentation/freshness diagnostics in [issue #801](https://github.com/peters/horizon/issues/801).
- Drive the isolated application through explicitly configured device CLI/MCP targets; native automation, when needed, must be scoped to the fixture display and exact PID. The Device panel is read-only for agents: its Interact toggle is a person's choice in the UI, and no agent tool can turn it on. Browser-page interaction remains on the `horizon-browser` skill's public `browser_*` tools; native VNC Device panels use the `horizon-device` skill and `device_panel`.
- For visible feature additions or behavior changes, capture a short video directly from the isolated test desktop, plus screenshots after launch and resize/fit. Use a recorder scoped to that display, start before the interaction, stop afterward, and play back or decode representative frames to verify the feature and movement were captured. A native Device panel does not itself record video. If recording is unavailable or stalls, report that lane as blocked; do not replace it with screenshots or noVNC. Keep finalized evidence private until publication is authorized, and record the candidate binary/commit and scenario with it.
- On macOS or Windows, use a dedicated test machine, VM or isolated desktop session exposed through VNC and viewed in the same native Device panel. The current viewer accepts numeric loopback endpoints only; use an explicitly authorized forwarding arrangement when needed, as the viewer does not create tunnels. Platform-specific graphics/input checks must still run on the actual target OS. If a suitable isolated environment or native viewing path is unavailable, report the blocked lane.
- Close the exact test window normally and clean up only fixture-owned processes and state. Close only the task-owned Device panel through `device_panel`; closing the viewer releases its connection but does not stop the application or fixture. Preserve the developer's Horizon, terminals and desktop throughout the test.
### Cross-Machine Smoke-Test Handoff
Multi-machine validation (e.g. macOS/Metal on one box, Linux/CUDA on another) is coordinated **through PR comments**, so agents on different machines can hand work to each other asynchronously:
- The implementing agent opens the PR, adds a smoke-test plan under `docs/testing/`, and posts a comment:
`SMOKE-TEST REQUEST <machine/os> — plan: docs/testing/<file>.md — scope: <which lanes>`
- The executing agent on the target machine checks out the PR branch, runs its lane of the plan, **fixes what it can and pushes those commits to the PR branch**, then replies with a report:
```
SMOKE-TEST REPORT (<machine/os>)
- <step>: pass | fail — short note
- ...
Summary: <what was fixed, what remains, anything the next agent must know>
SMOKE-TEST: DONE
```
The final line must be exactly `SMOKE-TEST: DONE` — waiting agents poll the PR for that marker. Never put the marker in a REQUEST comment; only in a completed REPORT.
- Any agent may then post a further `SMOKE-TEST REQUEST` (more scope, another machine) and the loop repeats — request → report → request → report — until every requested lane passes and the PR is ready.
- An agent resuming work on a PR must read the newest `SMOKE-TEST REPORT` comment before continuing.
### Smoke Test Reliability
- Treat motion-sensitive UI bugs as motion-sensitive validation problems: detached-window drag, resize feedback, hover jitter, and oscillation bugs are not reliably caught by still screenshots alone
- When a bug depends on live movement, capture a short video from the isolated test desktop; use high-frequency native window position traces for precise movement assertions and screenshots as supporting evidence. Scope any native recorder to the test desktop, never the developer's active desktop.
- If multiple Horizon processes may be running, scope automation and inspection to the exact PID under test. Do not target windows by application name alone; mixed-process sampling can produce false passes and false failures
- For detached-window smoke tests, verify the exact path being fixed. If the bug is in restore/relaunch behavior, seed or inspect the persisted runtime state and relaunch into that state instead of only testing fresh detach flows
- Prefer repeatable native-window assertions over visual guesswork: sample outer window position before, during, and after drag/resize to confirm the window moves monotonically and does not snap back, alternate between coordinates, or keep replaying a saved restore position
- Validate the exact branch or commit that will be pushed. If you use a merge-test worktree or disposable checkout for diagnosis, rerun the decisive smoke pass in the final branch/worktree before concluding the PR is fixed
- Window titles may lag behind state changes or differ across restore paths. When automating detached-window tests, identify the root and detached windows by PID plus non-root window membership rather than by title string alone when possible
### Speech-to-Text Smoke Testing
- On Linux with PipeWire/PulseAudio, prefer a deterministic synthetic-voice smoke over recording the room: select the monitor source for the output sink (for example `Monitor of <sink>`) as `features.speech.input_device`, then play a known phrase through that sink while holding the configured speech hotkey. Confirm the monitor source exists with `pactl list short sources` or CPAL device enumeration; do not change the user's default source or sink just to make the test work
- Use an isolated temporary config outside the repository and launch the exact candidate with `--config <temp-config> --ephemeral`. Unset `HORIZON` for the child process and define an explicit disposable terminal such as `/bin/bash --noprofile --norc`; otherwise a smoke launch from an active agent environment can inherit or resume an unrelated session
- Use the task-owned virtual display and lightweight window manager described in the native VNC testing rule above so focus and hotkey automation cannot interfere with the user's desktop. Allocate an unused display atomically, and stop only the display and window-manager processes created for the test
- Scope desktop automation to the exact Horizon PID and its non-root window. Focus only the disposable transcript terminal, press the configured hotkey down, run a blocking synthesizer command such as `spd-say -w "alpha bravo charlie, final marker zephyr"`, and release the hotkey only after synthesis finishes. Never target Horizon windows by application name when another Horizon process may be running
- Make the final words unique and assert that the inserted transcript includes the tail marker, not merely the beginning. Cover a short utterance, a normal 3–5 second utterance, an 8–10 second utterance, rapid release, and cancellation; the tail assertion is required because buffered capture regressions can look like successful partial transcription
- Capture proof of the recording state and the completed transcript, inspect the image, and correlate it with logs for the selected input device, negotiated sample rate/channel count, submitted duration, and transcription result. Treat output-device routing and audible playback as test preconditions rather than silently changing the user's audio setup
- Close the exact smoke-test window through its normal window-manager close path, verify that the candidate process exited, and remove the temporary config and proof artifacts. Never stop, signal, or reuse a pre-existing Horizon process for this test
### Windows Smoke Testing (Azure VMs / CI)
When creating an Azure VM for smoke testing, use **Standard_D4s_v3** with `MicrosoftVisualStudio:windowsplustools:base-win11-gen2:latest` as the current best-known disposable baseline. It worked in `northeurope` when DSv5 quota was unavailable. **Generate a fresh random password** for each VM — never hard-code or commit credentials to the repo.
- **Git and Git LFS are already present on the tested `windowsplustools` image**, but Horizon still needs `git lfs pull` after clone because icons and fonts are stored in LFS and `surge pack` depends on them
- **Install Visual Studio Build Tools before compiling if `C:\BuildTools\Common7\Tools\VsDevCmd.bat` is missing**:
```powershell
Invoke-WebRequest -Uri https://aka.ms/vs/17/release/vs_BuildTools.exe -OutFile C:\horizon-surge-smoke\vs_BuildTools.exe
C:\horizon-surge-smoke\vs_BuildTools.exe --quiet --wait --norestart --nocache --installPath C:\BuildTools --add Microsoft.VisualStudio.Workload.VCTools --add Microsoft.VisualStudio.Component.VC.Tools.x86.x64 --add Microsoft.VisualStudio.Component.Windows11SDK.22621
```
- **For iterative Windows smoke work, keep a warm VM and reuse it**. A reused VM avoids Azure provisioning, first boot, and the Build Tools install. If the helper is given an existing `--resource-group` and `--vm-name`, it should start and reuse that VM instead of creating a new one
- **Reuse the cached Surge toolchain whenever the Surge source commit has not changed**. `scripts/build-surge-toolchain.sh` stamps `.surge/toolchain-bin` with the source ref/commit and skips the rebuild when they still match
- **Use the released Surge tag as the default smoke path once it exists**. After `v1.0.0-beta.6`, the normal macOS/Linux/Windows smoke path is `./scripts/run-surge-filesystem-smoke.sh` or `./scripts/run-surge-azure-smoke.sh` with no override flags
- **Horizon ships `offline-gui` installers by default**. Do not switch `.surge/surge.yml` or the smoke harness back to `online-gui` unless the user explicitly wants network-at-install behavior
- **Use local or pinned Surge sources only for pre-merge validation**. Use `./scripts/run-surge-filesystem-smoke.sh --surge-path ../surge` for local unmerged smoke, or `./scripts/run-surge-azure-smoke.sh --surge-repo-url https://github.com/fintermobilityas/surge.git --surge-commit-sha <sha>` to validate an open Surge PR on Azure before merge
- **When overriding Surge for smoke, patch `surge-core` through a local `file://` Git source, not a raw crate path**. The Git source preserves Surge workspace dependency inheritance on Windows; raw crate-path overrides can fail to resolve `workspace = true` dependencies
- **On Windows, stop install-root processes before deleting `%LOCALAPPDATA%\\horizon` during repeated smoke runs**. A lingering `horizon.exe` or `surge-supervisor.exe` from the previous pass will otherwise make cleanup fail with `Device or resource busy`
- **After a headless installer run, stop the installer-launched `--surge-first-run` Horizon process before continuing the scripted smoke**. Leaving that managed-install app instance alive makes repeated Windows checks noisier and can hide later failures behind a still-running first-run session
- **Stream the guest-side smoke command output live in Azure instead of buffering it until the whole Bash command exits**. The Windows build/install path is too long to diagnose efficiently from a single final dump
- **Install prerequisites via winget only from a real user session** (winget is per-user and not available from SYSTEM):
```powershell
winget install Git.Git --accept-source-agreements --accept-package-agreements
winget install GitHub.GitLFS --accept-source-agreements --accept-package-agreements
winget install Microsoft.OpenSSH.Beta --accept-source-agreements --accept-package-agreements
winget install Rustlang.Rustup --accept-source-agreements --accept-package-agreements
winget install --id Microsoft.VisualStudio.2022.Community --source winget --force --override "--add Microsoft.VisualStudio.Component.VC.Tools.x86.x64 --add Microsoft.VisualStudio.Component.VC.Tools.ARM64 --add Microsoft.VisualStudio.Component.Windows11SDK.22621 --addProductLang En-us"
```
- **If you need SSH, add an SSH NSG rule alongside RDP and then switch to SSH once it works**. It is still faster and more reliable than repeated `az vm run-command` calls
- **Fix the OpenSSH firewall rule profile** — `Add-WindowsCapability` creates a rule for `Private` only, but Azure VMs use `Public`:
```powershell
Set-NetFirewallRule -Name "OpenSSH-Server-In-TCP" -Profile Any
```
- **Use debug builds** (`cargo build`) for smoke testing, not release — much faster compilation, sufficient for validating launch and crash behavior
- **`az vm run-command` runs as SYSTEM** — no desktop, no winget, single-command-at-a-time bottleneck. Use it for guest bootstrap, log collection, and starting scheduled tasks, not for the GUI smoke itself
- **Do not rely on the Startup folder alone to trigger the smoke**. On the tested image, autologon produced `explorer.exe` but the Startup launcher did not fire reliably. Force one autologon to create the console session, then start the actual smoke through a scheduled task with `LogonType Interactive`
- **Hyper-V Video + WARP** are the only GPU adapters on standard Azure VMs. GUI launch tests must run in an interactive user session (scheduled task with `Interactive` logon or RDP)
- **When SCP-ing manifest directories**, copy individual files — `scp -r` can create nested subdirectories that winget rejects with "Subdirectory not supported in manifest path"
### Performance Profiling
- Prefer repeatable workloads over ad-hoc observation: profile idle, panning, mouse-move, resize, and scroll as separate cases instead of treating "high CPU" as one bucket
- Start with `cargo build --profile profiling --features trace-profiling`, then use `HORIZON_TRACE_SPANS=1 RUST_LOG=info target/profiling/horizon` or `scripts/profile.sh <seconds>`; the script already falls back to traced spans when `perf` is blocked
- Keep before/after comparisons workload-matched: same board layout, same interaction script, same binary profile, same tracing mode, and the same isolated runtime state (`HOME`, config, and session inputs)
- Compare normalized span costs, not just total wall time: use per-call `avg_us` for key spans such as `horizon::app::update`, `horizon::app::lifecycle::render_active_view`, `horizon::app::panels::render_panels`, and `egui::context::pass` when frame counts differ between runs
- When profiling pointer-driven regressions, automate the interaction with native tooling on the host OS instead of relying on manual movement: `xdotool` on X11, AppleScript/Quartz or equivalent on macOS, and Win32/UIA tooling on Windows
- Scope automation and inspection to the exact PID under test when multiple Horizon processes may exist; avoid targeting windows by application name alone
- Keep interaction workloads separate from terminal-output workloads so PTY redraw noise does not mask pointer or layout regressions
- Treat short motion traces as provisional. If a resize or pointer regression only appears in a small sample, rerun it with a denser or longer scripted interaction before changing code
- For UI perf changes, capture a live screenshot after the profiled interaction as well as at launch; some "wins" are just incorrect culling
- Keep temporary perf harnesses, generated configs, and trace logs outside the repo unless the user explicitly asks to preserve them
### Typical Hotspots
- `crates/horizon-ui/src/app/panels.rs` — panel composition, titlebar chrome, history meter, hover work, and context-menu setup
- `crates/horizon-ui/src/terminal_widget/render.rs` — per-cell iteration, text batching, default-color conversion, decorations, and redundant fills
- `crates/horizon-ui/src/terminal_widget/input.rs` — per-panel event cloning/scanning and pointer-move fan-out
- `crates/horizon-ui/src/app/workspace.rs` and `crates/horizon-ui/src/app/canvas.rs` — workspace hover effects, cursor changes, and canvas redraws while moving the mouse
- `crates/horizon-ui/src/app/canvas.rs` and the minimap path — repeated static geometry such as dot grids, glows, and overview rects can become tessellation-heavy under pointer-only redraws
- Mouse-move spikes are often hover/redraw problems, not PTY/output problems; compare against an idle trace before touching terminal I/O
- Moving the pointer anywhere inside the Horizon window can trigger full-scene redraws; "empty canvas" is not a free case if visible panels still repaint
### Performance Guardrails
- Safe optimizations usually reduce repeated work on the common path: batch default terminal text, fast-path default colors, gate per-panel pointer processing behind actual interaction, and remove redundant paint passes
- Off-screen culling is generally safe; active-workspace culling is not. Do not skip panel body rendering just because a panel belongs to a non-active workspace if it can still be visible on screen
- If a perf change affects visibility, focus, hover, or selection behavior, treat it as correctness-sensitive and verify it live before keeping it
- Record the exact workload used for any reported perf gain in the commit message or PR notes so later agents can reproduce it
### UI Feature Perf Checklist
- Any new UI feature must identify its redraw surface up front: what pointer movement, hover state, animation, terminal output, or config changes will cause it to update
- Treat pointer-only frames as a first-class perf budget; if a feature adds hover behavior, compare idle vs scripted mouse-move before and after, including a pass over empty canvas outside any workspace and a pass over visible panel chrome
- Do not add unconditional per-frame work across every panel or workspace for convenience. Compute lazily on interaction, gate by on-screen visibility, or cache by stable keys
- Prefer cached meshes or cached shapes for repeated static decoration such as grids, minimaps, badges, panel chrome, and other geometry that does not semantically change every frame
- Avoid broad `request_repaint` or animation loops for passive UI polish. Repaint continuously only when there is active interaction, active animation, or real terminal/output change
- New menus, tooltips, badges, and summary labels should avoid eager text layout, string building, or list construction for every visible panel each frame
- If a feature needs per-panel pointer inspection, keep the hot path narrow: no cloning/scanning full input state for inactive panels unless the pointer is actually relevant to that panel
- When a feature adds a new always-visible overlay, include the overlay in the perf trace and live screenshot verification; overlays can dominate pointer redraw cost even when the cursor is elsewhere
- If a feature cannot be made lazy or cache-friendly, document why in the commit message or PR notes and include the measured cost on the target workload
### UI Launch Troubleshooting
- If Horizon "doesn't launch", first distinguish a crash from an unmapped window: `ps -C horizon` then `xwininfo -root -tree | rg Horizon`
- When `xwininfo -id <window-id> -stats` reports `Map State: IsUnMapped`, the process created a root window but the desktop never surfaced it; inspect first-frame UI/input code before blaming PTY startup
- When the map state is `IsViewable`, treat it as a focus, placement, or window-manager issue instead of a launch failure
## Architecture Notes
### Threading Model
- **Main thread:** egui event loop + rendering
- **Per-panel event loop thread:** alacritty_terminal `EventLoop` reads PTY output, parses VT sequences, and sends events via `mpsc::channel`
- **Input:** main thread writes to PTY via `EventLoopSender`
### Data Flow
```
Shell → PTY slave → PTY master → [alacritty EventLoop thread] → Term (VT parse) → channel → main thread → egui
Keyboard → main thread → EventLoopSender → PTY master → PTY slave → Shell
```
### Panel Lifecycle
1. `Board::create_panel()` opens a PTY via `alacritty_terminal::tty`, spawns `$SHELL`
2. alacritty `EventLoop` thread continuously reads PTY output and updates the `Term` grid
3. Each frame: drain event channel → render from `Term` state → egui
4. On resize: recalculate rows/cols → resize `Term` + PTY
5. On close: drop Panel (PTY handles cleaned up automatically)
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

