verify-pi-gui
minghinmatthewlam/pi-gui/.agents/skills/verify-pi-gui/SKILL.md
Verify pi-gui's core conversation flows through visible Electron UI with a real provider: sending, streaming, thread switching, tools, stop, drafts and restart. Use for desktop user-flow verification; includes a separate settings smoke and a prioritized feature map.
Skill1.2k starsChanged 8 days ago
- Installs packages
What's in it
- Verify pi-gui
- Launch
- Doctor
- Drive
- Evidence
- Cleanup
- Helpers
--- name: verify-pi-gui description: "Verify pi-gui's core conversation flows through visible Electron UI with a real provider: sending, streaming, thread switching, tools, stop, drafts and restart. Use for desktop user-flow verification; includes a separate settings smoke and a prioritized feature map." --- # Verify pi-gui Read [features/README.md](features/README.md) before choosing coverage. Conversation behavior is the primary proof. A passing settings/navigation smoke does not establish that the app can send or stream a message. Use the existing `apps/desktop/tests/helpers/electron-app.ts` launcher; the skill supplies UI recipes and evidence capture, not a second app implementation. ## Launch Run from the repository root with dependencies installed (Node >=22.19.0 <26, pnpm 10.25.0; setup is `pnpm install --frozen-lockfile`). The default command is the real-provider conversation proof: ```sh PI_APP_REAL_AUTH=1 \ PI_APP_REAL_AUTH_SOURCE_DIR="$HOME/.pi/agent" \ PI_GUI_PROVIDER=openai-codex \ PI_GUI_MODEL=gpt-5.6-luna \ .agents/skills/verify-pi-gui/scripts/prove.sh ``` Use a provider/model actually configured for this user; those example values are not a guarantee of usable authentication. The explicit environment opts into real requests and usage. Missing configuration exits 2 before build; invalid credentials fail the run. Never silently skip core proof or substitute fake auth. Custom endpoints that require `models.json` are not supported by this initial recipe; use a configured built-in provider or extend the credential/config setup deliberately. For the remaining mapped surfaces after the default conversation proof, use the same real-auth environment: ```sh PI_APP_REAL_AUTH=1 \ PI_APP_REAL_AUTH_SOURCE_DIR="$HOME/.pi/agent" \ PI_GUI_PROVIDER=openai-codex \ PI_GUI_MODEL=gpt-5.6-luna \ .agents/skills/verify-pi-gui/scripts/prove.sh --maintenance ``` That command starts threads from New thread, then checks Skills Try, pin while a run is going, Show more Today, a permanent worktree against `git worktree list`, and queued follow-up plus steer. Time grouping hides the folder row after those threads exist, so the worktree step switches Grouping to Workspace before it opens workspace actions. It is not a substitute for the default send/stream proof. In a Claude cloud session, the repository's SessionStart hook (`scripts/cloud/session-start.sh`) has already written `~/.pi/agent/auth.json` from the environment's `PI_AUTH_JSON_B64` variable, started Xvfb on `DISPLAY=:99`, and exported the real-auth variables above. Run `pnpm install --frozen-lockfile`, then `prove.sh` with no prefix. The environment must use `scripts/cloud/setup.sh` as its setup script and allow `chatgpt.com` and `auth.openai.com`. If the hook reports that auth is missing, or the provider rejects the credential, stop and ask the user for a fresh `PI_AUTH_JSON_B64`. Never print the auth file. For the secondary no-provider settings/navigation proof: ```sh .agents/skills/verify-pi-gui/scripts/prove.sh --smoke ``` These commands build first, show and focus Electron with `PI_APP_TEST_MODE` removed, and retain a unique `.artifacts/verify-pi-gui/run-XXXXXX/`. Build failure blocks launch; do not reuse stale output. On this host the full Xcode selection can block `git`/`swiftc` on its license. An already working Command Line Tools installation can be selected per invocation with `DEVELOPER_DIR=/Library/Developer/CommandLineTools`; verify `xcrun --find swiftc` under that environment first. Do not accept licenses or change global developer settings for the user. Each run uses a scratch workspace and isolated profile. The conversation proof copies only the selected provider credentials into a mode-0700 private temporary directory outside the evidence tree, with a mode-0600 auth file. It creates a minimal model configuration there once and reuses it across restart. It does not modify the source profile. Do not publish the private directory or credentials. Separate profiles prevent history collision, but build output and foreground input are shared: serialize runs and do not drive the user's installed app. ## Doctor On each launch, focus the owned app and perform the recipe's read-only identity checks: app path resolves to this checkout's `apps/desktop`, userData equals this run's profile, native window is visible and focused, test mode is absent, and main-process test hooks are absent. `firstWindow()` waits for DOM load and the preload bridge. Doctor JSON records actual paths and PID. Successful sending validates usable authentication; a configured provider name alone does not. If launch fails or auth is rejected, capture the error and stop that attempt. A map entry and a build are not runtime proof. ## Drive The default `conversation.spec.ts` drives [conversations](features/conversations.md) and [thread continuity](features/thread-continuity.md): 1. Click New thread, enter Alpha's prompt, click Start thread. Observe assistant text growing while the row reports running, type into a multiline draft while scrolling upward in small steps, verify the reading position, jump back to latest, then observe the final assistant marker and completion. 2. Create Bravo through the same UI. Its real shell tool writes a marker file and briefly waits. Select Alpha while Bravo runs; verify Alpha's text/draft and Bravo's continuing status. Let Bravo complete while Alpha is selected. 3. Select Bravo, inspect the assistant result and expanded tool output, and compare the actual file on disk. 4. Send a follow-up through the composer that starts a bounded 25-second shell command. Observe its running tool and start file, click Stop run, and require aborted output, idle controls, and no completion file before the natural deadline. Check distinct drafts by switching both ways. 5. Archive/restore Bravo through sidebar controls. Restart the app, select both conversations, and verify transcripts and drafts. `proof.spec.ts`, selected only by `--smoke`, clicks Settings sections, opens Skills and New thread, and verifies a preference after restart. `--maintenance` (`maintenance.spec.ts`) drives the remaining mapped surfaces in a visible session: Skills Try and slash aliases, pin/unpin while a run is going, New thread until Show more Today, Grouping switched to Workspace so the folder actions menu is visible, a permanent worktree checked with `git worktree list`, and a real-provider Enter-queue / modified-Enter steer. It does not create those threads through session IPC. The model configuration and scratch workspace are launch fixtures; they do not prove account onboarding or native folder opening. Playwright sends input to the real renderer; it does not move the desktop pointer. Native pickers, clipboard, permissions and OS switching need a native/Computer Use journey. The launcher uses the development Electron binary with freshly built app code, not a packaged installation. Read the feature map for other required entry points. [Queued follow-ups](features/follow-ups.md), worktrees, and native/package paths have separate coverage; do not imply they passed with the default proof. For product changes, use `apps/desktop/tests/AGENTS.md` and the current package scripts to choose additional regression lanes. Existing core/live specs may contain IPC fixtures or injected events; inspect them before treating them as real conversation proof. ## Evidence The helper prints the evidence directory. Conversation runs retain doctor JSON, screenshots and ARIA at each checkpoint, timestamped assistant streaming samples, scroll frame intervals, synthetic composer input/pre-close values, action traces, tool output file, Stop timing and cancellation evidence in `stop-proof.json`, build/run logs, exit code, progress/result JSON, and cleanup records. These recipes do not request Playwright Electron `recordVideo`. For a UI change verified in a Claude cloud session, record the whole display with the `record-screen` skill: start it just before `prove.sh` and stop it after, then attach the MP4 to the result. `prove.sh` does not record video itself. Inspect assistant-only content so a prompt containing the expected answer cannot make the test pass. Streaming requires observed growth during a run; a completed answer alone is insufficient. Persistence requires a second process using the same profile. Tool proof needs both visible output and the file side effect. Open the actual `conversation.zip` or `restart.zip` from the printed directory with `pnpm exec playwright show-trace`. The settings smoke instead uses `change.zip` and `restart.zip`. Maintenance uses `surfaces.zip` and `follow-ups.zip`. Inspect screenshots and traces as well as assertions. Keep credentials outside artifacts shared with reviewers; trace source includes test code and prompts, so use synthetic prompts only. This is not a dry-run: conversation mode contacts the configured provider, consumes usage, creates session files, and runs the specified scratch-workspace tool command. Report exact completed checkpoints and any failure. Never present a partial/blocked run as a full pass. ## Cleanup Both recipes close only their owned Electron applications in `finally`, including assertion/trace failures, and record the PIDs in `cleanup.json`. Confirm those processes exited and the proof artifacts still exist. For launch failure, inspect Playwright's teardown/error; terminate only a verified owned PID if stranded, never by process name. Root `AGENTS.md` prohibits deleting temporary artifacts without approval. Retain scratch state and the private profile; do not put that private credential-bearing directory into shared evidence. Request explicit cleanup authorization before deleting retained private state. Process teardown still completes on every attempted run. ## Helpers - `scripts/prove.sh [--conversation|--smoke|--maintenance]`: executable CLI; default is conversation. Builds, allocates artifacts, runs the chosen recipe, preserves exit status, checks retained proof. `--maintenance` requires the same real-auth environment as conversation. - `scripts/conversation.spec.ts`: default real-provider user journey. - `scripts/maintenance.spec.ts`: visible sweep of skills Try, pin while running, Show more Today, worktree plus Git, and queued follow-ups. - `scripts/proof.spec.ts`: secondary no-provider UI smoke. - `scripts/playwright.config.ts`: inherits repo defaults, selects these recipes, disables retries. Use `$maintain-verification-skill` to keep the map and live recipes aligned with the app.
More agent context in minghinmatthewlam/pi-gui
3 other files this repository gives its agents.
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.

