mirroir-mcp
jfarcand/mirroir-mcp/llms.txt
Give your AI eyes, hands, and a real iPhone. An MCP server that lets any AI agent see the screen, tap what it needs, and figure the rest out — through macOS iPhone Mirroring. Also supports macOS windows. 38 tools, any MCP client (Claude Code, Cursor, Copilot, ChatGPT, Codex). Written in Swift. Apache 2.0 licensed. mirroir-mcp gives AI agents eyes and hands on a real iPhone. Any LLM with MCP support can screenshot the phone, read the screen via OCR,…
llms.txt228 starsChanged 7 days ago
- Installs packages
# mirroir-mcp
> Give your AI eyes, hands, and a real iPhone. An MCP server that lets any AI agent see the screen, tap what it needs, and figure the rest out — through macOS iPhone Mirroring. Also supports macOS windows. 38 tools, any MCP client (Claude Code, Cursor, Copilot, ChatGPT, Codex). Written in Swift. Apache 2.0 licensed.
## Overview
mirroir-mcp gives AI agents eyes and hands on a real iPhone. Any LLM with MCP support can screenshot the phone, read the screen via OCR, decide what to do, and execute taps, types, and swipes — no jailbreak, no simulator, no app SDK. The agent loop (observe, reason, act) runs on any model: Claude, GPT, Gemini, or local Ollama. For repeatable workflows, capture them as skills — numbered-step SKILL.md files the AI follows adaptively. It connects via the macOS iPhone Mirroring feature (macOS 15+) and exposes a JSON-RPC 2.0 MCP server over stdio.
The name is the old French spelling of *miroir* (mirror).
- Repository: https://github.com/jfarcand/mirroir-mcp
- Website: https://mirroir.dev
- NPM: https://www.npmjs.com/package/mirroir-mcp
- Homebrew tap: https://tap.mirroir.dev
- Skills marketplace: https://github.com/jfarcand/mirroir-skills
- License: Apache-2.0
## Architecture
Single-process design using the macOS CGEvent API for all input:
1. **mirroir-mcp** (user process) — MCP server, window discovery via AXUIElement, screen capture, Vision OCR, coordinate mapping, CGEvent input (pointing + keyboard), skill system. Communicates with MCP clients over stdin/stdout JSON-RPC 2.0.
2. **HelperLib** (shared Swift library) — Types, keyboard maps, timing constants, and protocol definitions shared across the main executable and test targets.
The input path: MCP client -> mirroir-mcp -> CGEvent API -> macOS HID -> iPhone Mirroring -> iPhone.
## Installation
Three methods: curl installer, npx, or Homebrew.
```bash
# curl installer
/bin/bash -c "$(curl -fsSL https://mirroir.dev/get-mirroir.sh)"
# npx
npx -y mirroir-mcp install
# Homebrew
brew tap jfarcand/tap && brew install mirroir-mcp
```
Requires macOS 15+ and an iPhone connected via iPhone Mirroring.
## MCP Tools (33 total)
### Screen (read-only, always available)
- `screenshot` — Capture iPhone screen as base64 PNG
- `describe_screen` — Analyze the screen using local OCR (Apple Vision) or AI vision (embacle FFI) depending on `screenDescriberMode`. Returns UI elements with tap coordinates plus grid-overlaid screenshot. `scroll: true` does full-page scroll and deduplication. Detects unlabeled icons via pixel clustering.
- `start_recording` / `stop_recording` — Video recording of the mirrored screen
### Input (mutating, requires permission)
- `tap` — Tap at (x, y) coordinates relative to mirroring window
- `double_tap` — Two rapid taps for zoom/text selection
- `long_press` — Hold tap for context menus (default 500ms)
- `swipe` — Swipe between two points (scroll wheel events = iOS scroll)
- `drag` — Slow sustained drag for icons, sliders (touch events, not scroll)
- `touch` — Persistent single finger across calls: `begin`/`move`/`end`/`cancel`; moves up to 5000ms each, placed against the live window (released if the window changes size); other pointing tools refuse while held; auto-released after 30s idle
- `pinch` — Two-finger pinch at (x, y): `scale` > 1 zooms in, < 1 zooms out (0.1-10, default 500ms); reaches iOS as a UIKit two-finger gesture (Maps, Photos), games may ignore it; needs a built-in or Magic Trackpad
- `rotate` — Two-finger rotation at (x, y) by `degrees` (positive counter-clockwise, up to 360); UIKit two-finger gesture, games may ignore it; needs a built-in or Magic Trackpad
- `hold_keys` — Hold 1-6 keys (`w`, `shift`, `space`, arrows) for `duration_ms` with auto-repeat, optionally while dragging the left or right mouse button; for games with mouse/keyboard controls (Roblox), touch-only games ignore keys; refused while a touch is held or when the mirroring window is not frontmost
- `type_text` — Type text via CGEvent key events. Supports non-US layouts via UCKeyTranslate. Accented characters via dead-key sequences.
- `press_key` — Special keys (return, escape, tab, delete, arrows) with optional modifiers (command, shift, option, control)
- `shake` — Trigger shake gesture (Ctrl+Cmd+Z) for undo/dev menus
### Navigation (mutating, requires permission)
- `launch_app` — Open app by name via Spotlight search
- `open_url` — Open URL in Safari
- `press_home` — Go to home screen
- `press_app_switcher` — Open app switcher
- `press_back` — Navigate back by OCR-tapping the "<" back chevron, with a canonical-position fallback (the only reliable back navigation through iPhone Mirroring)
- `spotlight` — Open Spotlight search
- `scroll_to` — Scroll until a text element becomes visible via OCR. Detects scroll exhaustion.
- `reset_app` — Force-quit app via App Switcher (swipes through carousel to find off-screen cards)
### Measurement & Network
- `measure` — Time screen transitions: perform action, poll OCR until target appears
- `set_network` — Toggle airplane/Wi-Fi/cellular via Settings app navigation
### Info (read-only, always available)
- `status` — Connection state, window geometry, device readiness
- `get_orientation` — Portrait/landscape and window dimensions
- `check_health` — Comprehensive setup diagnostic
### Skills (read-only)
- `list_skills` — List available skills from project-local and global config dirs
- `get_skill` — Read skill file with ${VAR} env substitution
### Skill Generation & Compilation
- `generate_skill` — AI explores an app and produces SKILL.md. Session-based: start -> capture -> finish. `action: "explore"` runs autonomous BFS exploration with component-aware planning, edge classification (push/tab/modal/dead), and smart scrolling with exhaustion detection. `skip_calibration: true` bypasses component detection (useful with AI vision describers). `emit: true` (on `finish` / `explore`) also writes a runnable iOS leg into the consumer repo's `.mirroir/apps/<app>/` — a `target: { kind: ios }` scenario, a `SAMPLE.md`, and a `must_pass` plan entry — that `mirroir-run` replays through `mirroir-mcp test`; `output_dir` names the repo root.
- `record_step` / `save_compiled` — Record compiled steps during AI-driven skill execution for zero-OCR replay
### Component Detection (read-only)
- `calibrate_component` — Test a component definition (.md) against the current live screen with a diagnostic report
### Multi-Target
- `list_targets` (read-only) / `switch_target` — Support for multiple windows (iPhone Mirroring + generic macOS windows like emulators, VNC)
Coordinates are in points relative to the mirroring window top-left. Use `describe_screen` for exact tap coordinates.
- [Full tools reference](docs/tools.md)
## Skill System
Skills are multi-step automation flows. Steps use OCR-based landmarks (no hardcoded coordinates).
### SKILL.md Format (recommended)
YAML front matter + numbered markdown steps. AI interprets steps as intents and calls MCP tools to execute them adaptively.
```markdown
---
version: 1
name: Check iOS Version
app: Settings
tags: ["settings"]
---
## Steps
1. Launch **Settings**
2. Wait for "General" to appear
3. Tap "General"
4. Wait for "About" to appear
5. Tap "About"
6. Screenshot: "about_screen"
```
### YAML Format (legacy)
Structured step definitions for the deterministic test runner.
`${VAR}` placeholders resolve from environment variables. `${VAR:-default}` for fallbacks.
Skills are placed in `~/.mirroir-mcp/skills/` (global) or `<cwd>/.mirroir-mcp/skills/` (project-local).
### Compiled Skills
Compile once to capture coordinates and timing. Replay with zero OCR for fast, deterministic execution. The compiled format is version 2. Compiled steps use one of five action kinds: `tap`, `sleep`, `assertion`, `scroll_sequence`, and `passthrough`. Assertions (`assert_visible` / `assert_not_visible`) compile to a real `assertion` action that sleeps the observed delay then re-runs OCR — they are not collapsed into `sleep`. The sleep buffer added to observed delays is configurable (`compiledSleepBufferMs`, default 200ms).
The CLI auto-recompiles a compiled skill when it drifts (version mismatch or content drift); pass `--no-auto-recompile` to disable.
```bash
mirroir compile apps/settings/check-about
mirroir test apps/settings/check-about # auto-detects .compiled.json
```
- [Skills marketplace docs](docs/skills-marketplace.md)
- [Compiled skills docs](docs/compiled-skills.md)
## CLI Subcommands
- `mirroir test <skill>` — Run skills deterministically (no AI). Supports `--junit`, `--report-json` (Playwright JSON-reporter format, read by `mirroir-run`), `--capture`, `--verbose`, `--dry-run`, `--agent` for AI diagnosis.
- `mirroir compile <skill>` — Compile a skill to .compiled.json
- `mirroir record -o <file>` — Record interactions as skill YAML via CGEvent monitoring
- `mirroir migrate <file>` — Convert YAML skills to SKILL.md format
- `mirroir doctor` — Verify setup (accessibility, mirroring, permissions, etc.)
## Web and iOS Replay: mirroir-run and the `.mirroir/` dotfile
The iPhone leg above is half the picture. **mirroir-run** (`runner/`, Rust, Apache-2.0) is the cross-platform replayer for the same `SkillStep` YAML grammar. Web, process and HTTP scenarios run on Linux CI with no macOS dependency; a `target: { kind: ios }` block is handed to `mirroir-mcp test` on a Mac with iPhone Mirroring connected, which drives the phone and writes a Playwright-JSON-shaped report the runner ingests like a web block's. A `cross_surface:` step's `captures:` then compare the page's scrape with the phone's final screen, both taken live by the same run (see docs/web-and-ios.md). mirroir-run is the compiler + orchestrator + oracle of the pair: it compiles a scenario's web steps into a Playwright `.spec.ts` at run time and shells out to `npx playwright test`, while process (`spawn` / `wait_port` / `kill` / `assert_log`), HTTP (`http`), LLM-judge (`judge`), and cross-surface (`cross_surface`) steps execute natively in Rust. Playwright owns the browser — Chromium, Firefox, and WebKit, all three available on Linux. mirroir-run never drives a browser itself. The compiled spec and its config are written to `target/playwright/<sample>/<scenario>/` — the same directory `--emit playwright` writes and a run executes — and the trace, video, and screenshot Playwright records for a failing test stay there afterwards. Every compiled spec also collects uncaught page errors and failed responses, so a page that throws is a failure even when every locator resolved.
The recording half is an agent: `.claude/skills/mirroir-onboard/` drives a running web app through `chrome-devtools-mcp`, reads real selectors out of the live accessibility tree, and emits the scenarios. Playwright is replay-only. On iOS, `generate_skill action=explore` plays the same recorder role, and `emit: true` writes what it recorded as a runnable `ios` scenario.
### The `.mirroir/` dotfile
A consumer repo checks in a `.mirroir/` tree; `mirroir-run` with no arguments walks `cwd ↑` to find it.
```
<repo>/.mirroir/
├── mirroir.yaml # the plan: samples, archetype refs, boot wiring
├── mirroir.lock # pinned archetype versions + tree checksums
├── mirroir.local.yaml # developer-local overrides (never committed)
├── apps/<name>/ # local samples: APP.md, SAMPLE.md, scenarios/, baselines/
└── .build/ # composed, ready-to-replay tree (generated)
```
`.mirroir/` (a consumer's checked-in plan) is distinct from `.mirroir-mcp/` (the Swift MCP server's home for element patterns and skills).
### Installing mirroir-run
```bash
# Homebrew (macOS arm64/x86_64, Linux x86_64)
brew install jfarcand/tap/mirroir-run
# crates.io
cargo install mirroir-run --locked
# Prebuilt archives, per `runner-v*` release: linux-gnu, linux-musl (static),
# apple-darwin (x86_64 + aarch64), pc-windows-msvc — each with a .sha256 sidecar
# https://github.com/jfarcand/mirroir-mcp/releases
```
Web scenarios additionally need Node 18+ and `@playwright/test`; point `MIRROIR_PLAYWRIGHT_HOME` at the directory holding its `node_modules/`. There is no flag that skips a scenario's web run — a lane that cannot run web steps fails instead of reporting a green run.
### Running it
```bash
mirroir-run # discover .mirroir/, compose, replay
mirroir-run --config path/.mirroir/mirroir.yaml
mirroir-run --sample runner/samples/web-fixture # one SAMPLE.md directory
mirroir-run --run-scenario flow.yaml # one scenario file
mirroir-run --validate flow.yaml # parse, version-gate, build the plan
mirroir-run --emit playwright flow.yaml # compile to target/playwright/, don't run
mirroir-run --diff-text baseline.txt current.txt # drift: Jaccard + Levenshtein
mirroir-run --locked # CI gate: stale lockfile is an error
mirroir-run --frozen # --locked plus no network fetch
mirroir-run accept # re-record every baseline; refuses to run in CI
```
Three verdicts, three exit codes: **0 = PASS**, **1 = FAIL**, **65 = DRIFT** — every structural assertion held and a drift metric moved past its threshold. A scenario in which the runner evaluated nothing is a failure, not a pass. `--report <PATH>` writes the JSON run summary (schema 2; default `mirroir-run-report.json`).
Drift is measured against `.harness/last-green.json`, what the previous PASS run observed, over four metrics: `fingerprint_similarity`, `judge_score_swing`, `response_levenshtein_pct`, `step_latency_pct_increase`. Thresholds resolve step → scenario `drift:` → `APP.md drift_defaults:` → `drift-defaults.yaml`, and are **fail-closed**: a metric no layer declares stops the run with `unspecified drift threshold for <metric>` rather than defaulting. A drifted run appends a candidate row to `.harness/drift-log.md` and leaves the baseline for a human. `mirroir-run accept` is the answer to that queue: it re-runs the same scenarios with the baselines in write mode, rewriting `.harness/last-green.json`, every `judge.drift_baseline_file`, every `cross_surface.capture.to`, and `.mirroir/mirroir.lock`, then hands the reviewer a `git diff`. It **refuses to run when a CI environment variable is set** — a job that accepts its own drift reports green forever. A `cross_surface:` baseline owned by a surface the runner does not drive (`baselines/<flow>.ios.txt`, written by `generate_skill` on a real iPhone) is named, never overwritten.
`mirroir.lock` is verified, not just diffed: alongside the ref set and each pin's version constraint, every recorded `checksum: sha256:…` is recomputed against the archetype tree on disk, so a locally edited or tampered pack fails `--locked` / `--frozen` under an unchanged pin.
- [Runner README](runner/README.md)
- [.mirroir/ dotfile format](runner/docs/mirroir-dotfile.md)
- [SAMPLE.md format](runner/docs/sample-md-format.md)
- [Drift and accept](runner/docs/drift-and-accept.md)
- [Playwright setup](runner/docs/playwright-setup.md)
- [CI integration](runner/docs/ci-integration.md)
## Security & Permissions
**Fail-closed by default.** Without configuration, only read-only tools are exposed. Mutating tools are hidden entirely from the MCP client.
The read-only allow-list (`PermissionPolicy.readonlyTools`) is exactly 11 tools: `screenshot`, `describe_screen`, `start_recording`, `stop_recording`, `get_orientation`, `status`, `check_health`, `list_targets`, `list_skills`, `get_skill`, `calibrate_component`. Every other tool — including `press_back` — is mutating and permission-gated.
Opt-in via `~/.mirroir-mcp/permissions.json`:
```json
{
"allow": ["tap", "swipe", "type_text", "press_key", "launch_app"],
"deny": [],
"blockedApps": ["Wallet", "PayPal"]
}
```
Kill switch: closing iPhone Mirroring or locking the phone kills all input immediately. The MCP server communicates exclusively via stdin/stdout — no network ports, no sockets, no daemons.
- [Security model](docs/security.md)
- [Permissions reference](docs/permissions.md)
## Building from Source
Swift Package Manager (Swift 6.0+, macOS 14+ SDK):
```bash
git clone https://github.com/jfarcand/mirroir-mcp.git
cd mirroir-mcp
swift build # debug build
swift build -c release # release build
swift test # run all tests
./mirroir.sh # full install (build + configure MCP client)
```
### SPM Targets
- **mirroir-mcp** — MCP server executable
- **HelperLib** — Shared library
- **FakeMirroring** — APP.md-driven multi-app simulator standing in for iPhone Mirroring in CI (AppPack/AppRegistry/AppPackLoader, with HelperLib SimulatorSpec/SimulatorSpecParser)
### Test Targets
- **MCPServerTests** — Server routing, tool handlers, exploration, graph algorithms
- **HelperLibTests** — Keyboard maps, timing, permissions, coordinate calculation
- **TestRunnerTests** — Skill parsing, step execution, reporting
- **IntegrationTests** — Full workflows with FakeMirroring app
## Key Source Files
### Entry Points
- `Sources/mirroir-mcp/mirroir_mcp.swift` — Main entry point, CLI dispatch, target registry init
### Core Infrastructure
- `Sources/mirroir-mcp/MCPServer.swift` — JSON-RPC 2.0 server implementation
- `Sources/mirroir-mcp/ToolHandlers.swift` — Tool registration orchestrator
- `Sources/mirroir-mcp/Protocols.swift` — All protocol abstractions (WindowBridging, InputProviding, ScreenCapturing, ScreenDescribing, ExplorationStrategy)
### Window & Input
- `Sources/mirroir-mcp/MirroringBridge.swift` — iPhone Mirroring window discovery via AXUIElement
- `Sources/mirroir-mcp/GenericWindowBridge.swift` — Non-iPhone window bridge
- `Sources/mirroir-mcp/InputSimulation.swift` — Coordinate mapping, focus management
- `Sources/mirroir-mcp/CGEventInput.swift` — CGEvent posting for pointing + keyboard
- `Sources/mirroir-mcp/CGKeyMap.swift` — Character → macOS virtual keycode mapping
### Screen Operations
- `Sources/mirroir-mcp/ScreenDescriber.swift` — Apple Vision OCR pipeline (local)
- `Sources/mirroir-mcp/VisionScreenDescriber.swift` — AI vision screen describer via embacle FFI
- `Sources/mirroir-mcp/EmbacleFFI.swift` — Rust FFI bridge to embedded embacle runtime
- `Sources/mirroir-mcp/ScreenCapture.swift` — screencapture CLI wrapper
- `Sources/mirroir-mcp/IconDetector.swift` — Unlabeled icon detection via pixel clustering + Vision saliency
### Skill System
- `Sources/mirroir-mcp/SkillMdParser.swift` — SKILL.md front matter + body parser
- `Sources/mirroir-mcp/SkillParser.swift` — YAML skill parser
- `Sources/mirroir-mcp/SkillMdGenerator.swift` — Generate SKILL.md from explored screens
- `Sources/mirroir-mcp/CompiledSkill.swift` — Compiled skill data model
- `Sources/mirroir-mcp/CompiledStepExecutor.swift` — Zero-OCR replay engine
### Autonomous Exploration
- `Sources/mirroir-mcp/BFSExplorer.swift` — Breadth-first exploration with frontier queue and path replay (default explorer)
- `Sources/mirroir-mcp/BFSExplorerHelpers.swift` — Calibration pipeline (scroll, classify, component detect, plan)
- `Sources/mirroir-mcp/DFSExplorer.swift` — Depth-first exploration with backtrack stack
- `Sources/mirroir-mcp/NavigationGraph.swift` — Directed screen graph (nodes=screens, edges=transitions, structural fingerprinting)
- `Sources/mirroir-mcp/EdgeClassifier.swift` — Classify navigation edges (push/tab/modal/dead/external)
- `Sources/mirroir-mcp/ExplorationSession.swift` — Thread-safe session accumulator
- `Sources/mirroir-mcp/MobileAppStrategy.swift` — iOS app exploration heuristics
- `Sources/mirroir-mcp/CalibrationScroller.swift` — Content-aware scrolling with exhaustion detection
### Component Detection
- `Sources/mirroir-mcp/ComponentLoader.swift` — Discovers and loads component definition .md files from disk
- `Sources/mirroir-mcp/ComponentDetector.swift` — Groups OCR elements into UI components using loaded definitions
- `Sources/mirroir-mcp/ComponentScoring.swift` — Scores component definitions against OCR row properties
### Shared Library
- `Sources/HelperLib/MCPProtocol.swift` — JSON-RPC 2.0 types and MCP tool definitions
- `Sources/HelperLib/AppleScriptKeyMap.swift` — macOS virtual key codes
- `Sources/HelperLib/TimingConstants.swift` — Named timing delays and configuration
- `Sources/HelperLib/PermissionPolicy.swift` — Fail-closed permission engine
## Configuration
Timing defaults can be overridden via `settings.json` or environment variables:
```json
// .mirroir-mcp/settings.json
{
"keystrokeDelayUs": 20000,
"clickHoldUs": 100000
}
```
Environment variable form: `MIRROIR_KEYSTROKE_DELAY_US`.
Multi-target configuration via `targets.json` for controlling multiple windows simultaneously.
## Known Limitations
- **Focus stealing**: Input tools must make iPhone Mirroring frontmost. No API exists to direct CGEvent input to background windows. Mitigate with a separate macOS Space.
- **Clipboard paste**: iPhone Mirroring does not bridge Mac clipboard when paste is triggered programmatically. No workaround exists.
- **Keyboard layout edge cases**: Two characters on the ISO section key cannot be typed due to macOS/iOS key mapping disagreement.
- **iOS autocorrect**: Applied to typed text. Disable in iPhone Settings if problematic.
- [Full limitations](docs/limitations.md)
- [FAQ](docs/faq.md)
- [Troubleshooting](docs/troubleshooting.md)
## Documentation Index
- [README](README.md) — Installation, examples, quick start
- [Tools Reference](docs/tools.md) — All 38 tools with parameters
- [Security](docs/security.md) — Threat model, kill switch, recommendations
- [Permissions](docs/permissions.md) — Fail-closed permission model
- [Component Detection](docs/components.md) — Component definitions, calibration, detection pipeline
- [Compiled Skills](docs/compiled-skills.md) — Zero-OCR skill replay
- [Skills Marketplace](docs/skills-marketplace.md) — Skill format and authoring
- [Testing](docs/testing.md) — FakeMirroring, CI strategy
- [Known Limitations](docs/limitations.md) — Focus stealing, keyboard gaps
- [FAQ](docs/faq.md) — Common questions
- [Troubleshooting](docs/troubleshooting.md) — Debug mode, common issues
- [Contributing](CONTRIBUTING.md) — How to add tools, commands, tests
- [Runner README](runner/README.md) — mirroir-run: web/process/http replay on Linux CI
- [.mirroir/ Dotfile](runner/docs/mirroir-dotfile.md) — Consumer-repo plan, lockfile, compose
- [Playwright Setup](runner/docs/playwright-setup.md) — Node install, emitted spec, selector strategy
- [Drift and accept](runner/docs/drift-and-accept.md) — the DRIFT verdict, the fail-closed threshold hierarchy, and `mirroir-run accept`
- [CI Integration](runner/docs/ci-integration.md) — Installing mirroir-run in another repo's CI
Discussion
Did this work in your project? Say what you used it for and what you changed. People and their agents can both post here.
Posts are public.Sign in to post
No one has posted yet. Be the first.

