create-routine
Mentra-Community/MentraOS/.agents/skills/create-routine/SKILL.md
Create, edit or port a Mentra automated testing routine with English requirements, saved actions verified during authoring, and deterministic replay through the shared framework. To request routine runs or authoring from a PR, use select-pr-routines instead.
Skill2.4k starsChanged yesterday
What's in it
- Create, edit or port a routine
- Put the behavior in the routine
- Hold one session and verify the saved actions
- Shared lifecycle and modified miniapps
- Replay and publish
---
name: create-routine
description: Create, edit or port a Mentra automated testing routine with English requirements, saved actions verified during authoring, and deterministic replay through the shared framework. To request routine runs or authoring from a PR, use select-pr-routines instead.
---
# Create, edit or port a routine
Routines should be **fast, reliable and easy to create or edit**. Keep English
instructions, observable expectations and executable actions together in the private
[Mentra-Automated-Testing repository](https://github.com/Mentra-Community/Mentra-Automated-Testing).
Read its [porting guide](https://github.com/Mentra-Community/Mentra-Automated-Testing/blob/main/docs/ROUTINE-PORTING.md)
when migrating old coverage; it links the deleted source and explains what to reuse.
## Put the behavior in the routine
Inspect the selected harness revision's `routines/`, `framework/` and controller
schemas. Start from the closest routine for the platform and glasses. Preserve
proven product actions and fixtures; replace old executor/ownership wrappers.
| Location | Responsibility |
| --- | --- |
| `routines/<id>/routine.ts` | Export `createRoutine(state)` using `defineRoutine` and `step`; English metadata, product setup/steps/teardown and fixtures |
| `framework/drivers/` | Shared interactions and recorded observations; Mac uses `executeMacStep(action, context)` with the supplied `context.ui` |
| `framework/platforms/`, `framework/glasses/` | Composed platform and glasses lifecycle providers |
| `framework/authoring/`, `orchestration/` | Held sessions, jobs, lane ownership, repair and publication |
Declare platforms, entry (`home` or `sign-in`), account, requirements, fixtures,
stable ordered step IDs, `glasses.models` and required capability IDs in `requires`.
Keep device identities, secrets and tool paths in private lane configuration.
Confirm the installed lane offers those capabilities. Add a reusable provider once
for missing shared functionality; do not hide host setup in product steps.
Preflight the complete fixture contract before reserving hardware. Check tool roles,
not just executable hashes: the Mac UI driver and app launcher are distinct pins.
When a capability is missing, assign its shared provider work separately and use
the other lane or already installed routines while it is built.
No routine-name branches in workers, dispatch or catalog, and no hardcoded videos:
source enrollment discovers definitions; published passing runs supply examples.
## Hold one session and verify the saved actions
Start through the built-in machine `routine-work` create/edit job described in the
harness [job guide](https://github.com/Mentra-Community/Mentra-Automated-Testing/blob/main/docs/ROUTINE-WORK.md)
and [assigned-agent skill](https://github.com/Mentra-Community/Mentra-Automated-Testing/blob/main/.agents/skills/prepare-routine-work/SKILL.md).
The supervisor owns the workspace, machine agent, reservation and held session.
The following are inner operations for that assigned agent, using its provisioned
`MENTRA_TEST_CLIENT_CONFIG` and the current `mentra-test` CLI; they are not a
parallel coordinator authoring path:
```sh
bun run mentra-test lane request @reservation.json
bun run mentra-test lane wait @wait.json
bun run mentra-test author start @start.json
bun run mentra-test author command @command.json
bun run mentra-test author inspect @scope.json
bun run mentra-test lane give-back @give-back.json
```
Read harness `orchestration/README.md`, `orchestration/controller.ts` and
`framework/authoring/session.ts` for current schemas and held-session behavior;
`contracts/controller.ts` defines admission. Run the CLI from the harness checkout,
not MentraOS. Do not invent IDs or use an old standalone author CLI.
Use `ControllerClient`/these service endpoints for controller mutations, including
diagnostic attachments. Never open `ControllerStore` against the live database to
register evidence, change ownership or manufacture cleanup receipts. Read-only SQL
can help inspect state; a missing public operation is framework work to assign.
Reservation request supplies `requestId`, `laneId`, `purpose`, `admissionExpiresAt`.
Wait with `{reservationId, afterGeneration, timeoutMs}` until granted. Start supplies
the granted `reservationId`, `generation`, a stable `operationId`, selected `build`
and canonical editable `sourcePath`. Default start performs setup and starts the
original recorder; optional `setupMode: "manual"` exposes individual lifecycle actions.
Each author command carries `{reservationId, generation, operationId, command}`.
Nested commands use `op`: `steps`, `snapshot`, `step` with `stepId`, `actions` with
`phase: "setup" | "test" | "teardown"`, `action` with setup/teardown `phase` and
`actionId`, or `finish`. Use returned IDs and inspect `{reservationId, generation}`
until each operation settles; a settled operation may contain a failed assertion.
Use a new operation ID for each action; a lost response reuses its original ID to
reconcile that call. Direct driver calls must retain the supplied owned context.
For example, the inner product command is `{ "op": "step", "stepId": "saved-id" }`,
inside `command`, not a separate CLI verb.
Declare the complete flow before starting. Use computer use to discover controls,
save each action and assertion, then execute that saved action through the same
driver/helper replay will use. Traverse the **whole English flow** this way: a manual
click does not prove a different script written afterward. Prefer the simplest
supported interaction that works; verify outcomes rather than successful clicks.
Complete every saved product step and normal teardown before ordinary replay.
A partially successful held traversal or expired recording is not that boundary.
On a settled step failure, inspect the actual error, edit that existing action and
retry with a concrete `retryReason` from its current safe prerequisite state. Keep
the same owner, recorder and passing prefix; do not reinstall or restart setup for
an ordinary authoring mistake. Do not repeat an uncertain submission/firmware write.
If returning to a prerequisite needs an already passed product action, inspect
`{op: "actions", phase: "test"}` for eligibility and repeat that same saved action
with an explicit `retryReason` describing the observed prerequisite. The controller
must confirm its previous input settled; a source reload alone permits no repeat.
For a completed navigation tap, wait for its observable destination before the
next input. A delivered-but-rejected tap keeps its intent: reconcile the resulting
page without tapping again. Keep those checks in the same saved action for replay.
The held loader preserves `createRoutine(state)` state and original lifecycle while
reloading existing product steps. Changing step IDs/order, lifecycle callbacks or
metadata requires finishing the session first. Shared helper/native changes require
the updated installed revision and a fresh session. Fix a broken app control rather
than accumulating alternate input or focus algorithms.
Prepare and compile changed shared source off hardware while other work uses the
lanes. Once affected owners release, activate one frozen candidate; local iteration
may use reviewed source before merge while retaining the PR's review/CI merge gates.
Check the installed recorder's duration, byte limit and output allowance before a
long flow, including held editing and accepted operation settlement; step deadlines
do not extend capture. A Mac fixture with a recorded
browser window declares `external-window` with `fixture-data` and uses the shared
admitted policy for both recordings. A product update can continue after a recorder
or client deadline; inspect that original operation and settle it normally rather
than issuing another update or restarting the passing prefix.
Request authoring reservations before waiting for the current run to finish, so
the next queued job does not repeatedly displace ready authoring work.
Before ordinary dispatch, confirm the installed executor source and enrolled
definition revision agree; frozen requests do not change during a service upgrade.
Enroll the intended source before submitting new work. A stale local request that
never launched can be cancelled normally with a reason, then replaced with the
same saved actions on the intended source. Preserve launched/nightly requests.
Publish startup failures through normal evidence delivery without replay; inspect
the exact export error when publication stalls rather than repeating the test.
For UI transitions, verify the departing overlay disappears as well as the
new page appears. Home controls can remain visible behind a miniapp. Use bounded
postcondition observation; an acknowledged click is not a completed transition.
Inspect current controls rather than copying old labels blindly: Android's radio
icon can toggle while its label opens details, and an empty miniapp switcher opener
can remain present on idle Home. Require the actual state before sending input.
Before hardware, compare the old saved selector with the selected build's current
component. Several URL editors can coexist: preserve OTA's specific manifest
placeholder instead of selecting any editable field. After an uncertain typing
response, observe the exact requested value before clearing or typing again;
the empty placeholder disappears when input succeeded. Keep that observation
in the same saved action, not a separate replay technique.
Static headings may appear twice on a platform: require readable content, and use
exact IDs/counts for the actionable controls that must be unique.
On Android, use the supplied `ui.scroll(anchor, direction)` for a bounded gesture
inside the observed scroll view, then resnapshot. Check `checked` for toggles rather
than assuming a click changed them; public text replacement uses `clearText` before
`type`. Use `hideKeyboard` for the actual IME. The optional `systemUi` retains the
same ownership and permits only the enrolled system-dialog namespaces; normal `ui`
remains scoped to the Mentra App.
Reuse observations, not assumptions. Android can repeat a radio label on its
parent and text child; count actual checkable controls. A tall option group can
span viewports: accumulate all known checked/unchecked states under the same
foreground group, reject contradictions and scroll toward an unobserved option.
Validate visible preconditions again in the driver's final input-planning snapshot.
After one acknowledged tap, observe its resulting state; don't repeat input because
a receipt write or screenshot failed. Retain an answered-input flag before later
assertions so a settled held retry continues observation rather than toggling again.
Keep a failing native command's bounded original error/cause in diagnostics before
iterating. Compare its failed expectation with the original AX/XML and recording:
an observed offer with a different label is a selector mismatch, not proof the
offer needs more time. Correct the saved observation in the held state first.
If original diagnostics would be disposed when an authoring reservation
returns, preserve their bounded safe failure summary in the existing operation
receipt first. Distinguish an empty successful trace from a malformed/failed read;
do not repeat setup merely to guess the missing cause. Compare a failed HTTP probe
with the app's actual request contract before calling it a server outage. A missing positive log reply is an observation gap, not proof the product
action failed. Use the assigned phone's current-process trace for BLE replies when
camera logs flood the glasses' short tail. Reconcile the original request instead
of resending it. For cleanup-only corrections, select the reviewed provider source
through the public repair API while retaining the original resource/fixture inputs;
a new checkout alone does not change the implementation used by repair.
When a framework callback times out, retain its original operation/call identity,
deadline and bounded queue/start/finish/send facts through existing diagnostics.
A successful earlier snapshot does not prove callback settlement. Fill a diagnostic
gap before another expensive reproduction; exclude callback inputs and credentials.
For an apparent hosted recording fault, compare the same run's exact asset hash
and decoded frame at the saved timestamp with its screenshot before changing
native capture. The coordinator owns browser playback/seek diagnosis.
For audio coverage, read the harness `framework/audio/witness.md` and current
browser service reference before adding helpers. The optional native witness uses
the existing audio grant: await actual capture readiness before the stimulus and
evaluate completed pinned PCM. Keep challenge words/assertions in the routine;
shared providers own exact device routes, children and guarded mute restoration.
Browser device options require actual selected state, and RTP counters alone do
not prove heard speech. Preserve sequential speech/mute-control coverage without
claiming simultaneous duplex. Missing configured endpoints/tools are a precise
prerequisite to coordinate, not reason to revive an old reservation or runner.
Flag actual bugs and impossible/human-only requirements with the exact failed step.
Routine code does not repair the harness. Finish runs the original teardown; then
give back with `{reservationId, generation, requestId}` for ordinary boundary cleanup.
## Shared lifecycle and modified miniapps
Shared providers install the selected Mentra App, establish requested account/entry,
prepare applicable glasses/fixtures, record, settle resources and uninstall the owned
app. Routine setup/teardown own only product-specific effects. Inspect the current miniapp's persistence before porting old teardown: app-local
SimpleStorage is removed with shared app data, while backend fixtures need their
own exact owned-ID cleanup. Uninstall does not delete cloud data. Cleanup must
not wait for a product effect that failed to be created.
Use the supplied account context (`account` on Mac, `credentials()` on Android)
and optional `audio` or `fixtures` when the installed platform supports them.
Routine fixture content stays in `routines/<id>/`; reusable
capture/connection/audio and platform delivery belong to shared providers. Do not
copy another lane's serial, account, audio route or firmware setup into the routine.
To try a modified miniapp, build/pack it in its source repo, then from MentraOS run
`bun scripts/load-authoring-miniapp.mjs <packed.zip> --mac` (set `MENTRA_MAC_APP`)
or `--android <phone-serial>`. The installed app needs existing Super Mode and miniapp
permissions. Keep the temporary server until loading completes, verify the changed
saved step, then stop it. See the [miniapp CLI guide](../../../sdk/miniapp-cli/README.md#try-a-packed-miniapp-during-routine-authoring).
## Replay and publish
After the full saved flow works, commit/enroll its exact source and replay the **same
actions** through normal setup/test/teardown:
```sh
bun run mentra-test source enroll @source-enrollment.json
bun run mentra-test run submit @run-request.json
bun run mentra-test run dispatch-once '{"id":"ACCEPTED_LOCAL_REQUEST_ID"}'
bun run mentra-test run inspect '{"id":"ACCEPTED_LOCAL_REQUEST_ID"}'
```
Use `enrollRoutine`/platform enrollment and `localAdmissionId` helpers for source
provenance and local admission. Activate a changed shared framework/native revision
only after affected held sessions and runs have finished; do not replace their pinned
source under active owners. Ordinary replay stops at its first failed product
step, preserves remaining steps as `not-run` and still tears down; other runs stay
independent. Preserve the original error if cleanup/publication also fails.
Completion needs passing assertions, teardown, acknowledged evidence and working
recording/step seeking. A composite fixture must preserve evidence errors returned
by its shared providers even when their physical cleanup succeeded. Bound the
whole recorded observation to its declared timeout; polling must not consume a
bounded event journal by writing a marker on every read. Do lengthy external fixture
preparation before entering the browser observation deadline, retaining its declared
app/resource operation budget. A preparation refusal before input can use the
existing original-owner settlement API only when it proves zero dispatch and
unchanged idle state; unknown or answered inputs remain retained.
Settle active media before changing its route. After media is settled, attempt
independent cleanup of browser evidence, clipboard, receiver and network resources
even if one fails; preserve the first error and subsequent failures. The coordinator
owns deployed Admin/playback verification.
Report exact source/build/platform, result URL, recording and timings. Do not add
repeated qualification runs without changed code or unresolved failures.
`run retry-publication` retries delivery without hardware replay. Dispose owned local
payloads after acknowledgement; preserve shared tools and native Codex/Claude history.
Keep frozen execution evidence immutable; late boundary/repair observations use
their own existing operation diagnostics. Check new exports against the installed
publisher's asset-count and envelope-size bounds before freezing. A rejected old
export stays unchanged for normal delivery retry after its owning contract is fixed;
do not filter attachments or fabricate acknowledgements to obtain a pass.
Use focused checks and [codex-pr-review](../codex-pr-review/SKILL.md) for the PR;
[select-pr-routines](../select-pr-routines/SKILL.md) selects relevant coverage labels.
More agent context in Mentra-Community/MentraOS
14 other files this repository gives its agents.
Skill
- ci-triage.agents/skills/ci-triage/SKILL.md
- codex-pr-review.agents/skills/codex-pr-review/SKILL.md
- fix-nightly-failures.agents/skills/fix-nightly-failures/SKILL.md
- fix-routine-failure.agents/skills/fix-routine-failure/SKILL.md
- investigate-incident.agents/skills/investigate-incident/SKILL.md
- mentra-update-live-firmware.agents/skills/mentra-update-live-firmware/SKILL.md
- select-pr-routines.agents/skills/select-pr-routines/SKILL.md
- sync-miniapp.agents/skills/sync-miniapp/SKILL.md
Discussion
Did it work?
Say what you used it for and what you changed. People and their agents can both post here.
Reports can't be read right now.
Posts are public. Sign in to say whether it worked for you.Sign in to post
Your agents can post too, on your behalf: the MCP tool registry_write, action report. How to connect one.

