Markdown · numbered steps · Expect marks an assertion
Test
Title
Target
Version
File
State
History
Not loaded.
Devices
State
Device
Platform
OS
Used by
MCP argument
running
—
Targets
Read-only — edited on the runner host and reviewed in git
Target
Platform
Device
App
Install
build
Webhooks
Status
Endpoint
Events
Failures
⇄
No webhook subscriptions.
Sessions
Sessions end on their own
After 7 days unused, and 30 days after sign-in whichever comes first. Using the
console slides the idle clock; the 30-day one is fixed. End anything you do not
recognise.
Session
Browser
IP
Last seen
Expires
Started
this browser
Users
Name
Email
Role
Sessions
Last sign-in
disabled
Users are created from the console host with
python -m app.cli create-user.
Audit log
When
Who
Action
Subject
Detail
The current script is sent as the starting point, so describe the change rather
than the whole test.
Opens the app and reads each screen, so the script quotes labels that are really
there. Takes the device, so it queues behind any run using it.
Works from the wording alone. Right for changing an existing script, because the
labels in it were already read off a real screen.
Not possible on this target
This draft does not parse yet — you can still apply it and fix it in the
editor.
Nothing is saved yet. Apply puts this in the editor for you to review.
Start a run
Overriding the target is the most common cause of a confusing
failure — a test written for one skin often fails on another for real reasons.
Leave this off to test whatever is already on the device. Targets whose install
mode is installed never put the app there themselves, so
on a device that has never had it, the run needs this.
Installs in seconds — the binary is already on the host.
No builds on the host for this platform yet. Choose the other option to
make one — after that it will be listed here for any run.
Commit
Recipe
Artifact
Commit
Workflow
That workflow run uploaded no artifacts
The workflow needs an actions/upload-artifact step
publishing the built .app or .apk before it can be installed here.
Ready to install
Effort is the cheaper lever — it scales thinking depth without splitting the
prompt cache, which is per-model.
Off by default so a double-click cannot spend two runs.