Mobile testing for AI coding agents

Let your agent see the app.

Your agent tests and debugs your app on an iOS Simulator or an Android emulator, and brings back the evidence.

9:41 Welcome back Continue n5 · button
autonom ui wait --settled --timeout-ms 5000 {"ok": true, "settled": true, …} autonom ui tree {"ok": true, …, "nodes": […, {"ref": "n5", "role": "button", "text": null, "desc": "Continue", …}, …], …} autonom ui tap --desc "Continue" {"ok": true, "x": 196, "y": 605, "ref": "n5", …}

How it works

Look. Act. Prove.

The same verbs work on an iOS Simulator and an Android emulator. A node has the same shape on both, so one skill drives either platform.

Screen docs

  1. Look.

    Read the screen as a compact accessibility tree, with a screenshot for proof.

    autonom ui tree
  2. Act.

    Find and tap by what an element says, not by where it sits.

    autonom ui tap --desc "Continue"
  3. Prove.

    Every step lands in the session journal with its result and your notes.

    autonom journal

Before and after

Exit 0 is not proof.

A tap that returns ok does not prove the screen changed. Your agent compares the tree or the screenshot before and after.

Welcome back Continue
Before
Today
After

To keep a pair comparable, pin the status bar on both platforms, the keyboard and locale on iOS, and turn animations off on Android. The status bar and keyboard pins fix state the app does not own.

The evidence ladder

Climb it. Skip nothing.

Your agent starts at the code and works up to a before and after replay, one rung at a time.

Evidence docs

  1. Code
  2. Unit and widget tests
  3. Integration on an explicit target
  4. Screenshot, UI tree and logs
  5. Profile, memory and network
  6. Before and after replay

One target

Never a silent guess.

With more than one device ready, your agent names the one it means. Ambiguity is an error listing the candidates.

Sessions docs

{"ok": false, "error_code": "ambiguous_target", "error": "2 ready targets; pass --target (or --platform)", "hint": "Candidates: android:emulator-5554, ios:…", "candidates": […]} On stderr with exit code 2, so your agent branches on the code.

What it can do

One toolkit for the whole check.

Skills your agent loads and one command line that answers in JSON. No MCP server needed.

The screen, as structure.

UI and accessibility trees, screenshots and recordings, logs and crash reports. Mobile Canvas mirrors and lightly controls the target in your browser. On iOS, frames are polled screenshots.

Screen and evidence docs

Commands

  • ui tree
  • ui find
  • screenshot
  • record
  • logs
  • crash

Input by meaning.

Find and tap by label, then swipe, type into the focused field and press keys. ui wait --settled holds until two consecutive trees match.

Screen docs

Traffic, captured and mocked.

Watch HTTP and HTTPS traffic through a local proxy and swap in the response you need, fullest on Android. Capture starts only on an explicit flag, and exports to HAR. A screenshot taken during a mock is flagged as mocked.

Network docs

Flows you can replay.

Save a check as a strict YAML flow and replay it. A mistake is caught at its exact line, and a failed step can suggest close matches from its captured screen. Imports and exports flows within the Maestro Core Profile.

Flows docs

Reports

  • HTML
  • JUnit
  • Allure
  • agent JSON
  • CSV

Memory for each app.

Your agent keeps notes on each app in ~/.autonom/apps/. A flow approved after consecutive clean replays, three by default, can become an App Skill in your project.

App memory docs

Performance, measured.

Measure memory, CPU, frames and traces, fullest on Android. Performance skills cover Flutter, native Kotlin and Compose.

Performance docs

Safety

Capture is always opt-in.

Starting the proxy, attaching a device and installing a CA each take their own flag, every time. Credentials in captured traffic are masked before it is written.

Safety docs

Consent is never cached.

Capture needs --i-understand-mitm on every run, plus a typed phrase on an interactive terminal. No config file, environment variable or earlier grant stands in.

Flow secrets are never saved.

Flow secrets never enter artifacts, and the journal keeps only scrubbed arguments.

The proxy stays local.

It binds 127.0.0.1 with no flag to widen it, and host network settings are never changed.

Sessions live in ~/.autonom/sessions/, outside any repository, so a capture cannot be committed by accident.

Install

Ready in four steps.

Works with Claude Code, Codex, Grok and any agent that loads skills.

  1. Get the code.

    Clone the repository and step into it.

    git clone https://github.com/aiatsuk/autonom cd autonom
  2. Install it.

    A checklist asks which device tools to add. Tick your agents too, and each one gets the skills.

    ./install.sh
  3. Ask what this machine can do.

    Green means installed, not proven. Anything missing comes with the exact fix.

    autonom doctor
  4. Take the tour.

    Start a new agent session so it loads the skills. Then take the guided walk on your device.

    autonom tour

Or, without cloning.

Claude Code and Codex can add the skills as a plugin. Start a new session afterwards so it loads.

  • Claude Code
    claude plugin marketplace add aiatsuk/autonom claude plugin install autonom@autonom
  • Codex
    codex plugin marketplace add aiatsuk/autonom codex plugin add autonom@autonom

Needs git, Python 3.11 and Node 20.11 or later. iOS Simulator targets need a Mac with Xcode. Free under the MIT license.

Full install guide

Run your first check.

Point your agent at a Simulator or an emulator and get back a screenshot and a journal of every step.