
Natural-language Android automation agent that drives real devices via ADB, captures Logcat and screenshots, and exposes an MCP server for AI IDEs and CI test pipelines.
Let AI assistants and test suites use real phones like a human.
English • 中文文档 • Workflow Showcase • Quick Start • MCP for IDEs • Benchmarks • Discord Community
Live Demo: Setup driving routes and calculate total durations in Google Maps, then open YouTube to play a Coldplay song.
Antigravity uses ARTEMIS through MCP to turn a test request into a plan, device execution, and a diagnostic report:
Ensure an Android device (with USB Debugging enabled) or emulator is connected. The one-click startup script will automatically:
uv) dependencies.rules.md) into your AI IDEs (Antigravity, Cursor, Claude Code, Codex, Windsurf, VS Code, Cline/Roo, OpenClaw).# 1. Clone repo & navigate to directory
git clone https://github.com/google/artemis.git && cd artemis
# 2. One-click launch
./start.sh
# 1. Clone repo & navigate to directory
git clone https://github.com/google/artemis.git
cd artemis
# 2. One-click launch
.\start.bat
PowerShell does not search the current directory for executable scripts by default, so use
.\start.batwithout a trailing\. In Command Prompt (CMD), usestart.batinstead.
Tip: Opens
http://localhost:8000in your default browser with a device connection wizard, live screen mirroring, prompt sandbox, and execution replays. You can also run directly from CLI:uv run artemis run "Open Settings, find Battery and tell me current level" --profile flash.
ARTEMIS includes a native Model Context Protocol (MCP) server. Connect your real phone directly into AI IDEs:
Running ./start.sh (macOS/Linux) or .\start.bat (Windows PowerShell) will prompt you to configure global MCP and testing rules for detected IDEs (or you can install/update anytime later manually using the commands below):
# Auto-install MCP server & global rules for Antigravity / Jetski:
uv run artemis mcp --install antigravity
# Or install for all supported AI IDEs (including Codex):
uv run artemis mcp --install all
Tip: You can also configure MCP interactively during first-time setup via
uv run artemis init. Pro Tip: If you want to use theartemiscommand globally withoutuv runin any directory, runuv tool install -e .once in the project root.
Install the zero-runtime-dependency client on the development machine. ADB, agents, models, and image processing remain on the device host:
uv add "artemis-client @ git+https://github.com/google/artemis.git#subdirectory=packages/artemis-client"
import asyncio
from artemis_client import ArtemisClient
async def main():
client = ArtemisClient(
"http://artemis-host:8000",
device_serial="emulator-5554", # optional: target specific device serial
default_profile="flash", # "flash" (fast reactive) or "pro" (deep reasoning)
)
result = await client.run(
"Open System Settings, go to 'Battery', verify battery percentage is displayed, and check for any crash dialogs.",
)
assert result.succeeded, f"Test failed: {result.error or result.status}"
print(f"✅ Test Passed! Device: {result.device_serial} | Trace ID: {result.trace_id}")
if __name__ == "__main__":
asyncio.run(main())
Console Overview: ① View Switcher (Home / Workspace) · ② Model & Replay (Flash/Pro status & video replay) · ③ Live Agent Stream (Action perception, target coordinates & structured results) · ④ Prompt Dock (Natural language dispatch) · ⑤ Task Queue & Dashboard (Lifecycle & history)
uv run artemis ui): Real-time screen projection and interactive panel, supporting natural language test dispatch, live reasoning telemetry, action trajectories, and execution replay; manage server lifecycle anytime from any terminal using uv run artemis restart, uv run artemis stop, and uv run artemis status;uv run artemis run): Direct terminal execution for automated test cases, exploratory stability inspection, or AndroidWorld benchmarks with high-fidelity structured terminal output;The first task on a device installs the Artemis Accessibility Helper, a small
accessibility service that reads the screen layout without taking the
UiAutomation connection. Tools using UiAutomation can suppress the helper unless
they enable FLAG_DONT_SUPPRESS_ACCESSIBILITY_SERVICES. You will see
a collapsed "Artemis test helper is running" notification and a new entry under
Settings > Accessibility; both are that helper. It listens only on the phone
itself and sends nothing elsewhere.
uv run artemis helper installuv run artemis helper status / uv run artemis doctoruv run artemis helper uninstallARTEMIS_HIERARCHY_BACKEND=uiautomator in .envARTEMIS_HELPER_AUTO_INSTALL=false in .envIf the helper ever fails mid-task, ARTEMIS falls back to UIAutomator2 and says
so in the task timeline, in mobile_manage_task status, and in the final report.
Artemis achieved a 99%+ completion rate on AndroidWorld, Google Research's benchmark spanning 20+ apps and 100+ multi-step tasks.
ARTEMIS supports two execution profiles tailored for different automation requirements:
--profile flash): Fast and token-efficient reactive loop (~3–5s per step): one model observes the live screen, thinks, and acts, with no graph orchestration. Ideal for routine, deterministic UI tasks. The loop is unbounded by default (agent.flash.max_turns, 0 = unlimited) because history is compressed rather than capped: Flash shares the Pro session transcript ledger (session-relative T+mm:ss clock, screenshots folded into visual summaries, older steps chunked into eras and recallable on demand via search_history / replay_steps) and can query the session recording through video_analyzer. Transient UI (auto-fading control bars, toasts) is handled by chaining taps into one click_sequence. Limitations: No task plan or notes, no pre-execution safety net, no checkpoint verification or final report, and no ADB shell.--profile pro): A planning and verification workflow (~15–40s per step), built as a multi-agent graph. A Planner maintains a living Markdown task plan with milestones and verify / assert check items; the executes it with the full toolset (Explorer grounding whose / / tier is a user setting per profile — / in or — never chosen by the agent; notes, history recall, video analysis, ADB diagnostics). Every single action passes a pre-execution (XML-first, pixel fallback), while multi-action fire back to back to beat turn latency on transient UI. A blocked or failed action opens an that stays in the Operator's context until a later action succeeds, so recovery is handled by the Operator itself with no separate repair agent. A read-only verifies plan checkpoints and runs an exit final review against the original goal (: / (default) / / ), and plan milestone edits get an advisory review. Handles 100+ step long-horizon workflows, monitoring, and an optional written report.Contributions are warmly welcomed!
This project is licensed under the Apache License 2.0.
This project includes source code developed by Minitap, Inc..
|
1. Prompt Input (Task Dispatch) Describe your test scenario and target metrics in Antigravity
|
2. Test Plan Generation Formulates a step-by-step test plan & architecture for review
|
|
3. Autonomous Test Execution Drives real device, navigates UI, and profiles performance
|
4. Final Report Delivers structured audit findings, metric tables, and raw datasets
|
If you prefer to configure manually, run uv run artemis mcp --generate-config <client> (for example, codex or antigravity) to output the appropriate TOML or JSON snippet. Replace /path/to/artemis with your actual repo path and point command to your .venv Python executable:
~/.codex/config.toml):[mcp_servers.artemis]
command = "/path/to/artemis/.venv/bin/python"
args = ["-m", "mcp_server"]
cwd = "/path/to/artemis"
[mcp_servers.artemis.env]
PYTHONUNBUFFERED = "1"
PYTHONPATH = "/path/to/artemis"
~/.gemini/jetski/mcp_config.json):{
"mcpServers": {
"artemis": {
"command": "/path/to/artemis/.venv/bin/python",
"args": ["-m", "mcp_server"],
"cwd": "/path/to/artemis",
"env": {
"PYTHONUNBUFFERED": "1"
},
"tools": {
"mobile_run_task": { "eager": true },
"mobile_manage_task": { "eager": true },
"mobile_get_device_state": { "eager": true },
"mobile_inspect_trace": { "eager": true },
"mobile_diagnose": { "eager": true }
}
}
}
}
claude_desktop_config.json):{
"mcpServers": {
"artemis": {
"command": "/path/to/artemis/.venv/bin/python",
"args": ["-m", "mcp_server"],
"cwd": "/path/to/artemis"
}
}
}
To ensure your AI coding assistant acts with the rigor of a senior mobile test engineer and never hallucinates UI interactions, we provide a dedicated testing mindset rules file at mcp_server/rules.md (covering Active Exploration before coding, Flash vs. Pro routing strategy, Latency & Timing compensation, and the "Dynamic-First, Coordinate-Fallback" locator pattern).
You can mount or copy mcp_server/rules.md into your AI IDE's rule configuration:
rules.md to your Workspace Rules, Global Rules settings, or agent instructions.artemis mcp --install claude to install the rules to ~/.claude/rules/artemis.md (install to exactly one location — Claude Code loads both ~/.claude/CLAUDE.md and ~/.claude/rules/*.md, so duplicating the rules wastes context)..cursorrules or create a rule file at .cursor/rules/artemis.mdc.~/.codex/AGENTS.md (or the active AGENTS.override.md).For more details on the testing mindset and MCP architecture, see the MCP Server README.
In Codex, Antigravity, or Claude Code, simply prompt:
"Build the latest changes into an APK, install it on the connected device, open the login screen with a test account, verify if there are any unexpected popups after login, and return screenshots of the final page."
flashproultrapro.explorer.modeflash.explorer_modeconfig/artemis.jsonc--explorer-pro-mode--verification-levelofffinalcheckpointsstrict[Loop:continuous]