Test Result Schema¶
The test result schema defines the structure and format of test execution result files produced by the execute workflow (mauto guide execute, finalized with mauto result finalize). Every time a test scenario runs, it produces a result file conforming to this schema, capturing detailed execution metrics, assertion outcomes, and intelligent observations.
Overview¶
A test result is a comprehensive record of a single scenario execution, capturing: - Execution metadata — device, app version, environment, and timestamp - Step-by-step outcomes — status and duration of each test step - Assertion verdicts — whether each assertion passed or failed, with expected vs. actual values - Intelligent observations — regression detection, flakiness flags, and device/environment context - Variable captures — dynamic values extracted during execution
Results are stored as JSON files in mobile-automator/results/ with the naming pattern run_YYYYMMDD_HHMMSS.json.
Result files version differently from scenario files
Result files carry schema_version (no leading $), and its only valid value is "2.0". This is deliberate and not a mistake: scenario files use $schema_version (with the $), which accepts "2.0" or "2.1", while result files stay at schema_version: "2.0". A result produced from a $schema_version: "2.1" scenario still reports schema_version: "2.0".
When you run mauto result finalize, the result file is assembled and its typed observations (regression, flakiness, state context) are auto-harvested into cross-session memory — a best-effort step that never fails an otherwise-successful finalize.
Schema Structure¶
The result schema defines a single JSON object at the root level with required and optional fields:
Run Result
├─ Identifiers (run_id, scenario_id, schema_version)
├─ Metadata (device, app version, environment, timestamp)
├─ Status & Counters (status, assertions passed/failed/total)
├─ Duration (total execution time in seconds)
├─ Steps Array (step outcomes with retries, screenshots, errors)
├─ Assertions Array (assertion verdicts with expected/actual values)
├─ Observations Array (regression/flakiness/state context insights)
├─ Captured Variables (values extracted via capture_value steps)
└─ Summary (human-readable overall result)
Key Fields¶
Root Level Fields¶
| Field | Type | Required | Description |
|---|---|---|---|
run_id |
string | Yes | Unique run identifier in format run_YYYYMMDD_HHMMSS. Example: run_20260227_145230 |
scenario_id |
string | Yes | ID of the scenario that was executed |
schema_version |
string | No | Result-file schema version. Always "2.0" (the only valid value). Distinct from a scenario's $schema_version, which may be "2.0" or "2.1". |
status |
string | Yes | Overall result status: "passed", "failed", or "error" |
metadata |
object | Yes | Execution context including device, app version, environment, timestamp |
total_assertions |
integer | Yes | Total count of assertions in the execution |
passed_assertions |
integer | Yes | Count of assertions that passed |
failed_assertions |
integer | Yes | Count of assertions that failed |
duration_seconds |
number | Yes | Total execution time in seconds (minimum 0) |
captured_variables |
object | No | Variables captured during test execution via capture_value steps. Keys are variable names from the scenario. |
steps_executed |
array | Yes | Array of executed steps with outcomes and details |
assertion_results |
array | Yes | Array of assertion verdicts with expected vs. actual values |
observations |
array | Yes | Array of intelligent observations (regression, flakiness, state context) |
summary |
string | Yes | Human-readable overall summary of the result |
metadata Object¶
Contains contextual information about the execution environment:
| Field | Type | Required | Description |
|---|---|---|---|
app_version |
string | Yes | Version of the app at time of execution (e.g., "1.2.3") |
device_model |
string | Yes | Device used for execution (e.g., "Pixel 6", "iPhone 14 Pro") |
api_level |
string | Yes | Android API level or iOS version (e.g., "34" for Android, "17.2" for iOS) |
environment |
string | Yes | Target environment: "production", "staging", "development", etc. |
timestamp |
string | Yes | ISO-8601 datetime of execution (e.g., "2026-02-27T14:52:30Z") |
steps_executed[] Array Items¶
Each step execution object captures what happened during that test step:
| Field | Type | Required | Description |
|---|---|---|---|
step_id |
string | Yes | Step identifier. Snake_case string (e.g., "tap_login"). |
status |
string | Yes | Step execution status: "passed", "failed", "skipped", or "error" |
screenshot |
string | null | No | Path to captured screenshot for this step (relative to results directory) |
error_message |
string | null | No | Error details if step failed |
retried |
boolean | No | Whether this step was retried due to suspected flakiness (default: false) |
retry_count |
integer | No | Number of retry attempts made for this step (0 = no retries) |
step_duration_ms |
integer | No | Actual time taken to execute this step in milliseconds |
condition_evaluated |
boolean | null | No | Result of the step's condition evaluation. Null if no condition was set. |
sub_steps_executed |
array | No | Execution results for nested sub-steps if this step had sub_steps |
observations |
string | null | No | DEPRECATED — use the run-level observations array, which is typed. Retained for result files written before typed observations; no writer populates it. |
assertion_results[] Array Items¶
Each assertion result object captures whether an assertion passed or failed:
| Field | Type | Required | Description |
|---|---|---|---|
assertion_id |
string | Yes | Assertion identifier. Snake_case string (e.g., "assert_login_success"). |
status |
string | Yes | Assertion status: "passed" or "failed" |
expected |
string | null | No | What was expected (e.g., "Login button visible") |
actual |
string | null | No | What was actually found (e.g., "Button not found after 5 seconds") |
message |
string | Yes | Human-readable result description |
reference_screenshot |
string | null | No | Path to reference baseline screenshot (for screenshot_match assertions) |
actual_screenshot |
string | null | No | Path to screenshot captured during execution (for screenshot_match assertions) |
similarity_score |
number | null | No | Similarity score for screenshot_match assertions (0.0 to 1.0, where 1.0 is identical) |
observations[] Array Items¶
Intelligent observations detected during execution using observer traits:
| Field | Type | Required | Description |
|---|---|---|---|
type |
string | Yes | Observer trait type: "regression", "flakiness", or "state_context" |
step_id |
string | null | No | Related step if applicable, null if not step-specific. |
message |
string | Yes | Observation detail (human-readable explanation) |
Observation Types:
- regression — Visual or functional change detected beyond what assertions caught. Examples: "Button color changed from blue to red", "Element position shifted 10px left"
- flakiness — Timing issues, intermittent failures, or performance anomalies. Examples: "Step took 3.5s instead of typical 1.2s", "Loading indicator still visible on first attempt"
- state_context — Device/environment context that may affect test reliability. Examples: "Network was slow (3G detected)", "Device has low memory (512MB free)"
Field Validation Rules¶
- run_id — Must match pattern
^run_\d{8}_\d{6}$(format: run_YYYYMMDD_HHMMSS) - schema_version — Must be "2.0"
- status (root level) — One of: "passed", "failed", "error"
- status (steps) — One of: "passed", "failed", "skipped", "error"
- status (assertions) — One of: "passed", "failed"
- total/passed/failed_assertions — Non-negative integers
- duration_seconds — Non-negative number (can be 0 for very fast executions)
- step_duration_ms — Non-negative integer (0 is valid for instant operations)
- similarity_score — Range 0.0 to 1.0 (0 = completely different, 1.0 = identical)
- timestamp — Must be valid ISO-8601 datetime format
Examples¶
Passed Test Result¶
A successful execution of a login scenario with all steps and assertions passing:
{
"run_id": "run_20260227_145230",
"scenario_id": "login_flow_001",
"schema_version": "2.0",
"status": "passed",
"metadata": {
"app_version": "1.2.3",
"device_model": "Pixel 6",
"api_level": "34",
"environment": "staging",
"timestamp": "2026-02-27T14:52:30Z"
},
"total_assertions": 3,
"passed_assertions": 3,
"failed_assertions": 0,
"duration_seconds": 12.45,
"captured_variables": {
"auth_token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
"user_id": "user_12345"
},
"steps_executed": [
{
"step_id": "tap_login_button",
"status": "passed",
"screenshot": "steps/tap_login_button.png",
"step_duration_ms": 245,
"retry_count": 0
},
{
"step_id": "wait_for_username_field",
"status": "passed",
"screenshot": "steps/wait_for_username_field.png",
"step_duration_ms": 1200,
"retry_count": 0
},
{
"step_id": "type_credentials",
"status": "passed",
"screenshot": "steps/type_credentials.png",
"step_duration_ms": 890,
"retry_count": 0
},
{
"step_id": "submit_login",
"status": "passed",
"screenshot": "steps/submit_login.png",
"step_duration_ms": 5100,
"retry_count": 0,
"condition_evaluated": true
}
],
"assertion_results": [
{
"assertion_id": "assert_welcome_message",
"status": "passed",
"message": "Welcome message displayed correctly",
"expected": "Welcome, John Doe",
"actual": "Welcome, John Doe"
},
{
"assertion_id": "assert_profile_icon",
"status": "passed",
"message": "User profile icon visible in header",
"expected": "Profile icon present",
"actual": "Profile icon present"
},
{
"assertion_id": "assert_dashboard_loaded",
"status": "passed",
"message": "Dashboard screen fully loaded",
"expected": "All dashboard cards visible",
"actual": "All dashboard cards visible"
}
],
"observations": [
{
"type": "state_context",
"message": "Network conditions: WiFi (good signal strength)"
}
],
"summary": "Test passed successfully. All 3 assertions passed. Execution completed in 12.45 seconds."
}
Failed Test Result with Flakiness Detection¶
A failed execution where a step was retried and flakiness was detected:
{
"run_id": "run_20260227_145445",
"scenario_id": "checkout_flow_001",
"schema_version": "2.0",
"status": "failed",
"metadata": {
"app_version": "1.2.3",
"device_model": "iPhone 14 Pro",
"api_level": "17.2",
"environment": "production",
"timestamp": "2026-02-27T14:54:45Z"
},
"total_assertions": 4,
"passed_assertions": 2,
"failed_assertions": 2,
"duration_seconds": 18.75,
"captured_variables": {},
"steps_executed": [
{
"step_id": "navigate_to_checkout",
"status": "passed",
"screenshot": "steps/navigate_to_checkout.png",
"step_duration_ms": 890,
"retry_count": 0
},
{
"step_id": "wait_for_payment_form",
"status": "failed",
"screenshot": "steps/wait_for_payment_form_failed.png",
"step_duration_ms": 5250,
"retry_count": 2,
"retried": true,
"error_message": "Element not found after 5 retries: PaymentFormContainer"
},
{
"step_id": "enter_card_details",
"status": "skipped",
"error_message": "Skipped due to previous step failure"
},
{
"step_id": "confirm_payment",
"status": "skipped",
"error_message": "Skipped due to previous step failure"
}
],
"assertion_results": [
{
"assertion_id": "assert_checkout_screen",
"status": "passed",
"message": "Checkout screen displayed",
"expected": "Checkout header visible",
"actual": "Checkout header visible"
},
{
"assertion_id": "assert_price_summary",
"status": "passed",
"message": "Price summary matches cart total",
"expected": "Total: $99.99",
"actual": "Total: $99.99"
},
{
"assertion_id": "assert_payment_form",
"status": "failed",
"message": "Payment form failed to load",
"expected": "Card input field visible",
"actual": "Card input field not found"
},
{
"assertion_id": "assert_submit_button",
"status": "failed",
"message": "Submit button not visible due to form not loading",
"expected": "Submit button enabled",
"actual": "Submit button not found"
}
],
"observations": [
{
"type": "flakiness",
"step_id": "wait_for_payment_form",
"message": "Step failed initially but took longer on retries (first: 1.2s, final: 5.25s). Payment form may load slowly under production conditions."
},
{
"type": "state_context",
"message": "Network conditions: Cellular (3G, ~2Mbps). Payment endpoint may be slow."
},
{
"type": "regression",
"step_id": "wait_for_payment_form",
"message": "Payment form took significantly longer to appear than expected. Previous baseline: ~1.5s, actual: 5.25s."
}
],
"summary": "Test failed. 2 of 4 assertions failed. Payment form failed to load, causing subsequent steps to be skipped. Detected flakiness and potential performance regression. Execution completed in 18.75 seconds."
}
Error During Execution¶
An execution that encountered an unexpected error:
{
"run_id": "run_20260227_150000",
"scenario_id": "search_flow",
"schema_version": "2.0",
"status": "error",
"metadata": {
"app_version": "1.2.2",
"device_model": "Galaxy S24",
"api_level": "34",
"environment": "staging",
"timestamp": "2026-02-27T15:00:00Z"
},
"total_assertions": 5,
"passed_assertions": 2,
"failed_assertions": 1,
"duration_seconds": 8.32,
"steps_executed": [
{
"step_id": "enter_search",
"status": "passed",
"screenshot": "steps/enter_search.png",
"error_message": null
},
{
"step_id": "submit_search",
"status": "passed",
"screenshot": "steps/submit_search.png"
},
{
"step_id": "wait_for_results",
"status": "error",
"error_message": "mobile-mcp engine disconnected unexpectedly while executing the tap verb"
}
],
"assertion_results": [
{
"assertion_id": "assert_search_field",
"status": "passed",
"message": "Search field visible",
"expected": "Search field present",
"actual": "Search field present"
},
{
"assertion_id": "assert_search_text",
"status": "passed",
"message": "Entered search text correctly",
"expected": "Text: 'pizza'",
"actual": "Text: 'pizza'"
},
{
"assertion_id": "assert_results_loaded",
"status": "failed",
"message": "Search results not loaded due to execution error",
"expected": "Results displayed",
"actual": "Execution interrupted"
}
],
"observations": [
{
"type": "state_context",
"message": "MCP server lost connection. Device may have disconnected or server crashed."
}
],
"summary": "Test encountered an error during execution. MCP server disconnected during wait_for_results step. Completed 2 of 5 assertions before failure. Execution completed in 8.32 seconds."
}
Result File Locations & Naming¶
Result files are stored in the mobile-automator/results/ directory with the following naming convention:
mobile-automator/results/run_YYYYMMDD_HHMMSS.json
Examples:
- mobile-automator/results/run_20260227_145230.json
- mobile-automator/results/run_20260227_150000.json
Related screenshot files are stored in:
- mobile-automator/results/steps/<step_id>.png
Connecting Results to Scenarios¶
Each result references its source scenario via the scenario_id field. The corresponding scenario JSON is stored in:
mobile-automator/scenarios/<scenario_id>.json
Relationship:
- scenario_id in result → Find scenario at mobile-automator/scenarios/<scenario_id>.json
- The scenario defines: steps, assertions, variables, and preconditions
- The result captures: how those steps executed and whether assertions passed
Working with Results Programmatically¶
Reading Result Files¶
import json
# Load a result
with open("mobile-automator/results/run_20260227_145230.json") as f:
result = json.load(f)
# Access key information
print(f"Scenario: {result['scenario_id']}")
print(f"Status: {result['status']}")
print(f"Duration: {result['duration_seconds']}s")
print(f"Passed: {result['passed_assertions']}/{result['total_assertions']}")
# Iterate through failed assertions
for assertion in result['assertion_results']:
if assertion['status'] == 'failed':
print(f"\nAssertion {assertion['assertion_id']} failed")
print(f" Expected: {assertion['expected']}")
print(f" Actual: {assertion['actual']}")
# Check for flakiness
for obs in result['observations']:
if obs['type'] == 'flakiness':
print(f"\nFlakiness detected: {obs['message']}")
Filtering by Status¶
# Find all failed runs
import os
import json
failed_runs = []
for filename in os.listdir("mobile-automator/results/"):
if filename.endswith(".json"):
with open(f"mobile-automator/results/{filename}") as f:
result = json.load(f)
if result['status'] == 'failed':
failed_runs.append(result)
print(f"Found {len(failed_runs)} failed runs")
Related References¶
- Test Scenario Schema — The schema for defining test scenarios
- Assertion Types — Detailed documentation of all assertion types
- Execute Command Guide — How to run tests and generate results
- MCP Tools Reference — Mobile automation primitives available to test steps