Skip to content

Test Result Schema

The test result schema defines the structure and format of test execution result files produced by the execute workflow (mauto guide execute, finalized with mauto result finalize). Every time a test scenario runs, it produces a result file conforming to this schema, capturing detailed execution metrics, assertion outcomes, and intelligent observations.

Overview

A test result is a comprehensive record of a single scenario execution, capturing: - Execution metadata — device, app version, environment, and timestamp - Step-by-step outcomes — status and duration of each test step - Assertion verdicts — whether each assertion passed or failed, with expected vs. actual values - Intelligent observations — regression detection, flakiness flags, and device/environment context - Variable captures — dynamic values extracted during execution

Results are stored as JSON files in mobile-automator/results/ with the naming pattern run_YYYYMMDD_HHMMSS.json.

Result files version differently from scenario files

Result files carry schema_version (no leading $), and its only valid value is "2.0". This is deliberate and not a mistake: scenario files use $schema_version (with the $), which accepts "2.0" or "2.1", while result files stay at schema_version: "2.0". A result produced from a $schema_version: "2.1" scenario still reports schema_version: "2.0".

When you run mauto result finalize, the result file is assembled and its typed observations (regression, flakiness, state context) are auto-harvested into cross-session memory — a best-effort step that never fails an otherwise-successful finalize.

Schema Structure

The result schema defines a single JSON object at the root level with required and optional fields:

Run Result
├─ Identifiers (run_id, scenario_id, schema_version)
├─ Metadata (device, app version, environment, timestamp)
├─ Status & Counters (status, assertions passed/failed/total)
├─ Duration (total execution time in seconds)
├─ Steps Array (step outcomes with retries, screenshots, errors)
├─ Assertions Array (assertion verdicts with expected/actual values)
├─ Observations Array (regression/flakiness/state context insights)
├─ Captured Variables (values extracted via capture_value steps)
└─ Summary (human-readable overall result)

Key Fields

Root Level Fields

Field Type Required Description
run_id string Yes Unique run identifier in format run_YYYYMMDD_HHMMSS. Example: run_20260227_145230
scenario_id string Yes ID of the scenario that was executed
schema_version string No Result-file schema version. Always "2.0" (the only valid value). Distinct from a scenario's $schema_version, which may be "2.0" or "2.1".
status string Yes Overall result status: "passed", "failed", or "error"
metadata object Yes Execution context including device, app version, environment, timestamp
total_assertions integer Yes Total count of assertions in the execution
passed_assertions integer Yes Count of assertions that passed
failed_assertions integer Yes Count of assertions that failed
duration_seconds number Yes Total execution time in seconds (minimum 0)
captured_variables object No Variables captured during test execution via capture_value steps. Keys are variable names from the scenario.
steps_executed array Yes Array of executed steps with outcomes and details
assertion_results array Yes Array of assertion verdicts with expected vs. actual values
observations array Yes Array of intelligent observations (regression, flakiness, state context)
summary string Yes Human-readable overall summary of the result

metadata Object

Contains contextual information about the execution environment:

Field Type Required Description
app_version string Yes Version of the app at time of execution (e.g., "1.2.3")
device_model string Yes Device used for execution (e.g., "Pixel 6", "iPhone 14 Pro")
api_level string Yes Android API level or iOS version (e.g., "34" for Android, "17.2" for iOS)
environment string Yes Target environment: "production", "staging", "development", etc.
timestamp string Yes ISO-8601 datetime of execution (e.g., "2026-02-27T14:52:30Z")

steps_executed[] Array Items

Each step execution object captures what happened during that test step:

Field Type Required Description
step_id string Yes Step identifier. Snake_case string (e.g., "tap_login").
status string Yes Step execution status: "passed", "failed", "skipped", or "error"
screenshot string | null No Path to captured screenshot for this step (relative to results directory)
error_message string | null No Error details if step failed
retried boolean No Whether this step was retried due to suspected flakiness (default: false)
retry_count integer No Number of retry attempts made for this step (0 = no retries)
step_duration_ms integer No Actual time taken to execute this step in milliseconds
condition_evaluated boolean | null No Result of the step's condition evaluation. Null if no condition was set.
sub_steps_executed array No Execution results for nested sub-steps if this step had sub_steps
observations string | null No DEPRECATED — use the run-level observations array, which is typed. Retained for result files written before typed observations; no writer populates it.

assertion_results[] Array Items

Each assertion result object captures whether an assertion passed or failed:

Field Type Required Description
assertion_id string Yes Assertion identifier. Snake_case string (e.g., "assert_login_success").
status string Yes Assertion status: "passed" or "failed"
expected string | null No What was expected (e.g., "Login button visible")
actual string | null No What was actually found (e.g., "Button not found after 5 seconds")
message string Yes Human-readable result description
reference_screenshot string | null No Path to reference baseline screenshot (for screenshot_match assertions)
actual_screenshot string | null No Path to screenshot captured during execution (for screenshot_match assertions)
similarity_score number | null No Similarity score for screenshot_match assertions (0.0 to 1.0, where 1.0 is identical)

observations[] Array Items

Intelligent observations detected during execution using observer traits:

Field Type Required Description
type string Yes Observer trait type: "regression", "flakiness", or "state_context"
step_id string | null No Related step if applicable, null if not step-specific.
message string Yes Observation detail (human-readable explanation)

Observation Types:

  • regression — Visual or functional change detected beyond what assertions caught. Examples: "Button color changed from blue to red", "Element position shifted 10px left"
  • flakiness — Timing issues, intermittent failures, or performance anomalies. Examples: "Step took 3.5s instead of typical 1.2s", "Loading indicator still visible on first attempt"
  • state_context — Device/environment context that may affect test reliability. Examples: "Network was slow (3G detected)", "Device has low memory (512MB free)"

Field Validation Rules

  • run_id — Must match pattern ^run_\d{8}_\d{6}$ (format: run_YYYYMMDD_HHMMSS)
  • schema_version — Must be "2.0"
  • status (root level) — One of: "passed", "failed", "error"
  • status (steps) — One of: "passed", "failed", "skipped", "error"
  • status (assertions) — One of: "passed", "failed"
  • total/passed/failed_assertions — Non-negative integers
  • duration_seconds — Non-negative number (can be 0 for very fast executions)
  • step_duration_ms — Non-negative integer (0 is valid for instant operations)
  • similarity_score — Range 0.0 to 1.0 (0 = completely different, 1.0 = identical)
  • timestamp — Must be valid ISO-8601 datetime format

Examples

Passed Test Result

A successful execution of a login scenario with all steps and assertions passing:

{
  "run_id": "run_20260227_145230",
  "scenario_id": "login_flow_001",
  "schema_version": "2.0",
  "status": "passed",
  "metadata": {
    "app_version": "1.2.3",
    "device_model": "Pixel 6",
    "api_level": "34",
    "environment": "staging",
    "timestamp": "2026-02-27T14:52:30Z"
  },
  "total_assertions": 3,
  "passed_assertions": 3,
  "failed_assertions": 0,
  "duration_seconds": 12.45,
  "captured_variables": {
    "auth_token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9...",
    "user_id": "user_12345"
  },
  "steps_executed": [
    {
      "step_id": "tap_login_button",
      "status": "passed",
      "screenshot": "steps/tap_login_button.png",
      "step_duration_ms": 245,
      "retry_count": 0
    },
    {
      "step_id": "wait_for_username_field",
      "status": "passed",
      "screenshot": "steps/wait_for_username_field.png",
      "step_duration_ms": 1200,
      "retry_count": 0
    },
    {
      "step_id": "type_credentials",
      "status": "passed",
      "screenshot": "steps/type_credentials.png",
      "step_duration_ms": 890,
      "retry_count": 0
    },
    {
      "step_id": "submit_login",
      "status": "passed",
      "screenshot": "steps/submit_login.png",
      "step_duration_ms": 5100,
      "retry_count": 0,
      "condition_evaluated": true
    }
  ],
  "assertion_results": [
    {
      "assertion_id": "assert_welcome_message",
      "status": "passed",
      "message": "Welcome message displayed correctly",
      "expected": "Welcome, John Doe",
      "actual": "Welcome, John Doe"
    },
    {
      "assertion_id": "assert_profile_icon",
      "status": "passed",
      "message": "User profile icon visible in header",
      "expected": "Profile icon present",
      "actual": "Profile icon present"
    },
    {
      "assertion_id": "assert_dashboard_loaded",
      "status": "passed",
      "message": "Dashboard screen fully loaded",
      "expected": "All dashboard cards visible",
      "actual": "All dashboard cards visible"
    }
  ],
  "observations": [
    {
      "type": "state_context",
      "message": "Network conditions: WiFi (good signal strength)"
    }
  ],
  "summary": "Test passed successfully. All 3 assertions passed. Execution completed in 12.45 seconds."
}

Failed Test Result with Flakiness Detection

A failed execution where a step was retried and flakiness was detected:

{
  "run_id": "run_20260227_145445",
  "scenario_id": "checkout_flow_001",
  "schema_version": "2.0",
  "status": "failed",
  "metadata": {
    "app_version": "1.2.3",
    "device_model": "iPhone 14 Pro",
    "api_level": "17.2",
    "environment": "production",
    "timestamp": "2026-02-27T14:54:45Z"
  },
  "total_assertions": 4,
  "passed_assertions": 2,
  "failed_assertions": 2,
  "duration_seconds": 18.75,
  "captured_variables": {},
  "steps_executed": [
    {
      "step_id": "navigate_to_checkout",
      "status": "passed",
      "screenshot": "steps/navigate_to_checkout.png",
      "step_duration_ms": 890,
      "retry_count": 0
    },
    {
      "step_id": "wait_for_payment_form",
      "status": "failed",
      "screenshot": "steps/wait_for_payment_form_failed.png",
      "step_duration_ms": 5250,
      "retry_count": 2,
      "retried": true,
      "error_message": "Element not found after 5 retries: PaymentFormContainer"
    },
    {
      "step_id": "enter_card_details",
      "status": "skipped",
      "error_message": "Skipped due to previous step failure"
    },
    {
      "step_id": "confirm_payment",
      "status": "skipped",
      "error_message": "Skipped due to previous step failure"
    }
  ],
  "assertion_results": [
    {
      "assertion_id": "assert_checkout_screen",
      "status": "passed",
      "message": "Checkout screen displayed",
      "expected": "Checkout header visible",
      "actual": "Checkout header visible"
    },
    {
      "assertion_id": "assert_price_summary",
      "status": "passed",
      "message": "Price summary matches cart total",
      "expected": "Total: $99.99",
      "actual": "Total: $99.99"
    },
    {
      "assertion_id": "assert_payment_form",
      "status": "failed",
      "message": "Payment form failed to load",
      "expected": "Card input field visible",
      "actual": "Card input field not found"
    },
    {
      "assertion_id": "assert_submit_button",
      "status": "failed",
      "message": "Submit button not visible due to form not loading",
      "expected": "Submit button enabled",
      "actual": "Submit button not found"
    }
  ],
  "observations": [
    {
      "type": "flakiness",
      "step_id": "wait_for_payment_form",
      "message": "Step failed initially but took longer on retries (first: 1.2s, final: 5.25s). Payment form may load slowly under production conditions."
    },
    {
      "type": "state_context",
      "message": "Network conditions: Cellular (3G, ~2Mbps). Payment endpoint may be slow."
    },
    {
      "type": "regression",
      "step_id": "wait_for_payment_form",
      "message": "Payment form took significantly longer to appear than expected. Previous baseline: ~1.5s, actual: 5.25s."
    }
  ],
  "summary": "Test failed. 2 of 4 assertions failed. Payment form failed to load, causing subsequent steps to be skipped. Detected flakiness and potential performance regression. Execution completed in 18.75 seconds."
}

Error During Execution

An execution that encountered an unexpected error:

{
  "run_id": "run_20260227_150000",
  "scenario_id": "search_flow",
  "schema_version": "2.0",
  "status": "error",
  "metadata": {
    "app_version": "1.2.2",
    "device_model": "Galaxy S24",
    "api_level": "34",
    "environment": "staging",
    "timestamp": "2026-02-27T15:00:00Z"
  },
  "total_assertions": 5,
  "passed_assertions": 2,
  "failed_assertions": 1,
  "duration_seconds": 8.32,
  "steps_executed": [
    {
      "step_id": "enter_search",
      "status": "passed",
      "screenshot": "steps/enter_search.png",
      "error_message": null
    },
    {
      "step_id": "submit_search",
      "status": "passed",
      "screenshot": "steps/submit_search.png"
    },
    {
      "step_id": "wait_for_results",
      "status": "error",
      "error_message": "mobile-mcp engine disconnected unexpectedly while executing the tap verb"
    }
  ],
  "assertion_results": [
    {
      "assertion_id": "assert_search_field",
      "status": "passed",
      "message": "Search field visible",
      "expected": "Search field present",
      "actual": "Search field present"
    },
    {
      "assertion_id": "assert_search_text",
      "status": "passed",
      "message": "Entered search text correctly",
      "expected": "Text: 'pizza'",
      "actual": "Text: 'pizza'"
    },
    {
      "assertion_id": "assert_results_loaded",
      "status": "failed",
      "message": "Search results not loaded due to execution error",
      "expected": "Results displayed",
      "actual": "Execution interrupted"
    }
  ],
  "observations": [
    {
      "type": "state_context",
      "message": "MCP server lost connection. Device may have disconnected or server crashed."
    }
  ],
  "summary": "Test encountered an error during execution. MCP server disconnected during wait_for_results step. Completed 2 of 5 assertions before failure. Execution completed in 8.32 seconds."
}

Result File Locations & Naming

Result files are stored in the mobile-automator/results/ directory with the following naming convention:

mobile-automator/results/run_YYYYMMDD_HHMMSS.json

Examples: - mobile-automator/results/run_20260227_145230.json - mobile-automator/results/run_20260227_150000.json

Related screenshot files are stored in: - mobile-automator/results/steps/<step_id>.png

Connecting Results to Scenarios

Each result references its source scenario via the scenario_id field. The corresponding scenario JSON is stored in:

mobile-automator/scenarios/<scenario_id>.json

Relationship: - scenario_id in result → Find scenario at mobile-automator/scenarios/<scenario_id>.json - The scenario defines: steps, assertions, variables, and preconditions - The result captures: how those steps executed and whether assertions passed

Working with Results Programmatically

Reading Result Files

import json

# Load a result
with open("mobile-automator/results/run_20260227_145230.json") as f:
    result = json.load(f)

# Access key information
print(f"Scenario: {result['scenario_id']}")
print(f"Status: {result['status']}")
print(f"Duration: {result['duration_seconds']}s")
print(f"Passed: {result['passed_assertions']}/{result['total_assertions']}")

# Iterate through failed assertions
for assertion in result['assertion_results']:
    if assertion['status'] == 'failed':
        print(f"\nAssertion {assertion['assertion_id']} failed")
        print(f"  Expected: {assertion['expected']}")
        print(f"  Actual: {assertion['actual']}")

# Check for flakiness
for obs in result['observations']:
    if obs['type'] == 'flakiness':
        print(f"\nFlakiness detected: {obs['message']}")

Filtering by Status

# Find all failed runs
import os
import json

failed_runs = []
for filename in os.listdir("mobile-automator/results/"):
    if filename.endswith(".json"):
        with open(f"mobile-automator/results/{filename}") as f:
            result = json.load(f)
            if result['status'] == 'failed':
                failed_runs.append(result)

print(f"Found {len(failed_runs)} failed runs")

← Back to Reference Index