Screenshots, taps, UI events, app state, and logs are raw material. They become useful for evaluation only when a reviewer can tell what the agent was asked to do, which events belong to the same attempt, what changed, what evidence supports the outcome, and what the result can reasonably say.

As this series reaches its twentieth article, the useful step is to connect the earlier discussions: what a screenshot leaves out, what to record during a task, how raw runs become structured data, what an evaluation dataset needs, and how device profiles differ from run conditions. The thread through them is traceability: every judgment should lead back to a specific task attempt and its evidence.

A delayed order exposes the difference between action and outcome

Consider a controlled test in a shopping app. The task is to place exactly one order for a test item. The preconditions are explicit: a test account is signed in, the cart contains one item, and the order history contains no matching order. The completion criterion is also explicit: at a predeclared verification point after the attempt, exactly one matching order is present. The test plan should define that point in advance rather than choosing it after seeing the result.

Now suppose the agent taps Place order. The app shows a progress state but no confirmation. After waiting, the agent taps again. The app then shows an “Order placed” screen. A screenshot of that final screen suggests completion, but it does not answer whether the first request also created an order.

For this attempt, a useful trace distinguishes the events rather than compressing them into “order succeeded”:

  • Initial state: one item in the cart; no matching order in the test account’s history.
  • Action 1: the agent issues the first submit tap.
  • Observation 1: the app shows progress; no confirmation is visible yet.
  • Action 2: the agent issues a second submit tap after the wait.
  • Observation 2: the app displays an order confirmation.
  • Outcome check: if the app’s order history or an authorized test-state verifier is available, inspect it for matching orders.

If the check finds two matching orders, the attempt fails the stated “exactly one” criterion even though the final screen says the order was placed. If the test compares baseline and delayed-response conditions, keep the app build, account, cart, and device stable where practical, and vary the condition the question is about. That makes the comparison easier to interpret; it still does not establish causation if other relevant factors changed. The trace also shows a retry after an ambiguous response. If no reliable state check can distinguish one order from two, the outcome is unresolved; the UI confirmation alone does not justify upgrading it to success.

This is a hypothetical test design, not a report of a measured run. It illustrates why the task criterion, event order, observations, retry, and outcome evidence must remain connected. AndroidWorld provides a published example of task-specific initialization and success checks that inspect device state; that design supports reproducible task evaluation, but does not make one verifier suitable for every app or claim. AndroidWorld, ICLR 2025

A hypothetical order attempt: the agent retries after an ambiguous response, then checks order state against the task criterion.

Keep the evidence layers distinct

For the same attempt, several records may describe different things:

  • Agent action: what the agent issued, such as a tap or text entry.
  • UI observation: what the screen or accessibility hierarchy showed afterward.
  • Application or test state: what changed in the app’s underlying state, when that state can be inspected through an authorized test mechanism.
  • Task assessment: whether the verified state satisfies the predeclared criterion.

These layers can agree, but they are not interchangeable. A tap event does not prove the app received the action. A visible toast does not necessarily prove a durable side effect. An underlying state change does not prove that the agent followed the intended interaction policy. Record the evidence source for each claim and avoid presenting one layer as another.

Android’s UI Automator documentation describes interacting with user and system apps, waiting for UI conditions, and capturing screenshots. Those are useful UI-level testing capabilities; they are not, by themselves, a complete dataset format or an application’s authoritative state verifier. Android UI Automator documentation

Preserve enough structure to reconstruct an attempt

A record should let a reviewer reconstruct the attempt without guessing. At minimum, consider retaining:

  • A task definition, starting conditions, and an explicit completion criterion.
  • A stable attempt identifier and ordered events, with timestamps where available.
  • A distinction between agent actions and observations, including the source of each observation.
  • Relevant device, OS, app-build, account-state, and run-condition fields.
  • References to the evidence used by each outcome check, plus any missing or ambiguous evidence.
  • The resulting assessment and the rule used to assign it.

Not every study needs every field. Keep information that could affect execution, evidence, or interpretation for the stated question. Protect test-account data and avoid storing secrets or personal information in captured traces.

If network behavior is part of the question, record the condition that was configured separately from what was observed during the run. A configured delay is not a measurement of real connectivity. If the workflow cannot reliably configure or observe the relevant condition, document that limitation instead of implying it was controlled.

Make labels auditable, not merely convenient

Outcome labels are compact summaries. They become auditable when the dataset retains the rule and evidence behind them. For example, “success” could mean that exactly one matching order is visible in a specified verification source after the attempt. “Failure” could mean that the checked state violates that criterion. “Unresolved” could mean the available evidence cannot distinguish the two.

Do not silently place unresolved cases in the success or failure bucket. When reporting a rate, state how unresolved attempts were handled and show the denominator. Keep retries, repeated attempts, and distinct task conditions identifiable; several attempts on one starting state do not represent several independent app flows.

This is the difference between a score that can be interpreted and a number detached from its evidence. The earlier article on why task success or failure is not enough makes the same point from the perspective of outcome assessment.

Evaluate only the claim the sample can support

A dataset can be complete for a narrow question and still say little about other questions. If all attempts use one app build, one account state, and one device profile, the evaluation should say so. If several factors change together, report the observed association without claiming that one factor caused the difference. If a conclusion concerns broader transfer, the selection of apps, tasks, devices, and conditions must match that claim.

A useful evaluation report therefore pairs the score with the task definition, selection rule, attempt count, verification method, unresolved cases, and coverage limits. This echoes the distinction between coverage and comparability: more represented conditions may broaden coverage, while controlled comparisons help explain why outcomes differ. Neither property should be inferred from a single total count.

A practical readiness check

Before using a mobile interaction dataset for an evaluation, ask:

  • Can a reviewer state the claim and completion criterion without guessing?
  • Can each event be assigned to one attempt and put in order?
  • Are actions, UI observations, test-state evidence, and assessments distinguishable?
  • Can important outcome labels be traced to a rule and supporting evidence?
  • Are retries, missing observations, unresolved outcomes, and exclusions visible?
  • Are configured conditions separated from conditions observed during execution?
  • Does the conclusion stay within the documented task and environment coverage?

This is a practical review checklist, not a universal certification standard. Another question may require different evidence or make some fields unnecessary.

Infrastructure is one part of the chain

Real Android infrastructure can provide physical-device execution. It does not, by itself, define the evaluation question, preserve every required event, validate an outcome label, or establish that a dataset supports a particular claim. Those steps need to be specified and checked in the actual workflow.

Evaluation-ready data is not data that claims to answer every question. It is data whose task, trace, context, evidence, and limits are clear enough for another person to judge what the result means.