That does not mean collecting every possible signal. A useful evaluation dataset is built around a stated question. Its records should let another person inspect how the evidence supports the reported conclusion—and where the record stops.

Begin with the evaluation question

Before choosing fields, state what the evaluation is intended to compare or explain. Is it measuring task completion, robustness to a specific condition, interaction efficiency, or another defined property? The question determines which observations and context are relevant.

Define the task, completion criterion, and unit being counted. If one unit is an attempt, say whether retries are separate attempts or part of one episode. Record how many attempts were run and which attempts were included in the reported result. If a task has variants, describe how they differ.

Without those definitions, the same row can be interpreted in incompatible ways. A “success” label may refer to reaching a screen, satisfying a user goal, or passing a particular checker. The dataset should say which meaning applies.

Preserve the attempt, not only its final label

For each attempt, retain the goal and the initial conditions that matter to that task. A concise attempt record can identify the task variant, starting app or screen state, relevant device and software context, the agent or controller version, and the evaluation run.

Then preserve the ordered observations and actions needed to interpret the attempt. An observation might be a screenshot or another captured state. An action might be a tap, text entry, navigation command, or other issued input. Link each record to its artifacts and to the attempt it belongs to. State what timestamps mean and which clock produced them; do not assume that capture time, command time, and export time are interchangeable.

The goal is not to claim that every state change was caused by the preceding action. Background activity, network responses, or system behavior may also affect what appears. Keep the record precise about what was issued and what was observed.

The AndroidWorld benchmark offers a useful example of why task context matters: its tasks include dedicated initialization, success-checking, and tear-down logic to support reproducible evaluation. That is a benchmark design choice, not a requirement that every dataset use the same harness. It illustrates the value of making task setup and checks inspectable.

Connect outcomes to criteria and evidence

Store the completion criterion alongside the outcome assessment. Identify the checker or review method and the observation or state evidence used to support the conclusion. If a check could not be completed, preserve that status and its reason instead of silently mapping it to success or failure.

Consider a note task that asks an agent to save a specific title and body. Returning to the notes list is an observation, but may not establish that both fields persisted. Reopening the note and checking the requested fields provides evidence tied more directly to that criterion. The dataset should make clear which observation supported the assessment.

A note completion criterion linked to save and reopen actions, observed states, and an evidence-based outcome assessment.

Record conditions that affect interpretation

Keep context that can help explain differences between attempts: device model or configuration, Android and app versions, relevant permissions or settings, network conditions when the task depends on them, and agent, controller, or checker versions.

The useful context depends on the question. A network measurement may be unnecessary for an offline task; it may be essential when a task depends on a service response. State what was controlled, what was merely observed, and what was not collected. A configured condition is not the same as evidence that the condition behaved as expected.

Version the task definitions, schema, checkers, and transformations that affect reported results. If a checker changes, a score from the new checker may not be directly comparable with an earlier score. Keep the version history visible.

Make missing data and selection visible

Distinguish among “not collected,” “collection failed,” “not applicable,” and “not retained.” These states have different implications. A missing screenshot can limit review of an action; it does not automatically establish that the agent failed.

Document how attempts were selected, excluded, retried, or grouped. Keep related retries or steps together when separating them could leak nearly identical context across evaluation splits. For action prediction, do not place future observations or final assessments among the inputs available at the earlier step. They may remain in the source record while being excluded from that specific data view.

Report the denominator and how unresolved or invalid checks affect it. These are reporting recommendations, not universal category names. The important point is to make the counting rule explicit enough that a reader can reconstruct what the headline number includes.

Document the dataset as a release

A dataset release should identify its version, intended use, included records or selection rule, collection and processing methods, schema and checker versions, and known limitations. Include applicable access, privacy, and usage constraints so consumers can understand what may be inspected, shared, or reused.

Documentation is useful when it helps a reader make a decision, not when it becomes a checklist detached from the data. Datasheets for Datasets and Data Cards are examples of documentation approaches; neither implies that one fixed template fits every evaluation dataset.

A practical review asks: Can someone trace a reported outcome back to its criterion and evidence? Can they see the conditions and versions under which it was produced? Can they tell what is missing and how attempts entered the dataset? If those answers are available, the result is more inspectable. It still does not guarantee that the dataset is representative or that its conclusion generalizes beyond the evaluated conditions.

Real Android infrastructure can provide an execution environment for collecting device-based evidence. A complete evaluation dataset also depends on collection logic, checks, transformations, selection, and documentation. Confirm which components exist in a proposed setup; device access alone should not be read as proof that a complete dataset service is available.

The next article will examine the balance between repeatability in controlled conditions and the variation encountered in real mobile execution.