A useful task record preserves the context, sequence, and evidence needed to understand an attempt after it ends.
A mobile AI agent taps Save. The interface returns to a list. A few seconds later, the run stops.
What should the record contain so someone else can judge what happened? A screenshot may show the list. An action log may show the tap. A success label may state a conclusion. Without the connections between them, the reviewer still has to guess.
In the previous article, we described a mobile task trajectory as a connected account of one attempt. The next step is deciding what information should survive that attempt. The aim is to preserve enough evidence to answer specific questions, with the limits of that evidence visible.
Record the goal before interpreting the result
Start with the requested task, its relevant inputs, and the criterion used to assess completion. Give the attempt an identifier so its observations and actions can be distinguished from other runs.
Consider the same illustrative task as before: create a note titled “Site visit,” add “Bring the access badge” as its body, and check the saved contents. The criterion is about both fields. A title appearing in a list answers only part of the question.
Record the starting conditions that matter to that criterion: the app was already open, the attempt began on the notes list, and whether a matching note already existed was checked—or was not checked. An unchecked starting condition should not silently become an assumption.
This is a conceptual recording example, not an ARMARRAY product demonstration or a prescribed data format.
Preserve observations and their sources
Keep the observations used to interpret or select actions, along with the observations used to assess the result. These may include screenshots and, where available and appropriate, UI element descriptions.
Associate each observation with the attempt and the relevant step. Record its source and capture time. If a screenshot was resized or cropped before reaching the agent, retain enough information to identify which version it received and how its coordinates relate to the original view.
A screenshot and a UI description can complement each other, but neither should be presented as complete access to the app’s internal state. If an observation is unavailable, record that gap rather than substituting a nearby image without explanation.
Record issued actions separately from observed effects
For each action, preserve the operation and the information needed to interpret it: the target, coordinates or element reference where applicable, and relevant input parameters. Keep its relationship to the preceding observation explicit.
Distinguish the selected action, the command actually issued, and any execution response the controller reports. A controller accepting a command is not the same as the application achieving the requested change.
In the note example, “tap Save” belongs to the action record. “Returned to the list” belongs to a later observation. “Requested title and body were visible after reopening” belongs to the evidence used for assessment. Collapsing them into a single success event removes useful distinctions.

Make time mean something specific
A timestamp is most useful when its meaning is clear. Was it assigned when an image was captured, when it reached the controller, when an action was issued, or when a result was checked?
For steps where timing matters, preserve those distinct events rather than treating one timestamp as a substitute for all of them. An observation identifier can connect an action to its context even when events arrive out of order.
Suppose an image is captured before Save, but delivered after the tap. Sorting by delivery time could make that image look like evidence of the action’s effect. Capture time and an explicit association reveal the difference.
When records come from different clocks, document the clock basis and any known synchronization limits. Do not infer precise cross-system latency from timestamps whose relationship has not been established. Within an attempt, an explicit sequence and elapsed-time reference can help preserve order without claiming more timing accuracy than was measured.

Capture environment context that can change interpretation
Record the app and environment details needed to interpret or compare attempts. Depending on the task, these may include app version, Android version, emulator or physical-device configuration, display dimensions and orientation, language, permissions, and relevant starting app state.
Separate known configuration from observed runtime conditions. A device being configured for Wi-Fi does not establish that a remote request succeeded. A visible connectivity indicator does not measure latency or prove service availability.
For a task whose result depends on a network service, record the network observations or errors actually available, with their timing and source. If the task is entirely local, detailed network measurements may add little value. The recording scope should follow the question being investigated.
Some context can be recorded once per attempt, with changes recorded when they occur. Repeating an unchanged environment description at every step adds volume without necessarily adding evidence.
Keep failures, retries, and missing data visible
Preserve errors and recovery actions that affect the execution path. Include waits, timeouts, retries, and resets when they are relevant to understanding the attempt.
A retry after a save error may remain part of the same attempt under the collection rules. A reset that begins a new attempt needs a separate boundary. Keep the relationship between attempts explicit rather than joining their best-looking fragments.
Also distinguish a failed task from a failed recording. If screenshot capture stops, the app may continue running. That gap alone does not establish whether the task succeeded or failed. Recording limitations belong alongside task evidence so later users can assess what the trace supports.
Attach the result to its evidence
At the end, preserve the completion criterion, assessment method, result, and references to the evidence used. If an automated checker or human reviewer supplied the assessment, identify that source and the relevant checker version or review rule where applicable.
For the note task, reopening the saved note and observing both requested fields supports the defined visible-content check. It does not automatically establish synchronization to another device or future availability. Those claims would require different checks.
If the final observation is missing, an unresolved assessment may be more accurate than forcing a success or failure label. The record should also state why execution ended—for example, a stop decision, time limit, or interruption. An agent deciding it is done and a task being independently verified are different facts.
Research environments make this distinction concrete. AndroidWorld combines task initialization, agent execution, and task assessment in its workflow. Its design is a useful reference for separating execution evidence from the criterion used to judge a task. AndroidWorld project
Define a useful recording scope
A practical review can start with six questions:
- Which goal and attempt does this record describe?
- What information was available before each important action?
- What was issued, and what was subsequently observed?
- What do the event times and environment details actually establish?
- Where did retries, resets, or recording gaps occur?
- What evidence supports the final assessment?
Collect additional information when it helps answer a relevant question. Avoid treating every available log, screen, or identifier as necessary by default. Use test data where possible, and make any redaction or transformation explicit so the record remains interpretable.
These are recording design questions, not claims that a particular device environment automatically supplies every field. The same questions apply to emulator and physical-device attempts; available signals and collection methods may differ.
The next article will examine how to collect these connected observations and actions on real Android devices. Once the recording requirements are clear, the collection method can be evaluated against them.
