A folder contains screenshots, controller messages, and a final result file. Every file came from the same Android task. That does not yet tell a data consumer which image informed an action, whether a later image was captured before or after it, or how the final result was assessed.
The previous article described how to collect a reviewable attempt on a real Android device. The next step is transformation: give the collected evidence an explicit structure without silently changing its meaning.
A practical path is to preserve source artifacts, define record relationships, normalize documented conventions, validate the result, and build a dataset for a stated use. The examples below are design illustrations, not an ARMARRAY export specification or a measured device run.
Start with the intended use
Decide what the resulting data should support. Reviewing failures, learning from demonstrations, and evaluating task completion may use overlapping evidence, but they do not necessarily require the same examples or selection rules.
For an action-prediction example, identify the information available before the action and the action to be predicted. For an outcome review, retain the final evidence and the criterion used to judge it. Do not treat a single flattened row as automatically suitable for both.
Keep a sufficiently connected source record so these different views can be derived deliberately. If an observation was never collected, a new schema cannot recover it.
Preserve sources before transforming them
Keep the source artifacts needed to trace each derived record: images, issued-action records, available execution responses, and outcome evidence. Assign stable references and retain the original event identifiers where they exist.
Document what each timestamp means and which clock produced it. A filename containing a time is not enough if nobody knows whether it represents capture, receipt, or export.
For transformed artifacts, preserve the relationship to their source. A resized screenshot should have its own reference and the transformation details needed to interpret it. A checksum can help detect changed bytes; it does not establish that the image is a truthful or complete account of execution.
Keep retained sources under the applicable access and retention rules. A downstream dataset need not include every raw artifact, but its selected records should have a documented provenance path to the evidence that may be retained.
Define relationships before choosing a file format
The schema should describe entities and their relationships, not merely rename log fields. A useful starting model separates the attempt, observations, actions, artifacts, and assessments. Make record identifiers explain their type: for example, use OBS for an observation, ACT for an issued action, and OUT for an outcome assessment. A shared sequence number can show where an event belongs in one attempt, while the type prefix shows what kind of record it is. Spell out the prefixes the first time they appear; do not assume readers will infer them from a diagram.
The attempt identifies the goal, starting context, and execution boundary. An observation identifies what was captured and its artifact references. An action identifies what was issued, its parameters, and the observation used to select it. An assessment identifies a criterion, a conclusion, and its supporting evidence.
A subsequent observation can be linked to the relevant step without being described as proof that the action caused every visible change. Keep that distinction in the relationship names and documentation.
An existing episodic-data design can provide a useful comparison. RLDS organizes data into episodes and steps, with conventions for observations, actions, and episode boundaries. Its repository is archived, so it is a design reference here rather than a recommendation to adopt an actively maintained tool. RLDS repository
Your collection workflow may need a different representation. The important requirement is that readers can identify what each relationship means, including how the last observation is represented when there is no next action.
Map one concrete step without filling the gaps
Return to the illustrative note task: create “Site visit,” enter “Bring the access badge,” and inspect both saved fields. The sequence below names the events in order so each identifier can be matched to a point in the task. OBS means observation, ACT means issued action, and OUT means outcome assessment. The number after each prefix is the event’s position in this example attempt; it is not a product event ID.
OBS-01: the notes list before the task. 2.ACT-02: tap “New note,” usingOBS-01as the input observation. 3.OBS-03: the new note editor appears. 4.ACT-04: enter the title and body, usingOBS-03as context. 5.OBS-05: the editor visibly contains both requested fields. 6.ACT-06: tap “Save,” usingOBS-05as the input observation. 7.OBS-07: the app returns to the notes list. This shows a screen transition; by itself, it does not verify both saved fields. 8.ACT-08: reopen the target note, usingOBS-07to locate it. 9.OBS-09: the reopened note displays the title and body. 10.OUT-10: assess the criterion “both requested fields are visible after reopening,” citingOBS-09as supporting evidence.
The identifier prefixes separate record types: OBS-09 is an observation, while OUT-10 is the assessment made from evidence. They do not share an ambiguous single-letter O. The sequence numbers correspond to this illustrative attempt; they do not imply that every collection system must use this naming convention.

The diagram should preserve the same order and identifiers as the list. The number after each type prefix is the event’s position in this example attempt, so the sequence can be followed across observations, actions, and assessment. It is a map of the example, not a substitute for explaining the steps. If OBS-09 is missing, retain the gap and limit or leave OUT-10 unresolved. Do not create a replacement observation or infer success only from the Save command.
Normalize conventions without hiding changes
Choose and document common representations for action types, coordinate systems, time units, and missing values. Preserve the original value or a traceable source reference when conversion changes how a field is expressed.
For coordinates, record which image or device view the values refer to. A tap position in a resized image cannot be interpreted correctly without the relevant dimensions and mapping. Do not silently reinterpret it as a coordinate in the original view.
Treat “not collected,” “collection failed,” and “not applicable” as different states where that distinction affects use. A blank field or null value alone cannot explain why information is absent.
Version the schema and the transformation process separately. A field definition can remain unchanged while a converter bug is fixed. Conversely, the same converter version is not enough to explain a changed interpretation of an outcome label.
Validate structure, connections, and meaning
Validation needs more than a file opening successfully.
A schema check can test required properties, types, and permitted values. JSON Schema provides mechanisms for describing object properties and required fields. Those checks are useful for consistent records, but validating the shape of a record does not establish what happened on the device. JSON Schema object reference
Check relationships separately: referenced observations should exist, artifacts should resolve, and linked records should belong to the intended attempt. Where the workflow guarantees an ordering, check it against the appropriate sequence or clock basis.
Then review meaning. Does the selected observation actually correspond to the context used for the action? Does the assessment reference evidence appropriate to its criterion? A record can satisfy every required field while pointing to the wrong screenshot.
Keep validation findings and reasons for exclusion. A recording problem and a task failure should remain distinguishable, even if both make an example unsuitable for one particular dataset.

Build a dataset view with explicit selection rules
A dataset release is a documented selection of records and artifacts for a stated purpose. Record which attempts were included, what was filtered, which transformations were applied, and which schema and assessment versions were used.
Do not discard all unsuccessful attempts by default. Failure analysis may need them. Demonstration learning may use a narrower subset. An unresolved outcome might prevent use as a verified-success example while still supporting inspection of an earlier, well-recorded action.
Prevent inappropriate overlap between training and evaluation data. Avoid splitting adjacent steps of the same attempt across those sets when that would expose evaluation context during training. Related retries, duplicated artifacts, or near-identical task instances may require grouping too. Define the split around the generalization claim being tested.
For action prediction, future observations and final assessments should not accidentally appear among the inputs available to the model at that step. They may remain in the source record for review while being excluded from that training view.
A release manifest makes the selection inspectable. It should identify the dataset version, included records, relevant processing versions, and known limitations so another person can understand which data they received.
Keep structure separate from readiness
Structured data is easier to inspect and process. It is not automatically suitable for training or evaluation. Suitability still depends on coverage, evidence quality, labeling, selection, and the intended task.
Real Android infrastructure supplies an execution environment. Organizing runs into a reusable dataset also requires collection logic, transformation, validation, and documented selection. Confirm which parts are available and which require integration in a proposed setup.
Start by tracing one exported example back to its evidence. If its relationships and limits remain clear, apply the same checks to the larger collection.
The next article will examine why a simple success or failure label is not enough for mobile AI evaluation—and what evidence makes that label interpretable.
