A mobile task trajectory connects what was observed, what was done, what changed, and how a task ended—within one identifiable attempt.

A screenshot shows a moment. An action log lists operations. Neither necessarily explains how a mobile task unfolded.

In the previous article, we examined the gap between UI automation logs and usable AI data. The next question is how to organize execution into something a reviewer can follow. A useful starting point is the mobile task trajectory: a connected record of an attempt to achieve a goal.

The important word is connected. A folder of screenshots and a list of taps may contain the ingredients, but the relationships between them determine whether the attempt can be interpreted.

Start with one task and one attempt

Consider an illustrative task: create a note titled “Site visit,” add “Bring the access badge” as its body, and confirm that the saved note contains both.

The attempt begins with the notes app open on its list screen. The agent opens the editor, enters the requested text, saves, returns to the list, and opens the note to inspect it.

This is a conceptual example, not an ARMARRAY product demonstration. Its purpose is to show why a trajectory needs more than a successful final screenshot.

The task supplies the goal. The attempt supplies the execution boundary. Together, they let us ask which observations and actions belong to this particular effort—and what evidence supports its reported result.

Four concepts make the sequence readable

Observation is the information available at a particular point in the attempt. It might include a screenshot or an accessible description of the interface. An observation is a view of the environment, not necessarily its complete internal state.

Action is an operation selected or issued in that context: opening the editor, entering text, tapping Save, or navigating back to inspect the result. An issued action and a confirmed effect are different parts of the record.

State transition describes how the environment changes as execution proceeds. The record may reveal only part of that change. Seeing an editor replaced by a list screen establishes a visible transition; it does not, by itself, establish that the requested note was persisted correctly.

Outcome is the assessed result relative to the task goal. It needs a stated basis. In this example, reopening the note and observing the requested title and body supports a more specific conclusion than merely observing that the editor closed.

These concepts are an explanatory model, not a mandatory file format. A system may represent them differently while preserving the relationships needed to understand the attempt.

A notes list observation, an open-editor action, and the resulting empty editor.

Follow the evidence through the example

At the start, the observation shows the notes list. The next action opens a new note. A later observation shows an empty editor. That observation provides the context for entering the requested title and body.

After text entry, another observation shows the populated editor. The Save action follows. The next observation shows the list, including an entry titled “Site visit.” This is useful evidence, but the requested body has not yet been checked there.

The agent opens that entry. The resulting observation shows the title and body together. The outcome assessment can now refer to those visible contents and the criterion defined for this illustrative task.

This does not prove every possible property of storage, synchronization, or future availability. Those would require different criteria and evidence. A trajectory should support the claim actually being made, with its limits intact.

A reader can now trace the result backward: from the outcome assessment to its supporting observation, from that observation to the preceding action, and from that action to the context in which it was selected.

Sequence alone is not enough

Putting events in chronological order helps, but it does not automatically establish their relationships.

Suppose the screenshot attached to Save was captured before text entry finished. It is part of the same attempt, yet it provides the wrong context for interpreting that action. Or suppose the final screenshot came from a second attempt. The images may look consistent while describing different executions.

Mobile interfaces can also update while an agent is waiting. A dialog may appear, loading may continue, or an external event may alter the screen. The next observation should not automatically be described as an effect caused entirely by the immediately preceding action.

A useful trajectory preserves the observed sequence and makes uncertain relationships visible. Missing observations should remain missing; they should not be filled with an invented account of what must have happened.

Keep retries and endings visible

Now imagine that Save is followed by a loading indicator, then an error. The agent waits, observes the error, and tries again.

Removing those steps would make the attempt appear simpler than it was. Keeping them allows a reviewer to examine the recovery path and decide whether it is appropriate for the intended use of the data.

A retry can remain within the same attempt when that is how the collection process defines the boundary. If the environment is reset and a new attempt begins, that boundary should be explicit. Separate attempts should not be silently joined into one apparently continuous success.

An attempt can also end before its goal is verified. A timeout, interruption, or unavailable observation may leave the result unresolved. The record still describes a trajectory, although its completeness and suitability for training or evaluation need separate assessment.

One task attempt includes an error and retry; a reset starts a separate attempt.

A trajectory is evidence, not automatic training approval

A readable attempt helps people inspect decisions, locate a divergence, or compare execution paths. It does not make every included action a good training target. An unnecessary detour or an accidental tap remains part of what happened, even if it should not become an example of preferred behavior.

Likewise, a detailed trace does not replace an outcome criterion. AndroidWorld provides a concrete research example of this distinction: its task definitions include initialization and success-checking logic, alongside an agent interaction loop that gathers observations and executes actions. The execution sequence and the task assessment serve related but different purposes. AndroidWorld project

This reasoning applies to emulator and physical-device runs. Choosing an execution environment and designing an interpretable record are separate decisions. Physical execution alone does not establish that the trajectory is complete or that the task succeeded.

Next: What belongs in the record?

A mobile task trajectory gives an attempt a readable structure: a goal, connected observations and actions, evidence of change, and an ending whose meaning can be inspected.

The next practical question is what to capture so those connections survive beyond the run. The following article will examine the information worth recording during a mobile AI task—from screenshots and operations to timing, environment context, and result evidence.