A mobile agent taps Save and returns to a list. Did it complete the task? The answer depends on the task criterion. If the request was only to leave the editor, perhaps. If the task was to save a note with a specific title and body, the screen transition alone does not establish that both fields persisted.
The problem is not that success and failure are useless labels. The problem is asking one label to carry more meaning than the evidence supports.
Start with a criterion that can be checked
Consider a simple task: create a note titled “Site visit” with the body “Bring the access badge.” A useful evaluation begins by stating what counts as completion. For example: after saving, the note can be reopened and both requested fields are visible.
That criterion makes the result interpretable. “The agent tapped Save” describes an action. “The app returned to the list” describes an observed transition. Neither alone proves that the requested note was saved correctly. Reopening the note and checking both fields provides evidence that is directly connected to the criterion.
Benchmarks such as AndroidWorld make this relationship explicit through task-specific initialization and success-checking code. The benchmark paper also reports that task variations can materially change performance, a reminder that a score depends on which tasks and conditions were evaluated. AndroidWorld, ICLR 2025
Separate the outcome from the evidence
A report should make clear at least three things:
- Outcome: Did the attempt satisfy the stated task criterion?
- Evidence: Which observation or state check supports that conclusion?
- Evaluation status: Could the criterion be checked reliably for this attempt?
These are related, but they are not interchangeable. A checker can fail to run, a screenshot can be missing, or the final state can be ambiguous. Those cases do not establish task success. They also do not necessarily establish that the agent failed the task. Keep them unresolved or invalid for the relevant measurement, explain why, and show how they affect the denominator. Do not silently count them as passes or failures.
This distinction matters when the agent’s own message says “Done.” That message is part of the record; it is not independent proof that the app state meets the requested criterion. An evaluator should check the result at the level the task requires, using evidence that can be traced back to the attempt.

A single score can hide different failure modes
Two agents can receive the same failure label for very different reasons. One may choose the wrong control. Another may enter the right content but lose it during saving. A third may complete the task while the checker looks at the wrong screen. The headline rate collapses these cases into one number.
That is why a final success rate is useful for comparison but often weak for diagnosis. AgentBoard describes this limitation directly: final success rates alone reveal little about an agent’s process, motivating finer-grained progress analysis. AgentBoard, NeurIPS 2024
Additional measures should follow the question being asked. If the goal is to debug task completion, record the last verified state and the point where evidence diverges from the criterion. If the goal is to compare efficiency, actions or elapsed time may matter. If a task includes a constraint, evaluate that constraint explicitly. Partial progress can help explain an attempt, but it should not replace a required final-state check.
More metrics do not automatically make an evaluation better. Each extra measure needs a definition, a reason to include it, and a reliable way to compute it. A long dashboard of loosely defined numbers can obscure the result as easily as a single binary label.
Make the score comparable
A score is meaningful only in relation to the evaluation setup. Record the task definition and success criterion, initial state, app and relevant environment versions, agent or controller version, checker version, number of attempts, and the denominator used. If task variants or device conditions differ, state how they were grouped or controlled.
This does not mean every report needs a large benchmark harness. A small evaluation can still be clear: define a narrow set of tasks, retain the attempt evidence, use a stated checker, and report what the resulting score does and does not cover. When evidence cannot support a conclusion, say so.
A practical report can therefore present a headline success rate alongside a compact evidence summary: verified successes, verified failures, unresolved evaluations, the criteria used, and a small number of diagnostics chosen for the decision at hand. The exact categories and metrics should be documented rather than treated as a universal standard.
What this means for real-device evaluation
Physical Android devices can expose execution conditions that a benchmark or emulator may not represent. But running a task on real hardware does not, on its own, validate the outcome. Evaluation still depends on the task criterion, evidence collection, checker behavior, and the conditions under which attempts are compared.
ARMARRAY’s role is the real-device infrastructure layer. A complete evaluation workflow may require additional integration for task orchestration, trajectory recording, outcome assessment, and reporting. Those should be confirmed for the specific setup; real-device access alone should not be read as a claim that a complete evaluation product is already available.
Success and failure remain useful. Treat them as the beginning of a result, then show the criterion, evidence, and evaluation limits that let someone understand what the score actually says.
The next article looks at what a Mobile AI evaluation dataset needs to retain so these results can be inspected and compared responsibly.
