“Add more diversity” sounds like an obvious way to improve mobile AI data. But diversity is not a single dial. A broader app list, different account and system states, additional locales, and varied network conditions can each change a different part of an evaluation. Adding them without a question can make results harder to interpret without making them more useful.
A better approach is to define the claim first, identify the conditions that could affect it, and record which parts of that space the data actually covers.
Start with the claim, not the inventory
Consider an agent asked to update a delivery address and place an order. If the claim is that it can complete this task reliably across supported shopping apps, app-specific checkout flows and account states may change the path: an address may already exist, a sign-in prompt may appear, or an extra confirmation step may be required. Locale can change address formats and labels; a slow or interrupted connection can make the submission state ambiguous; keyboard layout, screen size, or a system prompt can affect whether a control remains visible.
Those differences do not all belong in every study. If the question is whether the agent recovers from a delayed order submission, hold the app, account, locale, and device profile steady where practical, then vary the network condition and preserve evidence of the final order state. If the claim is transfer across apps, select distinct checkout flows and define the same task-level completion criterion for each. Device models become relevant when the claim concerns layout, input, or hardware-dependent behavior—not simply because more models make the inventory look broader.
The same dataset can be diverse in one dimension and narrow in another. State the intended claim before counting what is included.

Dimensions that may matter—and why
The examples below are useful lenses, not a complete or mutually exclusive taxonomy. A study may need other dimensions—such as app build, input method, accessibility settings, time zone, or notification state—and some factors interact. Choose dimensions because they could change execution, evidence, or interpretation for the claim at hand.
App diversity asks whether behavior transfers across applications, interfaces, and interaction conventions. A dozen tasks in one app do not establish cross-app coverage. Record the app and relevant build or version, and separate app families when their flows differ materially.
Task diversity asks whether the dataset covers different user goals and interaction patterns. Rephrasing the same sequence does not create a new task class. Useful variation may include search, form entry, editing, navigation, multi-step completion, and recovery—but select categories that match the intended use rather than treating a longer task list as proof of breadth.
Network diversity asks how behavior changes under specified connectivity conditions, such as latency, limited throughput, or interruption. Keep configured network profiles distinct from observed connectivity. A simulated profile is evidence about that configured condition, not a measurement of a carrier network. Android’s Emulator documentation describes controls for simulating network speeds; those settings help define a repeatable test condition, but do not stand in for every real connection.
System-state diversity asks whether the task is robust to relevant device and app states: permission prompts, dialogs, loading states, keyboard visibility, orientation, background/foreground transitions, or an existing signed-in session. List only states that could alter the path or the outcome check. Randomly varying state makes comparisons difficult; deliberately selecting it makes the coverage legible.
Locale diversity asks whether language, formatting, or regional conventions affect the task. Translated strings alone may not cover date, number, address, or right-to-left layout differences. Record the locale and the specific behavior being tested. Android’s app language guidance describes per-app language support, one example of why “language” should be specified rather than treated as a generic label.
Account-state diversity asks whether behavior depends on authentication, permissions, personalization, or account history. Signed-out, new, and established accounts can expose different screens and data. Protect privacy: use appropriate test accounts and avoid placing secrets or personal information in captured evidence.
Device-model diversity asks whether hardware and software differences affect the intended behavior: screen size, OS release, manufacturer customizations, input methods, or available sensors. A list of models is not automatically a representative sample. Describe the selection rule and the properties those devices are meant to cover. Android’s device testing guidance recommends combining virtual devices with testing on selected physical devices according to the question being tested.
These dimensions can interact. A locale change may alter the layout; a different OS version may change a permission prompt; a weak connection may expose a loading state only in one app. When the goal is to explain a difference, vary one factor at a time where practical. When the goal is to observe realistic combinations, preserve the combinations and report them as such.
Breadth and comparability need separate views
A coverage matrix can show which combinations were tested, but a filled cell does not mean equal evidence. For each cell, keep the task definition, starting conditions, attempts, observed states, outcome criterion, and missing evidence. A single run on a device model is not equivalent to repeated runs across that model; five locale labels do not show that locale-sensitive tasks were exercised.
Keep two views of the results:
- Coverage: which apps, tasks, conditions, account states, locales, and device profiles are represented, and which are missing.
- Comparison: which conditions were held constant, which changed, how many attempts were observed, and whether the outcome criterion was applied consistently.
This distinction matters because broad coverage and clean causal comparison are different goals. A broad sample can reveal where behavior may vary. A controlled comparison can help explain a variation. Neither should be presented as the other.

Choose diversity in proportion to the claim
For a narrow regression check, use a small, controlled set of tasks and configurations, then add the exact variation implicated by the change. For a cross-app capability claim, select distinct app flows and task patterns, document the selection method, and test enough examples to expose variation within each chosen group. For a claim about physical-device behavior, name the device and software scope, and avoid generalizing from a convenience sample to all Android users. The dimensions listed above are starting points; include additional factors when the intended claim depends on them.
A practical selection sequence is:
- Write the claim the evaluation is meant to support.
- Identify the dimensions that could change task execution or interpretation.
- Choose a baseline and define the variations to compare.
- Record configured conditions separately from observed conditions.
- Apply the same completion criterion and evidence requirements across comparable runs.
- Report the covered scope, selection limits, attempts, and unresolved checks next to the result.
This sequence does not prescribe a universal number of apps, tasks, accounts, or devices. The right sample depends on the claim, the cost of a missed failure, and the variation expected in the intended use.
Diversity is a coverage decision, not a quality score
A dataset is not strong simply because it contains many labels. It is useful when its coverage matches the question, its evidence supports the stated conclusion, and its limits remain visible. More diversity can reveal failure modes; without clear task definitions and comparable evidence, it can also obscure why a result changed.
For Mobile AI data, the practical question is not “How many kinds of things did we include?” It is “Which differences did we test, why do they matter, and what can the resulting evidence support?”
Real Android infrastructure can provide a physical execution environment for device-based work. Dataset design, task orchestration, capture, evaluation checks, and reporting still need to be specified for the study; access to devices alone does not demonstrate that a complete data or evaluation platform is present.
