Repeatability and real-world variation are often presented as competing goals. For Mobile AI evaluation, they answer different questions.

A controlled run helps isolate a defined condition and compare attempts. A physical-device run can reveal how an interaction behaves on actual hardware and a particular software configuration. Neither result is automatically more valid. The right environment depends on the claim the evaluation needs to support.

Repeatability is a method, not an emulator property

To compare attempts, define what should stay the same: the task and completion criterion, starting app state, app and operating-system versions, agent and checker versions, permissions, and relevant network conditions. Record what was configured and what was actually observed.

Even with a fixed setup, mobile applications perform asynchronous work. A screen may still be loading, a background operation may finish at a different time, or an external service may respond differently. Android’s testing guidance identifies synchronization and unknown background work as sources of flaky results, and cautions against arbitrary fixed delays as a general solution. A repeatable protocol therefore needs observable wait conditions, recorded attempts, and explicit handling of unresolved checks—not just the same device profile.

Retries should remain visible as attempts. If a run is repeated after a timeout or an ambiguous observation, retain the first result and the reason for retrying rather than silently replacing it with the later outcome.

A virtual device can control useful conditions

An emulator is not merely a generic stand-in. Android’s documentation describes virtual device profiles and API levels, along with controls for conditions such as location, network speed, rotation, and sensor input. These controls can make it practical to rerun a scenario or vary a specified input while keeping other settings stable.

That is useful when the question is about a defined behavior: for example, whether an agent reaches the same requested state under a chosen network profile or screen configuration. The result establishes what happened under those configured conditions. It does not establish how every physical device, carrier, app build, or service will behave.

Simulation should be described precisely. Name the profile and the conditions changed. Do not report a simulated setting as though it were a measurement of a real device or network.

A physical-device run adds context, not automatic proof

Some questions depend on the particular hardware and software execution context. A team may need to inspect behavior across device models, OS versions, app builds, permissions, or physical inputs. Android’s device guidance recommends emulator use for platform versions and screen sizes alongside testing on real devices; its broader testing strategy places different test scopes on different environments, including emulators and selected phones or other devices.

Physical runs also vary. Network state, background activity, device configuration, timing, and app state can differ between attempts. If those conditions are not recorded, a result may be harder to reproduce even though it came from real hardware. “Real device” describes the execution environment; it does not by itself validate the task outcome or explain why two runs differed.

Pair the environments around one evaluation question

Start with one task and a checkable completion criterion. Run a controlled set of attempts where the goal is to compare a defined change. Then select physical-device runs for the conditions that matter to the intended use. Keep the task definition, evidence format, and outcome check aligned where possible so the results can be compared meaningfully.

Suppose the question is whether an agent can submit a form when connectivity is slow. A virtual-device run can repeat the flow with a specified network profile. Selected physical-device runs can then show how the same task behaves on particular devices and connections. Record the configured profile separately from observed connectivity and the evidence that the form was actually submitted.

For each attempt, preserve the device or virtual profile, OS and app versions, starting state, agent and checker versions, configured conditions, observations, actions, timestamps, outcome evidence, and any retries or missing data. Where practical, vary one condition at a time to make differences easier to interpret. When several conditions change together, report that clearly rather than attributing the outcome to one of them.

The same slow-connectivity form task compared across repeatable controlled runs and selected physical-device runs, with conditions interpreted in context.

If a controlled run passes and a physical-device run does not, neither environment should be dismissed by default. Trace both outcomes back to the criterion and evidence. Was the initial state different? Did the same action reach a different screen? Did a permission, network response, or checker behave differently? Does that difference matter to the evaluation question?

Report what each result can support

Separate the questions in the report. A controlled comparison can show whether attempts differ under specified settings. A selected physical-device set can reveal behavior across those device and software conditions. Neither alone proves that the result represents all users, devices, or real-world use.

State the tested scope, selection rules, number of attempts, unresolved checks, and known gaps. Distinguish settings that were assigned from conditions that were observed. If a result depends on a particular device group or app version, keep that boundary next to the conclusion.

The goal is not to maximize realism or repeatability in isolation. It is to make clear which evidence supports which claim—and which evidence is still missing.

Real Android infrastructure can provide a physical execution environment for device-based work. The collection logic, task orchestration, evaluation checks, and reporting required for a particular study should be confirmed separately; device access alone is not evidence that a complete evaluation service is present.