Ten Android phones on a desk may be enough to explore an app flow or run a small, manually supervised study. The number itself tells us little about whether the setup can produce data that another person can interpret, reproduce, or audit.

The important question is not whether the phones are physically real. It is whether each run can be connected to the device that performed it, the conditions under which it ran, and the evidence behind its outcome. A phone collection becomes useful infrastructure when those relationships survive routine interruptions: a cable moves, an app stalls, a device reboots, or an attempt must be repeated.

A device count is not a run record

Consider a small study asking an Android agent to update a profile field and verify that the saved value appears afterward. Ten phones are connected to a workstation. The operator labels them by USB port, starts runs, and saves screenshots into folders named after those ports.

During one run, a device disconnects. The operator reconnects a phone to a different port and retries the task. The screenshot shows the expected value, but the folder name now identifies the port rather than the physical phone. If the app was also restarted, the screenshot may not reveal whether it belongs to the first attempt or the retry. The task may have completed, yet the record no longer supports a confident statement about which device and attempt produced the evidence.

This is not a failure of having only ten devices. It is an identity and traceability failure. If the lab had stable device identifiers, attempt identifiers, and a record of reassignment and retry, ten devices could support a well-run study. A much larger rack with ambiguous device-to-run mapping could not.

A useful unit of work is therefore not “phone 7” or “the 3:00 p.m. batch.” It is a specific task attempt bound to a known device profile, starting conditions, ordered events, and an outcome check. That is the connection the earlier article on turning raw device runs into structured AI data asks readers to preserve.

The operating system around a device matters

A repeatable device workflow has several responsibilities. They can be handled manually at small scale, provided the procedure is explicit and the records remain complete.

Identify the device. Keep a stable identifier for each physical unit and record relevant properties such as model, Android build, and app build. A USB port, shelf position, or temporary network address is a location, not a durable device identity. If a unit is replaced, preserve the change in the run record.

Define the starting point. Specify what must be true before the task begins: app and account state, permissions, locale, orientation, and other conditions relevant to the question. A reset procedure should say what it resets and what it leaves intact. “Restarted” is not enough to establish that two attempts began from comparable states.

Assign each attempt. Give every attempt a unique identifier and bind it to the task definition, device, software profile, and run conditions. If an attempt is interrupted, record that attempt as interrupted and assign a new identifier to a retry. Do not silently turn two attempts into one successful run.

Preserve evidence and outcome rules. Keep actions, observations, timestamps where available, and the source used to check the result together. State the completion criterion before execution. A screenshot can show a visible state; it may not prove a durable application change. When an authorized and reliable state check is unavailable, label the outcome unresolved rather than guessing.

Recover without erasing history. Devices will become unavailable. A useful workflow records why a run stopped, what recovery action was taken, whether the device returned to a known state, and whether the next attempt is comparable. Retain failed and incomplete attempts when they explain gaps in the dataset.

A traceable device run connects device identity and run context to separately recorded attempts, evidence, and an outcome check.

Power, thermal state, and network are part of the run context

A device pool also depends on conditions outside the app. Record the power and connection arrangement well enough to explain an interruption: which unit or hub connection was involved, whether charging or a cable change affected availability, and what recovery followed. These details do not need to become elaborate telemetry for every study. They matter when they affect whether a device can start, continue, or return to a known state.

For sustained or performance-sensitive runs, thermal state may also matter. Android exposes thermal status, and its documentation explains that heat can lead to throttling and affect performance. That does not make thermal telemetry mandatory for every UI task; it makes thermal conditions a candidate variable when duration or device performance is part of the question. See the Android Thermal API documentation.

Network state needs similar care. Android documents that network capabilities and transports can change during execution. A configured network profile describes what was intended; it is not proof of the connection observed throughout an attempt. Where connectivity matters, retain the configured condition separately from observations or connectivity changes available to the workflow. See Android documentation on reading network state.

A central control screen can make a device pool easier to operate, but a green “available” indicator is useful only if its meaning is defined. It might mean the device is reachable, the app is prepared, or a task can start; those are different states. Whatever the interface, the run record should preserve which state was checked and when.

The platform examples support the distinction, not a universal schema

Google’s Firebase Test Lab documentation describes test matrices in terms of selected devices and configurations such as Android API level, orientation, and locale, and describes reviewing test results and artifacts. Android’s Test Orchestrator documentation explains how separate test invocations can reduce shared state and isolate crashes, while noting the added restart time.

These are examples from app testing. They do not prescribe a Mobile AI data schema, and they do not mean every study needs a cloud test matrix or an orchestrator. They do illustrate a useful engineering principle: execution configuration and attempt boundaries are explicit parts of a test, not details that can be reconstructed from a device count after the fact.

For a small collection, a careful operator and a simple manifest may be sufficient. As concurrency and device variation increase, manual bookkeeping becomes harder to keep consistent. Automation can help with inventory, assignment, status reporting, capture, and recovery, but each automated step still needs observable records and clear failure behavior. A button that says “run” does not by itself establish that an attempt started, completed, or passed its outcome check.

Measure the workflow, not just the shelf

When describing a device collection, report more than how many phones are present. Useful operational details include which devices were eligible for a run, how many attempts were started and completed, which were interrupted or retried, what evidence was captured, and how outcomes were checked. Keep the denominator visible when reporting rates, and separate “device unavailable,” “execution failed,” “criterion not met,” and “evidence insufficient.” These states answer different questions.

The appropriate detail depends on the claim. A study about a single app flow may need only a narrow device profile and careful attempt records. A transfer claim across Android versions or manufacturers needs a deliberate selection method and evidence that the chosen combinations were actually exercised. Adding phones without changing the represented profiles or conditions may increase throughput, but it does not automatically broaden coverage.

What ten phones can be

Ten phones can be a prototype lab, a supervised test bench, or a dependable small device pool. They can also be a pile of hardware that produces screenshots without reliable context. The difference is not a threshold number or a rack design. It is whether the workflow can answer, for every result: which task and attempt was this, which device and conditions were involved, what happened, what evidence supports the label, and what remains unknown?

Real Android infrastructure provides the physical execution substrate. The task model, device-to-run assignment, data capture, outcome verification, recovery rules, and reporting still have to fit the intended evaluation. Device access is necessary for some questions, but device count alone is not a data collection system.