A device API can tell a system how to query a phone, change a setting, or launch an app. That is useful, but it does not by itself make a data collection workflow programmable.
A workflow also needs to say what work was accepted, which device or devices ran it, what happened after submission, how an uncertain run can be recovered, and what evidence supports the result. Without those connections, automation can issue commands while leaving the surrounding run difficult to inspect or reproduce.
The practical distinction is between controlling a device and exposing a bounded unit of work that software can submit, follow, and assess.
Start with a run contract
Consider a hypothetical evaluation: an Android agent changes an app’s notification preference, then verifies that the preference remains changed after leaving and reopening the relevant screen. The useful question is not simply whether a tap command was sent. The run needs a defined starting context, an eligible device, a completion condition, and a way to inspect what the device showed afterward.
A run contract can make those expectations explicit. It may identify the requested task, device or device profile, app and software context, starting conditions, completion criteria, timeout behavior, evidence to retain, and cleanup expectations. The exact fields depend on the work. The point is to make the request and its boundaries legible to the systems that submit and execute it.
For a multi-device workflow, the system may also need to coordinate eligibility, reservation, concurrency, and release. Those are orchestration concerns around device execution; they need not all live in one service or be implemented by a central scheduler.
Give asynchronous work a trackable identity
Physical-device runs take time. A submission may be accepted before a device is ready, while an app is launching, or while observations are being collected. A caller therefore needs a way to distinguish “the request was received” from “the work finished.”
One common API pattern is to return an operation reference for long-running work and let the caller retrieve progress or a final result. Google’s API design guidance describes this pattern, while leaving the details to the API’s own contract. It is a useful design reference, not a universal schema. Google AIP-151
Identity should also match the level being tracked. An operation ID can identify the overall request; an attempt ID can identify one execution on a device. If a request fans out to several devices, each attempt can have its own state and evidence while remaining linked to the parent operation. That makes partial completion visible instead of flattening a mixed result into one ambiguous status.
The lifecycle vocabulary should be documented and stable. A system might distinguish accepted, assigned, running, completed, failed, and canceled, for example. Whatever terms are chosen, callers need to know which transitions are possible and what each state guarantees.
For example, a command waiting for device discovery returns IDs you can use to follow it:
POST /api/devices/:uuid/command
{
"status": "queued_discovery",
"command_id": "[redacted]",
"discovery_id": "[redacted]"
}
Make retries describe intent
A network timeout does not prove that a device action never ran. The device may have completed the action while the response was lost. Blindly resubmitting can perform the action twice; refusing to retry can leave a run stranded.
A client-supplied request identifier can help a service recognize a repeated submission and avoid treating it as new work. Google’s API guidance describes request IDs as a pattern for idempotency and auditing. The exact behavior must still be defined by the service. Google AIP-155
It is useful to distinguish two cases: retrying delivery of the same request after an uncertain response, and deliberately starting a new attempt because the prior execution failed or needs comparison. They have different intent and should not be collapsed into one generic “retry” button. For state-changing actions, the interface should explain whether a retry resumes, deduplicates, or starts fresh.
Return evidence as well as status
“Completed” answers whether the execution reached an end state. It does not necessarily answer whether the requested outcome was achieved. In the notification example, a successful command response is weaker evidence than observing the setting after reopening the screen.
A useful result links the requested criterion to what was observed. Depending on the task, that may include timestamps, device and software context, action history, screenshots or other observation references, and an outcome assessment. The result should also make uncertainty visible: an execution can finish while its evidence remains incomplete or its outcome cannot be determined.
Errors benefit from similar precision. A failure before assignment, a device disconnect during execution, and a completed run with missing evidence call for different recovery decisions. Multi-device jobs may also end partially: some attempts can complete while others fail or remain unresolved. Preserving per-attempt results helps downstream evaluation decide what can be used, retried, or excluded.

Keep device control and task orchestration distinct
At the device-control layer, ARMARRAY supports device discovery and status, configuration, commands, WebSocket-based registration and control, remote sessions, MQTT, and related connectivity. A data collection workflow brings these capabilities together with task submission, resource allocation, operation and attempt tracking, observations, artifacts, and outcome criteria. How these parts are organized can vary with the project; what matters is that callers can follow a requested run through device execution to its evidence-backed assessment.
Questions that make an interface reviewable
When reviewing a proposed integration, ask how a caller identifies a request and an execution attempt; what acceptance and completion mean; where progress, errors, and partial results appear; how duplicate submissions and intentional retries differ; which evidence is returned or retained; and how cancellation, cleanup, and device release work. These questions are a practical starting point, not a fixed checklist for every system.
A real-device data collection system becomes programmable when software can express a bounded request, follow its execution, recover from uncertainty, and connect the final assessment to inspectable evidence. Device commands are part of that system. The surrounding contract is what lets teams operate it reliably.
This builds on the infrastructure question discussed in Why 10 Phones on a Desk Are Not a Data Collection Infrastructure: once devices form an operational pool, the next question is how software can work with that pool through explicit, observable interfaces.
