Test Automation

Real device coverage: which phones you actually need

Every mobile testing conversation reaches the same question eventually: how many real devices do we actually need? The honest answer is fewer than a device farm will sell you and more than nothing, and the number depends on a market rather than a best practice.

What follows is how we pick, and why emulators stop being enough earlier in a banking app than in almost anything else.

What an emulator cannot reproduce

An emulator runs your code. It does not run your code on a phone, and the gap between those two things is where a particular class of defect lives.

  • Biometric hardware. A fingerprint sensor on a mid-range Android handset behaves differently from the emulator’s simulated prompt, and differently again between manufacturers. Authentication flows that pass on a simulator fail on a real sensor with a wet thumb.
  • SMS and OTP delivery. Autofill, notification interception, the race between the code arriving and the field accepting it. None of this exists on an emulator, and all of it is load-bearing for a bank.
  • Manufacturer skins. One UI, MIUI, HyperOS and ColorOS each change permission dialogs, background execution limits and notification behaviour. Aggressive battery management on some skins kills a background sync your emulator runs happily.
  • Attestation and root detection. Play Integrity, SafetyNet’s successor checks, and your own root detection all treat an emulator as compromised — so the paths that matter most to a financial app are the paths an emulator cannot exercise at all.
  • Thermal and memory pressure. A device that has been in a pocket in Karachi in June throttles. Low-RAM handsets kill backgrounded activities. Neither happens on a workstation with 32GB.
  • Real networks. Not a throttled profile — an actual handover between cells, a captive portal, a connection that is present but useless.

Why a bank is the worst case

Most apps degrade gracefully when one of those things misbehaves. A financial app does not, because the features that depend on device hardware are the features that cannot be skipped.

Device binding ties an account to a handset. Screenshot blocking is a regulatory expectation in several markets. Root and emulator detection deliberately refuse to run in exactly the environment your test suite lives in. Biometric login is the primary authentication path for most users, not a convenience.

So in a consumer app the emulator covers perhaps ninety per cent of the surface. In a banking app it covers the ninety per cent nobody was worried about.

Which devices you actually need

Not the flagship you own. The method we use is dull and it works:

  1. Pull your own analytics first. Device model and OS version for the last ninety days of real sessions. Not market share reports — your users. The distribution is almost never what the team assumes.
  2. Cover to the long tail, not the mode. Take models in descending order until you reach roughly 80% of sessions. In Pakistan that usually means Samsung mid-range, Infinix, Tecno, Xiaomi and Vivo — and very rarely the handset anyone in the room is carrying.
  3. Add the floor deliberately. The oldest OS version and the lowest RAM configuration you still support. This is where crashes concentrate, and it is the device nobody volunteers to test on.
  4. Add one of each skin you are exposed to. Skin behaviour differs more than Android version does, which is the part teams consistently get wrong.
  5. Keep an iOS pair. Current and current-minus-two, because iOS fragmentation is real even if it is narrow.

That lands most teams between eight and twelve devices. A hundred-device matrix is a procurement decision rather than a testing one.

Cloud farm or your own shelf

Both, and for different jobs.

A cloud device farm is right for breadth: the regression run across twenty models overnight, the one-off reproduction on a handset nobody owns. It is wrong for anything involving SMS to a real number, a SIM, biometrics that need enrolling, or a flow where latency to a remote device changes the behaviour you are measuring.

A small physical shelf — the eight to twelve from the method above — is right for the authentication paths, the OTP flows, and anything a regulator may ask you to demonstrate. It is also the only way to hand someone a phone during a UAT session.

Teams that buy only the farm discover the gap the first time an OTP flow breaks. Teams that buy only the shelf cannot run breadth. The split is usually obvious once you have the device list.

What this looks like in a suite

Real-device coverage is not a separate test pass. It is a tier in the same suite: the fast, broad regression runs on emulators and the farm; the narrow set of hardware-dependent journeys runs on physical devices, on a schedule, with the results in the same report.

The failure mode we see most often is a mobile suite that is green on CI and has never once run the biometric login path, because that path cannot run where CI lives. If your pipeline cannot execute your primary authentication flow, your pipeline is not testing your app — it is testing the part of your app that happens to be convenient.

We build and run suites on this basis as part of our QA and test automation work, and we wrote about what else gives way under real conditions in testing fintech apps.

Building something that has to hold up?

Bring the process, not a specification. We will tell you honestly whether an agent is the right answer.