Four techniques, and what each one cannot see
Static analysis, dynamic analysis, network interception and runtime instrumentation each answer a different question about a mobile app - and each is blind in a different direction. What the core techniques actually find, and what to ask for in the report.
There is a disagreement that turns up on almost every mobile assessment, and it is worth starting with because it explains why the work takes the shape it does.
Read the code and you will find a session token being written to NSUserDefaults. That looks like a finding, and on paper it is one. Run the app and you will never see that write happen, because the branch that performs it only fires once the user has upgraded from an earlier version. Put a proxy in the path and the same token turns up in a request header, in the clear, on every call the app makes.
All three of those observations are true. Only the third belongs in a report, and the first is the one that usually gets written.
This is why mobile testing uses four techniques instead of one good one. They answer different questions, and each of them is blind in a direction the others are not.
Reading the app without running it
Static analysis means working on the shipped artefact: the IPA or APK, the decompiled bytecode, the manifest or Info.plist, the bundled assets and native libraries. jadx will give you readable Java back out of a DEX file. Ghidra or Hopper will take on the Mach-O and the .so files, though you will spend most of that time doing something closer to archaeology than reading. For everything else there is strings, entropy analysis and pattern matching, which is unglamorous and finds a surprising amount.
What you are hunting for is anything the developer left switched on at build time. Hardcoded keys and tokens. android:allowBackup at its default, which means anyone with the handset and a USB cable can pull the app's private data. A debuggable flag that shipped. NSAllowsArbitraryLoads set in the Info.plist, quietly disabling App Transport Security. Entitlements wider than the app needs. An exported component with no permission guarding it. Weak or misused crypto primitives. And the full inventory of third-party SDKs that came along for the ride, which is usually longer than the development team expects.
The trouble with all of it is that static analysis cannot tell you whether a line of code runs. A token written to unencrypted storage inside a function nothing ever calls is a true observation and a worthless finding, and keeping those two apart is most of the skill.
It also degrades badly on modern frameworks. A React Native app ships a JavaScript bundle, a Flutter app ships a compiled libapp.so, and a Java decompiler has no idea what to do with either. Both routinely contain the strings you are looking for. When a tool cannot read the artefact, the honest thing to write down is "not read", and it is remarkable how often that gets recorded as "not found" instead.
Running it and watching what it does
Dynamic analysis means putting the app on a device or an emulator and watching. What it writes to disk, what lands in the keychain or the keystore, what it puts in the logs, what it leaves in the app switcher screenshot, what it puts on the clipboard, which permissions it asks for and at what point in the flow.
The evidence here is more physical than in the static pass. You pull the app container off a rooted or jailbroken device, you read logcat on Android or os_log on iOS, you take a filesystem snapshot before and after a flow and compare the two. This is the pass that catches a token still valid after logout, or personal data cached somewhere world-readable.
What limits it is coverage. You only see the paths you actually drive. Anything behind a paid tier, a particular account state, a region, or a second-run branch stays invisible, and automation tends to fall over the moment there is a login, an MFA prompt or any anti-automation on the way in. A dynamic run is evidence about the flows you drove. It is not evidence about the app.
There is also a confound worth knowing about before you draw conclusions. Getting at the app container means rooting or jailbreaking the device, and a well-built app notices. If its detection fires and it starts degrading, you are now testing a different program from the one your client's customers are running.
The proxy in the middle
For the network pass you put Burp or ZAP between the app and its backend, install the CA certificate, and read the traffic.
The highest-severity findings in most mobile assessments come out of this pass, and usually not because of the cryptography. It is where you get to test the API. Whether an object identifier in a request can be swapped for someone else's and comes back with their data. Whether authorisation is actually enforced server-side or just assumed because the client only ever offers you your own records. Whether a request can be replayed. Whether tokens are turning up in query strings, which puts them in every access log along the path. Whether anything sensitive is going over plain HTTP.
What it cannot reach is the traffic the app encrypts itself, or pins. On Android 7.0 and later an app no longer trusts user-installed CAs by default, so getting the proxy in means a system-level certificate or a repackaged build. Certificate pinning defeats it outright, which is the entire point of pinning. And none of this touches traffic that never leaves the device, or anything that is not HTTP.
Hooking the process
The last technique is to attach to the running app with Frida and hook its own methods.
This is the only one that answers whether a control is real. You can read a key out of memory at the moment the app uses it, rather than hunting through a binary for it. You can look at a payload after the app has finished decrypting it. You can skip root and jailbreak detection, step past a debugger check, and call an internal function directly to see what it would have done.
It is also the easiest technique to misreport. Instrumentation proves what is possible for someone who has already decided to reverse your client's app and has the tooling to do it. That is not the same risk as a key sitting in plaintext in a binary anyone can download, and writing up a Frida bypass at the severity of a leaked credential is the fastest way to lose the reader. The finding is what the bypass gets you, not the fact that it worked.
What actually gets tested
The techniques are how you test. These are the things being tested, and each one is only reachable by the right technique.
- Storage. Where tokens, credentials and personal data land, and whether the platform's protected store was used. Static tells you which API was called, dynamic tells you what is on disk afterwards.
- Cryptography. Which primitives, which modes, where the keys come from, whether any of them are hardcoded. Static for the algorithm, instrumentation for the key.
- Authentication and sessions. Token lifetime, refresh behaviour, whether logout invalidates anything, whether a biometric gate can be stepped around by calling past it.
- Network. TLS validation, pinning, cleartext endpoints, and the authorisation logic sitting behind every call.
- Platform interaction. Exported components, deep links and URL schemes, IPC, WebViews with JavaScript bridges, and what another app on the same device can reach.
- Code and configuration. Debug flags, backup settings, sensitive values in logs, and the dependency inventory.
- Resilience. Root, jailbreak, emulator and debugger detection, integrity checks, obfuscation. You test these by trying to defeat them.
- Privacy. What personal data leaves the device, to whom, and through which SDK.
Where reports go wrong
Static observations written up as runtime facts. This is the most common defect in a mobile report, and it is a claim about behaviour that was only ever inferred from source. If a path was never executed, that belongs in the report as a limitation.
Interception results written up as coverage. A proxy only sees what the app chose to send over HTTP. An empty intercept log is evidence of pinning or of some other transport, and treating it as evidence of a clean app is how a report ends up saying nothing while sounding thorough.
Dynamic results written up as completeness. An afternoon of tapping through an app is a test of that afternoon.
And instrumentation findings written up as though the impact were self-evident. A bypass proves a control can be removed by someone holding the device and the tooling. Say who that is and what they walk away with.
What to ask for
- Which technique produced each finding. "The token is stored insecurely" is not a finding. "Read from the app container after login" is.
- What was not tested. Every engagement has a boundary somewhere: an account tier, a platform, a pinned endpoint nobody repackaged. The boundary is part of the result and should be stated as plainly as the findings.
- Whether the artefact was actually decompiled. A bundle that could not be read is not a clean bundle, and a report that treats the two the same is not worth much.
- The reproduction. A request that can be replayed, a hook that can be re-run, a file that can be pulled off the device.
- And treat everything client-side as a speed bump. Root detection, obfuscation, integrity checks, even pinning: these raise the cost of an attack, and that is a real benefit, but none of them creates trust. Anything that genuinely must not be bypassed has to be enforced on the server, because that is the only place the attacker does not control.
None of this is exotic; the four techniques have been standard practice for well over a decade. What separates a useful report from a decorative one is usually just whether the author kept track of which technique produced each line, and whether they were honest about the parts they never got to.
