What 116 findings look like when you have to triage them
One scan produced 116 findings. Eight rows described a single missing capability, and 93 of them landed in the tier everyone skips.
One production iOS app. One scan. 116 findings.
HIGH 10
MEDIUM 11
INFO 93
LOW 2
Read that as a to-do list and you have 116 problems and no idea where to start. Read it properly and you have a few root causes plus a long tail of facts. Nearly all the work that gets you from the first reading to the second happens after collection, and it's the part most scanning tools skip.
Eight rows, one missing capability
Eight of that scan's eleven MEDIUM findings:
✗ Missing: Certificate pinning
✗ Missing: Jailbreak detection
✗ Missing: Debugger detection
✗ Missing: Anti-swizzling
✗ Missing: Integrity verification
✗ Missing: Code signing verification
✗ Missing: Keychain protection
✗ No code obfuscation detected
All eight carry the same CWE, CWE-693. All eight sit under the same control, MASVS-RESILIENCE. Seven close with the same line: across 32 first-party source files. A team could file eight tickets off that list. A person reading it sees one sentence: this build has no client-side hardening in it.
That gap matters more than it sounds. Eight rows means one root cause spread across eight owners, and eight rows is how you teach somebody to skim. Enumeration is what a scanner is good at. It isn't triage.
So why didn't we merge them? Because the tool can't honestly do it. Those rows state eight different facts: pinning is absent, jailbreak detection is absent, integrity verification is absent. A rule loose enough to fuse those would also fuse "pinning is absent" with "pinning is present but weak" on the next app, and that's a wrong answer rather than a tidier one. What the tool owes the reader is the material to group them. Same CWE, same control, same scope line, repeated down the page. The grouping stays a judgement, and it stays the reader's.
What we do merge, and what we don't
Worth being exact about this, because "our tool deduplicates" is easy to claim and easy to disprove. Three things get merged, and you can check all three in the output.
The same finding from several analyzers. MobSF, Androguard and our manifest analyzer each report an app's exported components in their own words. They arrive as one row whose source reads androguard + manifest_analyzer + mobsf, with a note recording how many analyzers agreed.
Two wordings of one fact. "Clear text traffic is Enabled For App[android:usesCleartextTraffic=true]" and "Cleartext traffic allowed" are the same attribute of the same <application> element. Both map to a single label before grouping, so a report never charges you twice for one flag.
The same title across three or more files. Those collapse into one row carrying an affected_files count and the list of sites it covers. One detector firing in eleven files is one finding with a count, not eleven findings.
What we don't merge is two findings that merely share a root cause. That line is deliberate. Fusing the eight absences above would mean substituting our judgement for the evidence, and a report that does that is worth less than one that hands you the pieces.
The tier everyone skips
93 of that scan's 116 findings were filed INFO. It is the tier everyone skips, and it should not be.
84 of those were discovered endpoints, lifted out of the binary one string at a time. Not one of them is a vulnerability on its own. A URL is a fact until something reads it and works out what it is looking at, and most of these were never worth the second look: font licences, Adobe XMP namespaces, Apple's own certificate servers.
That is not an argument for promoting them, and it is not an argument for deleting the tier either. INFO says nothing about how much a finding matters. It says what has been proven so far. Discovery and analysis are different jobs, and a tool that only escalates what already looks scary will miss exactly the findings that needed somebody to interpret them.
The tier has to stay honest in the other direction too. Most of what a scanner extracts genuinely is informational: platform artefacts, library licences, framework internals. Report that as risk and you've taught your users to ignore you.
The noise you have to throw away
A large chunk of any mobile scan is third-party code, and findings in it are close to worthless to the customer. Nobody can fix a dependency's internal logging call, and that finding will be back on every scan until the library updates.
So library paths get filtered before they reach the report. Tools are weighted by how much their evidence actually supports a claim. A manifest declaration is close to certain. A regex hit against obfuscated code is a guess. They don't deserve the same confidence, so they don't get it. Findings that contradict a tool's known behaviour get suppressed by pattern instead of resurfacing on every scan.
None of it is glamorous. All of it decides whether your report gets read or stops getting opened.
What the number should mean
The question worth asking about a scan isn't "how many findings". It's "how many things do I have to change". Those are different numbers, and the distance between them is where reporting quality lives.
Three habits keep it honest:
- Group by root cause before you count. If two findings have the same fix, treat them as one finding with a count. Then check which direction that claim runs. A tool merging them for you is asserting they are one thing, and on the evidence alone it usually can't know.
- Keep the informational tier. Not everything needs escalating. Everything needs recording. The finding you wave off this quarter is the one you'll want a record of next quarter.
- State confidence per finding. A deterministic string pulled out of a binary and a regex hit on obfuscated code both produce "findings". Present them at the same confidence and you've built noise. Say which is which and you've built evidence.
