VAPT
VAPT of Android Applications
An Android application runs on hardware the attacker owns. Every control implemented on the device — root detection, certificate pinning, obfuscation, local authentication — is a speed bump on someone else's machine, and the only question worth asking is how long it holds. Testing therefore covers the app, the device it runs on, and the APIs it talks to, because a finding in any one of them is a finding in the product.
Methodology
- 01
Reconnaissance and build analysis
The APK or AAB is unpacked and the manifest, SDK levels, permissions, exported components, network security configuration and signing scheme are recorded. This establishes what the app is permitted to do before any testing of what it actually does.
- 02
Static analysis (SAST)
Decompilation to Java and Smali, then review for hardcoded secrets, weak cryptography, insecure storage paths, debug flags, backup settings, exported providers and unsafe WebView configuration. Automated scanning is the first pass, not the finding.
- 03
Device and storage assessment
The app is exercised on a rooted device and every artefact it writes is examined — shared preferences, SQLite databases, Realm stores, the keystore, cache, logs and external storage — for credentials, tokens, PII and session material left in the clear.
- 04
Dynamic analysis and traffic interception
Traffic is proxied with a user-installed CA. Where certificate pinning is present it is bypassed at runtime, because the assessment question is not whether pinning exists but whether the API behind it is safe once it is defeated.
- 05
Runtime manipulation
Frida and Objection are used to hook methods, defeat root and emulator detection, force branch outcomes, and test whether authorisation and business logic are enforced on the server or merely on the client.
- 06
API and backend testing
Every endpoint the app calls is tested in its own right for authentication, authorisation, IDOR, mass assignment and rate limiting. Most severe mobile findings are server-side findings reached through the app.
- 07
Exploitation and evidence
Confirmed issues are exploited to demonstrate real impact, and captured as reproducible steps with screenshots, request and response pairs and, where relevant, a proof-of-concept build.
- 08
Reporting and retest
Technical report and executive summary are issued together, followed by a retest after remediation that evidences closure rather than assuming it.
Approach to testing
- Testing runs against a real device and an emulator. Emulator-only testing misses hardware-backed keystore behaviour and device-specific storage; device-only testing makes instrumentation slower and some conditions harder to force.
- Both a rooted and an unrooted device are used. Rooted answers what an attacker can reach; unrooted answers what a normal user is exposed to. Reporting only the rooted result overstates risk and reporting only the unrooted result hides it.
- The most recent supported Android release and the app's declared minimum are both covered, because a control introduced in a newer API level may simply be absent on the oldest version the app still allows.
- Grey box by default. CSS is given a working build and test credentials for at least two roles, which is what makes authorisation testing between roles possible at all.
- Client-side controls are treated as findings only where they are the sole control. Where a server-side equivalent exists, the client-side bypass is recorded as context rather than as an issue in its own right.
Types of assessment
Black box
Production APK only, no credentials, no source. Mirrors an attacker who has downloaded the app from the store. Cheapest, and finds the least.
Grey box (default)
Working build, test credentials for multiple roles, API documentation. The best coverage per unit of effort, and the only way to test authorisation properly.
White box
Full source, build pipeline and backend access. Adds code-level review of cryptography and authorisation logic, and finds classes of issue that runtime testing cannot reach.
Resilience assessment
MASVS-RESILIENCE only: root detection, tamper detection, obfuscation and anti-instrumentation, measured as time-to-defeat rather than present or absent. Relevant where the app handles payments or licensed content.
Frameworks and standards
- OWASP MASVS 2.x
- The verification standard the assessment is scored against. MASVS-STORAGE, CRYPTO, AUTH, NETWORK, PLATFORM, CODE and RESILIENCE.
- OWASP MASTG
- The corresponding test procedures, so each MASVS control maps to a documented test rather than to tester preference.
- OWASP Mobile Top 10
- Used for executive communication, because it is the list a board is likely to have heard of.
- NIST SP 800-115
- The technical testing lifecycle the engagement runs on — planning, discovery, attack, reporting.
- PTES
- Execution standard for the exploitation and post-exploitation phases.
- CWE / CVSS v3.1
- Weakness classification and severity scoring, so findings are comparable between engagements and between vendors.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
MobSF
First-pass static and dynamic analysis, and the manifest baseline.
jadx / apktool
Decompilation to Java and Smali, and repackaging for patched builds.
Frida
Runtime hooking: pinning bypass, root detection bypass, method tracing and forced returns.
Objection
Frida-backed runtime exploration, keystore dumping and storage inspection without bespoke scripts.
Burp Suite Professional
Interception, request manipulation and API testing behind the app.
Drozer
IPC and exported component testing — activities, services, broadcast receivers and content providers.
adb
Device interaction, logcat review, backup extraction and file system access.
Semgrep
Rule-based source review where the engagement is white box.
apksigner / apkleaks
Signing scheme verification and secret discovery in packaged builds.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Storage (MASVS-STORAGE)
- Credentials, tokens or PII in shared preferences, SQLite or Realm
- Sensitive data written to external storage or cache
- android:allowBackup and the resulting extractable data set
- Keystore usage, key invalidation on biometric change, hardware backing
- Sensitive data in logcat, crash reports and third-party analytics
- Screenshot and task-switcher exposure of sensitive screens
Cryptography (MASVS-CRYPTO)
- Hardcoded keys, IVs or salts in the package
- ECB mode, static IVs, or a cipher chosen without authentication
- Custom or home-rolled cryptographic routines
- Random number generation from a predictable source
Authentication and authorisation (MASVS-AUTH)
- Session token lifetime, rotation and invalidation on logout
- Biometric and local authentication enforced server-side, not only in the UI
- Role separation tested between at least two accounts
- IDOR across every object identifier the app sends
- Step-up authentication for sensitive operations
Network (MASVS-NETWORK)
- TLS version, cipher suites and certificate validation
- Certificate pinning presence, and behaviour once defeated
- Cleartext traffic permitted by the network security configuration
- Sensitive data in URLs, query strings or headers
Platform interaction (MASVS-PLATFORM)
- Exported activities, services, receivers and content providers
- Deep link and app link handling, including unvalidated parameters
- WebView configuration: JavaScript, file access, addJavascriptInterface
- Permission set against actual functional need
- Clipboard, keyboard cache and accessibility service exposure
Code quality and resilience (MASVS-CODE, MASVS-RESILIENCE)
- Debuggable flag and debug artefacts in a production build
- Third-party SDK inventory and known vulnerable versions
- Root, emulator and hooking detection, measured as time-to-defeat
- Obfuscation coverage of security-relevant code paths
- Integrity verification and response to a repackaged build
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
A single tester works one hypothesis at a time. Parallel agents drive storage, network, IPC and runtime tracks simultaneously against the same build, so an artefact written during one interaction is caught while another track is still exercising the UI.
Transient state — a token written to cache during a failed login, a key held in memory only between two screens — exists for seconds and is routinely missed by sequential manual testing.
Deep link and exported component space is combinatorial. Enumerating it exhaustively is machine work; judging which results matter is not.
Scanner output and manual findings are cross-checked against each other, so a MobSF result nobody reproduced is not reported, and a manual finding a scanner missed is not assumed unique.
Every agent finding is validated by a human tester before it reaches the report. The swarm decides what to look at; it does not decide what is true. Anything not reproduced by hand is discarded.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Exploited, not scanned
A MobSF warning is a starting point. Findings reach the report only when they have been exploited on a device with reproducible steps and evidence attached. Where a control cannot be bypassed, that is stated as a positive result rather than quietly dropped.
Both reports, always
A technical assessment report for the engineers and an executive summary for the people funding the fix, issued together. Most vendors produce one and let the other audience make do with it.
Fixation, not a backlog
CSS works remediation with the development team — not only what is wrong, but the code-level change — and retests to evidence closure. An assessment that ends at the report has moved the risk to a spreadsheet, not reduced it.
The API is in scope
Many mobile assessments stop at the client. The endpoints behind the app are tested as first-class targets, because that is where the severe findings usually are.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Scope, build hash, device and OS matrix, and the exact test window
- Every finding with CVSS v3.1 vector, CWE reference and MASVS control
- Reproduction steps precise enough for a developer to follow without the tester present
- Evidence: screenshots, request and response pairs, decompiled excerpts, Frida scripts used
- Code-level remediation guidance, not 'implement certificate pinning'
- Controls tested that held, so the report records what is working
- Retest results appended against each original finding
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- Risk position in business terms — what an attacker could do to the business, not to the binary
- Severity distribution and how it compares to the previous assessment
- The three things that most need funding, with an indication of effort
- Regulatory and compliance exposure where relevant (DPDP Act, PCI DSS, RBI directions)
- Remediation timeline and the retest date
- One page. If it runs to five it will not be read by the person who signs the budget.
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A retail bank's customer app, roughly 2 million installs, tested grey box before a payments release.
- Finding
- Certificate pinning was implemented and defeated in under an hour with a standard Frida script. Behind it, the transaction endpoint authorised on a customer identifier supplied by the client, so changing it returned another customer's balance and transaction history.
- Recommendation
- Move authorisation server-side and derive the customer identifier from the session rather than the request body. Treat pinning as an attacker-delay control, not an access control, and stop relying on it.
- Outcome
- The IDOR was fixed before release and verified by retest. Pinning was retained but reclassified internally, which changed how the team reasoned about every subsequent client-side control.
A logistics operator's driver app, around 8,000 field devices on managed hardware.
- Finding
- The app wrote an OAuth refresh token to shared preferences in cleartext, and android:allowBackup was left enabled. Any device with USB debugging on — common across the fleet for support reasons — allowed the token to be extracted with adb and reused from anywhere.
- Recommendation
- Move token storage to the hardware-backed keystore, disable backup for the app, and bind refresh tokens to a device identifier so an extracted token is useless off the device.
- Outcome
- All three implemented. The device binding proved the more valuable of the changes, because it also closed token reuse from a lost or stolen handset.
A healthcare provider's patient app, tested against DPDP Act obligations.
- Finding
- An exported content provider, present to support an internal companion app that had been retired two years earlier, returned appointment records to any application installed on the device without a permission check.
- Recommendation
- Remove the provider. Where an export is genuinely required, protect it with a signature-level permission so only applications signed with the same key can reach it, and audit exported components on every release.
- Outcome
- Provider removed. The audit step was added to the release checklist, which subsequently caught a second exported service before it reached production.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.