VAPT
VAPT of Web Applications
Most web application scanners find the same injection and header issues on every target. The findings that actually cost an organisation money are authorisation flaws and business logic abuse — a user reaching another user's data, a workflow completed out of order, a price modified in transit. Those require an authenticated tester who understands what the application is for, which is what this assessment is built around.
Methodology
- 01
Reconnaissance and mapping
Full crawl of the authenticated and unauthenticated surface, technology fingerprinting, endpoint and parameter inventory, and identification of every role and state the application supports.
- 02
Threat modelling against the business
What the application is worth, what an attacker would want from it, and which workflows would hurt if abused. The test plan is written against this rather than against a generic checklist.
- 03
Automated discovery
Authenticated scanning to clear the known classes quickly, so tester time goes to what a scanner cannot reach. Scanner output is triage input, never a finding.
- 04
Authentication and session testing
Registration, login, MFA, password reset, session fixation, concurrent sessions, token entropy and invalidation, and every path that issues or accepts credentials.
- 05
Authorisation testing
Every object identifier and every function tested horizontally between peer accounts and vertically between privilege levels. This is the single highest-yield phase and it needs at least two accounts per role.
- 06
Business logic abuse
Workflows driven out of sequence, steps skipped, quantities and prices manipulated, race conditions on limited resources, and negative or boundary values where the application assumes good faith.
- 07
Injection and client-side testing
SQL, NoSQL, command, template and LDAP injection; XSS in all three contexts; SSRF, XXE, deserialisation, and file upload handling.
- 08
Exploitation and impact
Confirmed issues chained to demonstrate real business impact — not 'reflected XSS exists' but 'this is how an account is taken over'.
- 09
Reporting and retest
Technical report and executive summary together, then a retest that evidences closure.
Approach to testing
- Grey box by default, with at least two accounts per role. Authorisation testing between roles is impossible without them, and it is where most severe findings live.
- Testing runs against a staging environment that mirrors production. Where only production is available, destructive tests are agreed in writing beforehand and run in a defined window.
- Rate limiting and WAF behaviour are tested, then testing continues with the tester allowlisted. Both answers matter: what an attacker faces, and what is actually behind the protection.
- Findings that a scanner reported but a tester could not reproduce are discarded, not reported with a caveat.
- Chained findings are reported as one issue with its real severity, rather than three medium findings that together amount to account takeover.
Types of assessment
Black box
No credentials, no documentation. Mirrors an internet-based attacker with no foothold. Useful as a perimeter check, weak as an application assessment.
Grey box (default)
Credentials for every role, plus documentation. The best coverage per unit of effort and the only way to test authorisation properly.
White box
Source and architecture access alongside runtime testing. Reaches authorisation logic and cryptographic implementation that black box testing cannot.
Continuous / retest cycle
Assessment repeated each release against a fixed scope, so regression is caught rather than discovered at the next annual test.
Frameworks and standards
- OWASP ASVS 4.x
- The verification standard the application is scored against, at the level agreed for its risk tier.
- OWASP Top 10
- Used in the executive summary, because it is the taxonomy a board recognises.
- OWASP WSTG
- Test procedures, so each control maps to a documented test rather than tester preference.
- PTES
- Execution standard across exploitation and post-exploitation.
- NIST SP 800-115
- The overall technical testing lifecycle.
- CWE / CVSS v3.1
- Classification and severity scoring, comparable across engagements and vendors.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Burp Suite Professional
The primary tool: interception, repeater, intruder, and custom extensions for application-specific logic.
OWASP ZAP
Secondary automated pass, for cross-checking Burp's scanner rather than replacing it.
sqlmap
Confirming and exploiting injection points already identified by hand.
ffuf / feroxbuster
Content and parameter discovery across the authenticated surface.
Nuclei
Templated checks for known CVEs in identified components and versions.
Semgrep
Source review where the engagement is white box.
Autorize / Authz plugins
Systematic horizontal and vertical authorisation comparison across every request.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Authentication
- Credential policy, lockout and anti-automation
- MFA enrolment, bypass and recovery paths
- Password reset token entropy, expiry and single use
- Session token generation, rotation and invalidation
- Concurrent session handling and logout completeness
Authorisation
- Horizontal access between peer accounts on every object identifier
- Vertical escalation between privilege levels
- Forced browsing to functions absent from the UI
- Authorisation on the API as well as the rendered page
- Multi-tenant isolation where the application is shared
Business logic
- Workflow steps skipped or completed out of order
- Quantity, price and discount manipulation in transit
- Race conditions on limited or single-use resources
- Negative, zero and boundary values
- Replay of one-time operations
Injection and input handling
- SQL, NoSQL, command, template and LDAP injection
- Cross-site scripting in HTML, attribute and JavaScript contexts
- SSRF against internal services and cloud metadata
- XXE and unsafe deserialisation
- File upload type, size and content validation
Configuration and exposure
- Security headers and cookie attributes
- TLS configuration and certificate validity
- Verbose errors, stack traces and debug endpoints
- Exposed administrative interfaces and backup files
- Component inventory against known CVEs
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Authorisation testing is combinatorial: every object identifier against every role. A tester samples it; agents enumerate it exhaustively and flag only the differences that matter.
Race conditions need many requests fired within milliseconds of each other. That is machine work, and a manual tester will rarely catch a race that only opens for 40ms.
Long multi-step workflows have state that a single tester loses track of. Parallel agents hold several states at once and try transitions between them.
Scanner findings and manual findings are cross-checked against each other, so nothing unreproduced reaches the report and nothing a scanner missed is assumed unique.
Every agent finding is validated by a human tester before it reaches the report. The swarm decides what to look at; it does not decide what is true.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Exploited, not scanned
Findings are proven by exploitation with reproducible steps and evidence. A scanner alert that cannot be reproduced does not appear in the report at all.
Both reports, always
Technical assessment report and executive summary issued together, written separately for two audiences rather than being one document at two lengths.
Fixation, not a backlog
CSS works remediation with the development team and retests to evidence closure. An assessment that ends at the report has moved risk to a spreadsheet.
Chained, not itemised
Three medium findings that together produce account takeover are reported as one critical with the chain shown, rather than as three mediums a triage process will deprioritise.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Scope, environment, credentials used and the exact test window
- Every finding with CVSS v3.1 vector, CWE reference and ASVS control
- Reproduction steps precise enough to follow without the tester present
- Evidence: request and response pairs, screenshots, proof-of-concept payloads
- Attack chains shown end to end where findings combine
- Controls tested that held, so the report records what is working
- Retest results appended against each original finding
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- Risk position in business terms — what an attacker could do to the business
- Severity distribution and movement since the previous assessment
- The three things that most need funding, with effort indicated
- Regulatory exposure where relevant (DPDP Act, PCI DSS, RBI, SEBI)
- Remediation timeline and retest date
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A B2B SaaS platform serving around 400 corporate tenants.
- Finding
- Tenant identifier was accepted from a request header and trusted. Changing it returned another tenant's user list, invoices and uploaded documents, from any authenticated account on any tenant.
- Recommendation
- Derive tenant context from the session server-side and never from client-supplied input. Add a tenant-isolation test to the regression suite so the class cannot return.
- Outcome
- Fixed within a week and verified by retest. The regression test caught a similar issue on a new endpoint four months later.
An insurance portal handling policy purchase and claims.
- Finding
- The premium was calculated client-side and submitted with the order. Intercepting and reducing it produced a valid policy at the attacker's price, because the server re-validated the policy but not the amount.
- Recommendation
- Recalculate every monetary value server-side at the point of transaction and treat client-submitted amounts as display data only.
- Outcome
- Recalculation moved server-side. A reconciliation check against historical orders found no prior abuse, which is what made the finding closable rather than an incident.
A public sector grievance portal, roughly 90,000 registered citizens.
- Finding
- Password reset tokens were derived from a timestamp and were six digits. With no rate limit on the reset endpoint, any account could be taken over in under ten minutes of automated guessing.
- Recommendation
- Cryptographically random tokens of at least 128 bits, single use, 15-minute expiry, with rate limiting and alerting on the reset endpoint.
- Outcome
- Implemented before public disclosure was necessary. Rate limiting was extended to authentication endpoints generally, which also reduced credential-stuffing traffic.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.