Skip to content
CyberSmithSECURE
Under Attack

VAPT

VAPT of APIs

An API has no user interface to constrain what a caller does, which removes the accidental protection a web front end provides. Almost every severe API finding is an authorisation failure — an object identifier that is trusted, a function that checks authentication but not permission, a field that should never have been writable. This assessment is built around those, because scanners are poor at them and they are what gets exploited.

Methodology

  1. 01

    Specification and surface mapping

    OpenAPI, GraphQL schema or WSDL ingested where available; where it is not, the surface is discovered from client traffic and fuzzing. Every endpoint, method, parameter and object type is inventoried.

  2. 02

    Authentication model review

    How tokens are issued, scoped, refreshed and revoked. OAuth flows, JWT signing and validation, API key handling, and whether any endpoint is reachable unauthenticated.

  3. 03

    Object-level authorisation testing

    Every identifier in every request substituted across peer accounts. This is the highest-yield phase and is run exhaustively rather than sampled.

  4. 04

    Function-level authorisation testing

    Every endpoint called from every role, including endpoints absent from the documentation and from the client application.

  5. 05

    Property-level testing

    Mass assignment and excessive data exposure: fields accepted that should be server-controlled, and fields returned that the caller should never see.

  6. 06

    Business flow abuse

    Automated abuse of flows designed for humans — bulk enumeration, repeated redemption, and operations replayed beyond their intended limit.

  7. 07

    Injection and resource testing

    Injection where input reaches a query or command, SSRF where the API fetches a caller-supplied URL, and resource exhaustion through unbounded queries or deeply nested GraphQL.

  8. 08

    Reporting and retest

    Technical report and executive summary together, then a retest evidencing closure.

Approach to testing

  • At least two accounts per role. Object-level authorisation cannot be tested with one account, and it is where the severe findings are.
  • Undocumented endpoints are treated as in scope. An endpoint absent from the specification is often absent from the security review too.
  • GraphQL is tested for introspection, query depth and cost, batching abuse and per-resolver authorisation — a single authorisation check at the query root is a common and severe error.
  • Rate limiting is tested per endpoint rather than globally. A limit on login that does not apply to password reset is not a limit.
  • Findings are reported against the API itself, not against a client that happens to call it, because the same interface is usually reachable from several clients.

Types of assessment

Black box

No specification, no credentials. Surface discovered from observed traffic. Realistic for a public API, weak for anything authenticated.

Grey box (default)

Specification plus credentials for every role. The only way to test object and function level authorisation properly.

White box

Source access alongside runtime testing, which reaches authorisation logic and reveals endpoints never exposed in any specification.

Continuous / pipeline

Contract and authorisation tests run per deployment, so a new endpoint without an authorisation check fails the build rather than the next annual assessment.

Frameworks and standards

OWASP API Security Top 10 (2023)
The primary taxonomy — BOLA, broken authentication, BOPLA, unrestricted resource consumption, BFLA and the rest.
OWASP ASVS 4.x
Verification controls where the API backs a web or mobile application.
OWASP WSTG
Test procedures adapted for interfaces without a rendered UI.
NIST SP 800-115 / PTES
Testing lifecycle and exploitation standard.
CWE / CVSS v3.1
Classification and severity scoring.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

Burp Suite Professional

Primary interception, repeater and intruder work, with extensions for authorisation comparison.

Postman / Insomnia

Specification-driven request construction and role-by-role replay.

Autorize

Systematic horizontal and vertical authorisation comparison across every observed request.

GraphQL Voyager / InQL

Schema exploration, introspection abuse and per-resolver authorisation mapping.

ffuf

Endpoint and parameter discovery beyond the documented surface.

Nuclei

Templated checks for known issues in identified API gateways and frameworks.

jwt_tool

JWT signing, algorithm confusion and claim manipulation testing.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Object-level authorisation (BOLA)

  • Every identifier substituted across peer accounts
  • Sequential and guessable identifiers enumerated
  • Nested and relational objects reached indirectly
  • Identifiers in headers, cookies and body as well as the path

Authentication

  • JWT algorithm, signature validation and claim trust
  • Token expiry, refresh and revocation on logout
  • OAuth flow, redirect URI validation and scope enforcement
  • API key scoping and rotation
  • Endpoints reachable with no authentication at all

Property-level authorisation (BOPLA)

  • Mass assignment of server-controlled fields
  • Excessive data returned beyond what the caller needs
  • Internal fields exposed in error responses
  • Field-level permissions on update operations

Function-level authorisation (BFLA)

  • Every endpoint called from every role
  • Administrative endpoints reachable by standard users
  • HTTP method substitution on restricted endpoints
  • Undocumented and deprecated endpoints still live

Resource consumption

  • Rate limiting per endpoint, not only globally
  • GraphQL query depth, complexity and batching limits
  • Pagination limits and unbounded result sets
  • File upload size and processing limits

Inventory and configuration

  • Old API versions still reachable
  • Non-production environments exposed to the internet
  • Verbose errors disclosing stack traces or internals
  • CORS configuration and credential exposure
  • TLS configuration on every published host

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • BOLA testing is the product of every identifier and every role. A tester samples it; agents enumerate it exhaustively and surface only the responses that differ in a way that matters.

  • Undocumented and deprecated endpoints are found by patient enumeration rather than insight, which is precisely what parallel agents are good at.

  • GraphQL resolver authorisation has to be checked per field, not per query. The combination space is large enough that manual testing reliably samples rather than covers it.

  • Agent output is cross-checked against manual findings so nothing unreproduced is reported and nothing a tool missed is assumed unique.

Every agent finding is validated by a human tester before it reaches the report. The swarm decides what to look at; it does not decide what is true.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Exploited, not scanned

API scanners are particularly weak on authorisation because they do not know which objects belong to whom. Findings here are demonstrated with two accounts and captured request pairs.

Both reports, always

Technical report and executive summary together, written for two audiences.

Fixation, not a backlog

Remediation worked with the engineering team, then retested to evidence closure.

Exhaustive, not sampled, on authorisation

Most vendors sample object-level authorisation because doing it exhaustively by hand is impractical. CSS enumerates it, which is the single most common source of severe API findings.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Scope, specification version, environments and accounts used, and the test window
  • Every finding with CVSS v3.1 vector, CWE reference and API Top 10 category
  • Reproduction as complete request and response pairs for both accounts involved
  • Evidence of impact — the data actually returned, redacted
  • Endpoint inventory including undocumented endpoints discovered
  • Controls tested that held
  • Retest results appended against each original finding

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • Risk position in business terms
  • Severity distribution and movement since the last assessment
  • The three things that most need funding
  • Regulatory exposure where the API carries regulated data
  • Remediation timeline and retest date
  • One page

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A payments aggregator's partner API, around 120 integrated merchants.

Finding
Merchant identifier was taken from the request body rather than the authenticated key. Any merchant could retrieve settlement reports, transaction records and payout details for every other merchant on the platform.
Recommendation
Derive merchant context from the API key server-side. Add an integration test asserting that a key cannot reach another merchant's objects, and run it per deployment.
Outcome
Fixed in eight days and verified by retest. The integration test is now a release gate, which is what stopped the class returning when a new reporting endpoint shipped.

A logistics platform's GraphQL API used by web, mobile and partner clients.

Finding
Authorisation was enforced once at the query root. Nested resolvers inherited no check, so a shipment query could traverse to the customer object and return contact details and addresses for accounts the caller had no relationship with.
Recommendation
Enforce authorisation at every resolver rather than at the entry point, disable introspection in production, and apply query depth and cost limits.
Outcome
Per-resolver authorisation implemented. Depth limiting also removed a denial-of-service path that had not been reported, because deeply nested queries had been timing out in production for months.

A healthtech provider's REST API backing a patient portal.

Finding
A deprecated v1 API remained live alongside v2. It had no rate limiting and an older authorisation model, and returned full patient records where v2 returned a filtered view. Nothing referenced it except a retired client.
Recommendation
Decommission v1. Where a version must stay live for compatibility, bring it under the same authorisation model and include it in scope for every assessment.
Outcome
v1 withdrawn after a 30-day partner notice. An inventory step was added to the release process so a version cannot be silently left running.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.