Skip to content
CyberSmithSECURE
Under Attack

Vulnerability Fixation

Verification Retest

A finding marked closed without verification is an assumption. In practice a meaningful share of fixes either do not close the issue, close it in one place and not others, or introduce a new problem. Retest exists to establish which, and to produce the evidence a customer, auditor or regulator will ask for.

Methodology

  1. 01

    Scope agreement

    Which findings are ready for retest, what change was made for each, and where. Retesting a finding whose fix has not deployed to the tested environment wastes both parties' time.

  2. 02

    Fix review

    The change itself examined before testing, where source or configuration is available. A fix that is obviously incomplete can be returned without a test cycle.

  3. 03

    Original reproduction attempt

    The exact original reproduction steps re-run first. If they still work, the fix has failed and no further analysis is needed.

  4. 04

    Variant testing

    Where the original no longer works, variants are attempted — the fix may address one input or path while leaving others open, which is the most common form of partial remediation.

  5. 05

    Instance coverage check

    For findings affecting multiple hosts, endpoints or code paths, every instance is checked rather than the one originally reported.

  6. 06

    Regression assessment

    Whether the fix has introduced a new issue, which happens often enough that omitting this step is negligent.

  7. 07

    Closure determination

    Each finding marked closed, partially closed or still open, with evidence for the determination.

  8. 08

    Reporting

    A retest report appended to the original assessment, so the two read as one record rather than as separate documents.

Approach to testing

  • The original reproduction steps are run first and exactly. Testing a fix by a different route is how a partially closed finding gets recorded as closed.
  • Variants are attempted on every finding that passes the original test, because fixing the reported input while leaving three others open is the most common partial remediation.
  • Every instance is checked for multi-instance findings. A fix deployed to nine of eleven servers is not closed, and it is frequently reported as closed.
  • Regression is assessed. A fix that closes a finding and opens a new one has not improved anything.
  • Partial closure is reported as partial. Rounding up to closed is how the same finding reappears in next year's assessment with everyone surprised.

Types of assessment

Standard retest (default)

All remediated findings from a prior assessment, tested and evidenced. Usually included in the original engagement.

Rolling retest

Findings tested as they are fixed rather than in a single cycle at the end. Suits continuous remediation programmes and gives faster feedback.

Targeted retest

A specific finding or small set, usually because a customer or regulator has asked for evidence on that item.

Full re-assessment

Where remediation was extensive or architectural, a fresh assessment rather than a retest, because the attack surface has genuinely changed.

Frameworks and standards

Original assessment methodology
Retest follows whatever standard the original engagement used — OWASP ASVS, MASVS, PTES — so results are directly comparable.
CVSS v3.1
Residual severity where a finding is partially closed.
PCI DSS / ISO 27001 evidence requirements
Where the retest exists to satisfy an audit obligation, the evidence is prepared to that standard.
NIST SP 800-115
Testing lifecycle consistency with the original engagement.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

Original engagement artefacts

The same scripts, payloads and requests used in the original test, so the comparison is exact.

Burp Suite

Application finding reproduction and variant testing.

Nuclei

Templated verification across every affected host for multi-instance findings.

Frida / Objection

Mobile finding reproduction on device.

Client ticketing system

Closure status written back to where the client's teams actually track work.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Preparation

  • Findings confirmed as deployed to the environment being tested
  • Change description obtained for each finding
  • Original reproduction steps and artefacts retrieved
  • Test credentials and access still valid
  • Environment matches the original test target

Verification

  • Original reproduction steps re-run exactly
  • Variant inputs and alternative paths attempted
  • Every affected instance checked, not a sample
  • Fix reviewed at source or configuration level where available
  • Compensating controls verified where the underlying issue remains

Regression

  • New issues introduced by the fix
  • Adjacent functionality affected
  • Error handling changes exposing new information
  • Performance or availability impact from the change

Determination

  • Closed, partially closed or open recorded explicitly
  • Residual severity scored for partial closures
  • Evidence captured for each determination
  • Findings the client has accepted as risk recorded with owner and review date

Reporting

  • Retest results appended to the original report
  • Closure evidence suitable for audit
  • Register updated in the client's own system
  • Outstanding findings carried forward with revised timelines

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • Multi-instance verification across hundreds of hosts or endpoints is exhaustive checking, and sampling is precisely how a partially deployed fix gets recorded as complete.

  • Variant generation around a fixed input — encoding, alternative parameters, adjacent endpoints — explores far more of the space than a tester re-running the original payload.

  • Regression checking across adjacent functionality is breadth work that retest budgets rarely stretch to manually.

  • Comparing the fix against every other instance of the same root cause in the codebase finds the call sites the developer missed.

Closure determination is always made by a human tester. Agents verify at scale and generate variants; whether a finding is closed is a judgement that carries audit consequences and is not delegated. A wrongly closed finding stops being looked at, which makes it worse than an open one.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Partial closure is reported as partial

The uncomfortable answer is the useful one. Rounding a partially fixed finding up to closed is how the same issue reappears next year with everyone surprised.

Every instance, not a sample

Multi-instance findings are verified across every affected host or endpoint, because a fix deployed to most of them is a fix that has not closed the finding.

Regression assessed

Fixes introduce new issues often enough that not checking is negligent. Retest covers what the change broke as well as what it fixed.

Evidence suitable for audit

Closure is evidenced rather than asserted, in a form a customer, auditor or regulator will accept without further questions.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Retest scope, environment and date, with reference to the original engagement
  • Per finding: original severity, change made, test performed, determination
  • Evidence for each determination, including failed reproduction attempts
  • Partially closed findings with residual severity and what remains
  • Variants that still succeed where the original no longer does
  • Instance coverage for multi-instance findings
  • Regression findings introduced by remediation
  • Updated overall risk position against the original assessment

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • How many findings are closed, partially closed and still open
  • Movement in overall risk position since the original assessment
  • Fixes that did not work, and why that matters more than the count suggests
  • Outstanding items with revised timelines
  • Whether the evidence package meets the audit or customer requirement it was produced for
  • One page

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A fintech retesting 62 findings from a web application assessment.

Finding
41 were fully closed. 14 were partially closed — mostly input validation applied to the reported parameter while adjacent parameters on the same endpoint remained vulnerable. 5 had not been fixed at all despite being marked complete. 2 regressions were introduced, including a new authorisation gap created by a refactor intended to fix a different finding.
Recommendation
Return the partial fixes with the specific remaining variants, and treat the regression as a new critical finding rather than folding it into the original.
Outcome
The second retest closed all but three. The client changed its internal process to require the developer to test the variant set rather than only the reported case.

A manufacturer retesting infrastructure findings across 340 hosts.

Finding
A configuration fix had been deployed by a script that failed silently on hosts running an older operating system. Verification across every affected host found 47 still vulnerable, against an internal record showing 100% deployment.
Recommendation
Re-run deployment with verification per host, and add a post-deployment check to the change process so completion is measured rather than assumed.
Outcome
All 47 remediated within two weeks. The verification step has since caught two further partial deployments.

A healthcare provider preparing evidence for a customer security assessment.

Finding
The client needed evidenced closure on 18 findings. Retest confirmed 16 closed. Two had been accepted as risk internally but recorded as closed, which the customer's questionnaire would have treated as a misrepresentation.
Recommendation
Record the two as accepted risks with owner, rationale and review date, and present them to the customer as such rather than as closed.
Outcome
The customer accepted both with the documented rationale. The client's view was that the honest categorisation was less damaging than the discovery would have been.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.