Vulnerability Fixation
Verification Retest
A finding marked closed without verification is an assumption. In practice a meaningful share of fixes either do not close the issue, close it in one place and not others, or introduce a new problem. Retest exists to establish which, and to produce the evidence a customer, auditor or regulator will ask for.
Methodology
- 01
Scope agreement
Which findings are ready for retest, what change was made for each, and where. Retesting a finding whose fix has not deployed to the tested environment wastes both parties' time.
- 02
Fix review
The change itself examined before testing, where source or configuration is available. A fix that is obviously incomplete can be returned without a test cycle.
- 03
Original reproduction attempt
The exact original reproduction steps re-run first. If they still work, the fix has failed and no further analysis is needed.
- 04
Variant testing
Where the original no longer works, variants are attempted — the fix may address one input or path while leaving others open, which is the most common form of partial remediation.
- 05
Instance coverage check
For findings affecting multiple hosts, endpoints or code paths, every instance is checked rather than the one originally reported.
- 06
Regression assessment
Whether the fix has introduced a new issue, which happens often enough that omitting this step is negligent.
- 07
Closure determination
Each finding marked closed, partially closed or still open, with evidence for the determination.
- 08
Reporting
A retest report appended to the original assessment, so the two read as one record rather than as separate documents.
Approach to testing
- The original reproduction steps are run first and exactly. Testing a fix by a different route is how a partially closed finding gets recorded as closed.
- Variants are attempted on every finding that passes the original test, because fixing the reported input while leaving three others open is the most common partial remediation.
- Every instance is checked for multi-instance findings. A fix deployed to nine of eleven servers is not closed, and it is frequently reported as closed.
- Regression is assessed. A fix that closes a finding and opens a new one has not improved anything.
- Partial closure is reported as partial. Rounding up to closed is how the same finding reappears in next year's assessment with everyone surprised.
Types of assessment
Standard retest (default)
All remediated findings from a prior assessment, tested and evidenced. Usually included in the original engagement.
Rolling retest
Findings tested as they are fixed rather than in a single cycle at the end. Suits continuous remediation programmes and gives faster feedback.
Targeted retest
A specific finding or small set, usually because a customer or regulator has asked for evidence on that item.
Full re-assessment
Where remediation was extensive or architectural, a fresh assessment rather than a retest, because the attack surface has genuinely changed.
Frameworks and standards
- Original assessment methodology
- Retest follows whatever standard the original engagement used — OWASP ASVS, MASVS, PTES — so results are directly comparable.
- CVSS v3.1
- Residual severity where a finding is partially closed.
- PCI DSS / ISO 27001 evidence requirements
- Where the retest exists to satisfy an audit obligation, the evidence is prepared to that standard.
- NIST SP 800-115
- Testing lifecycle consistency with the original engagement.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Original engagement artefacts
The same scripts, payloads and requests used in the original test, so the comparison is exact.
Burp Suite
Application finding reproduction and variant testing.
Nuclei
Templated verification across every affected host for multi-instance findings.
Frida / Objection
Mobile finding reproduction on device.
Client ticketing system
Closure status written back to where the client's teams actually track work.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Preparation
- Findings confirmed as deployed to the environment being tested
- Change description obtained for each finding
- Original reproduction steps and artefacts retrieved
- Test credentials and access still valid
- Environment matches the original test target
Verification
- Original reproduction steps re-run exactly
- Variant inputs and alternative paths attempted
- Every affected instance checked, not a sample
- Fix reviewed at source or configuration level where available
- Compensating controls verified where the underlying issue remains
Regression
- New issues introduced by the fix
- Adjacent functionality affected
- Error handling changes exposing new information
- Performance or availability impact from the change
Determination
- Closed, partially closed or open recorded explicitly
- Residual severity scored for partial closures
- Evidence captured for each determination
- Findings the client has accepted as risk recorded with owner and review date
Reporting
- Retest results appended to the original report
- Closure evidence suitable for audit
- Register updated in the client's own system
- Outstanding findings carried forward with revised timelines
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Multi-instance verification across hundreds of hosts or endpoints is exhaustive checking, and sampling is precisely how a partially deployed fix gets recorded as complete.
Variant generation around a fixed input — encoding, alternative parameters, adjacent endpoints — explores far more of the space than a tester re-running the original payload.
Regression checking across adjacent functionality is breadth work that retest budgets rarely stretch to manually.
Comparing the fix against every other instance of the same root cause in the codebase finds the call sites the developer missed.
Closure determination is always made by a human tester. Agents verify at scale and generate variants; whether a finding is closed is a judgement that carries audit consequences and is not delegated. A wrongly closed finding stops being looked at, which makes it worse than an open one.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Partial closure is reported as partial
The uncomfortable answer is the useful one. Rounding a partially fixed finding up to closed is how the same issue reappears next year with everyone surprised.
Every instance, not a sample
Multi-instance findings are verified across every affected host or endpoint, because a fix deployed to most of them is a fix that has not closed the finding.
Regression assessed
Fixes introduce new issues often enough that not checking is negligent. Retest covers what the change broke as well as what it fixed.
Evidence suitable for audit
Closure is evidenced rather than asserted, in a form a customer, auditor or regulator will accept without further questions.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Retest scope, environment and date, with reference to the original engagement
- Per finding: original severity, change made, test performed, determination
- Evidence for each determination, including failed reproduction attempts
- Partially closed findings with residual severity and what remains
- Variants that still succeed where the original no longer does
- Instance coverage for multi-instance findings
- Regression findings introduced by remediation
- Updated overall risk position against the original assessment
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- How many findings are closed, partially closed and still open
- Movement in overall risk position since the original assessment
- Fixes that did not work, and why that matters more than the count suggests
- Outstanding items with revised timelines
- Whether the evidence package meets the audit or customer requirement it was produced for
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A fintech retesting 62 findings from a web application assessment.
- Finding
- 41 were fully closed. 14 were partially closed — mostly input validation applied to the reported parameter while adjacent parameters on the same endpoint remained vulnerable. 5 had not been fixed at all despite being marked complete. 2 regressions were introduced, including a new authorisation gap created by a refactor intended to fix a different finding.
- Recommendation
- Return the partial fixes with the specific remaining variants, and treat the regression as a new critical finding rather than folding it into the original.
- Outcome
- The second retest closed all but three. The client changed its internal process to require the developer to test the variant set rather than only the reported case.
A manufacturer retesting infrastructure findings across 340 hosts.
- Finding
- A configuration fix had been deployed by a script that failed silently on hosts running an older operating system. Verification across every affected host found 47 still vulnerable, against an internal record showing 100% deployment.
- Recommendation
- Re-run deployment with verification per host, and add a post-deployment check to the change process so completion is measured rather than assumed.
- Outcome
- All 47 remediated within two weeks. The verification step has since caught two further partial deployments.
A healthcare provider preparing evidence for a customer security assessment.
- Finding
- The client needed evidenced closure on 18 findings. Retest confirmed 16 closed. Two had been accepted as risk internally but recorded as closed, which the customer's questionnaire would have treated as a misrepresentation.
- Recommendation
- Record the two as accepted risks with owner, rationale and review date, and present them to the customer as such rather than as closed.
- Outcome
- The customer accepted both with the documented rationale. The client's view was that the honest categorisation was less damaging than the discovery would have been.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.