Configuration Review
Configuration Review of Backup and Recovery
Every organisation has backups. Far fewer have backups a ransomware operator cannot reach, and fewer still have tested that a full restoration completes inside the time the business can survive. Modern ransomware targets the backup infrastructure first and deliberately, so this review treats the backup system as a primary attack target rather than as an IT utility.
Methodology
- 01
Architecture and coverage review
What is backed up, what is not, where copies live, and whether coverage matches the business impact assessment rather than what was easy to configure.
- 02
Immutability and isolation assessment
Whether at least one copy is genuinely immutable or offline, and whether the credentials that manage backups are reachable from the production domain an attacker would compromise first.
- 03
Backup infrastructure hardening
The backup servers, appliances and consoles assessed as high-value targets: authentication, MFA, network exposure, patch currency and administrative separation.
- 04
Retention and RPO/RTO validation
Retention against stated recovery point objectives, and whether the schedule actually achieves them for the systems that matter.
- 05
Restoration testing review
Evidence of restoration tests, their scope, and whether anyone has ever tested a full system recovery rather than a single file.
- 06
Ransomware scenario walkthrough
A tabletop against a realistic scenario: domain compromise, backup console access attempted, production encrypted. What survives, and how long recovery takes.
- 07
Reporting and retest
Technical report and executive summary together, with a gap list against 3-2-1-1-0, then a retest confirming closure.
Approach to testing
- The backup system is assessed as an attack target, not as infrastructure. The question is what a domain administrator could delete, because that is who the attacker will be.
- Immutability claims are verified rather than accepted. Vendor marketing and configured reality differ frequently, and a retention lock that an administrator can shorten is not immutability.
- Restoration is assessed on evidence. 'We test restores' without documented test records, timings and scope is treated as untested.
- RTO is measured against the business's stated tolerance, not against the backup product's specification sheet. Restoring 40 TB over a link that takes nine days is a finding regardless of the RPO.
- The review covers SaaS data explicitly — Microsoft 365, Google Workspace and Salesforce are not backed up by the provider in the sense most organisations assume.
Types of assessment
Architecture and configuration review (default)
Coverage, immutability, isolation and hardening. Read-only, no disruption.
Restoration validation
An observed restoration of a nominated system, timed end to end, so RTO is measured rather than estimated.
Ransomware resilience assessment
Scenario-driven: assume domain compromise, then assess what the attacker can reach and what survives.
SaaS data protection review
Focused on Microsoft 365, Google Workspace and similar, where the provider's retention is commonly mistaken for backup.
Frameworks and standards
- 3-2-1-1-0 rule
- Three copies, two media, one offsite, one immutable or offline, zero errors on verification. The structural baseline.
- NIST SP 800-34
- Contingency planning guidance, used for RPO/RTO structure and plan assessment.
- ISO/IEC 27031
- ICT readiness for business continuity, where the client maintains an ISMS.
- CIS Controls v8 Control 11
- Data recovery control assessment and maturity.
- Vendor hardening guides
- Veeam, Commvault, Rubrik, Veritas and equivalent — their own guidance on protecting the backup platform.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Vendor consoles (read-only)
Configuration, job history, retention and immutability settings direct from the platform.
BloodHound
Determining whether backup service accounts and consoles are reachable from a compromised domain identity.
Nmap
Confirming network isolation of backup infrastructure from production segments.
Custom job history parsers
Analysing months of job logs for silent failures and the systems that never complete successfully.
Storage vendor tooling
Verifying object lock, retention lock and WORM configuration at the storage layer rather than the application layer.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Coverage
- Systems in the business impact assessment against systems actually backed up
- SaaS data: Microsoft 365, Google Workspace, Salesforce and similar
- Endpoint and remote worker data
- Configuration and infrastructure-as-code, not only data
- Databases backed up consistently rather than as open files
- Systems excluded from backup, and whether anyone decided that
Immutability and isolation
- At least one copy immutable or physically offline
- Immutability verified, not assumed, including whether an administrator can shorten it
- Backup network segmentation from production
- Backup credentials separate from the production domain
- MFA on backup console access
- Air-gapped or logically isolated copy for critical systems
Infrastructure hardening
- Backup server patch currency
- Console exposure and authentication
- Service account privilege on production systems
- Vendor hardening guide compliance
- Backup repository access control and encryption
Recovery capability
- Documented RPO and RTO per system tier
- Evidence of restoration testing with dates, scope and timings
- Full system recovery tested, not only file-level
- Bare-metal and cloud recovery paths
- Recovery runbooks and whether they are current
- Restoration bandwidth against data volume
Monitoring and verification
- Job failure alerting and who receives it
- Silent failures: jobs reporting success with incomplete data
- Backup verification and integrity checking
- Retention compliance against policy and regulation
- Capacity trending against growth
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Job history across months and hundreds of systems hides the ones that have quietly failed. Correlating every job against the protected system inventory finds what aggregate success rates conceal.
Coverage gaps are found by diffing the asset inventory against the backup inventory, which is mechanical comparison and is almost never done by hand at full scale.
Whether a backup credential is reachable from a compromised domain account is a graph question over the directory, not a configuration setting.
Immutability claims are verified per repository rather than per platform, because one misconfigured repository is where the attacker will go.
Every agent finding is validated by a human reviewer. The assessment is read-only: no job is modified, no restore triggered, and nothing deleted. Restoration validation, where in scope, is performed by the client's own team with CSS observing.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Assessed as an attack target
Most backup reviews check schedules and retention. CSS asks what a domain administrator can delete, because that is who the attacker will be by the time backups matter.
Immutability verified, not accepted
Retention lock is tested against whether an administrator can shorten or remove it. Configured immutability and real immutability differ more often than vendors suggest.
RTO measured, not quoted
Where restoration validation is in scope, recovery is timed end to end against real data volumes and real link speeds, rather than taken from a specification sheet.
Both reports, always
Technical report and executive summary together, plus a ransomware scenario walkthrough the board can follow.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Scope: platforms, repositories, protected systems and the review date
- Coverage analysis: protected systems against the asset and business impact inventory
- Immutability and isolation assessment per repository, with verification evidence
- Backup infrastructure hardening findings against vendor guidance
- Job history analysis including silent and recurring failures
- RPO and RTO assessment per system tier, with measured figures where validated
- Ransomware scenario outcome: what survives and how long recovery takes
- Retest results appended against each original finding
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- If ransomware encrypted production tonight, what the organisation could restore and how long it would take
- Position against 3-2-1-1-0, stated as which elements are met and which are not
- The three changes that most improve survivability
- Regulatory and insurance position, since cyber insurers increasingly require immutability evidence
- Remediation timeline and retest date
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A distribution business of roughly 900 staff, running a well-regarded backup platform.
- Finding
- Backups ran nightly with 30-day retention to a disk repository, replicated offsite. The backup service account was a domain administrator, the console authenticated against the same domain with no MFA, and the repository was mounted as a share reachable from production. A domain compromise would have reached every copy.
- Recommendation
- Move the backup platform to a separate identity boundary, enforce MFA on the console, convert the repository to a hardened immutable target, and remove domain administrator from the service account.
- Outcome
- Identity separation and MFA completed in three weeks; the immutable repository followed in the next budget cycle. The retest confirmed a simulated domain administrator could no longer delete or alter any backup copy.
A professional services firm that had migrated to Microsoft 365 four years earlier.
- Finding
- No backup existed for Microsoft 365. The assumption was that Microsoft handled it. Retention policies covered accidental deletion for 30 to 93 days depending on the workload, and nothing covered a malicious administrator, a ransomware event in OneDrive, or a departed user's mailbox after licence removal.
- Recommendation
- Deploy a third-party Microsoft 365 backup with independent retention and immutability, covering Exchange, SharePoint, OneDrive and Teams.
- Outcome
- Backup deployed over a month. Within the year it was used to recover a SharePoint site a user had deleted seven months earlier, which was well outside Microsoft's retention.
A manufacturer with 40 TB of production data and a stated four-hour RTO.
- Finding
- Restoration validation timed a full recovery of the primary ERP system at 71 hours, against a stated RTO of four. The bottleneck was a 200 Mbps link to the offsite repository, which nobody had measured against the data volume. Restore tests had only ever covered individual files.
- Recommendation
- Hold a local copy of tier-one systems for fast recovery, keep the offsite copy for disaster scenarios, and restate the RTO honestly for systems where four hours is not achievable.
- Outcome
- A local immutable copy was added for the top six systems, bringing measured RTO to under five hours. The honest restatement of RTO for lower tiers was, per the client, the more valuable finding.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.