Skip to content
CyberSmithSECURE
Under Attack

Configuration Review

Configuration Review of Backup and Recovery

Every organisation has backups. Far fewer have backups a ransomware operator cannot reach, and fewer still have tested that a full restoration completes inside the time the business can survive. Modern ransomware targets the backup infrastructure first and deliberately, so this review treats the backup system as a primary attack target rather than as an IT utility.

Methodology

  1. 01

    Architecture and coverage review

    What is backed up, what is not, where copies live, and whether coverage matches the business impact assessment rather than what was easy to configure.

  2. 02

    Immutability and isolation assessment

    Whether at least one copy is genuinely immutable or offline, and whether the credentials that manage backups are reachable from the production domain an attacker would compromise first.

  3. 03

    Backup infrastructure hardening

    The backup servers, appliances and consoles assessed as high-value targets: authentication, MFA, network exposure, patch currency and administrative separation.

  4. 04

    Retention and RPO/RTO validation

    Retention against stated recovery point objectives, and whether the schedule actually achieves them for the systems that matter.

  5. 05

    Restoration testing review

    Evidence of restoration tests, their scope, and whether anyone has ever tested a full system recovery rather than a single file.

  6. 06

    Ransomware scenario walkthrough

    A tabletop against a realistic scenario: domain compromise, backup console access attempted, production encrypted. What survives, and how long recovery takes.

  7. 07

    Reporting and retest

    Technical report and executive summary together, with a gap list against 3-2-1-1-0, then a retest confirming closure.

Approach to testing

  • The backup system is assessed as an attack target, not as infrastructure. The question is what a domain administrator could delete, because that is who the attacker will be.
  • Immutability claims are verified rather than accepted. Vendor marketing and configured reality differ frequently, and a retention lock that an administrator can shorten is not immutability.
  • Restoration is assessed on evidence. 'We test restores' without documented test records, timings and scope is treated as untested.
  • RTO is measured against the business's stated tolerance, not against the backup product's specification sheet. Restoring 40 TB over a link that takes nine days is a finding regardless of the RPO.
  • The review covers SaaS data explicitly — Microsoft 365, Google Workspace and Salesforce are not backed up by the provider in the sense most organisations assume.

Types of assessment

Architecture and configuration review (default)

Coverage, immutability, isolation and hardening. Read-only, no disruption.

Restoration validation

An observed restoration of a nominated system, timed end to end, so RTO is measured rather than estimated.

Ransomware resilience assessment

Scenario-driven: assume domain compromise, then assess what the attacker can reach and what survives.

SaaS data protection review

Focused on Microsoft 365, Google Workspace and similar, where the provider's retention is commonly mistaken for backup.

Frameworks and standards

3-2-1-1-0 rule
Three copies, two media, one offsite, one immutable or offline, zero errors on verification. The structural baseline.
NIST SP 800-34
Contingency planning guidance, used for RPO/RTO structure and plan assessment.
ISO/IEC 27031
ICT readiness for business continuity, where the client maintains an ISMS.
CIS Controls v8 Control 11
Data recovery control assessment and maturity.
Vendor hardening guides
Veeam, Commvault, Rubrik, Veritas and equivalent — their own guidance on protecting the backup platform.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

Vendor consoles (read-only)

Configuration, job history, retention and immutability settings direct from the platform.

BloodHound

Determining whether backup service accounts and consoles are reachable from a compromised domain identity.

Nmap

Confirming network isolation of backup infrastructure from production segments.

Custom job history parsers

Analysing months of job logs for silent failures and the systems that never complete successfully.

Storage vendor tooling

Verifying object lock, retention lock and WORM configuration at the storage layer rather than the application layer.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Coverage

  • Systems in the business impact assessment against systems actually backed up
  • SaaS data: Microsoft 365, Google Workspace, Salesforce and similar
  • Endpoint and remote worker data
  • Configuration and infrastructure-as-code, not only data
  • Databases backed up consistently rather than as open files
  • Systems excluded from backup, and whether anyone decided that

Immutability and isolation

  • At least one copy immutable or physically offline
  • Immutability verified, not assumed, including whether an administrator can shorten it
  • Backup network segmentation from production
  • Backup credentials separate from the production domain
  • MFA on backup console access
  • Air-gapped or logically isolated copy for critical systems

Infrastructure hardening

  • Backup server patch currency
  • Console exposure and authentication
  • Service account privilege on production systems
  • Vendor hardening guide compliance
  • Backup repository access control and encryption

Recovery capability

  • Documented RPO and RTO per system tier
  • Evidence of restoration testing with dates, scope and timings
  • Full system recovery tested, not only file-level
  • Bare-metal and cloud recovery paths
  • Recovery runbooks and whether they are current
  • Restoration bandwidth against data volume

Monitoring and verification

  • Job failure alerting and who receives it
  • Silent failures: jobs reporting success with incomplete data
  • Backup verification and integrity checking
  • Retention compliance against policy and regulation
  • Capacity trending against growth

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • Job history across months and hundreds of systems hides the ones that have quietly failed. Correlating every job against the protected system inventory finds what aggregate success rates conceal.

  • Coverage gaps are found by diffing the asset inventory against the backup inventory, which is mechanical comparison and is almost never done by hand at full scale.

  • Whether a backup credential is reachable from a compromised domain account is a graph question over the directory, not a configuration setting.

  • Immutability claims are verified per repository rather than per platform, because one misconfigured repository is where the attacker will go.

Every agent finding is validated by a human reviewer. The assessment is read-only: no job is modified, no restore triggered, and nothing deleted. Restoration validation, where in scope, is performed by the client's own team with CSS observing.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Assessed as an attack target

Most backup reviews check schedules and retention. CSS asks what a domain administrator can delete, because that is who the attacker will be by the time backups matter.

Immutability verified, not accepted

Retention lock is tested against whether an administrator can shorten or remove it. Configured immutability and real immutability differ more often than vendors suggest.

RTO measured, not quoted

Where restoration validation is in scope, recovery is timed end to end against real data volumes and real link speeds, rather than taken from a specification sheet.

Both reports, always

Technical report and executive summary together, plus a ransomware scenario walkthrough the board can follow.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Scope: platforms, repositories, protected systems and the review date
  • Coverage analysis: protected systems against the asset and business impact inventory
  • Immutability and isolation assessment per repository, with verification evidence
  • Backup infrastructure hardening findings against vendor guidance
  • Job history analysis including silent and recurring failures
  • RPO and RTO assessment per system tier, with measured figures where validated
  • Ransomware scenario outcome: what survives and how long recovery takes
  • Retest results appended against each original finding

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • If ransomware encrypted production tonight, what the organisation could restore and how long it would take
  • Position against 3-2-1-1-0, stated as which elements are met and which are not
  • The three changes that most improve survivability
  • Regulatory and insurance position, since cyber insurers increasingly require immutability evidence
  • Remediation timeline and retest date
  • One page

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A distribution business of roughly 900 staff, running a well-regarded backup platform.

Finding
Backups ran nightly with 30-day retention to a disk repository, replicated offsite. The backup service account was a domain administrator, the console authenticated against the same domain with no MFA, and the repository was mounted as a share reachable from production. A domain compromise would have reached every copy.
Recommendation
Move the backup platform to a separate identity boundary, enforce MFA on the console, convert the repository to a hardened immutable target, and remove domain administrator from the service account.
Outcome
Identity separation and MFA completed in three weeks; the immutable repository followed in the next budget cycle. The retest confirmed a simulated domain administrator could no longer delete or alter any backup copy.

A professional services firm that had migrated to Microsoft 365 four years earlier.

Finding
No backup existed for Microsoft 365. The assumption was that Microsoft handled it. Retention policies covered accidental deletion for 30 to 93 days depending on the workload, and nothing covered a malicious administrator, a ransomware event in OneDrive, or a departed user's mailbox after licence removal.
Recommendation
Deploy a third-party Microsoft 365 backup with independent retention and immutability, covering Exchange, SharePoint, OneDrive and Teams.
Outcome
Backup deployed over a month. Within the year it was used to recover a SharePoint site a user had deleted seven months earlier, which was well outside Microsoft's retention.

A manufacturer with 40 TB of production data and a stated four-hour RTO.

Finding
Restoration validation timed a full recovery of the primary ERP system at 71 hours, against a stated RTO of four. The bottleneck was a 200 Mbps link to the offsite repository, which nobody had measured against the data volume. Restore tests had only ever covered individual files.
Recommendation
Hold a local copy of tier-one systems for fast recovery, keep the offsite copy for disaster scenarios, and restate the RTO honestly for systems where four hours is not achievable.
Outcome
A local immutable copy was added for the top six systems, bringing measured RTO to under five hours. The honest restatement of RTO for lower tiers was, per the client, the more valuable finding.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.