Red Teaming
Assumed Breach Assessment
Given enough attempts, someone will click something. Perimeter strength is therefore a delaying control, not a determining one, and the question that decides whether an incident is an inconvenience or a catastrophe is what an attacker reaches afterwards. An assumed breach assessment skips the way in entirely, spending the whole engagement on escalation, movement and detection.
Methodology
- 01
Foothold definition
The starting position agreed with the client: a standard user account, a managed laptop, a compromised web server, or a supplier VPN connection. Each answers a different question.
- 02
Local reconnaissance
From the foothold: what the account can see, what is cached on the host, what shares and services are reachable, and what the local security controls permit.
- 03
Privilege escalation
Local escalation on the host, then domain or platform escalation, with every path recorded rather than only the successful one.
- 04
Credential access
Harvesting from memory, disk, configuration files, shares and scripts — credentials in a file share are still among the most reliable escalation routes.
- 05
Lateral movement
Movement towards crown jewel systems, with segmentation tested as an actual constraint rather than as documentation.
- 06
Objective assessment
What the attacker can reach: the finance system, the customer database, the source repository, the backup platform. Reported as reachability rather than as a list of vulnerabilities.
- 07
Detection assessment
What the client's tooling recorded at each stage, and whether anyone responded.
- 08
Reporting and retest
Technical report, executive summary and reachability map, then a retest confirming closure.
Approach to testing
- The foothold is given, not earned. Time that a full red team spends on initial access goes to escalation and movement instead, which is where the findings that change an incident's cost live.
- The blue team may or may not be informed, and this is agreed explicitly. Informed produces better learning; uninformed produces a truer detection measurement.
- Segmentation is tested from the foothold outward, because a zone boundary that holds against a scanner often does not hold against an authenticated user with a legitimate reason to be on the network.
- Crown jewels are agreed in advance. Without that, the assessment measures how far a tester got rather than whether the business's important systems are reachable.
- Every action is logged with timestamps so detection can be reconciled during the debrief.
Types of assessment
Standard user foothold (default)
A normal domain account on a managed device. Models the phishing outcome, which is how most incidents begin.
Compromised server
Starts from a web or application server in a DMZ. Models exploitation of an internet-facing service.
Supplier or contractor access
Starts from third-party VPN or remote access credentials. Models supply chain compromise, and frequently finds the weakest segmentation.
Insider simulation
A legitimate employee acting maliciously, with their real entitlements. Different from a compromised account, because the actions look authorised.
Frameworks and standards
- MITRE ATT&CK
- Technique mapping for both the attack narrative and the detection coverage comparison.
- PTES
- Execution standard for post-exploitation.
- Microsoft Enterprise Access Model
- Tiering reference where the environment is Microsoft-centric, since tier violations are the usual escalation route.
- NIST SP 800-115
- Overall technical testing lifecycle.
- Cyber Kill Chain
- Executive narrative structure, focused on the post-intrusion phases.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
BloodHound
Attack path enumeration from the foothold principal to crown jewel systems.
Impacket
SMB, Kerberos and RPC interaction during movement.
Seatbelt / winPEAS
Host situational awareness and local escalation enumeration.
CrackMapExec / NetExec
Credential validation and reachability mapping across host sets.
Snaffler
File share enumeration for credentials and sensitive data, which is consistently high yield.
Custom C2
Where detection measurement is in scope and off-the-shelf tooling would be caught on signature alone.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Host and local
- Local privilege escalation from the foothold account
- Cached credentials and tokens on the host
- Local security control effectiveness: EDR, application control, script restrictions
- Host firewall and its effect on movement
- Sensitive data accessible locally
Credential access
- LSASS and memory credential extraction
- Credentials in file shares, scripts and configuration
- Service account credentials and their scope
- Kerberoastable and AS-REP roastable accounts
- Credential reuse across hosts and tiers
- Password manager and browser stored credentials
Escalation paths
- Every path from foothold to domain or platform administrator
- Tier violations where a tiering model exists
- Delegation and ACL abuse routes
- Certificate services escalation
- Cloud identity escalation where hybrid
Lateral movement and segmentation
- Hosts reachable from the foothold
- Zone boundaries crossed and whether they were intended to be crossable
- Administrative protocols permitted between zones
- Jump host effectiveness
- Crown jewel system reachability
Detection
- Actions that generated telemetry
- Actions that generated an alert
- Alerts investigated versus closed
- Time to detection per attack stage
- Containment capability if detection had occurred
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Path enumeration from the foothold is a graph problem across the whole directory, and the shortest path is rarely the one a tester finds first by intuition.
File share enumeration across terabytes for credentials and sensitive data is patient search — high yield, and something a tester under time pressure samples.
Reachability mapping from the foothold to every host and service is exhaustive scanning that defines the true blast radius, rather than the documented one.
Detection correlation across every logged action and the client's telemetry is a join across two large datasets.
Every agent finding is validated by a human operator and all exploitation stays under human control. On a production network the blast radius of an automated action is the business. The swarm enumerates and correlates; the operator acts.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Measures blast radius, not entry
The perimeter will eventually be crossed. This assessment quantifies what that costs, which is the number that actually determines incident severity and is missing from most testing programmes.
Every path, not the first
Escalation paths are enumerated exhaustively, because closing one route in a graph of many produces a retest that fails.
Both reports, always
Technical report and executive summary together, plus a reachability map showing what the foothold reached.
Detection assessed alongside
The report states what the client's monitoring saw at each stage. A finding that nothing was detected is frequently more valuable than the escalation path itself.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Foothold definition, rules of engagement and the engagement window
- Complete timeline of actions with timestamps for telemetry reconciliation
- Every escalation path from foothold to privileged access, with length and preconditions
- Reachability map: which crown jewel systems were reached and by which route
- Segmentation findings against the documented design
- MITRE ATT&CK technique mapping with detection outcome per technique
- Findings with severity expressed in terms of reachability
- Retest results appended against each original finding
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- Starting from one compromised laptop, what an attacker reaches and how long it takes
- Which crown jewel systems are reachable, and which are not
- Whether the organisation would have noticed, and at which stage
- The three changes that most reduce blast radius
- Remediation timeline and retest date
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A retail group, starting from a standard head office user account.
- Finding
- A file share reachable by all staff contained a deployment script with a service account password in clear. The account was a local administrator across 600 hosts including the point-of-sale management servers. Time from foothold to POS management access was under two hours.
- Recommendation
- Remove credentials from scripts in favour of managed service accounts, restrict the share, and remove blanket local administrator rights.
- Outcome
- Credentials rotated immediately and the share restricted the same day. The managed service account migration ran over a quarter, and the retest could not reach POS management from a standard account.
A financial services firm with a documented tiering model, starting from a tier-two workstation.
- Finding
- The tiering model was sound on paper. In practice, three tier-zero administrators used their privileged accounts to log into tier-two workstations for support, leaving credentials in memory. Harvesting one gave direct tier-zero access.
- Recommendation
- Enforce tier separation with authentication policy silos so a tier-zero credential cannot be used on a tier-two host, and provide administrators with a supported alternative support workflow.
- Outcome
- Authentication policy silos deployed over two months. The alternative workflow was the harder part, since the behaviour existed because support was genuinely difficult otherwise.
A manufacturer, starting from a supplier VPN account.
- Finding
- Supplier access was intended to reach one application server. Network rules permitted the whole server VLAN, and from there the engineering network was reachable because a historical route had never been removed. The supplier connection reached production control systems in four hops.
- Recommendation
- Restrict supplier VPN to the single required host and port, remove the historical route, and review every third-party access path against its stated business need.
- Outcome
- Supplier access restricted within a week. The wider review covered eleven third-party connections and narrowed seven of them.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.