Red Teaming
Purple Team Exercises
A red team tells you whether you were caught. A purple team tells you why not, and fixes it the same afternoon. Techniques are executed openly with the security operations team watching their own console, so every gap is diagnosed immediately — was there no telemetry, no rule, or a rule that fired and was ignored? Those three have completely different remedies and a covert engagement cannot distinguish them.
Methodology
- 01
Technique selection
An ATT&CK technique set chosen for the client's threat profile and current coverage, agreed in advance so the exercise is targeted rather than exhaustive.
- 02
Baseline coverage assessment
Existing detection rules and log sources reviewed against the selected techniques, producing an expected-coverage map before anything is executed.
- 03
Execution with observation
Each technique executed in a controlled way with the blue team watching. Timestamped so telemetry can be located precisely.
- 04
Immediate triage
For each technique: did telemetry exist, did a rule fire, did an analyst see it, did they act. The gap is classified rather than simply recorded as a miss.
- 05
Rule development
Where a gap is a missing rule, a detection is written and tested during the exercise, then re-run to confirm it fires.
- 06
Telemetry gap remediation
Where a gap is missing telemetry, the logging change is identified and scoped — usually a longer piece of work than a rule.
- 07
Re-execution and verification
Every technique re-run after remediation, so the exercise ends with measured improvement rather than a to-do list.
- 08
Reporting
Coverage before and after, with the rules developed handed over as artefacts.
Approach to testing
- Collaborative and overt throughout. The value is in the diagnosis, and hiding from the defenders destroys it.
- Gaps are classified into three kinds — no telemetry, no rule, no action — because they have different owners and different costs, and a report that says 'not detected' helps nobody.
- Detection rules are written during the exercise and verified by re-execution, not recommended for later.
- Coverage is measured before and after, so the deliverable is an improvement figure rather than an assessment.
- Technique selection is deliberately narrow. Twenty techniques closed properly beats a hundred and fifty recorded as gaps.
Types of assessment
Standard purple team (default)
An agreed technique set executed over several days with the blue team present, rules written as gaps are found.
Detection engineering sprint
Focused on building coverage for a specific tactic — credential access, lateral movement, exfiltration — rather than breadth.
Post-red-team purple
Run after a covert red team, working through the techniques that went undetected. The most effective sequence, because the gaps are already known.
Continuous validation
A recurring subset of techniques re-run monthly or quarterly to confirm detections still work after platform and rule changes.
Frameworks and standards
- MITRE ATT&CK
- The technique catalogue and the coverage model the whole exercise is scored against.
- Atomic Red Team
- Standardised, bounded technique implementations, so execution is reproducible and safe.
- MITRE D3FEND
- Mapping defensive countermeasures to the techniques they address.
- Detection Engineering maturity models
- Used to frame where the client's detection function sits and what to build next.
- Sigma
- Rules are written in Sigma where possible so they are portable between SIEM platforms.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Atomic Red Team
Bounded, documented technique execution with defined cleanup.
Caldera
Automated adversary emulation for longer chains, where a scripted sequence is more repeatable than manual execution.
Sigma / SIEM native rule languages
Writing detections during the exercise, portable where the platform allows.
Client SIEM and EDR consoles
The exercise is run against the client's own tooling, sitting beside their analysts.
VECTR
Tracking technique outcomes and coverage before and after, and producing the comparison view.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Initial access and execution
- Phishing attachment and link execution
- Malicious macro and script execution
- Living-off-the-land binary usage
- PowerShell and command-line logging coverage
- Application control bypass techniques
Persistence and privilege escalation
- Registry run keys and startup folder
- Scheduled task creation
- Service creation and modification
- Token manipulation and UAC bypass
- Accessibility feature abuse
Credential access
- LSASS memory access
- Credential dumping from registry hives
- Kerberoasting and AS-REP roasting
- Browser and credential manager access
- Cached credential retrieval
Discovery and lateral movement
- Domain and network enumeration
- Remote service execution: WMI, WinRM, PsExec
- Pass-the-hash and pass-the-ticket
- Remote desktop and admin share usage
- Scheduled task creation on remote hosts
Collection and exfiltration
- Archive creation and staging
- Exfiltration over HTTPS, DNS and cloud storage
- Large-volume data transfer detection
- Clipboard and screen capture
Response quality
- Alert fired versus alert seen versus alert actioned
- Time from execution to analyst acknowledgement
- Investigation quality and escalation decisions
- Out-of-hours coverage
- Containment action availability and speed
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Technique variation matters more than technique coverage: a rule that catches one implementation of credential dumping frequently misses three others. Agents generate variants so the rule is tested rather than the tool.
Telemetry correlation across SIEM, EDR and platform logs for each execution is a join across large datasets, and doing it by hand slows the exercise to a crawl.
Coverage mapping against the full ATT&CK matrix, including sub-techniques, is bookkeeping that should not consume analyst time during a live exercise.
Re-execution after rule development is repetitive by design and benefits from automation, which is what makes same-day verification possible.
Execution stays under operator control and every technique is agreed in advance with the blue team present. Agents handle variation, correlation and bookkeeping; they do not decide what to run on a client network.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Gaps are diagnosed, not just recorded
Every miss is classified as no telemetry, no rule, or no action. Those have different owners and different costs, and 'not detected' on its own is not an actionable finding.
Rules written and verified during the exercise
Detections are developed and re-tested the same day rather than recommended for later. The client ends the engagement with working coverage, not a backlog.
Improvement is measured
Coverage before and after, with the same techniques re-run. The deliverable is a figure, not an opinion.
Both reports, always
Technical report and executive summary together, plus the rule artefacts in a portable format.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Technique set, selection rationale and the exercise window
- Coverage before and after, per technique and per tactic
- For each technique: telemetry present, rule fired, analyst actioned — with timings
- Gap classification with owner and estimated effort for each
- Rules developed during the exercise, in Sigma or the client's native format
- Telemetry gaps requiring logging changes, scoped
- Response quality observations, including alerts closed incorrectly
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- Detection coverage before and after, as two figures
- Which attack stages the organisation can and cannot see
- The three investments that would most improve detection
- Whether the gap is tooling, configuration or staffing — they are usually not the same
- Recommended cadence for revalidation
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A financial services firm with a SIEM and a recently hired detection engineer.
- Finding
- Of 32 techniques executed, 19 produced no alert. Classification showed only four were genuinely missing telemetry; eleven had the telemetry and no rule, and four fired rules that analysts had suppressed months earlier because of noise.
- Recommendation
- Write the eleven missing rules during the exercise, re-tune the four suppressed rules rather than leaving them off, and scope the four telemetry gaps as a logging project.
- Outcome
- Coverage moved from 41% to 84% within the exercise week. The suppressed-rule finding prompted a standing review of every suppression, which found nine more.
A manufacturer with an outsourced security operations provider.
- Finding
- Alerts fired correctly for 22 of 26 techniques. The provider acknowledged 19 and escalated 3. Median time from execution to acknowledgement was 47 minutes during business hours and over nine hours overnight, against a contracted 15 minutes.
- Recommendation
- Raise the response findings at the contract level with the evidence, and add out-of-hours coverage to the service agreement.
- Outcome
- The contract was renegotiated with measured evidence rather than impressions. The exercise is now repeated biannually as a contractual verification mechanism.
A healthcare provider building detection in-house after moving from a managed service.
- Finding
- Credential access techniques were entirely undetected. LSASS access telemetry was available in the EDR but was not forwarded to the SIEM, and the EDR's own detections were configured to report rather than block or alert.
- Recommendation
- Forward EDR telemetry to the SIEM, switch credential access detections to alerting, and build correlation rules combining EDR and directory telemetry.
- Outcome
- Forwarding configured during the exercise and rules written the same week. Credential access coverage went from zero to full for the tested technique set.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.