Skip to content
CyberSmithSECURE
Under Attack

Phishing Simulation & Awareness

Spear Phishing and Business Email Compromise Simulation

Mass phishing measures the organisation. Spear phishing measures the twelve people who can move money, approve access or release data — and those are the people a real attacker will research for a fortnight before sending anything. The simulation therefore begins with genuine open-source reconnaissance and produces messages that are plausible because they are specific.

Methodology

  1. 01

    Target selection

    Agreed with a small control group: finance approvers, executives, IT administrators, HR and anyone else who can authorise something consequential.

  2. 02

    Open-source reconnaissance

    Public sources only — company filings, LinkedIn, conference material, press coverage, breach corpora. The reconnaissance findings are themselves a deliverable, because most clients do not know their exposure.

  3. 03

    Pretext development

    Scenarios built from what the reconnaissance found: a real supplier, a genuine project, an actual conference the target attended. Plausibility comes from specificity.

  4. 04

    Infrastructure preparation

    Lookalike domains registered, mail authentication configured so messages pass basic checks, and landing pages matched to the client's real systems.

  5. 05

    Campaign execution

    Low volume, individually crafted, sent over a window rather than in a batch. Batch sending is the pattern filters catch.

  6. 06

    MFA relay testing

    Where in scope, credential capture with real-time relay to test whether MFA actually stops an attacker who is present at the moment of authentication.

  7. 07

    Post-compromise demonstration

    Where credentials are captured and authorised, demonstrating what the account reaches — mailbox rules, delegate access, financial approval.

  8. 08

    Debrief and reporting

    Individual debriefs with targets handled carefully, then technical report and executive summary.

Approach to testing

  • Targets are individuals and the debrief is personal. Being fooled by a well-researched spear phish is not a failing, and the debrief says so — otherwise the programme teaches executives to hide incidents.
  • Reconnaissance is public-source only. No pretexting of staff to gather information, no contact with real suppliers, and no use of breach data beyond confirming exposure.
  • MFA relay is tested where authorised, because 'we have MFA' is the most common and most misplaced reassurance in this area.
  • Volume stays low and delivery is staggered. A spear phishing simulation that looks like a batch is testing the mail filter, not the person.
  • Executive impersonation is agreed with the impersonated executive beforehand. Simulating the CEO without telling the CEO ends badly.

Types of assessment

Targeted spear phishing (default)

A small set of named individuals with researched pretexts. Measures the people who matter most.

Business email compromise simulation

Finance-focused: fraudulent payment requests, supplier bank detail changes, invoice redirection. Tests process as much as people.

Executive impersonation

Messages appearing to come from named executives to their reports. Tests whether authority overrides process.

MFA relay assessment

Credential capture with live relay, demonstrating whether the MFA implementation resists an attacker present at authentication.

Frameworks and standards

MITRE ATT&CK T1566.002 / T1534
Spear phishing link and internal spear phishing technique references.
OSINT frameworks
Structured reconnaissance methodology so the exposure assessment is repeatable.
NIST SP 800-50
Awareness programme structure for the training that follows.
FBI IC3 BEC typologies
Real business email compromise patterns, so scenarios reflect what actually happens rather than what is easy to simulate.

Tools used

Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.

Evilginx

Reverse-proxy credential capture with real-time MFA relay, where authorised.

Gophish

Campaign tracking and landing page delivery for non-relay scenarios.

Maltego / SpiderFoot

Structured open-source reconnaissance and relationship mapping.

theHarvester / Hunter

Email address discovery and format confirmation.

Breach corpus lookup

Confirming which target credentials appear in public breach data — a check the client can and should run themselves.

Custom lookalike infrastructure

Domains, certificates and mail configuration matched to the pretext.

Checklist approach

The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.

Exposure

  • Executive and finance staff discoverable from public sources
  • Email address format confirmable externally
  • Organisational structure inferable from public profiles
  • Supplier and partner relationships publicly visible
  • Target credentials present in breach corpora
  • Out-of-office and travel information leaking externally

Technical controls

  • DMARC enforcement preventing exact-domain spoofing
  • Lookalike domain detection and monitoring
  • External sender warnings and whether they are noticed
  • Impersonation protection for named executives
  • Link rewriting and time-of-click analysis
  • Mail filtering against low-volume targeted messages

Authentication resilience

  • MFA method and its resistance to real-time relay
  • Number matching or phishing-resistant factors for privileged users
  • Conditional access limiting where a session can be used from
  • Session token lifetime after authentication
  • Detection of impossible travel and anomalous sign-in

Process controls

  • Payment change process and out-of-band verification
  • Whether authority can override process
  • Dual approval for consequential financial actions
  • Supplier bank detail change procedure
  • Escalation path when a request feels wrong

Response

  • Whether targets reported, and how quickly
  • Reporting behaviour among executives specifically
  • Detection of the sign-in following credential capture
  • Containment speed once compromise was evident

How findings are scored

Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.

Critical
Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
High
Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
Medium
Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
Low
Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
Informational
A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.

Scan types selected

  • Safe Checks
  • Standard / OWASP Top 10
  • Destructive
  • SANS Top 25
  • Business Logic Vulnerability Testing

Standard toolset by stage

OSINT
Datasploit, Google Dorks, Shodan
Enumeration & Scanning
Nmap, Wfuzz, Unicornscan
Domain Enumeration
Nikto, DnsRecon, Knock
Crawling & Fuzzing
Burp Suite, Acunetix, Netsparker
Vulnerability Analysis
OpenSSL, sqlmap, CVE-Details
Exploitation
Metasploit, Netcat, Exploit-DB

How CSS tests

A unified swarm of agents, for blind spot detection

AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.

  • Open-source reconnaissance across filings, social platforms, conference material and breach corpora is breadth work, and the useful detail is usually one item among thousands.

  • Relationship mapping between staff, suppliers and partners builds the graph that makes a pretext plausible, and it is tedious rather than clever.

  • Lookalike domain permutation across character substitution, homoglyphs and TLD variants is exhaustive generation, then filtered by a human for the ones that would actually fool someone.

  • Breach corpus correlation across target addresses and historical dumps is lookup at scale.

Reconnaissance is assisted; pretext writing and target interaction are not. Every message is written and reviewed by a human and approved by the client's control group before sending. No agent contacts any person at any point.

Why this differs

What CSS does that most vendors do not

Every one of these is checkable. Ask any vendor for the same and compare the answers.

Real reconnaissance, not a template

Pretexts are built from what is actually discoverable about the target. A generic template measures template recognition; a researched one measures judgement under plausibility.

The reconnaissance is a deliverable

Clients receive the full OSINT findings, because most have never seen their own public exposure laid out and it usually changes what they publish.

MFA tested against relay

'We have MFA' is the most common false reassurance in this area. Where authorised, CSS demonstrates whether the specific implementation resists an attacker present at authentication.

Debriefs are not punitive

Targets are debriefed individually and supportively. A programme that embarrasses executives produces executives who conceal real incidents.

Reporting

Two documents, two audiences

Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.

Technical assessment report

For the engineers who will fix it

  • Disclaimer, and Limitations on Disclosure and Use
  • Risk Level & Description — the five levels above, scored on CVSS 3.1
  • Scan Type — which of the five assessment types were selected
  • Assessment Scope — the control areas covered
  • Assessment Date — the exact testing window
  • Objective of the Assessment — objectives listed against completion status
  • Tools Utilization — manual and automated tooling by stage
  • Summary of the Assessment
  • Overall Recommendations, split into Must Have and Should Have
  • Vulnerability Overall Classifications as per Organization
  • Security Issues Highlighted
  • The Key Findings — each with evidence and detailed recommendation
  • Summary of Findings & Conclusion

For this assessment specifically

  • Targets, pretexts and the reconnaissance that produced them
  • Full open-source exposure findings per target
  • Campaign results: opened, clicked, submitted, reported, with timings
  • MFA relay outcome where in scope, with the specific weakness identified
  • Technical control findings: DMARC, impersonation protection, filtering
  • Process findings where a financial or access request was acted on
  • Post-compromise reach demonstrated

Executive summary

For the people who will fund the fix

  • Objectives, each against a completion status
  • Overall Finding of the Assessment — total threats identified, broken down by component and severity
  • Summary of the Assessment
  • Artefacts of the Assessment — the key findings as a numbered register with severity
  • Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
  • Overall Recommendation, including a Business Enabling Recommendation sequence
  • Must Have and Should Have actions

For this assessment specifically

  • Whether a targeted attack on the people who matter would succeed today
  • Whether MFA would have stopped it
  • Whether process would have caught it after authentication succeeded
  • The three changes that most reduce exposure
  • Public exposure findings the organisation may wish to reduce
  • One page

Case studies

What this finds in practice

Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.

A manufacturing group, targeted campaign against eight finance staff.

Finding
A pretext built from a genuine supplier relationship visible in a public case study requested a bank detail change. Three of eight engaged, and one initiated the change. The finance process required verbal confirmation, but the caller ID was spoofed to the supplier's published number and the confirmation call was accepted.
Recommendation
Verify bank changes by calling the number held on file, not the number provided or displayed, and require dual approval for any supplier bank detail change.
Outcome
Process changed within two weeks. Four months later the same control stopped a genuine attempt that had reached the same stage.

A financial services firm with MFA enforced on all accounts.

Finding
A relay campaign against twelve senior staff captured credentials from four, and all four approved the resulting push notification. Session tokens were valid for eight hours and usable from any location, so MFA provided no protection against an attacker present at authentication.
Recommendation
Move privileged and finance users to phishing-resistant authentication, enable number matching in the interim, and apply conditional access restricting session use by location and device compliance.
Outcome
Number matching enabled within days; hardware keys rolled out to 40 high-risk users over a quarter. The retest captured no usable session.

A healthcare provider, executive impersonation against departmental heads.

Finding
A message appearing to come from the CEO requesting an urgent staff list for a board paper was actioned by six of nine recipients within two hours, four of whom sent personal data including home addresses. External sender warnings were present but rendered at the bottom of the message on mobile.
Recommendation
Move external sender warnings to the top of the message body, add impersonation protection for named executives, and establish a standing rule that staff data requests go through HR regardless of who asks.
Outcome
Warning placement and impersonation protection implemented within a month. The HR routing rule was the change staff cited as most useful, because it removed the judgement call entirely.

Next

Scope this assessment

Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.