Phishing Simulation & Awareness
Mass Phishing Simulation
A single phishing campaign produces a number and no learning. The point of simulation is the trend — whether the organisation gets better, which departments do not, and whether reporting improves alongside clicking. This runs as a programme with a baseline, targeted training and repeat campaigns, and the reporting rate is treated as the more important metric than the click rate.
Methodology
- 01
Programme design
Campaign frequency, population, difficulty progression and the metrics that will be tracked, agreed with HR and internal communications before anything is sent.
- 02
Baseline campaign
A first campaign at moderate difficulty, sent without warning, establishing the starting position for click, submit and report rates.
- 03
Result analysis
Results broken down by department, role, tenure and location — the aggregate figure hides the population that actually needs attention.
- 04
Targeted training
Training assigned based on behaviour rather than to everybody. Those who reported get acknowledgement; those who submitted credentials get the most.
- 05
Repeat campaigns
Subsequent campaigns at varying difficulty and on varying themes, so the programme measures capability rather than familiarity with one template.
- 06
Reporting channel assessment
Whether the report button works, where reports go, how quickly they are triaged, and whether a real phish reported by a user would be acted on.
- 07
Trend reporting
Longitudinal reporting showing movement over quarters, which is the only output that tells the organisation whether the investment works.
Approach to testing
- Reporting rate is the primary metric, not click rate. An organisation where 40% click and 60% report is in better shape than one where 10% click and nobody reports, because the second has no early warning.
- No naming and shaming. Individual results are confidential to the individual and their training assignment; management sees aggregates. Programmes that punish clicking suppress reporting, which is the opposite of the goal.
- Difficulty is declared and progressive. Reporting a 2% click rate on an obvious template tells the board something false.
- Campaigns are agreed with HR and internal communications beforehand, and a disclosure plan exists for the people who will be upset.
- Themes avoid genuine distress. Simulated redundancy notices and bonus announcements produce excellent click rates and lasting resentment.
Types of assessment
Baseline campaign
A single campaign establishing the starting position. Useful once; not a programme.
Continuous programme (default)
Quarterly or monthly campaigns with behaviour-driven training and trend reporting. The only form that changes behaviour.
Difficulty-graded programme
Campaigns at declared difficulty levels so the trend is comparable and improvement is real rather than an artefact of easier templates.
Reporting-focused programme
Explicitly aimed at raising the report rate, with the report button, triage process and user feedback loop as the primary scope.
Frameworks and standards
- NIST SP 800-50 / 800-16
- Security awareness and training programme structure.
- SANS Security Awareness Maturity Model
- Used to describe where the programme sits and what the next stage requires.
- ISO/IEC 27001 Annex A.6.3
- Awareness, education and training, where the client maintains an ISMS.
- MITRE ATT&CK T1566
- Phishing technique reference, used when results feed a detection programme.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Gophish
Campaign delivery, landing pages and result tracking where a self-hosted platform is preferred.
Client awareness platform
Where the client already runs KnowBe4, Proofpoint or similar, CSS designs and analyses rather than replacing the tool.
Custom landing pages
Templates matched to the client's own systems, because generic templates measure template recognition rather than judgement.
Mail authentication tooling
Verifying that simulated mail is delivered on the same footing as real external mail, not allowlisted into the inbox.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Campaign design
- Population coverage including contractors and shared mailboxes
- Difficulty declared and consistent between campaigns
- Theme appropriateness and distress avoidance
- Allowlisting configured so delivery is realistic but not bypassing all controls
- Landing page behaviour and data capture minimisation
Measurement
- Click rate, credential submission rate and report rate
- Time to first click and time to first report
- Breakdown by department, role, tenure and location
- Repeat clickers across campaigns
- Comparison against declared difficulty
Technical controls
- Whether the simulated mail was filtered, and at which layer
- SPF, DKIM and DMARC behaviour for lookalike domains
- Link rewriting and sandboxing effectiveness
- Attachment handling policy
- Web filtering of the landing page
Reporting process
- Report button present and functional across mail clients
- Where reports go and who triages them
- Time from report to triage
- User feedback after reporting
- Whether a genuine phish would be actioned differently
Training
- Training assigned by behaviour rather than universally
- Completion rates and follow-up for non-completion
- Content relevance to the campaign that triggered it
- Role-specific content for high-risk functions
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Template variation matters: measuring against one template measures familiarity with it. Generating variants across theme, sender and urgency produces a susceptibility figure that means something.
Segmentation analysis across department, role, tenure and location surfaces the population that needs attention, which a single organisational figure hides entirely.
Repeat-clicker identification across campaigns over quarters is longitudinal correlation, and it identifies the small group that drives most of the risk.
Delivery verification — whether each message actually landed in an inbox or was filtered — is per-recipient checking that sampling gets wrong.
Every campaign is reviewed and approved by a human before sending, and by the client before that. Agents assist with template variation and result analysis; no message is sent to any employee without explicit client sign-off on the exact content.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Reporting rate is the headline
Most vendors lead with click rate because it makes a better chart. The rate at which staff report is what gives the security team early warning of a real campaign, and CSS reports it first.
No naming and shaming
Individual results stay with the individual. Punitive programmes suppress reporting, which makes the organisation less safe while making the numbers look better.
Difficulty declared
A 2% click rate on an obvious template is not an achievement. Campaigns carry a declared difficulty so trends are honest and comparable.
Both reports, always
Technical report on controls and delivery, executive summary on behaviour and trend.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Campaign design, population, templates used and declared difficulty
- Click, submission and report rates with breakdowns by segment
- Delivery analysis: which messages were filtered and at which control
- Technical control findings: mail authentication, link rewriting, sandboxing
- Reporting process assessment with measured triage times
- Repeat clicker analysis across campaigns
- Trend comparison against previous campaigns at equivalent difficulty
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- Report rate and click rate, in that order, with the trend
- Which parts of the organisation need attention, without naming individuals
- Whether technical controls or human behaviour is the larger gap
- The three changes that would most improve resilience
- Recommended programme cadence
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A professional services firm of around 800 staff running its first programme.
- Finding
- The baseline returned a 31% click rate and a 4% report rate. The report button existed in Outlook but sent reports to a shared mailbox nobody monitored; three genuine phishing emails reported by staff in the preceding quarter had gone unread.
- Recommendation
- Route reports to the security team with a triage SLA, acknowledge every report to the reporter, and make report rate the headline metric in management reporting.
- Outcome
- Report rate reached 38% within three quarters and click rate fell to 12%. The acknowledgement loop was, per staff feedback, the single change that made people bother.
A manufacturer with plant and office staff.
- Finding
- The organisational click rate of 14% concealed a 6% office rate and a 41% rate among plant supervisors, who used shared terminals, had minimal email experience and had never received training because the programme assumed desk-based staff.
- Recommendation
- Build role-specific training for shared-terminal users, and report by population rather than as a single organisational figure.
- Outcome
- Supervisor click rate fell to 17% over two quarters. Reporting by population is what made the gap visible after it had been invisible in the aggregate for two years.
A financial services firm that had run simulations for three years with improving numbers.
- Finding
- Click rate had fallen from 22% to 4%, which the board saw as success. Review found every campaign had used the same platform's default templates, and staff recognised the sender pattern. A campaign using a bespoke template matched to the client's own systems returned 27%.
- Recommendation
- Vary templates and senders, declare difficulty, and treat the previous three years of trend data as measuring template familiarity rather than resilience.
- Outcome
- The programme was redesigned with graded difficulty. Restating the trend honestly to the board was uncomfortable and led to funding for role-based training that the false trend had made look unnecessary.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.