Red Teaming
Red Teaming of AI and LLM Systems
An LLM application is a system that takes untrusted input and hands it to something with permissions. The model is rarely the interesting target; the interesting target is what it is connected to — the retrieval corpus, the tool it can call, the database it can query, the email it can send. Testing therefore concentrates on where instructions cross a trust boundary, which is a class of flaw most application security programmes have no existing coverage for.
Methodology
- 01
Architecture and trust boundary mapping
Every input source, retrieval corpus, tool, function and downstream system the model can reach, and the privilege each carries. The map is the assessment's foundation because the risk is in the connections.
- 02
Direct prompt injection
System prompt extraction, instruction override, role manipulation and guardrail bypass tested systematically rather than by trying jailbreaks found online.
- 03
Indirect prompt injection
Instructions planted in content the model will later ingest — documents, web pages, emails, database fields — testing whether data becomes instruction. This is the highest-severity class in most deployments.
- 04
Tool and agent abuse
Whether the model can be induced to call tools outside its intended purpose, chain calls to reach unintended systems, or pass attacker-controlled arguments to a privileged function.
- 05
Data exposure testing
Training and retrieval data leakage, cross-tenant retrieval in shared corpora, and whether the model returns content the requesting user is not entitled to.
- 06
Output handling testing
Whether model output is treated as trusted downstream — rendered as HTML, executed as code, passed to a shell or used in a query without validation.
- 07
Resource and cost abuse
Token exhaustion, recursive agent loops, and whether an attacker can impose unbounded cost.
- 08
Reporting and retest
Technical report and executive summary together, then a retest confirming closure — with the caveat that model behaviour is probabilistic and closure is evidenced across repeated trials.
Approach to testing
- Findings are demonstrated across repeated trials, not once. A model is probabilistic, and a bypass that works one time in twenty is still a finding — but it is reported with its success rate rather than as a binary.
- Guardrail bypass alone is not treated as high severity. The severity comes from what the bypass reaches: a model that can be made to say something rude is a brand issue, one that can be made to call a payment tool is a security issue.
- Indirect injection is prioritised over direct, because the user is usually not the attacker — the attacker is whoever controls a document the model will read.
- Agent systems are tested for privilege at the tool level, since the model's effective permission is the union of every tool it can call.
- Model updates invalidate results. The report states the exact model version and date, because the same prompt against a newer version may behave differently.
Types of assessment
LLM application assessment (default)
A chat or completion application with retrieval. Covers injection, data exposure and output handling.
Agentic system assessment
Systems where the model calls tools and takes actions. Higher risk and a different methodology, because the blast radius is real.
RAG pipeline assessment
Focused on the retrieval corpus: poisoning, cross-tenant leakage and whether document-level permissions survive into retrieval.
Model integration review
Architecture-level review of trust boundaries, tool permissions and output handling, without adversarial testing. Appropriate pre-launch.
Frameworks and standards
- OWASP Top 10 for LLM Applications
- The primary taxonomy — prompt injection, insecure output handling, excessive agency, and the rest.
- MITRE ATLAS
- Adversarial threat landscape for AI systems, used for technique mapping and reporting.
- NIST AI Risk Management Framework
- Governance context where the client needs to demonstrate a managed approach to AI risk.
- OWASP ASVS
- The conventional application controls around the model, which are frequently weaker than the model controls.
- EU AI Act
- Where the client's system falls in scope, obligations are noted alongside technical findings.
Tools used
Tooling is where testing starts, not where it ends. Every automated result is reproduced by hand before it reaches a report.
Garak
Automated probing across known LLM vulnerability classes as a first pass.
PyRIT
Microsoft's risk identification toolkit for structured adversarial testing at scale.
Promptfoo
Regression testing so a fixed bypass can be verified across repeated trials and future model versions.
Burp Suite
The application around the model, which is still a web application with conventional flaws.
Custom injection corpora
Payloads written for the client's specific tools and retrieval sources, because generic jailbreaks test the model vendor, not the client.
Checklist approach
The checklist is the floor, not the ceiling. It guarantees coverage so nothing standard is missed; the findings that matter usually come from what a tester does after it is complete.
Prompt injection
- System prompt extraction and disclosure
- Instruction override and role manipulation
- Guardrail bypass, with success rate across trials
- Indirect injection through retrieved documents
- Indirect injection through user-supplied content stored and later read
- Multi-turn injection building across a conversation
Excessive agency
- Complete tool inventory and the privilege each carries
- Tools callable outside their intended context
- Chained tool calls reaching unintended systems
- Attacker-controlled arguments passed to privileged functions
- Human-in-the-loop requirements for consequential actions
- Whether the model can modify its own instructions or configuration
Data exposure
- Retrieval returning content outside the user's entitlement
- Cross-tenant leakage in shared vector stores
- Training data extraction where the model is fine-tuned on client data
- Sensitive data in system prompts
- Conversation history isolation between users
Output handling
- Model output rendered as HTML without sanitisation
- Output used in SQL, shell or code execution
- Output passed to downstream systems as trusted input
- Markdown and link rendering enabling exfiltration
- Structured output validated against a schema
Resource and availability
- Token and request rate limiting per user
- Recursive agent loop termination
- Cost controls and alerting
- Context window exhaustion handling
Governance
- Model version pinning and change management
- Logging of prompts and completions, and its privacy implications
- Human review for consequential actions
- Incident process for model misbehaviour
How findings are scored
Every finding is scored on CVSS 3.1 and placed in one of five levels. The executive summary adds a sixth band — Compliant — so components that passed appear on the same chart as those that did not.
- Critical
- Immediate measures must be taken. These vulnerabilities can allow an attacker to take complete control of the application or server — stealing user data, tricking users into supplying sensitive information, or defacing the site.
- High
- Maximum risk associated with a specific vulnerability instance. May enable an attacker to compromise the application and its data, partially or completely, or to modify application behaviour beyond its intended purpose. To be handled with utmost priority.
- Medium
- Considerable risk. May enable an attacker to exploit the application to a particular level, gaining low-level information that can be used to craft more specific attacks.
- Low
- Lowest risk. May allow an attacker to gain some information about the application that was not intended to be known, without an exploitation technique currently available at that instance.
- Informational
- A functionality or component is missing best-practice implementation. Not a risk today, but may become one as the application changes or as exploitation techniques, policy or legal requirements evolve.
Scan types selected
- Safe Checks
- Standard / OWASP Top 10
- Destructive
- SANS Top 25
- Business Logic Vulnerability Testing
Standard toolset by stage
- OSINT
- Datasploit, Google Dorks, Shodan
- Enumeration & Scanning
- Nmap, Wfuzz, Unicornscan
- Domain Enumeration
- Nikto, DnsRecon, Knock
- Crawling & Fuzzing
- Burp Suite, Acunetix, Netsparker
- Vulnerability Analysis
- OpenSSL, sqlmap, CVE-Details
- Exploitation
- Metasploit, Netcat, Exploit-DB
How CSS tests
A unified swarm of agents, for blind spot detection
AI agents drive several testing tracks against the same target at once, then cross-check each other. A single tester works one hypothesis at a time; parallel agents cover the space a sequential pass leaves behind.
Model behaviour is probabilistic, so a single trial proves nothing. Agents run each payload hundreds of times and report a success rate, which is the only honest way to characterise a bypass.
Injection payload space is effectively unbounded. Parallel agents explore mutation and combination far beyond what a tester types by hand, and the successful payload is rarely the obvious one.
Multi-turn injection builds across a conversation; exploring conversation trees is combinatorial and is exactly what parallel agents are for.
Indirect injection requires planting content in every ingestible source and waiting to see what surfaces — a wide, patient search rather than a clever one.
Agents generate and evaluate payloads; a human assessor judges severity and confirms impact. This is the one capability where the swarm is testing a system of the same kind, and the report is explicit that automated results are filtered by a human before publication — including the many that look alarming and are not.
Why this differs
What CSS does that most vendors do not
Every one of these is checkable. Ask any vendor for the same and compare the answers.
Severity from reach, not from output
Most AI red teaming reports jailbreaks. CSS reports what the bypass reaches — a rude response and a callable payment tool are not the same finding, and conflating them wastes the client's attention.
Success rates, not anecdotes
Every bypass is reported with a measured success rate across repeated trials, because a probabilistic system cannot honestly be described with a single example.
Indirect injection prioritised
The user is usually not the attacker. Testing concentrates on content the model ingests from elsewhere, which is where real compromise comes from and where most assessments are thin.
Regression suite delivered
Findings ship as a Promptfoo suite the client can run against future model versions, because a model update can silently reopen a closed finding.
Reporting
Two documents, two audiences
Both are produced for every engagement. They are not the same document at two lengths — they answer different questions and are written separately. The structure below is the one CSS actually issues.
Technical assessment report
For the engineers who will fix it
- Disclaimer, and Limitations on Disclosure and Use
- Risk Level & Description — the five levels above, scored on CVSS 3.1
- Scan Type — which of the five assessment types were selected
- Assessment Scope — the control areas covered
- Assessment Date — the exact testing window
- Objective of the Assessment — objectives listed against completion status
- Tools Utilization — manual and automated tooling by stage
- Summary of the Assessment
- Overall Recommendations, split into Must Have and Should Have
- Vulnerability Overall Classifications as per Organization
- Security Issues Highlighted
- The Key Findings — each with evidence and detailed recommendation
- Summary of Findings & Conclusion
For this assessment specifically
- Scope: application, model version and date, tools, retrieval sources and the test window
- Findings mapped to OWASP LLM Top 10 and MITRE ATLAS
- Every bypass with its measured success rate across trials
- Trust boundary map showing what each finding reaches
- Reproduction payloads verbatim, with the conversation context required
- Tool inventory with effective privilege per tool
- Explicit statement that results are bound to the tested model version
- Retest results appended, evidenced across repeated trials
Executive summary
For the people who will fund the fix
- Objectives, each against a completion status
- Overall Finding of the Assessment — total threats identified, broken down by component and severity
- Summary of the Assessment
- Artefacts of the Assessment — the key findings as a numbered register with severity
- Observation of the Assessment — the major attacks the organisation should be prepared for, given what was found
- Overall Recommendation, including a Business Enabling Recommendation sequence
- Must Have and Should Have actions
For this assessment specifically
- What an attacker could make this system do, in business terms
- Whether the risk is reputational, data exposure, or unauthorised action — these are very different and are usually conflated
- The three changes that most reduce blast radius
- Regulatory position where the EU AI Act or sector guidance applies
- The model version dependency, stated plainly, because it affects how long the assessment remains valid
- One page
Case studies
What this finds in practice
Representative engagement patterns. Sector and scale only — no client is named, and no detail is included that could identify one.
A financial services firm's customer support assistant with retrieval over internal knowledge base articles.
- Finding
- A knowledge base article editable by any support agent was used to plant instructions. When the assistant retrieved that article, it followed them — disclosing its system prompt and, in 34% of trials, including content from other customers' conversation summaries held in the same vector store.
- Recommendation
- Treat retrieved content as data rather than instruction through prompt structure and delimiters, partition the vector store per tenant, and restrict knowledge base editing to a reviewed workflow.
- Outcome
- Partitioning and editorial review implemented before launch. The regression suite runs in CI and caught a reoccurrence when the retrieval prompt was refactored.
A logistics company's internal agent with tools for querying shipments and sending customer emails.
- Finding
- A shipment note field, populated from customer-supplied data, contained injected instructions. Processing that shipment caused the agent to send an email to an attacker-supplied address containing data from the query results. The email tool had no recipient allowlist and no human approval step.
- Recommendation
- Restrict the email tool to verified recipient domains, require human approval for any outbound message, and sanitise customer-supplied fields before they enter the model context.
- Outcome
- Recipient allowlist and approval step implemented within two weeks. The finding reframed the client's whole approach: tools now start with no privilege and gain it by justification.
A healthtech startup's clinical summarisation feature.
- Finding
- Model output was rendered as Markdown in the clinician interface without sanitisation. An injected instruction caused the model to emit an image tag with a URL containing summarised patient data, which the browser fetched automatically — exfiltrating clinical data to an external server on render.
- Recommendation
- Sanitise model output before rendering, disable automatic remote resource loading in the interface, and apply a content security policy restricting outbound requests.
- Outcome
- All three implemented before the feature reached production. The client extended output sanitisation to every model-facing surface after the finding.
Next
Scope this assessment
Most scopes are settled in one call. Tell us what the application does and who uses it, and we will tell you what testing it properly involves.