What an AI red-team assessment covers
An assessment tests the whole AI feature, not just the model. The serious failures usually happen where the model meets your system: the system prompt, the retrieval layer, the tools it can call and the data it can see. I agree the scope with you in writing, then test against a threat model built for your product and your users.
Where you already have guardrails or moderation in place, I test whether they hold, not just whether they exist.
- Direct and indirect prompt injection, including instructions hidden in documents, emails, web pages and retrieved content
- Jailbreaks and policy bypasses, in single and multi-turn conversations
- Data leakage: system prompts, secrets, personal data and other users’ records surfacing in responses
- Unsafe tool use and excessive permissions in agents and function calling
- Model output that reaches HTML, SQL, shell commands or downstream APIs without checks
- Abuse that drives runaway token use, loops or denial of service
Prompt injection and jailbreak testing
Prompt injection is the failure I test for first, because almost every LLM application is exposed to it. Direct injection comes from a user typing instructions into the chat. Indirect injection is harder to spot: the instructions arrive inside content the model reads, such as a support ticket, an uploaded PDF, a web page or a record pulled in by RAG, and the model treats them as commands.
I test both, using attacks written for your product rather than only a generic list. Jailbreak testing runs over multiple turns, because a model that refuses a harmful request in one message will often comply after a few turns of role-play, reframing or gradual escalation. Every successful attack is logged with the exact inputs, the response and the conditions needed to reproduce it, so your engineers can confirm it and test the fix.
AI agent and tool-use security testing
Agents raise the stakes. When a model can send emails, query databases, call internal APIs or run code, a successful injection is no longer just a bad reply: it becomes an action taken with your credentials.
I review agent trajectories step by step: what the agent was asked, what it decided, which tools it called with which arguments, and whether each step stayed within what that user was allowed to do. It is the same trajectory review I carry out on autonomous coding agents in evaluation work.
- Tools and permissions broader than the task needs
- Actions taken without a human confirmation step where one is needed
- Cross-user or cross-tenant access through shared tools, memory or retrieval
- Agents that report success when the task did not succeed
- Loops and chained calls that run up cost or hit rate limits
Mapping findings to the OWASP Top 10 for LLM Applications
The OWASP Top 10 for LLM Applications gives security teams, auditors and enterprise buyers a shared vocabulary for LLM risk, and many now ask for it by name. I map every finding to the relevant OWASP category, such as prompt injection, sensitive information disclosure or excessive agency, and give it a severity score based on likelihood and impact in your context.
That gives your security team a familiar structure, makes the report easy to use in customer security questionnaires and in a SOC 2 or ISO 27001 risk register, and shows clearly which categories were tested and which were out of scope. The framework is a floor, not a ceiling: findings that don’t fit a category neatly are still reported.
Why frontier-lab evaluation experience matters
Alongside consulting, I’ve done evaluation, red-teaming and training work on frontier LLMs and autonomous coding agents through Mercor, Turing, Uber AI Solutions, DataAnnotation, Mindrift, AfterQuery and Terac. It includes adversarial quality control, trajectory review and multi-turn reasoning audits, plus auditing synthetic repositories, test suites and Docker runtimes to catch reward hacking, answer leakage and false-positive verifications.
In practice, that means I know how models fail when someone is trying to make them fail, and how an evaluation can mislead. An agent that passes because the test was weak is something I look for by default. Because I design LLM evaluation and benchmarking frameworks too, I can leave you with an automated test suite that catches regressions, not just a one-off report.
The security grounding comes from an MSc in Cyber Security and Forensic Information Technology at the University of Portsmouth. I’m also a Databricks Certified Generative AI Engineer and was named Cybersecurity Engineer of the Year 2026 in the Corporate LiveWire Innovation & Excellence Awards.
What you get
- A written threat model and agreed test scope for your AI feature
- Risk-scored findings, each mapped to the OWASP Top 10 for LLM Applications
- Reproduction steps for every finding: inputs, conversation turns, tool calls and observed output
- A specific fix for each finding, from prompt and guardrail changes to permission and architecture changes
- A plain-English executive summary for leadership, customers and auditors
- A reusable adversarial test set your team can run in CI before each release
- One retest of fixed findings, with an updated report
How it works
Step 01 · Before testing
Scope and threat model
NDA first, then a working session on how the feature works, who uses it, what it can reach and what would hurt most if it went wrong. You get a written scope and a fixed or capped price before any testing starts.
Step 02 · Weeks 1–2
Adversarial testing
Hands-on testing in staging or with test accounts: prompt injection, jailbreaks, data leakage, tool use and agent trajectories. Anything critical is reported to you the same day, not held back for the report.
Step 03 · End of week 2
Report and walkthrough
Risk-scored findings, reproduction steps and fixes, followed by a call with your engineers to agree priorities and owners.
Step 04 · After your fixes
Retest
I rerun the original attacks, plus variants, against the fixed version and issue an updated report you can share with customers or auditors.
Proof
“Steady progress, predictable delivery, and code that’s easy to review and integrate.”
Sean Metcalf
Founder, getKaivo
“He has been instrumental in leading key security initiatives that have significantly strengthened our company’s overall security posture.”
Sam Tszho Ho
Head of AI and Platform
“A strong blend of expertise in cybersecurity, digital forensics tools, and software development — reliable and consistently high standard.”
Soraya Harding
Senior Lecturer, Cybersecurity Intelligence & Digital Forensics
Price
From £3,500. Fixed scope, with a fixed or capped price agreed in writing before testing starts. Every engagement starts with a free 1-hour intro call, and you get a written scope with a fixed or capped price before any work starts.
FAQ
How much does AI red teaming cost in the UK?
Assessments start from £3,500 for a fixed, agreed scope. The final price depends on how many features, user roles and tools are in scope, and whether agents can take real actions. You get a written scope and a fixed or capped price before any testing starts.
Is this the same as a standard penetration test?
No. A standard penetration test looks at your network, infrastructure and web application. This is specialist adversarial testing of the AI layer: prompts, retrieval, tools, memory and agent behaviour. I’m not a CREST- or CHECK-accredited tester, so if you need an accredited test as well, the two reports sit well side by side.
Do you need access to our source code or model weights?
No. Most assessments are black-box or grey-box: I test through the same interfaces your users have, with a test account for each role. Read access to the system prompt, tool definitions and retrieval setup makes testing faster and deeper, and I handle it under NDA.
We use OpenAI, Claude or Gemini. Isn’t the model provider responsible for safety?
The provider is responsible for the base model. You are responsible for your system prompt, the data you put in context, the tools you connect and what your application does with the output. That layer is where prompt injection turns into data leakage or unsafe actions, and no provider can test it for you.
Can you test an AI agent that takes real actions safely?
Yes. I test agents in staging or with sandboxed tools and test data, so nothing real is sent, deleted or charged during the assessment. If a production-only integration has to be tested, we agree the limits in writing first.
How often should we red-team an LLM application?
Before launch, and again after any significant change: a new model version, new tools, new data sources or a major rewrite of the system prompt. Between assessments, the regression test set I hand over lets your team rerun known attacks in CI on every release.