Advisory · AI Red Teaming · Partner-Led · Continuous

Red team your AI before attackers do.

Prompt-injection-focused adversarial testing of your AI systems: chat, RAG, agents, copilots, and multimodal apps. Prompt Shields scopes and coordinates, and vetted offensive-security partners run the attacks. You get findings, severity-ranked vectors, and a remediation plan in two to four weeks, wired into your CI/CD so it keeps working past your next model update.

The Architecture Problem · The Root Cause

LLMs read instructions and data on the same channel.

The model can't reliably tell “content to process” from “commands to obey.” That single fact is the root cause behind every question a security-minded buyer asks about LLMs. Prompt injection sits at #1 on the OWASP LLM Top 10 (2025) for the second edition running, precisely because it's architectural rather than a bug you patch.

The moment a model can act, whether it calls tools, hits APIs, or runs skills, that confusion stops being a content problem and becomes an authorisation problem. The whole conversation is really one question: how do you keep an attacker's words from becoming your system's actions?

Plain LLM / chatbot

Blast radius is informational: a wrong answer, leaked text, offensive output, a hallucinated number. Mitigation lives at the input/output layer.

Risk = what it says

Agentic system

The model's output becomes an action: a tool call, an API write, a file sent, a database row changed. A successful injection redirects a real decision, well beyond bad text.

Risk = what it's allowed to do

  • Excessive agency (LLM06) — every tool you grant is new attack surface.
  • Persistence — memory and RAG poisoning survive across sessions; data poisoning becomes a runtime concern.
  • Chained trust — multi-step, multi-tool, sometimes multi-agent. A compromise at step 2 quietly steers steps 3 through 10.
For a chatbot, security is about what it says. For an agent, security is about what it's allowed to do — and prompt injection turns into privilege escalation.

OWASP now ships a separate Top 10 for Agentic Applications (first edition Dec 2025) for exactly this reason. The chatbot list doesn't cover the surface.

Where The Risk Lives

The vendor secured the model. You own the wrapper.

Organisations treat “we wrapped GPT or Claude” as low-risk because the model is someone else's. The risk didn't go away. It moved into the wrapper, and the wrapper is the part you own. Here are six things teams underestimate when they build on top of commercial models.

Your glue code is the trust boundary

The vendor secured the model. You own the RAG sources, connectors, tool definitions, and system prompts. That's where injection actually lands.

Excessive agency by default

Teams wire up broad API scopes and service accounts to get the prototype working, then never tighten them. A compromised agent inherits all of it.

Indirect injection via your own data

Citizen submissions, uploaded PDFs, emails, case documents, scraped web. Any of it can carry instructions the model reads as commands.

System prompt leakage (LLM07)

Teams hide secrets, policy, and business logic in the system prompt assuming it's private. It isn't. Extraction is reliable.

Tool & MCP supply chain

A malicious or poisoned tool description or skill file can compromise the agent before any user even types. 2025 saw coding agents fully compromised through MCP tool descriptions.

No eval / regression harness

Ship without a test suite, and every model update or prompt tweak quietly changes behaviour with nothing to catch it. This is the number-one quiet incident pattern.

The model is the safest part of your stack. The risk is in the wrapper you built around it — and you own all of it.

What Teams Get Wrong

The honest frame, then the mistakes.

Prompt injection cannot be fully solved at the model level today. OpenAI, Anthropic, and Google DeepMind all said so explicitly in 2025. Any defence expressed as a prompt instruction can itself be overridden by another instruction. So the first mistake is treating this as something you can win with a cleverly worded system prompt.

1

Adding a system-prompt instruction telling the model to ignore malicious instructions

Prompt-level guardrails are bypassable by construction. They might be necessary, but they're never enough on their own. This is the single most common false sense of security.

2

Defending only against direct injection, ignoring indirect

Teams test the chat box but forget the model also reads emails, documents, web pages, RAG chunks, and tool outputs, all of which an attacker can control. The real production incidents are indirect. EchoLeak / CVE-2025-32711, a zero-click exfil from M365 Copilot, is the canonical example.

3

No separation of trusted vs. untrusted content

Untrusted data gets concatenated straight into the same context as instructions, with no provenance tracking. The model can't tell them apart because you didn't either.

4

Over-permissioned agents

Broad tool and API scopes give a successful injection somewhere to go. The injection is the match, and excessive agency is the gasoline.

5

No egress control

Even when you can't stop the injection itself, you can often stop the data from leaving. But teams leave open image fetches, rendered URLs, markdown links, and arbitrary HTTP tools that become exfil channels.

6

Treating a guardrail classifier as a silver bullet

Input / output filters help as defence in depth, but they have false negatives. Relying on them alone just moves the arms race.

7

No human-in-the-loop on high-consequence actions

Irreversible, state-changing, or money-or-data-out actions that auto-execute with no confirmation. A successful injection writes itself a cheque.

The Reframe · Defence In Depth

Stop trying to win at the model layer. Build four pillars around it.

Provenance + separation

Track what's untrusted. Keep retrieved and tool data structurally separate from instructions so the model never has to guess which is which.

Least privilege

Minimise each agent's tools and scopes so a successful injection has very little reach. Constrain credentials per task and per user.

Deterministic policy outside the LLM

Push real control-flow decisions into code the model can't talk its way past. Patterns like CaMeL and FIDES enforce policy outside the prompt loop.

Constrained egress

Block image fetches, arbitrary URL rendering, markdown link rendering, and outbound HTTP. Injection may succeed, but the data still can't leave.

The one heuristic to remember

Meta's Rule of Two

In any single operation, an agent should have at most two of three of the following properties. All three at once = exploitable.

  1. 1Processes untrusted input
  2. 2Has access to sensitive systems or data
  3. 3Can change external state or send data out

This single heuristic catches most of the dangerous designs we see in production agent architectures.

Injection is inevitable; impact is a design choice.

What We Test

Six prompt-injection attack categories

We calibrate to your surface (chat, RAG, agents, copilots, multimodal), and every engagement covers at least these six. Multilingual coverage is included by default, because English-tuned filters routinely miss non-English payloads.

Direct prompt injection

Overriding the system prompt with user-supplied instructions. Think the classic 'ignore all previous instructions' family, plus its modern variants.

Indirect prompt injection

Hostile content inside retrieved documents, RAG sources, tool outputs, emails, or web pages that the model treats as authoritative instructions. This is the real production-incident surface.

Jailbreaks

Role-play, encoding tricks, multi-turn priming, multilingual payloads, and stylistic chains that bypass safety filters or alignment guardrails.

System-prompt extraction

Coaxing the model to reveal its system prompt, tool catalogue, or hidden knobs. It's the LLM equivalent of source disclosure.

Tool & agent abuse

Chaining function calls, code interpreters, and browsing tools to read or move data the user should never have reached. This is the bridge from injection to incident.

Multimodal injection

Payloads embedded in images, PDFs, or audio that vision and audio models read as instructions. It's the dominant new attack surface on agents.

How We Test

In layers — cheap and fast to deep and expensive.

Direct injection, indirect injection and data exfiltration are three distinct things, and we test them distinctly. We check whether the data could actually get out, not just whether the model followed the bad instruction. The injection is the symptom; exfiltration is the incident.

Curated attack suite

A versioned library of known payloads run as deterministic regression tests: instruction overrides, jailbreaks, encoded and obfuscated text, hidden CSS, payloads in PDFs, images, and HTML, and your users' languages. English-tuned filters routinely miss non-English.

Indirect-specific scenarios

Plant payloads in the sources the agent actually ingests: a malicious instruction inside a document, a calendar invite, an email body, a web page, a RAG entry. Run a normal user task and watch for hijack. It's the test almost everyone skips.

Exfil canaries

Unique secret tokens placed in the agent's context or its connected data. Watch every egress channel: tool args, generated URLs, image src, network calls. If the canary leaves, you have a real exfil path, no debate.

Automated / generative red teaming

Tools that generate and mutate attacks at scale rather than hand-writing each one: PromptFoo, DeepEval / DeepTeam, Giskard, Microsoft PyRIT, Garak. Findings map cleanly to OWASP categories so coverage is auditable.

Tool / agent-level probing

For each tool integration, test for escalation: can the agent be steered into calling it with attacker-chosen arguments? Benchmarks like Agent Security Bench (ASB) formalise this.

Human red team

Periodic expert adversarial testing for the creative attacks automation misses. It complements the automated layers rather than replacing them, and findings feed back into the suite.

Scoring

We track attack success rate per category over time. The goal isn't zero percent on a fixed set, since you can overfit to your own tests. The goal is a declining success rate against an evolving set, plus zero successful exfiltrations on your canaries.

Continuous, Not One-Time

Red teaming as a pipeline, not an annual event.

A point-in-time red team before go-live is already stale by the next model update. Three things change underneath you constantly: the foundation model (vendors ship updates weekly), your own prompts, tools, and data, and the attacker playbook. Static testing can't keep up.

The mature teams we see are moving red teaming from an annual report to a CI/CD gate, applying the same discipline as unit tests to adversarial prompt-injection and exfil suites. Here's what that looks like in five steps.

Step 1

Adversarial tests in CI/CD

Your prompt-injection and exfil suite runs on every change as a release gate: prompt edit, model bump, new tool, data refresh. It's the same discipline as unit tests, and PromptFoo and DeepEval plug straight into pipelines.

Step 2

Regression gating

A known attack that starts passing again blocks the deploy. That's how you catch drift before it reaches production, and it's something an annual pentest will never do for you.

Step 3

Runtime monitoring

A background adversarial process continuously probes the live system with evolving payloads. Runtime guardrails and egress monitoring watch real traffic for injection and exfil attempts.

Step 4

Feedback loop

Every new attack found, whether in production, in research, or on disclosure lists, gets added to the suite. Coverage compounds instead of decaying.

Step 5

Periodic deep human red team

Roughly quarterly, on top of the automated programme, for novel creative attacks. Findings feed straight back into the automated suite, closing the loop.

Regulatory Tailwind

EU AI Act is making this a compliance expectation.

High-risk system obligations land August 2026. Adversarial-testing duties for general-purpose AI under Article 55 are already in effect. Continuous adversarial testing is becoming a compliance requirement rather than a nice-to-have, which usually reframes the budget conversation internally.

Red teaming as an event tests the system you had. Red teaming as a pipeline tests the system you actually shipped this morning. Your model changes weekly — your testing has to move at the same cadence.
Our Model · Partner-Led

Why we run this with security partners — not on our own.

AI red teaming sits at the intersection of two crafts: deep offensive-security experience (decades of breaking real systems for a living) and applied LLM expertise (how prompts, embeddings, tools, and guardrails actually fail). Few teams have both. Pretending otherwise is how engagements turn into glossy reports with no findings that matter.

Our model is honest: Prompt Shields brings the AI domain expertise — what to test, what to score it against, and what the findings mean for your governance posture. Our vetted security partners, offensive-security firms with proven AI / LLM red-teaming practices, run the attacks. You get one engagement, one report, and one accountable lead from our side.

How An Engagement Runs

Four phases, two to four weeks

We calibrate to scope. Most engagements run three weeks end-to-end with a retest in the fourth. After that, the suite lives in your CI/CD.

Phase 1

Scope

We map your AI surface (models, RAG sources, tools, user roles, data sensitivity) and agree the rules of engagement with you in writing.

Phase 2

Recon

Partners profile your hosted model and integration surface: system prompts, tool inventory, retrieval pipelines, and deployment context.

Phase 3

Attack

Hand-crafted and tooled prompt-injection campaigns across the agreed surface, recorded as reproducible proof-of-concept transcripts.

Phase 4

Remediation

Written findings, severity-ranked vectors, and a hardening plan your team can ship, with a follow-up retest included.

What You Walk Away With

Findings your team can actually act on

Findings report

Every successful attack with reproducible proof-of-concept transcripts, conditions and impact.

Severity-ranked vectors

Each finding rated for exploitability and business impact, mapped to OWASP LLM Top 10 and MITRE ATLAS.

Remediation roadmap

Concrete, sequenced controls at the prompt layer, the gateway layer, and in your guardrails, with quick wins called out.

Board-ready summary

A one-page narrative leadership can act on without translation, suitable for board or audit-committee packs.

Findings mapped to OWASP LLM Top 10 · OWASP Agentic Top 10 · MITRE ATLAS · NIST AI RMF · ISO 42001

Scoping call

Find out what your AI breaks under real attack.

Book a 30-minute scoping call. We'll walk your team through the methodology, map your specific surface against Meta's Rule of Two, propose a partner from our roster matched to your stack, and agree the rules of engagement before you commit.

Book a discovery call