Advisory · AI Red Teaming · Partner-Led · Continuous
Red team your AI before attackers do.
Prompt-injection-focused adversarial testing of your AI systems: chat, RAG, agents, copilots, and multimodal apps. Prompt Shields scopes and coordinates, and vetted offensive-security partners run the attacks. You get findings, severity-ranked vectors, and a remediation plan in two to four weeks, wired into your CI/CD so it keeps working past your next model update.
The Architecture Problem · The Root Cause
LLMs read instructions and data on the same channel.
The model can't reliably tell “content to process” from “commands to obey.” That single fact is the root cause behind every question a security-minded buyer asks about LLMs. Prompt injection sits at #1 on the OWASP LLM Top 10 (2025) for the second edition running, precisely because it's architectural rather than a bug you patch.
The moment a model can act, whether it calls tools, hits APIs, or runs skills, that confusion stops being a content problem and becomes an authorisation problem. The whole conversation is really one question: how do you keep an attacker's words from becoming your system's actions?
Plain LLM / chatbot
Blast radius is informational: a wrong answer, leaked text, offensive output, a hallucinated number. Mitigation lives at the input/output layer.
Risk = what it says
Agentic system
The model's output becomes an action: a tool call, an API write, a file sent, a database row changed. A successful injection redirects a real decision, well beyond bad text.
Risk = what it's allowed to do
- Excessive agency (LLM06) — every tool you grant is new attack surface.
- Persistence — memory and RAG poisoning survive across sessions; data poisoning becomes a runtime concern.
- Chained trust — multi-step, multi-tool, sometimes multi-agent. A compromise at step 2 quietly steers steps 3 through 10.
For a chatbot, security is about what it says. For an agent, security is about what it's allowed to do — and prompt injection turns into privilege escalation.
OWASP now ships a separate Top 10 for Agentic Applications (first edition Dec 2025) for exactly this reason. The chatbot list doesn't cover the surface.
Where The Risk Lives
The vendor secured the model. You own the wrapper.
Organisations treat “we wrapped GPT or Claude” as low-risk because the model is someone else's. The risk didn't go away. It moved into the wrapper, and the wrapper is the part you own. Here are six things teams underestimate when they build on top of commercial models.
Your glue code is the trust boundary
The vendor secured the model. You own the RAG sources, connectors, tool definitions, and system prompts. That's where injection actually lands.
Excessive agency by default
Teams wire up broad API scopes and service accounts to get the prototype working, then never tighten them. A compromised agent inherits all of it.
Indirect injection via your own data
Citizen submissions, uploaded PDFs, emails, case documents, scraped web. Any of it can carry instructions the model reads as commands.
System prompt leakage (LLM07)
Teams hide secrets, policy, and business logic in the system prompt assuming it's private. It isn't. Extraction is reliable.
Tool & MCP supply chain
A malicious or poisoned tool description or skill file can compromise the agent before any user even types. 2025 saw coding agents fully compromised through MCP tool descriptions.
No eval / regression harness
Ship without a test suite, and every model update or prompt tweak quietly changes behaviour with nothing to catch it. This is the number-one quiet incident pattern.
The model is the safest part of your stack. The risk is in the wrapper you built around it — and you own all of it.
What Teams Get Wrong
The honest frame, then the mistakes.
Prompt injection cannot be fully solved at the model level today. OpenAI, Anthropic, and Google DeepMind all said so explicitly in 2025. Any defence expressed as a prompt instruction can itself be overridden by another instruction. So the first mistake is treating this as something you can win with a cleverly worded system prompt.
Adding a system-prompt instruction telling the model to ignore malicious instructions
Prompt-level guardrails are bypassable by construction. They might be necessary, but they're never enough on their own. This is the single most common false sense of security.
Defending only against direct injection, ignoring indirect
Teams test the chat box but forget the model also reads emails, documents, web pages, RAG chunks, and tool outputs, all of which an attacker can control. The real production incidents are indirect. EchoLeak / CVE-2025-32711, a zero-click exfil from M365 Copilot, is the canonical example.
No separation of trusted vs. untrusted content
Untrusted data gets concatenated straight into the same context as instructions, with no provenance tracking. The model can't tell them apart because you didn't either.
Over-permissioned agents
Broad tool and API scopes give a successful injection somewhere to go. The injection is the match, and excessive agency is the gasoline.
No egress control
Even when you can't stop the injection itself, you can often stop the data from leaving. But teams leave open image fetches, rendered URLs, markdown links, and arbitrary HTTP tools that become exfil channels.
Treating a guardrail classifier as a silver bullet
Input / output filters help as defence in depth, but they have false negatives. Relying on them alone just moves the arms race.
No human-in-the-loop on high-consequence actions
Irreversible, state-changing, or money-or-data-out actions that auto-execute with no confirmation. A successful injection writes itself a cheque.
The Reframe · Defence In Depth
Stop trying to win at the model layer. Build four pillars around it.
Provenance + separation
Track what's untrusted. Keep retrieved and tool data structurally separate from instructions so the model never has to guess which is which.
Least privilege
Minimise each agent's tools and scopes so a successful injection has very little reach. Constrain credentials per task and per user.
Deterministic policy outside the LLM
Push real control-flow decisions into code the model can't talk its way past. Patterns like CaMeL and FIDES enforce policy outside the prompt loop.
Constrained egress
Block image fetches, arbitrary URL rendering, markdown link rendering, and outbound HTTP. Injection may succeed, but the data still can't leave.
The one heuristic to remember
Meta's Rule of Two
In any single operation, an agent should have at most two of three of the following properties. All three at once = exploitable.
- 1Processes untrusted input
- 2Has access to sensitive systems or data
- 3Can change external state or send data out
This single heuristic catches most of the dangerous designs we see in production agent architectures.
Injection is inevitable; impact is a design choice.
What We Test
Six prompt-injection attack categories
We calibrate to your surface (chat, RAG, agents, copilots, multimodal), and every engagement covers at least these six. Multilingual coverage is included by default, because English-tuned filters routinely miss non-English payloads.
Direct prompt injection
Overriding the system prompt with user-supplied instructions. Think the classic 'ignore all previous instructions' family, plus its modern variants.
Indirect prompt injection
Hostile content inside retrieved documents, RAG sources, tool outputs, emails, or web pages that the model treats as authoritative instructions. This is the real production-incident surface.
Jailbreaks
Role-play, encoding tricks, multi-turn priming, multilingual payloads, and stylistic chains that bypass safety filters or alignment guardrails.
System-prompt extraction
Coaxing the model to reveal its system prompt, tool catalogue, or hidden knobs. It's the LLM equivalent of source disclosure.
Tool & agent abuse
Chaining function calls, code interpreters, and browsing tools to read or move data the user should never have reached. This is the bridge from injection to incident.
Multimodal injection
Payloads embedded in images, PDFs, or audio that vision and audio models read as instructions. It's the dominant new attack surface on agents.
How We Test
In layers — cheap and fast to deep and expensive.
Direct injection, indirect injection and data exfiltration are three distinct things, and we test them distinctly. We check whether the data could actually get out, not just whether the model followed the bad instruction. The injection is the symptom; exfiltration is the incident.
Curated attack suite
A versioned library of known payloads run as deterministic regression tests: instruction overrides, jailbreaks, encoded and obfuscated text, hidden CSS, payloads in PDFs, images, and HTML, and your users' languages. English-tuned filters routinely miss non-English.
Indirect-specific scenarios
Plant payloads in the sources the agent actually ingests: a malicious instruction inside a document, a calendar invite, an email body, a web page, a RAG entry. Run a normal user task and watch for hijack. It's the test almost everyone skips.
Exfil canaries
Unique secret tokens placed in the agent's context or its connected data. Watch every egress channel: tool args, generated URLs, image src, network calls. If the canary leaves, you have a real exfil path, no debate.
Automated / generative red teaming
Tools that generate and mutate attacks at scale rather than hand-writing each one: PromptFoo, DeepEval / DeepTeam, Giskard, Microsoft PyRIT, Garak. Findings map cleanly to OWASP categories so coverage is auditable.
Tool / agent-level probing
For each tool integration, test for escalation: can the agent be steered into calling it with attacker-chosen arguments? Benchmarks like Agent Security Bench (ASB) formalise this.
Human red team
Periodic expert adversarial testing for the creative attacks automation misses. It complements the automated layers rather than replacing them, and findings feed back into the suite.
Scoring
We track attack success rate per category over time. The goal isn't zero percent on a fixed set, since you can overfit to your own tests. The goal is a declining success rate against an evolving set, plus zero successful exfiltrations on your canaries.
Continuous, Not One-Time
Red teaming as a pipeline, not an annual event.
A point-in-time red team before go-live is already stale by the next model update. Three things change underneath you constantly: the foundation model (vendors ship updates weekly), your own prompts, tools, and data, and the attacker playbook. Static testing can't keep up.
The mature teams we see are moving red teaming from an annual report to a CI/CD gate, applying the same discipline as unit tests to adversarial prompt-injection and exfil suites. Here's what that looks like in five steps.
Adversarial tests in CI/CD
Your prompt-injection and exfil suite runs on every change as a release gate: prompt edit, model bump, new tool, data refresh. It's the same discipline as unit tests, and PromptFoo and DeepEval plug straight into pipelines.
Regression gating
A known attack that starts passing again blocks the deploy. That's how you catch drift before it reaches production, and it's something an annual pentest will never do for you.
Runtime monitoring
A background adversarial process continuously probes the live system with evolving payloads. Runtime guardrails and egress monitoring watch real traffic for injection and exfil attempts.
Feedback loop
Every new attack found, whether in production, in research, or on disclosure lists, gets added to the suite. Coverage compounds instead of decaying.
Periodic deep human red team
Roughly quarterly, on top of the automated programme, for novel creative attacks. Findings feed straight back into the automated suite, closing the loop.
Regulatory Tailwind
EU AI Act is making this a compliance expectation.
High-risk system obligations land August 2026. Adversarial-testing duties for general-purpose AI under Article 55 are already in effect. Continuous adversarial testing is becoming a compliance requirement rather than a nice-to-have, which usually reframes the budget conversation internally.
Red teaming as an event tests the system you had. Red teaming as a pipeline tests the system you actually shipped this morning. Your model changes weekly — your testing has to move at the same cadence.
Why we run this with security partners — not on our own.
AI red teaming sits at the intersection of two crafts: deep offensive-security experience (decades of breaking real systems for a living) and applied LLM expertise (how prompts, embeddings, tools, and guardrails actually fail). Few teams have both. Pretending otherwise is how engagements turn into glossy reports with no findings that matter.
Our model is honest: Prompt Shields brings the AI domain expertise — what to test, what to score it against, and what the findings mean for your governance posture. Our vetted security partners, offensive-security firms with proven AI / LLM red-teaming practices, run the attacks. You get one engagement, one report, and one accountable lead from our side.
How An Engagement Runs
Four phases, two to four weeks
We calibrate to scope. Most engagements run three weeks end-to-end with a retest in the fourth. After that, the suite lives in your CI/CD.
Scope
We map your AI surface (models, RAG sources, tools, user roles, data sensitivity) and agree the rules of engagement with you in writing.
Recon
Partners profile your hosted model and integration surface: system prompts, tool inventory, retrieval pipelines, and deployment context.
Attack
Hand-crafted and tooled prompt-injection campaigns across the agreed surface, recorded as reproducible proof-of-concept transcripts.
Remediation
Written findings, severity-ranked vectors, and a hardening plan your team can ship, with a follow-up retest included.
What You Walk Away With
Findings your team can actually act on
Findings report
Every successful attack with reproducible proof-of-concept transcripts, conditions and impact.
Severity-ranked vectors
Each finding rated for exploitability and business impact, mapped to OWASP LLM Top 10 and MITRE ATLAS.
Remediation roadmap
Concrete, sequenced controls at the prompt layer, the gateway layer, and in your guardrails, with quick wins called out.
Board-ready summary
A one-page narrative leadership can act on without translation, suitable for board or audit-committee packs.
Findings mapped to OWASP LLM Top 10 · OWASP Agentic Top 10 · MITRE ATLAS · NIST AI RMF · ISO 42001
Find out what your AI breaks under real attack.
Book a 30-minute scoping call. We'll walk your team through the methodology, map your specific surface against Meta's Rule of Two, propose a partner from our roster matched to your stack, and agree the rules of engagement before you commit.
Book a discovery call