AI red-teaming is a practice that adversarially probes a model for harmful outputs, misuse pathways, and safety failures before and after deployment. The discipline borrows its name from cybersecurity practice, but the targets are different: a tester is not looking for an exposed S3 bucket, but for a jailbreak that bypasses a chat assistant's safety policy, an indirect prompt injection that hijacks an agent through a poisoned web page, or a chain of prompts that elicits chemical, biological, radiological, and nuclear (CBRN) uplift. Major labs treat AI red-teaming as a core release gate, but they define and run it differently. Anthropic frames it as adversarial testing to identify vulnerabilities; Google describes it as a less structured complement to development evaluations; OpenAI commits to multi-faceted testing regimes for major public releases. NIST has begun publishing technical work on agent hijacking that points to where the next gap is. The sections below cover what AI red-teaming actually is, the attack classes testers probe, how the three frontier labs run their programs, where automation fits, and why mandated external testing keeps surfacing in policy debates.
What AI Red-Teaming Is

AI red-teaming is the practice of adversarially probing an AI system to expose vulnerabilities, misuse pathways, and safety failures that standard development evaluations do not catch. Anthropic defines the discipline as adversarial testing of a technological system to identify potential vulnerabilities, and calls it a critical tool for improving the safety and security of AI systems (Anthropic, Challenges in Red Teaming AI Systems). Google frames safety testing as a required, iterative step in the deployment pipeline (Google AI for Developers, Safety guidance) and frames the exercise itself as adversarial testing where specialist groups launch attacks on an AI system (Google AI for Developers, Responsible AI evaluation). The distinction from a standard eval matters. Benchmarks measure performance against fixed datasets and known failure categories; adversarial probing searches for the unknown failure that no fixed dataset captured. For the broader safety stack that surrounds this practice, see how AI safety and alignment keeps models on track, and for the model class itself, how modern generative AI systems work.
- Adversarial testing: probing a model with inputs designed to make it fail, not inputs designed to measure baseline quality.
- Vulnerability: a reproducible failure pathway that produces harmful, prohibited, or unsafe output under realistic attacker conditions.
- Safety testing: the wider set of evaluations a vendor runs before release; AI red-teaming is its adversarial, less structured branch.
- Misuse pathway: a sequence of prompts or tool calls that turns a permitted capability into a prohibited one, such as CBRN uplift or cyber offense.
Types of AI Red-Teaming
AI red-teaming covers several distinct methods, each suited to a different threat surface: domain-specific expert testing, policy vulnerability testing, automated LLM-assisted generation, and multilingual or multicultural probing. Anthropic describes domain-specific expert testing as collaborating with subject matter experts to identify and assess potential vulnerabilities within their area of expertise, and runs a frontier program focused on CBRN, cybersecurity, and autonomous AI risks (Anthropic, Challenges in Red Teaming AI Systems). The same source introduces Policy Vulnerability Testing (PVT) as in-depth, qualitative testing conducted with external subject matter experts on policy topics covered under Anthropic's Usage Policy. PVT differs from a domain-specific expert sweep in scope: a frontier-risk biologist probes whether the model gives uplift on a pathogen; a PVT collaborator probes whether the model's responses on, say, election integrity or self-harm align with the published policy. Anthropic also flags multilingual and multicultural testing because the majority of this work takes place in English from a United States perspective, which leaves non-English failure modes underexplored. Each of these is a different lens on the same target model, and labs typically run several in parallel rather than treat them as substitutes. For the related practice of fixed-benchmark evaluation, see how AI model evaluation benchmarks measure model capability and safety.
- Domain-specific expert testing: a subject matter expert probes one risk area in depth, per Anthropic, with frontier work concentrated on CBRN, cyber, and autonomous AI risks.
- Policy vulnerability testing (PVT): qualitative external testing against the vendor's published usage policy, introduced by Anthropic.
- Automated adversarial generation: a language model generates adversarial inputs at scale against the target model.
- External testing: independent domain experts probe the model outside the vendor's organization, recommended by Google and OpenAI.
- Multilingual testing: probing in non-English languages and non-US cultural contexts to surface failure modes English-only testing misses.
Attack Techniques Red Teams Use
AI red-teaming exercises map to a consistent set of attack classes: jailbreaks that bypass system instructions, direct and indirect prompt injection, societal-harm elicitation, and frontier-risk probing in domains such as CBRN and cybersecurity. A jailbreak is a prompt or sequence of prompts that overrides the model's safety policy, often by role-play framing or instruction laundering. Direct prompt injection sits in the user message; indirect prompt injection is more dangerous because the malicious instruction arrives inside a document, web page, or tool output the model is asked to process. NIST's technical blog on agent hijacking documents the latter class in production: AI agents have been shown vulnerable to agent hijacking, a type of indirect prompt injection where an attacker inserts malicious instructions into data ingested by the agent and causes unintended, harmful actions (NIST, Strengthening AI Agent Hijacking Evaluations). Frontier-risk probing targets the highest-stakes categories. Anthropic's frontier program concentrates on CBRN, cybersecurity, and autonomous AI risks (Anthropic, Challenges in Red Teaming AI Systems), and OpenAI's GPT-4o system card lists some of the risk categories evaluated, including speaker identification, unauthorized voice generation, the potential generation of copyrighted content, ungrounded inference, and disallowed content (OpenAI, GPT-4o System Card). Different products surface different attack categories, which is why a single benchmark cannot replace adversarial testing.
- Jailbreak: a prompt or prompt chain that bypasses a model's safety policy, often through role-play, hypothetical framing, or obfuscation.
- Direct prompt injection: a malicious instruction placed in the user prompt itself, telling the model to ignore its system instructions.
- Indirect prompt injection: a malicious instruction hidden in third-party content (a web page, email, PDF, tool output) that the model later ingests.
- Agent hijacking: per NIST, an indirect prompt injection variant that redirects an AI agent's tool use toward harmful actions.
- Societal-harm elicitation: probing for biased, discriminatory, or otherwise harmful content covered under the vendor's usage policy.
- Frontier-risk probing: CBRN uplift, autonomous AI risk, and cybersecurity uplift attempts on the model's most dangerous capabilities.
- Disallowed content elicitation: per OpenAI's GPT-4o system card, attempts to draw out content prohibited by the model's policy.
How Anthropic, OpenAI, and Google Run Red-Teaming Programs
AI red-teaming programs differ across labs: Anthropic runs domain-specific expert testing and policy vulnerability testing with external subject-matter experts; OpenAI publishes a Preparedness Framework and ran pre-deployment evaluations for GPT-4o and DALL-E; Google describes its adversarial work as less structured than development evaluations and supplements it with independent external evaluations. Anthropic centers its program on PVT and domain-specific expert testing, with frontier work focused on CBRN, cyber, and autonomous AI risks, and the company explicitly flags the absence of standardized practices as a comparability problem across vendors (Anthropic, Challenges in Red Teaming AI Systems). OpenAI publishes evaluation results inside system cards: the GPT-4o system card shares the model's Preparedness Framework evaluations and lists the categories evaluated end to end (OpenAI, GPT-4o System Card), and DALL-E underwent pre-deployment safety evaluation as part of OpenAI's frontier-risk process (OpenAI, Our Approach to Frontier Risk). Google's responsible AI evaluation guidance positions the exercise as less structured than its development evaluations and points to external evaluations by independent domain experts as the complement that catches limitations internal groups miss (Google AI for Developers, Responsible AI evaluation).
| Lab | Named program elements | Public artifact | External experts |
|---|---|---|---|
| Anthropic | Domain-specific expert testing, Policy Vulnerability Testing, frontier program (CBRN, cyber, autonomous AI) | Challenges in Red Teaming AI Systems | External subject matter experts on policy and frontier risks |
| OpenAI | Preparedness Framework evaluations, system cards, pre-deployment evaluation (GPT-4o, DALL-E) | GPT-4o System Card; Our Approach to Frontier Risk | Independent domain experts for major public releases |
| Adversarial testing as less structured probing, plus development evaluations and external evaluations | Responsible AI evaluation guidance | Independent external domain experts |
Automated Red Teaming: Using AI to Test AI

AI red-teaming can itself be automated: a language model generates adversarial inputs at scale, while classifiers such as Llama Guard or Microsoft PyRIT score whether the target model responds with harmful content. Anthropic describes using language models as the attacker as using AI systems to automatically generate adversarial examples and test the robustness of other AI models, which raises coverage and reduces the manual effort of human-only testing (Anthropic, Challenges in Red Teaming AI Systems). The mechanic is straightforward in principle: an attacker model is prompted to produce adversarial inputs across a defined threat taxonomy, those inputs are fed to the target model, and a classifier or judge model scores each response. Meta's Llama Guard family and Microsoft's PyRIT (Python Risk Identification Toolkit) are two open-source pieces that fit into this loop, the first as a safety classifier and the second as an orchestration framework for adversarial-prompt generation. OpenAI has likewise written about combining human testers with AI-generated adversarial inputs in its red-teaming work (OpenAI, Advancing red teaming with people and AI). Automation is not a replacement for domain experts. It scales the easy cases, surfaces patterns at volume, and frees humans to chase the hardest attack paths, while flagged failures route to mitigation work.
- Define a threat taxonomy: the attack classes the run will probe, such as jailbreaks, prompt injection, and disallowed content.
- Generate adversarial inputs: an attacker language model produces prompts across each category, often varied across languages and framings.
- Run against the target model: each adversarial prompt is sent to the model under test, with system instructions and tools matching production.
- Score the responses: a safety classifier such as Llama Guard, or an orchestration framework such as Microsoft PyRIT, flags harmful or policy-violating outputs.
- Triage and feed back: confirmed failures route to mitigation work, and successful attack patterns seed the next automated cycle.
Gaps and Standardization Challenges
AI red-teaming lacks a single canonical standard: Anthropic has noted the absence of standardized practices makes it hard to compare safety results across models, and NIST's work on agent hijacking evaluations illustrates how even well-scoped test suites require continuous probing as new attack methods emerge. Anthropic states directly that the lack of standardized practices complicates the situation and the comparability of vendor safety claims (Anthropic, Challenges in Red Teaming AI Systems). The NIST technical blog adds an operational gap: AI agents have been shown vulnerable to agent hijacking, a type of indirect prompt injection, and ongoing testing reveals new weaknesses even as new systems address previously known attacks (NIST, Strengthening AI Agent Hijacking Evaluations). NIST's Center for AI Standards and Innovation tested AgentDojo's baseline attack methods together with novel methods developed jointly with the UK AI Security Institute, which produced fresh attack variants the original benchmark did not include. MITRE ATLAS, the public knowledge base of adversarial tactics against AI systems, plays a complementary cataloging role, but no taxonomy substitutes for live adversarial probing and the downstream mitigation work it triggers. Multilingual probing remains underdeveloped, since most published work is English-first. For where these gaps hit hardest in deployment, see how autonomous AI agents work and where they introduce new attack surfaces.
- No common standard: Anthropic flags standardization gaps that make cross-vendor safety comparisons unreliable.
- Agent hijacking surface: NIST documents indirect prompt injection against agents and the need for continuous testing as systems evolve.
- Catalogs, not tests: MITRE ATLAS catalogs adversarial tactics but does not replace live adversarial probing.
- Language and culture coverage: Anthropic notes the bulk of this work runs in English from a US perspective, leaving multilingual probing underdeveloped.
- Eval lifespan: benchmarks like AgentDojo decay as attacks evolve, which is why NIST and the UK AI Security Institute produced novel attack methods through joint testing.
The Case for Mandated External Red-Teaming
AI red-teaming is increasingly treated as a governance requirement, not a voluntary quality step: Anthropic recommends making external testing a precondition for releasing advanced AI systems, and OpenAI has committed to multi-faceted testing regimes drawing on independent domain experts for all major public releases. Anthropic's policy paper proposes mandating external testing, either through a centralized third party such as NIST or in a decentralized manner via researcher API access, to standardize adversarial evaluation as a precondition for releasing advanced AI systems (Anthropic, Charting a Path to AI Accountability). OpenAI's governance commitments mirror that direction: companies should commit to internal and external testing of models in areas including misuse, societal risks, and national security concerns, and should develop a multi-faceted, specialized, and detailed regime drawing on independent domain experts for all major public releases (OpenAI, Moving AI Governance Forward). The same source frames thorough adversarial evaluation as essential for product success and for guarding against significant national security threats. None of this is binding regulation. It is normative guidance from the labs themselves, and the open question is whether a regulator like NIST or an analog body will codify a baseline that all frontier developers have to meet before release. For related governance reading, see AI safety standards and certification frameworks and algorithmic accountability frameworks for AI systems.
- Anthropic's recommendation: external testing as a precondition for releasing advanced AI systems, run centrally through NIST or decentralized via researcher API access.
- OpenAI's commitment: internal and external evaluation across misuse, societal risks, and national security concerns, with a multi-faceted regime for every major public release.
- Independent domain experts: both labs converge on bringing in external subject matter experts rather than relying on internal teams alone.
- National security framing: OpenAI links thorough adversarial evaluation directly to guarding against significant national security threats, raising the policy stakes.
- Governance gap: none of this is binding; codifying a baseline AI red-teaming standard remains an open governance question.
References
- Anthropic, Challenges in Red Teaming AI Systems
- Anthropic, Charting a Path to AI Accountability
- OpenAI, Moving AI Governance Forward
- OpenAI, GPT-4o System Card
- OpenAI, Our Approach to Frontier Risk
- OpenAI, Advancing red teaming with people and AI
- Google AI for Developers, Responsible AI evaluation
- Google AI for Developers, Safety guidance
- NIST, Strengthening AI Agent Hijacking Evaluations
Further reading
Frequently Asked Questions
What is the difference between AI red-teaming and standard safety testing?
AI red-teaming is adversarial and less structured than development evaluations: testers actively try to make the model fail in unexpected ways. Standard safety testing runs fixed benchmarks against known failure categories, while adversarial probing surfaces novel attack paths those benchmarks do not cover, per Google's responsible AI evaluation guidance.
What attack categories do AI red teams typically probe?
Testers probe jailbreaks, direct and indirect prompt injection, harmful content elicitation, and frontier risks such as CBRN and cybersecurity uplift. Anthropic says its frontier program focuses on chemical, biological, radiological, nuclear, and autonomous AI risks specifically.
Is external red-teaming required before releasing an AI model?
No binding global regulation mandates it today, but Anthropic recommends external testing as a precondition for releasing advanced AI systems. OpenAI has committed to a multi-faceted regime drawing on independent domain experts for all major public releases.
Can AI be used to automate red-teaming of other AI models?
Yes. Automated adversarial testing uses a language model to generate inputs at scale against a target model. Anthropic describes this as using AI systems to automatically generate adversarial examples and test the robustness of other AI models, reducing the manual effort of human-only testing while increasing coverage.









