Adversarial and red-team testing for generative AI
Finds the ways a generative or general-purpose model can be made to produce harmful, deceptive or dangerous output before users or attackers do, and feeds the findings back into mitigations.
- Duties satisfied
- 1
- done properly, does the work
- Duties supported
- 7
- contributes; the duty needs more
- Jurisdictions
- 5
- Evidence items
- 3
How is it implemented?
A red team with a mix of security, domain and safety expertise, drawn partly from outside the development group, attacks the model and the product around it using scenario libraries that cover prompt injection, jailbreaks, data leakage, harmful-capability elicitation, bias provocation and misuse of tools or agents. Exercises are scoped and rules of engagement agreed in advance; findings are rated, tracked to closure and regression-tested in the next round. For the largest models the programme includes structured evaluations of dangerous capabilities and the organisation's mitigations are recorded in the safety framework. Results are summarised for governance and, where required, for authorities or downstream integrators.
Which legal duties does it serve?
Satisfies means the control, operated properly, does the work the duty asks for. Supports means it contributes but the duty needs more. The official text decides; open it before relying on either.
European Union 2 duties
-
satisfies Legal requirementManage systemic risk for high-impact general-purpose models
EU AI Act · Articles 51, 52 and 55 · applies from 2 Aug 2025
Model evaluation including adversarial testing.
-
supports Legal requirementAchieve appropriate accuracy, robustness and cybersecurity
EU AI Act · Article 15 · applies from 2 Aug 2026
Adversarial exercises cover poisoning, evasion and extraction attempts.
California (United States) 2 duties
-
supports Legal requirementLarge frontier developers must publish a frontier AI framework
California SB 53 · Business and Professions Code, Chapter 25.1 (as added by SB 53) · applies from 1 Jan 2026
Catastrophic-risk assessment evidence.
-
supports Legal requirementLarge frontier developers must send periodic summaries of catastrophic-risk assessments to the state
California SB 53 · Business and Professions Code Section 22757.12 (as added by SB 53) · applies from 1 Jan 2026
Dangerous-capability evaluations feed the assessment.
South Korea 1 duty
-
supports Legal requirementOperators of AI above the compute threshold must run lifecycle risk management and report safety results
Framework Act on the Development of Artificial Intelligence and Establishment of a Foundation for Trust · Article 32 · applies from 22 Jan 2026
Evidence for the risk identification step.
Texas (United States) 2 duties
-
supports Legal requirementDevelopers and deployers must not use AI to incite self-harm, harm to others or crime
Texas Responsible AI Governance Act (TRAIGA) · Business and Commerce Code Section 551.052 · applies from 1 Jan 2026
Tests for harmful-incitement behaviour in generative systems.
-
supports Legal requirement confidence highDevelopers and distributors must not build AI intended to produce child sexual abuse material or unlawful sexual deepfakes
Texas Responsible AI Governance Act (TRAIGA) · Business and Commerce Code Section 551.057 · applies from 1 Jan 2026
Tests that safeguards block sexual content involving minors.
United States 1 duty
-
supports Voluntary confidence highMeasure and test trustworthiness characteristics (Measure)
NIST AI RMF · MEASURE function
Red-teaming and independent review for generative AI.
What evidence shows it is operating?
| Evidence | Type | What it shows |
|---|---|---|
| Red-team exercise report | Evaluation or test report | Scope, methods, findings by severity and remediation status for one exercise. |
| Red-team rules of engagement and scenario library | Procedure or standard operating process | |
| Adversarial findings tracker | Risk register |
Owner: AI security or safety lead. Frequency: at launch and on material change.
Which risks does it address?
Subdomains of the MIT AI Risk Repository, with the incidents the AI Incident Database has recorded under each. Counts are live; they say how often a risk has materialised, not how well this control prevents it.
- 2.2 AI system security vulnerabilities and attacks Privacy & Security24 incidents · 112 risk entries
- 4.2 Cyberattacks, weapon development or use, and mass harm Malicious actors15 incidents · 82 risk entries
- 4.3 Fraud, scams, and targeted manipulation Malicious actors412 incidents · 77 risk entries
- 1.2 Exposure to toxic content Discrimination & Toxicity91 incidents · 116 risk entries
- 7.2 AI possessing dangerous capabilities AI system safety, failures, & limitations0 incidents · 78 risk entries
Which standards clauses does it correspond to?
Clause numbers only. A reference means the standard asks for overlapping work, so evidence may be reusable; it never means the standard discharges a legal duty.
| Framework | Reference | Note | Confidence |
|---|---|---|---|
| NIST AI RMF | MEASURE 2.6, 2.7; MANAGE 2.2 | Testing safety, security and resilience, including structured adversarial exercises for generative AI. | high |
| ISO/IEC 42001 | Annex A.6.2.4 | medium | |
| OWASP LLM Top 10 | LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM05 Improper Output Handling, LLM06 Excessive Agency, LLM07 System Prompt Leakage | high | |
| MITRE ATLAS | AML.M0020 Generative AI Guardrails, AML.M0022 Generative AI Model Alignment | medium |
Cite this record
AIPolicyTracker (2026). “Adversarial and red-team testing for generative AI”. https://aipolicytracker.org/controls/genai-red-team-testing (accessed 24 September 2026). Data licensed CC BY 4.0.
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.
Frequently asked questions
- Which legal duties does "Adversarial and red-team testing for generative AI" satisfy?
- It is recorded as satisfying 1 and supporting 7 duties across European Union, California (United States), South Korea, Texas (United States) and United States. A mapping means the control, operated properly, does the work the duty asks for; the official text decides whether it is enough.
- What evidence shows this control is operating?
- Red-team exercise report, Red-team rules of engagement and scenario library and Adversarial findings tracker. Owner: AI security or safety lead. Frequency: at launch and on material change.