AIPolicyTracker
Technical measure Owner: AI security or safety lead At launch and on material change

Adversarial and red-team testing for generative AI

Finds the ways a generative or general-purpose model can be made to produce harmful, deceptive or dangerous output before users or attackers do, and feeds the findings back into mitigations.

Duties satisfied
1
done properly, does the work
Duties supported
7
contributes; the duty needs more
Jurisdictions
5
Evidence items
3

How is it implemented?

A red team with a mix of security, domain and safety expertise, drawn partly from outside the development group, attacks the model and the product around it using scenario libraries that cover prompt injection, jailbreaks, data leakage, harmful-capability elicitation, bias provocation and misuse of tools or agents. Exercises are scoped and rules of engagement agreed in advance; findings are rated, tracked to closure and regression-tested in the next round. For the largest models the programme includes structured evaluations of dangerous capabilities and the organisation's mitigations are recorded in the safety framework. Results are summarised for governance and, where required, for authorities or downstream integrators.

Which legal duties does it serve?

Satisfies means the control, operated properly, does the work the duty asks for. Supports means it contributes but the duty needs more. The official text decides; open it before relying on either.

European Union 2 duties

California (United States) 2 duties

Texas (United States) 2 duties

United States 1 duty

What evidence shows it is operating?

Evidence this control produces
EvidenceTypeWhat it shows
Red-team exercise reportEvaluation or test reportScope, methods, findings by severity and remediation status for one exercise.
Red-team rules of engagement and scenario libraryProcedure or standard operating process
Adversarial findings trackerRisk register

Owner: AI security or safety lead. Frequency: at launch and on material change.

Which risks does it address?

Subdomains of the MIT AI Risk Repository, with the incidents the AI Incident Database has recorded under each. Counts are live; they say how often a risk has materialised, not how well this control prevents it.

Which standards clauses does it correspond to?

Clause numbers only. A reference means the standard asks for overlapping work, so evidence may be reusable; it never means the standard discharges a legal duty.

Framework references
FrameworkReferenceNoteConfidence
NIST AI RMFMEASURE 2.6, 2.7; MANAGE 2.2Testing safety, security and resilience, including structured adversarial exercises for generative AI.high
ISO/IEC 42001Annex A.6.2.4medium
OWASP LLM Top 10LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM05 Improper Output Handling, LLM06 Excessive Agency, LLM07 System Prompt Leakagehigh
MITRE ATLASAML.M0020 Generative AI Guardrails, AML.M0022 Generative AI Model Alignmentmedium

Cite this record

AIPolicyTracker (2026). “Adversarial and red-team testing for generative AI”. https://aipolicytracker.org/controls/genai-red-team-testing (accessed 24 September 2026). Data licensed CC BY 4.0.

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.

Frequently asked questions

Which legal duties does "Adversarial and red-team testing for generative AI" satisfy?
It is recorded as satisfying 1 and supporting 7 duties across European Union, California (United States), South Korea, Texas (United States) and United States. A mapping means the control, operated properly, does the work the duty asks for; the official text decides whether it is enough.
What evidence shows this control is operating?
Red-team exercise report, Red-team rules of engagement and scenario library and Adversarial findings tracker. Owner: AI security or safety lead. Frequency: at launch and on material change.