# Adversarial and red-team testing for generative AI

- **Record type**: Control
- **Kind**: Technical measure
- **Owner**: AI security or safety lead
- **Frequency**: At launch and on material change
- **Duties served**: 8

## What the control achieves

Finds the ways a generative or general-purpose model can be made to produce harmful, deceptive or dangerous output before users or attackers do, and feeds the findings back into mitigations.

## How it is typically implemented

A red team with a mix of security, domain and safety expertise, drawn partly from outside the development group, attacks the model and the product around it using scenario libraries that cover prompt injection, jailbreaks, data leakage, harmful-capability elicitation, bias provocation and misuse of tools or agents. Exercises are scoped and rules of engagement agreed in advance; findings are rated, tracked to closure and regression-tested in the next round. For the largest models the programme includes structured evaluations of dangerous capabilities and the organisation's mitigations are recorded in the safety framework. Results are summarised for governance and, where required, for authorities or downstream integrators.

## Evidence it produces

- Red-team exercise report (evaluation_report): Scope, methods, findings by severity and remediation status for one exercise.
- Red-team rules of engagement and scenario library (procedure)
- Adversarial findings tracker (risk_register)

## Legal duties this control serves

- Achieve appropriate accuracy, robustness and cybersecurity — EU AI Act, European Union (supports): https://aipolicytracker.org/obligations/eu-ai-act-accuracy-robustness-cybersecurity
- Manage systemic risk for high-impact general-purpose models — EU AI Act, European Union (satisfies): https://aipolicytracker.org/obligations/eu-ai-act-gpai-systemic-risk
- Operators of AI above the compute threshold must run lifecycle risk management and report safety results — Framework Act on the Development of Artificial Intelligence and Establishment of a Foundation for Trust, South Korea (supports): https://aipolicytracker.org/obligations/south-korea-ai-basic-act-art-32-safety-measures-for-high-performance-ai
- Large frontier developers must publish a frontier AI framework — California SB 53, California (United States) (supports): https://aipolicytracker.org/obligations/us-california-sb-53-frontier-ai-framework
- Large frontier developers must send periodic summaries of catastrophic-risk assessments to the state — California SB 53, California (United States) (supports): https://aipolicytracker.org/obligations/us-california-sb-53-catastrophic-risk-assessment-summaries
- Developers and deployers must not use AI to incite self-harm, harm to others or crime — Texas Responsible AI Governance Act (TRAIGA), Texas (United States) (supports): https://aipolicytracker.org/obligations/us-texas-responsible-ai-governance-act-traiga-manipulation-prohibition
- Developers and distributors must not build AI intended to produce child sexual abuse material or unlawful sexual deepfakes — Texas Responsible AI Governance Act (TRAIGA), Texas (United States) (supports): https://aipolicytracker.org/obligations/us-texas-responsible-ai-governance-act-traiga-sexual-content-and-csam-prohibition
- Measure and test trustworthiness characteristics (Measure) — NIST AI RMF, United States (supports): https://aipolicytracker.org/obligations/us-nist-ai-rmf-measure

## Standards clauses it corresponds to (clause numbers only)

- NIST AI RMF 1.0: MEASURE 2.6, 2.7; MANAGE 2.2 — Testing safety, security and resilience, including structured adversarial exercises for generative AI.
- ISO/IEC 42001:2023: Annex A.6.2.4
- OWASP Top 10 for LLM Applications: LLM01 Prompt Injection, LLM02 Sensitive Information Disclosure, LLM05 Improper Output Handling, LLM06 Excessive Agency, LLM07 System Prompt Leakage
- MITRE ATLAS: AML.M0020 Generative AI Guardrails, AML.M0022 Generative AI Model Alignment

## MIT AI Risk Repository subdomains addressed

2.2, 4.2, 4.3, 1.2, 7.2

## Provenance

- **Record page**: https://aipolicytracker.org/controls/genai-red-team-testing
- **Official source**: none recorded — this record is incomplete, see https://aipolicytracker.org/gaps
- **Review status**: pending review
- **Confidence**: high
- **Facts last confirmed**: never confirmed against the official source
- **Retrieved**: 2026-09-24
- **Licence**: https://creativecommons.org/licenses/by/4.0/

> This record is a structured summary with a link to the official text. It is not legal advice. Open the official source before relying on any date or duty. How current each record type must be is published at https://aipolicytracker.org/verification; what a record must carry at all is published at https://aipolicytracker.org/coverage.
