AI Red-Team and Evaluation Test Plan
In brief
The AI Red-Team and Evaluation Test Plan is a free XLSX and DOCX assessment for EU AI Act and NIST AI RMF. A test plan and case log for adversarial and evaluation testing: accuracy, robustness, prompt injection, bias, harmful content, privacy leakage and misuse, with severity and retest tracking. It is licensed CC BY 4.0 and is not legal advice.
- Format
- XLSX and DOCX · Assessment
- Version
- v1, built 28 Sep 2026
- Duties cited
- 8 from 5 instruments
- Rows from the records
- 8
- Frameworks
- EU AI Act, NIST AI RMF
- Written for
- General-purpose AI model provider, Provider / developer, Deployer / user organisation
- Price and licence
- Free · CC BY 4.0
What's inside
- Test cases sheet with area, scenario, expected and observed behaviour, outcome and severity
- Plan document: scope, independence, method, exit criteria
- Testing duties sheet: safety-testing and robustness duties on record
Preview
The sheets and sections of version v1, as built. Columns marked ▾ have a dropdown; ƒ is a formula.
Sheet: Test cases
| Test ID | Area ▾ | Scenario | Expected behaviour | Observed | Outcome ▾ | Severity if failed ▾ | Remediation | Retested |
|---|---|---|---|---|---|---|---|---|
| Rows are yours to fill; the dropdowns, formulas and colour rules are already in place. | ||||||||
One row per test case. Plan the cases before the run; record what was observed, not what was hoped for.
Sheet: Testing duties
| Duty | Category | Instrument | Jurisdiction | Who it binds | Nature | Source reference | Applies from | What it requires | Evidence a reviewer expects | ISO/IEC 42001 | NIST AI RMF | Verification | Record |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Large frontier developers must publish a frontier AI framework | Safety testing and evaluation | California SB 53 | California (United States) | General-purpose AI model provider | Legal requirement | Business and Professions Code, Chapter 25.1 (as added by SB 53) | 2026-01-01 | Large frontier developers must publish and maintain a framework describing how they incorporate national and international standards, assess catastrophic risk, | Published frontier AI framework | GOVERN 1.x; NIST AI 600-1 | Source-linked | https://aipolicytracker.org/obligations/us-california-sb-53-frontier-ai-framework | |
| Large frontier developers must send periodic summaries of catastrophic-risk assessments to the state | Safety testing and evaluation | California SB 53 | California (United States) | General-purpose AI model provider | Legal requirement | Business and Professions Code Section 22757.12 (as added by SB 53) | 2026-01-01 | A large frontier developer must transmit to the California Office of Emergency Services, on the periodic schedule the statute sets, a summary of any assessment | Internal-use catastrophic risk assessment summary sent to the Office of Emergency Services | Clause 9.1; Annex A.8.3 | MEASURE 2.6, MANAGE 1.2, GOVERN 4.3 | Verified against the official source 26 Sep 2026 | https://aipolicytracker.org/obligations/us-california-sb-53-catastrophic-risk-assessment-summaries |
| Achieve appropriate accuracy, robustness and cybersecurity | Accuracy, robustness and cybersecurity | EU AI Act | European Union | Provider / developer | Legal requirement | Article 15 | 2027-12-02 | High-risk AI systems must achieve an appropriate level of accuracy, robustness and cybersecurity and perform consistently throughout their lifecycle. Accuracy l | Accuracy metrics and test evidence; AI security assessment | Annex A controls on AI system verification and validation | MEASURE 2.5, 2.6, 2.7 | Source-linked | https://aipolicytracker.org/obligations/eu-ai-act-accuracy-robustness-cybersecurity |
| Manage systemic risk for high-impact general-purpose models | Safety testing and evaluation | EU AI Act | European Union | General-purpose AI model provider | Legal requirement | Articles 51, 52 and 55 | 2025-08-02 | A general-purpose model is presumed to have systemic risk when the cumulative compute used for training exceeds 10^25 floating-point operations, or when the Com | Model evaluation and red-teaming reports; Commission notification record | MEASURE 2.x; NIST AI 600-1 | Source-linked | https://aipolicytracker.org/obligations/eu-ai-act-gpai-systemic-risk | |
| Providers of systemic-risk GPAI models must secure the model and its infrastructure | Accuracy, robustness and cybersecurity | EU AI Act | European Union | General-purpose AI model provider | Legal requirement | Article 55(1)(d) | 2025-08-02 | Providers of general-purpose AI models with systemic risk must ensure an adequate level of cybersecurity protection for the model and for the physical infrastru | Model and infrastructure security assessment; Weight access control records | MEASURE 2.7, MANAGE 2.2 | Verified against the official source 26 Sep 2026 | https://aipolicytracker.org/obligations/eu-ai-act-art-55-systemic-risk-cybersecurity | |
| Employers and employment agencies must obtain an independent bias audit before using an automated employment decision tool | Accuracy, robustness and cybersecurity | NYC Local Law 144 (automated employment decision tools) | New York (United States) | Deployer / user organisation | Legal requirement | NYC Administrative Code Section 20-871(a)(1); 6 RCNY Section 5-301 | 2023-07-05 | An automated employment decision tool may not be used to screen candidates or employees for hiring or promotion in New York City unless it has been the subject | Independent bias audit report; Audit data extract and category mapping | Annex A.6.2.4; Clause 9.2 | MEASURE 2.11, MEASURE 1.3 | Verified against the official source 26 Sep 2026 | https://aipolicytracker.org/obligations/us-new-york-city-local-law-144-bias-audit |
Document outline (DOCX)
- AI red-team and evaluation test plan
- Scope
- Team and independence
- Method
- Exit criteria
- The duties this plan serves
- Large frontier developers must publish a frontier AI framework
- Large frontier developers must send periodic summaries of catastrophic-risk assessments to the state
- Achieve appropriate accuracy, robustness and cybersecurity
- Manage systemic risk for high-impact general-purpose models
- Providers of systemic-risk GPAI models must secure the model and its infrastructure
- Employers and employment agencies must obtain an independent bias audit before using an automated employment decision tool
- Operators of AI above the compute threshold must run lifecycle risk management and report safety results
- Measure and test trustworthiness characteristics (Measure)
How to use it
- 1Request the files. Enter your name, company and work email in the form on this page. The XLSX and DOCX download links arrive by email and work for 7 days.
- 2Read the README page. It states the version (v1), the dataset it was built from and the licence, so anyone reviewing your copy knows which records it reflects.
- 3Fill in your rows. Complete the "Test cases" sheet for your own systems. Dropdowns, formulas and colour rules are already set.
- 4Check the duties against your situation. The "Testing duties" sheet lists the recorded duties with their source references. Mark which apply to you and follow each link to the official text.
- 5Complete the document. Work through the DOCX sections (AI red-team and evaluation test plan, The duties this plan serves) and replace each placeholder with your organisation's answer.
- 6Keep the evidence and watch for new versions. Link each completed row to the evidence that supports it. When the law on record changes, this template gets a new version and a changelog on this page.
Duties this template covers (8)
Each is cited in the file with its source reference and a link back to the record.
- Large frontier developers must publish a frontier AI framework
- Large frontier developers must send periodic summaries of catastrophic-risk assessments to the state
- Achieve appropriate accuracy, robustness and cybersecurity
- Manage systemic risk for high-impact general-purpose models
- Providers of systemic-risk GPAI models must secure the model and its infrastructure
- Employers and employment agencies must obtain an independent bias audit before using an automated employment decision tool
- Operators of AI above the compute threshold must run lifecycle risk management and report safety results
- Measure and test trustworthiness characteristics (Measure)
Legal basis
Version history
| Version | Built | Dataset | What changed |
|---|---|---|---|
| v1 | bb068ecd9dad | First version, built from dataset bb068ecd9dad. |
Only the latest version is served. A rebuild that changes the content adds a version; a rebuild that does not is skipped.
Frequently asked questions
What is in the AI Red-Team and Evaluation Test Plan?
Test cases sheet with area, scenario, expected and observed behaviour, outcome and severity. Plan document: scope, independence, method, exit criteria. Testing duties sheet: safety-testing and robustness duties on record.
Which duties does it cite?
8 recorded duties from California SB 53, EU AI Act, NYC Local Law 144 (automated employment decision tools) and Framework Act on the Development of Artificial Intelligence and Establishment of a Foundation for Trust, including Business and Professions Code, Chapter 25.1 (as added by SB 53), Business and Professions Code Section 22757.12 (as added by SB 53), Article 15, Articles 51, 52 and 55, Article 55(1)(d) and NYC Administrative Code Section 20-871(a)(1); 6 RCNY Section 5-301. Each row links to the record, and the record to the official source.
Who is it for?
The duties it cites fall on general-purpose ai model provider, provider / developer and deployer / user organisation. Whoever owns AI governance for those roles usually completes it, with the system owner supplying the facts.
Is it free?
Yes. Request the XLSX and DOCX with your work email on this page; the download links arrive by email, valid for 7 days. No account and no charge. Licensed CC BY 4.0. You may use, adapt and share this template, including commercially, with attribution to aipolicytracker.org.
How will I know when it changes?
Version v1 was built on 28 September 2026. The library is rebuilt daily; when a change to the records reaches this template it gets the next version, a changelog below and an entry in the templates feed.
Does completing it make us compliant?
No. It is an informational resource, not legal advice; it helps produce the evidence a regulator, customer or auditor asks for. Whether a duty applies to you is a judgement the template cannot make.
Disclaimer: informational only, not legal advice. Verify every claim against the linked official sources and consult a qualified lawyer before acting.