AIPolicyTracker
AssessmentFree · sent to your work emailEU AI ActNIST AI RMF

AI Red-Team and Evaluation Test Plan

Formats: XLSX and DOCX · Version v1 · Built from dataset bb068ecd9dad · CC BY 4.0. You may use, adapt and share this template, including commercially, with attribution to aipolicytracker.org.

In brief

The AI Red-Team and Evaluation Test Plan is a free XLSX and DOCX assessment for EU AI Act and NIST AI RMF. A test plan and case log for adversarial and evaluation testing: accuracy, robustness, prompt injection, bias, harmful content, privacy leakage and misuse, with severity and retest tracking. It is licensed CC BY 4.0 and is not legal advice.

Format
XLSX and DOCX · Assessment
Version
v1, built 28 Sep 2026
Duties cited
8 from 5 instruments
Rows from the records
8
Written for
General-purpose AI model provider, Provider / developer, Deployer / user organisation
Price and licence
Free · CC BY 4.0

What's inside

  • Test cases sheet with area, scenario, expected and observed behaviour, outcome and severity
  • Plan document: scope, independence, method, exit criteria
  • Testing duties sheet: safety-testing and robustness duties on record

Preview

The sheets and sections of version v1, as built. Columns marked ▾ have a dropdown; ƒ is a formula.

Sheet: Test cases · 9 columns · blank, 150 rows ready to fill
First rows of the Test cases sheet
Test IDArea ▾ScenarioExpected behaviourObservedOutcome ▾Severity if failed ▾RemediationRetested
Rows are yours to fill; the dropdowns, formulas and colour rules are already in place.

One row per test case. Plan the cases before the run; record what was observed, not what was hoped for.

Sheet: Testing duties · 14 columns · 8 rows from the records
First rows of the Testing duties sheet
DutyCategoryInstrumentJurisdictionWho it bindsNatureSource referenceApplies fromWhat it requiresEvidence a reviewer expectsISO/IEC 42001NIST AI RMFVerificationRecord
Large frontier developers must publish a frontier AI frameworkSafety testing and evaluationCalifornia SB 53California (United States)General-purpose AI model providerLegal requirementBusiness and Professions Code, Chapter 25.1 (as added by SB 53)2026-01-01Large frontier developers must publish and maintain a framework describing how they incorporate national and international standards, assess catastrophic risk, Published frontier AI frameworkGOVERN 1.x; NIST AI 600-1Source-linkedhttps://aipolicytracker.org/obligations/us-california-sb-53-frontier-ai-framework
Large frontier developers must send periodic summaries of catastrophic-risk assessments to the stateSafety testing and evaluationCalifornia SB 53California (United States)General-purpose AI model providerLegal requirementBusiness and Professions Code Section 22757.12 (as added by SB 53)2026-01-01A large frontier developer must transmit to the California Office of Emergency Services, on the periodic schedule the statute sets, a summary of any assessment Internal-use catastrophic risk assessment summary sent to the Office of Emergency ServicesClause 9.1; Annex A.8.3MEASURE 2.6, MANAGE 1.2, GOVERN 4.3Verified against the official source 26 Sep 2026https://aipolicytracker.org/obligations/us-california-sb-53-catastrophic-risk-assessment-summaries
Achieve appropriate accuracy, robustness and cybersecurityAccuracy, robustness and cybersecurityEU AI ActEuropean UnionProvider / developerLegal requirementArticle 152027-12-02High-risk AI systems must achieve an appropriate level of accuracy, robustness and cybersecurity and perform consistently throughout their lifecycle. Accuracy lAccuracy metrics and test evidence; AI security assessmentAnnex A controls on AI system verification and validationMEASURE 2.5, 2.6, 2.7Source-linkedhttps://aipolicytracker.org/obligations/eu-ai-act-accuracy-robustness-cybersecurity
Manage systemic risk for high-impact general-purpose modelsSafety testing and evaluationEU AI ActEuropean UnionGeneral-purpose AI model providerLegal requirementArticles 51, 52 and 552025-08-02A general-purpose model is presumed to have systemic risk when the cumulative compute used for training exceeds 10^25 floating-point operations, or when the ComModel evaluation and red-teaming reports; Commission notification recordMEASURE 2.x; NIST AI 600-1Source-linkedhttps://aipolicytracker.org/obligations/eu-ai-act-gpai-systemic-risk
Providers of systemic-risk GPAI models must secure the model and its infrastructureAccuracy, robustness and cybersecurityEU AI ActEuropean UnionGeneral-purpose AI model providerLegal requirementArticle 55(1)(d)2025-08-02Providers of general-purpose AI models with systemic risk must ensure an adequate level of cybersecurity protection for the model and for the physical infrastruModel and infrastructure security assessment; Weight access control recordsMEASURE 2.7, MANAGE 2.2Verified against the official source 26 Sep 2026https://aipolicytracker.org/obligations/eu-ai-act-art-55-systemic-risk-cybersecurity
Employers and employment agencies must obtain an independent bias audit before using an automated employment decision toolAccuracy, robustness and cybersecurityNYC Local Law 144 (automated employment decision tools)New York (United States)Deployer / user organisationLegal requirementNYC Administrative Code Section 20-871(a)(1); 6 RCNY Section 5-3012023-07-05An automated employment decision tool may not be used to screen candidates or employees for hiring or promotion in New York City unless it has been the subject Independent bias audit report; Audit data extract and category mappingAnnex A.6.2.4; Clause 9.2MEASURE 2.11, MEASURE 1.3Verified against the official source 26 Sep 2026https://aipolicytracker.org/obligations/us-new-york-city-local-law-144-bias-audit

Document outline (DOCX)

  1. AI red-team and evaluation test plan
  2. Scope
  3. Team and independence
  4. Method
  5. Exit criteria
  6. The duties this plan serves
  7. Large frontier developers must publish a frontier AI framework
  8. Large frontier developers must send periodic summaries of catastrophic-risk assessments to the state
  9. Achieve appropriate accuracy, robustness and cybersecurity
  10. Manage systemic risk for high-impact general-purpose models
  11. Providers of systemic-risk GPAI models must secure the model and its infrastructure
  12. Employers and employment agencies must obtain an independent bias audit before using an automated employment decision tool
  13. Operators of AI above the compute threshold must run lifecycle risk management and report safety results
  14. Measure and test trustworthiness characteristics (Measure)

How to use it

  1. 1Request the files. Enter your name, company and work email in the form on this page. The XLSX and DOCX download links arrive by email and work for 7 days.
  2. 2Read the README page. It states the version (v1), the dataset it was built from and the licence, so anyone reviewing your copy knows which records it reflects.
  3. 3Fill in your rows. Complete the "Test cases" sheet for your own systems. Dropdowns, formulas and colour rules are already set.
  4. 4Check the duties against your situation. The "Testing duties" sheet lists the recorded duties with their source references. Mark which apply to you and follow each link to the official text.
  5. 5Complete the document. Work through the DOCX sections (AI red-team and evaluation test plan, The duties this plan serves) and replace each placeholder with your organisation's answer.
  6. 6Keep the evidence and watch for new versions. Link each completed row to the evidence that supports it. When the law on record changes, this template gets a new version and a changelog on this page.

Duties this template covers (8)

Each is cited in the file with its source reference and a link back to the record.

Legal basis

Version history

Versions of AI Red-Team and Evaluation Test Plan
VersionBuiltDatasetWhat changed
v1bb068ecd9dadFirst version, built from dataset bb068ecd9dad.

Only the latest version is served. A rebuild that changes the content adds a version; a rebuild that does not is skipped.

Frequently asked questions

What is in the AI Red-Team and Evaluation Test Plan?

Test cases sheet with area, scenario, expected and observed behaviour, outcome and severity. Plan document: scope, independence, method, exit criteria. Testing duties sheet: safety-testing and robustness duties on record.

Which duties does it cite?

8 recorded duties from California SB 53, EU AI Act, NYC Local Law 144 (automated employment decision tools) and Framework Act on the Development of Artificial Intelligence and Establishment of a Foundation for Trust, including Business and Professions Code, Chapter 25.1 (as added by SB 53), Business and Professions Code Section 22757.12 (as added by SB 53), Article 15, Articles 51, 52 and 55, Article 55(1)(d) and NYC Administrative Code Section 20-871(a)(1); 6 RCNY Section 5-301. Each row links to the record, and the record to the official source.

Who is it for?

The duties it cites fall on general-purpose ai model provider, provider / developer and deployer / user organisation. Whoever owns AI governance for those roles usually completes it, with the system owner supplying the facts.

Is it free?

Yes. Request the XLSX and DOCX with your work email on this page; the download links arrive by email, valid for 7 days. No account and no charge. Licensed CC BY 4.0. You may use, adapt and share this template, including commercially, with attribution to aipolicytracker.org.

How will I know when it changes?

Version v1 was built on 28 September 2026. The library is rebuilt daily; when a change to the records reaches this template it gets the next version, a changelog below and an entry in the templates feed.

Does completing it make us compliant?

No. It is an informational resource, not legal advice; it helps produce the evidence a regulator, customer or auditor asks for. Whether a duty applies to you is a judgement the template cannot make.

Disclaimer: informational only, not legal advice. Verify every claim against the linked official sources and consult a qualified lawyer before acting.