Accuracy, robustness, fairness and security testing
Verifies before release, and again after material change, that a system performs to declared accuracy levels, behaves consistently across groups and conditions, and resists the attacks specific to AI.
- Duties satisfied
- 6
- done properly, does the work
- Duties supported
- 9
- contributes; the duty needs more
- Jurisdictions
- 9
- Evidence items
- 3
How is it implemented?
A test plan defines the metrics, datasets, thresholds and acceptance criteria for each trustworthiness property the risk assessment flagged: accuracy and calibration, performance across demographic and operational subgroups, resilience to noisy or out-of-distribution input, and security against data and model poisoning, evasion and extraction. Tests are run by people independent of the developers where practical, results are compared against the criteria and failures block release until fixed or formally accepted. Declared performance figures in user documentation are drawn directly from the latest report, and the plan is re-run on retraining or a change in operating context.
Which legal duties does it serve?
Satisfies means the control, operated properly, does the work the duty asks for. Supports means it contributes but the duty needs more. The official text decides; open it before relying on either.
European Union 4 duties
-
satisfies Legal requirement confidence highAchieve appropriate accuracy, robustness and cybersecurity
EU AI Act · Article 15 · applies from 2 Aug 2026
Declared accuracy metrics, robustness and AI-specific security testing come from the release test programme.
-
satisfies Legal requirementProviders of systemic-risk GPAI models must secure the model and its infrastructure
EU AI Act · Article 55(1)(d) · applies from 2 Aug 2025
Security testing of the model and its serving infrastructure.
-
supports Legal requirement confidence highEstablish a risk management system for high-risk AI
EU AI Act · Article 9 · applies from 2 Aug 2026
Testing confirms that mitigations work.
-
supports Legal requirementProviders of generative AI must mark synthetic output as artificially generated in a machine-readable way
EU AI Act · Article 50(2) · applies from 2 Aug 2026
Tests that the mark is robust and reliable.
New York (United States) 2 duties
-
satisfies Legal requirement confidence highEmployers and employment agencies must obtain an independent bias audit before using an automated employment decision tool
NYC Local Law 144 (automated employment decision tools) · NYC Administrative Code Section 20-871(a)(1); 6 RCNY Section 5-301 · applies from 5 Jul 2023
Disparate-impact testing performed by an independent auditor on the statutory metrics.
-
supports Legal requirement confidence highEmployers and employment agencies must publish a summary of the bias audit results
NYC Local Law 144 (automated employment decision tools) · NYC Administrative Code Section 20-871(a)(2); 6 RCNY Section 5-302 · applies from 5 Jul 2023
Produces the results being summarised.
United Kingdom 2 duties
-
satisfies VoluntaryEnsure AI systems are safe, secure and robust throughout their lifecycle
UK AI regulation framework · Principle 1, Part 3
Security and robustness evidence to show a regulator.
-
satisfies VoluntaryUse AI in ways that are fair and do not discriminate unlawfully
UK AI regulation framework · Principle 3, Part 3
Disparate-outcome testing across protected characteristics before and after deployment.
United States 2 duties
-
satisfies VoluntaryMeasure and test trustworthiness characteristics (Measure)
NIST AI RMF · MEASURE function
Metrics and test methods for each trustworthiness characteristic.
-
supports Legal requirementApply minimum risk-management practices to high-impact AI
OMB M-25-21 · Section 4
Pre-deployment test evidence agencies can reuse.
Australia 1 duty
-
supports Voluntary confidence highTest and monitor systems, enable human control, and be transparent with users (guardrails 4 to 6)
Australian Voluntary AI Safety Standard · Guardrails 4, 5 and 6
Guardrail 3: pre-deployment testing.
Colorado (United States) 1 duty
-
supports Legal requirement confidence highDevelopers must use reasonable care to avoid algorithmic discrimination
Colorado AI Act · C.R.S. 6-1-1702(1) · applies from 30 Jun 2026
Disparate-impact testing across protected classes.
India 1 duty
-
supports Legal requirementImplement reasonable security safeguards and notify breaches
India DPDP Act · Section 8(5) and 8(6); DPDP Rules on breach intimation
Reasonable security safeguards evidenced by testing.
Singapore 1 duty
-
supports VoluntaryManage data quality, model development and monitoring across the lifecycle
Singapore Model AI Governance Framework · Second edition, Part on operations management
Repeatability and robustness evidence before release.
Texas (United States) 1 duty
-
supports Legal requirementDevelopers and deployers must not use AI with the intent to unlawfully discriminate against a protected class
Texas Responsible AI Governance Act (TRAIGA) · Business and Commerce Code Section 551.056 · applies from 1 Jan 2026
Fairness testing evidence supports the absence of intent.
What evidence shows it is operating?
| Evidence | Type | What it shows |
|---|---|---|
| Pre-release test report | Evaluation or test report | Metrics, subgroup results, robustness and security findings against acceptance criteria. |
| Test plan and acceptance criteria | Procedure or standard operating process | |
| Release test sign-off | Approval or sign-off record |
Owner: Quality or testing lead. Frequency: at launch and on material change.
Which risks does it address?
Subdomains of the MIT AI Risk Repository, with the incidents the AI Incident Database has recorded under each. Counts are live; they say how often a risk has materialised, not how well this control prevents it.
- 7.3 Lack of capability or robustness AI system safety, failures, & limitations305 incidents · 126 risk entries
- 1.3 Unequal performance across groups Discrimination & Toxicity34 incidents · 17 risk entries
- 2.2 AI system security vulnerabilities and attacks Privacy & Security24 incidents · 112 risk entries
- 1.1 Unfair discrimination and misrepresentation Discrimination & Toxicity118 incidents · 83 risk entries
Which standards clauses does it correspond to?
Clause numbers only. A reference means the standard asks for overlapping work, so evidence may be reusable; it never means the standard discharges a legal duty.
| Framework | Reference | Note | Confidence |
|---|---|---|---|
| ISO/IEC 42001 | Annex A.6.2.4 | Verification and validation. | high |
| NIST AI RMF | MEASURE 2.5, 2.6, 2.7, 2.11 | high | |
| ISO/IEC 27001 | Clause 8.1; Annex A 8.29 Security testing in development | medium | |
| MITRE ATLAS | AML.M0008 Validate ML Model, AML.M0003 Model Hardening, AML.M0015 Adversarial Input Detection | medium | |
| OWASP LLM Top 10 | LLM04 Data and Model Poisoning | medium |
Cite this record
AIPolicyTracker (2026). “Accuracy, robustness, fairness and security testing”. https://aipolicytracker.org/controls/accuracy-robustness-security-testing (accessed 24 September 2026). Data licensed CC BY 4.0.
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.
Frequently asked questions
- Which legal duties does "Accuracy, robustness, fairness and security testing" satisfy?
- It is recorded as satisfying 6 and supporting 9 duties across European Union, New York (United States), United Kingdom, United States, Australia, Colorado (United States), India, Singapore and Texas (United States). A mapping means the control, operated properly, does the work the duty asks for; the official text decides whether it is enough.
- What evidence shows this control is operating?
- Pre-release test report, Test plan and acceptance criteria and Release test sign-off. Owner: Quality or testing lead. Frequency: at launch and on material change.