AIPolicyTracker
Technical measure Owner: Quality or testing lead At launch and on material change

Accuracy, robustness, fairness and security testing

Verifies before release, and again after material change, that a system performs to declared accuracy levels, behaves consistently across groups and conditions, and resists the attacks specific to AI.

Duties satisfied
6
done properly, does the work
Duties supported
9
contributes; the duty needs more
Jurisdictions
9
Evidence items
3

How is it implemented?

A test plan defines the metrics, datasets, thresholds and acceptance criteria for each trustworthiness property the risk assessment flagged: accuracy and calibration, performance across demographic and operational subgroups, resilience to noisy or out-of-distribution input, and security against data and model poisoning, evasion and extraction. Tests are run by people independent of the developers where practical, results are compared against the criteria and failures block release until fixed or formally accepted. Declared performance figures in user documentation are drawn directly from the latest report, and the plan is re-run on retraining or a change in operating context.

Which legal duties does it serve?

Satisfies means the control, operated properly, does the work the duty asks for. Supports means it contributes but the duty needs more. The official text decides; open it before relying on either.

European Union 4 duties

New York (United States) 2 duties

United Kingdom 2 duties

United States 2 duties

Australia 1 duty

Colorado (United States) 1 duty

India 1 duty

Singapore 1 duty

Texas (United States) 1 duty

What evidence shows it is operating?

Evidence this control produces
EvidenceTypeWhat it shows
Pre-release test reportEvaluation or test reportMetrics, subgroup results, robustness and security findings against acceptance criteria.
Test plan and acceptance criteriaProcedure or standard operating process
Release test sign-offApproval or sign-off record

Owner: Quality or testing lead. Frequency: at launch and on material change.

Which risks does it address?

Subdomains of the MIT AI Risk Repository, with the incidents the AI Incident Database has recorded under each. Counts are live; they say how often a risk has materialised, not how well this control prevents it.

Which standards clauses does it correspond to?

Clause numbers only. A reference means the standard asks for overlapping work, so evidence may be reusable; it never means the standard discharges a legal duty.

Framework references
FrameworkReferenceNoteConfidence
ISO/IEC 42001Annex A.6.2.4Verification and validation.high
NIST AI RMFMEASURE 2.5, 2.6, 2.7, 2.11high
ISO/IEC 27001Clause 8.1; Annex A 8.29 Security testing in developmentmedium
MITRE ATLASAML.M0008 Validate ML Model, AML.M0003 Model Hardening, AML.M0015 Adversarial Input Detectionmedium
OWASP LLM Top 10LLM04 Data and Model Poisoningmedium

Cite this record

AIPolicyTracker (2026). “Accuracy, robustness, fairness and security testing”. https://aipolicytracker.org/controls/accuracy-robustness-security-testing (accessed 24 September 2026). Data licensed CC BY 4.0.

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.

Frequently asked questions

Which legal duties does "Accuracy, robustness, fairness and security testing" satisfy?
It is recorded as satisfying 6 and supporting 9 duties across European Union, New York (United States), United Kingdom, United States, Australia, Colorado (United States), India, Singapore and Texas (United States). A mapping means the control, operated properly, does the work the duty asks for; the official text decides whether it is enough.
What evidence shows this control is operating?
Pre-release test report, Test plan and acceptance criteria and Release test sign-off. Owner: Quality or testing lead. Frequency: at launch and on material change.