Data governance and dataset documentation
Makes sure the data used to train, tune, validate and test an AI system is fit for purpose, understood and documented, so that quality and bias problems are found in the data before they surface in decisions.
- Duties satisfied
- 3
- done properly, does the work
- Duties supported
- 5
- contributes; the duty needs more
- Jurisdictions
- 5
- Evidence items
- 3
How is it implemented?
Each dataset used in building a system gets a documentation sheet describing where it came from, how it was collected and labelled, what it is meant to represent, known gaps and the cleaning and transformation steps applied. Data owners check representativeness against the intended population and run bias and quality checks before a dataset is approved for use. Access, retention and change control apply to datasets in the same way as to code, and the documentation is versioned alongside the model that consumed it. Deployers who supply their own input data apply a lighter version of the same checks to confirm relevance.
Which legal duties does it serve?
Satisfies means the control, operated properly, does the work the duty asks for. Supports means it contributes but the duty needs more. The official text decides; open it before relying on either.
European Union 3 duties
-
satisfies Legal requirement confidence highApply data governance and quality criteria to training, validation and testing data
EU AI Act · Article 10 · applies from 2 Aug 2026
Dataset documentation, representativeness and bias checks per dataset.
-
satisfies Legal requirementDeployers must ensure input data they control is relevant and representative
EU AI Act · Article 26(4) · applies from 2 Aug 2026
Applied to operational input data rather than training data.
-
supports Legal requirement confidence highDraw up technical documentation before placing a high-risk system on the market
EU AI Act · Article 11 and Annex IV · applies from 2 Aug 2026
Dataset sheets supply the data sections.
Singapore 2 duties
-
satisfies VoluntaryManage data quality, model development and monitoring across the lifecycle
Singapore Model AI Governance Framework · Second edition, Part on operations management
Data lineage and dataset bias minimisation.
-
supports Legal requirementIdentify consent or an applicable PDPA exception before using personal data in AI
PDPC AI advisory guidelines · Advisory guidelines, sections on consent, business improvement and research exceptions
Dataset sheets record the legal basis alongside provenance.
India 1 duty
-
supports Legal requirementProcess personal data only with valid consent or a legitimate use, after notice
India DPDP Act · Sections 4 to 7
Dataset documentation records consent status per source.
New York (United States) 1 duty
-
supports Legal requirementEmployers and employment agencies must disclose the data collected and their retention policy for the tool
NYC Local Law 144 (automated employment decision tools) · NYC Administrative Code Section 20-871(b)(3); 6 RCNY Section 5-303 · applies from 5 Jul 2023
Source of the data inventory disclosed.
United Kingdom 1 duty
-
supports VoluntaryUse AI in ways that are fair and do not discriminate unlawfully
UK AI regulation framework · Principle 3, Part 3
Bias checks on training data.
What evidence shows it is operating?
| Evidence | Type | What it shows |
|---|---|---|
| Dataset documentation sheet | Dataset documentation | Provenance, collection, labelling, representativeness, limitations and preparation steps for one dataset. |
| Data quality and bias check report | Evaluation or test report | |
| Dataset approval for use | Approval or sign-off record |
Owner: Data governance lead. Frequency: once per ai system.
Which risks does it address?
Subdomains of the MIT AI Risk Repository, with the incidents the AI Incident Database has recorded under each. Counts are live; they say how often a risk has materialised, not how well this control prevents it.
- 1.1 Unfair discrimination and misrepresentation Discrimination & Toxicity118 incidents · 83 risk entries
- 1.3 Unequal performance across groups Discrimination & Toxicity34 incidents · 17 risk entries
- 7.3 Lack of capability or robustness AI system safety, failures, & limitations305 incidents · 126 risk entries
- 2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information Privacy & Security88 incidents · 80 risk entries
Which standards clauses does it correspond to?
Clause numbers only. A reference means the standard asks for overlapping work, so evidence may be reusable; it never means the standard discharges a legal duty.
| Framework | Reference | Note | Confidence |
|---|---|---|---|
| ISO/IEC 42001 | Annex A.7.2, A.7.3, A.7.4, A.7.6 | high | |
| NIST AI RMF | MAP 2.3; MEASURE 2.1, 2.2, 2.11 | medium | |
| MITRE ATLAS | AML.M0007 Sanitize Training Data | medium | |
| OWASP LLM Top 10 | LLM04 Data and Model Poisoning | medium |
Cite this record
AIPolicyTracker (2026). “Data governance and dataset documentation”. https://aipolicytracker.org/controls/data-governance-and-dataset-documentation (accessed 24 September 2026). Data licensed CC BY 4.0.
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.
Frequently asked questions
- Which legal duties does "Data governance and dataset documentation" satisfy?
- It is recorded as satisfying 3 and supporting 5 duties across European Union, Singapore, India, New York (United States) and United Kingdom. A mapping means the control, operated properly, does the work the duty asks for; the official text decides whether it is enough.
- What evidence shows this control is operating?
- Dataset documentation sheet, Data quality and bias check report and Dataset approval for use. Owner: Data governance lead. Frequency: once per ai system.