# Training-data provenance and copyright register

- **Record type**: Control
- **Kind**: Process
- **Owner**: Model development lead
- **Frequency**: At launch and on material change
- **Duties served**: 4

## What the control achieves

Records the origin, licence and rights status of every source used to train or fine-tune a model so that the organisation can honour opt-outs, answer downstream questions and publish a truthful summary of training content.

## How it is typically implemented

Model developers keep a register of training sources that captures the source, the acquisition route, the licence or legal basis relied on, whether machine-readable rights reservations were checked and honoured, and the date of collection. A written copyright and rights policy sets out how crawls respect opt-out signals, how licensed and public-domain material is distinguished and how takedown or objection requests are handled. From the register the team generates the public summary of training content and the information pack that downstream integrators need. The register is updated on every new training run and reviewed by legal before a model is released.

## Evidence it produces

- Training source register (register_entry): Source, acquisition route, licence or legal basis, opt-out check and collection date per source.
- Copyright and rights-reservation policy (policy_document)
- Public summary of training content (disclosure_notice)

## Legal duties this control serves

- Apply data governance and quality criteria to training, validation and testing data — EU AI Act, European Union (supports): https://aipolicytracker.org/obligations/eu-ai-act-data-governance
- Meet general-purpose AI model provider obligations — EU AI Act, European Union (satisfies): https://aipolicytracker.org/obligations/eu-ai-act-gpai-provider-obligations
- Collect and use personal information only with consent and for the stated purpose — Nepal Privacy Act 2075, Nepal (supports): https://aipolicytracker.org/obligations/nepal-privacy-act-consent-and-purpose
- Developers must document high-risk systems and disclose known risks — Colorado AI Act, Colorado (United States) (supports): https://aipolicytracker.org/obligations/us-colorado-developer-documentation-and-disclosure

## Standards clauses it corresponds to (clause numbers only)

- ISO/IEC 42001:2023: Annex A.7.3, A.7.5 — Data acquisition and provenance records.
- NIST AI RMF 1.0: MAP 4.1; GOVERN 6.1
- MITRE ATLAS: AML.M0025 Maintain AI Dataset Provenance

## MIT AI Risk Repository subdomains addressed

6.3, 2.1, 6.5

## Provenance

- **Record page**: https://aipolicytracker.org/controls/training-data-provenance-and-copyright-register
- **Official source**: none recorded — this record is incomplete, see https://aipolicytracker.org/gaps
- **Review status**: pending review
- **Confidence**: medium
- **Facts last confirmed**: never confirmed against the official source
- **Retrieved**: 2026-09-24
- **Licence**: https://creativecommons.org/licenses/by/4.0/

> This record is a structured summary with a link to the official text. It is not legal advice. Open the official source before relying on any date or duty. How current each record type must be is published at https://aipolicytracker.org/verification; what a record must carry at all is published at https://aipolicytracker.org/coverage.
