MIT AI Risk Repository · Risk Category · 59.15.00
Inappropriate data splitting
Description
"In data-driven AI development, the annotated data set is commonly split into training, validation, and test sets, whereby it is essential that the latter is not used for development but only for evaluation. Using the test set for training manipulates the testing strategy, which is the basis of the system’s quality assurance."
From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.0
- Causal entity
- Human
- Intent
- Other
- Timing
- Pre-deployment
How other frameworks describe this risk
Other entries from Schnitzer2024
- Inadequate specification of ODD
- Inappropriate degree of automation
- Inadequate planning of performance requirements
- Insufficient AI development documentation
- Inappropriate degree of transparency to end users
- Missing requirements for the implemented hardware
- Choice of untrustworthy data source
- Lack of data understanding
- Discriminative data bias
- Harming users’ data privacy
- Incorrect data labels
- Data poisoning