MIT AI Risk Repository · Risk Category · 59.15.00

Inappropriate data splitting

Description

"In data-driven AI development, the annotated data set is commonly split into training, validation, and test sets, whereby it is essential that the latter is not used for development but only for evaluation. Using the test set for training manipulates the testing strategy, which is the basis of the system’s quality assurance."

From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).

Classification

Subdomain
7.0
Causal entity
Human
Intent
Other

How other frameworks describe this risk

Other entries from Schnitzer2024