MIT AI Risk Repository · Risk Sub-Category · 45.01.09
Risks from data (Risks of unregulated training data annotation)
Category: AI's inherent safety risks
Description
"Issues with training data annotation, such as incomplete annotation guidelines, incapable annotators, and errors in annotation, can affect the accuracy, reliability, and effectiveness of models and algorithms. Moreover, they can introduce training biases, amplify discrimination, reduce generalization abilities, and result in incorrect outputs."
From AI Safety Governance Framework (TC2602024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.3 Lack of capability or robustness
- Causal entity
- Human
- Intent
- Unintentional
- Timing
- Pre-deployment
Subdomain definition: AI systems that fail to perform reliably or effectively under varying conditions, exposing them to errors and failures that can have significant consequences, especially in critical applications or areas that require moral reasoning.
Real-world incidents in this subdomain
- Purported AI Name-Reading System Reportedly Skipped and Misannounced Graduates at Arizona's Glendale Community College Commencement
- PocketOS Production Database Was Reportedly Deleted by Cursor AI Agent Running Claude Opus 4.6
- Baidu Apollo Go Robotaxis Stopped in Traffic During Reported System Failure in Wuhan, Stranding Some Passengers
- Purportedly AI-Enabled Targeting System Was Reportedly Implicated in Deadly U.S. Strike on Iranian Primary School
- Claude Code Agent Reportedly Deleted DataTalks.Club Production Infrastructure, Database, and Snapshots via Terraform
- Purportedly AI-Generated Sepsis Alert Reportedly Prompted Potentially Inappropriate IV Fluid Administration for a Dialysis Patient, Averted by Clinician Intervention
How other frameworks describe this risk
Other entries from TC2602024
- AI's inherent safety risks
- Risks from models and algorithms (Risks of explainability)
- Risks from models and algorithms (Risks of bias and discrimination)
- Risks from models and algorithms (Risks of robustness)
- Risks from models and algorithms (Risks of stealing and tampering)
- Risks from models and algorithms (Risks of unreliable output)
- Risks from models and algorithms (Risks of adversarial attack)
- Risks from data (Risks of illegal collection and use of data)
- Risks from data (Risks of improper content and poisoning in training data)
- Risks from data (Risks of data leakage)
- Risks from AI systems (Risks of exploitation through defects and backdoors)
- Risks from AI systems (Risks of computing infrastructure security)