MIT AI Risk Repository · Risk Sub-Category · 45.01.01
Risks from models and algorithms (Risks of explainability)
Category: AI's inherent safety risks
Description
"AI algorithms, represented by deep learning, have complex internal workings. Their black-box or grey-box inference process results in unpredictable and untraceable outputs, making it challenging to quickly rectify them or trace their origins for accountability should any anomalies arise."
From AI Safety Governance Framework (TC2602024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Other
Subdomain definition: Challenges in understanding or explaining the decision-making processes of AI systems, which can lead to mistrust, difficulty in enforcing compliance standards or holding relevant actors accountable for harms, and the inability to identify and correct errors.
Real-world incidents in this subdomain
- Uber Launched Opaque Algorithm That Changes Drivers' Payments in the US
- Israeli Tax Authority Reportedly Used an Opaque Automated System to Issue a Fine, Declining to Explain or Disclose the Underlying Calculation
- Uber Allegedly Wrongfully Accused Drivers of Fraud via Automated Systems
- Houston ISD's EVAAS Teacher-Evaluation System Reportedly Put Teachers' Jobs at Risk Through Unverifiable Scores
- Dutch City Court Defended Home Value Generated by Black-Box Algorithm
How other frameworks describe this risk
- Intelligibility
- Attributing the responsibility for AI's failures
- Model Evaluations (Interpretability/Explainability)
- Model outputs inconsistent with chain-of-thought reasoning
- Lack of understanding of in-context learning in language models
- General Evaluations (Difficulty of identification and measurement of capabilities)
- Lack of transparency and interpretability
- Opacity (the black box problem)
Other entries from TC2602024
- AI's inherent safety risks
- Risks from models and algorithms (Risks of bias and discrimination)
- Risks from models and algorithms (Risks of robustness)
- Risks from models and algorithms (Risks of stealing and tampering)
- Risks from models and algorithms (Risks of unreliable output)
- Risks from models and algorithms (Risks of adversarial attack)
- Risks from data (Risks of illegal collection and use of data)
- Risks from data (Risks of improper content and poisoning in training data)
- Risks from data (Risks of unregulated training data annotation)
- Risks from data (Risks of data leakage)
- Risks from AI systems (Risks of exploitation through defects and backdoors)
- Risks from AI systems (Risks of computing infrastructure security)