MIT AI Risk Repository · Risk Category · 56.06.00
Lack of transparency and interpretability
Description
"Today's Frontier AI is difficult to interpret and lacks transparency. Contextual understanding of the training data is not explicitly embedded within these models. They can fail to capture perspectives of underrepresented groups or the limitations within which they are expected to perform without fine tuning or reinforcement learning with human feedback (RLHF)."
From Future Risks of Frontier AI (GOS2023), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Pre-deployment
Subdomain definition: Challenges in understanding or explaining the decision-making processes of AI systems, which can lead to mistrust, difficulty in enforcing compliance standards or holding relevant actors accountable for harms, and the inability to identify and correct errors.
Real-world incidents in this subdomain
- Uber Launched Opaque Algorithm That Changes Drivers' Payments in the US
- Israeli Tax Authority Reportedly Used an Opaque Automated System to Issue a Fine, Declining to Explain or Disclose the Underlying Calculation
- Uber Allegedly Wrongfully Accused Drivers of Fraud via Automated Systems
- Houston ISD's EVAAS Teacher-Evaluation System Reportedly Put Teachers' Jobs at Risk Through Unverifiable Scores
- Dutch City Court Defended Home Value Generated by Black-Box Algorithm
How other frameworks describe this risk
- Intelligibility
- Attributing the responsibility for AI's failures
- General Evaluations (Difficulty of identification and measurement of capabilities)
- Lack of understanding of in-context learning in language models
- Model outputs inconsistent with chain-of-thought reasoning
- Model Evaluations (Interpretability/Explainability)
- Opacity (the black box problem)
- Transparency - Explainability
Other entries from GOS2023
- Discrimination
- Inequality
- Environmental impacts
- Amplification of biases
- Harmful responses
- Intellectual property rights
- Providing new capabilities to a malicious actor
- Misapplication by a non-malicious actor
- Poor performance of a model used for its intended purpose, for example leading to biased decisions
- Unintended outcomes from interactions with other AI systems
- Impacts resulting from interactions with external societal, political, and economic systems
- Loss of human control and oversight, with an autonomous model then taking harmful actions