MIT AI Risk Repository · Risk Category · 05.13.00
Transparency - Explainability
Description
Being a multifaceted concept, the term 'transparency' is both used to refer to technical explainability as well as organizational openness. Regarding the former, papers underscore the need for mechanistic interpretability and for explaining internal mechanisms in generative models. On the organizational front, transparency relates to practices such as informing users about capabilities and shortcomings of models, as well as adhering to documentation and reporting requirements for data collection processes or risk evaluations.
From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
Subdomain definition: Challenges in understanding or explaining the decision-making processes of AI systems, which can lead to mistrust, difficulty in enforcing compliance standards or holding relevant actors accountable for harms, and the inability to identify and correct errors.
Real-world incidents in this subdomain
- Uber Launched Opaque Algorithm That Changes Drivers' Payments in the US
- Israeli Tax Authority Reportedly Used an Opaque Automated System to Issue a Fine, Declining to Explain or Disclose the Underlying Calculation
- Uber Allegedly Wrongfully Accused Drivers of Fraud via Automated Systems
- Houston ISD's EVAAS Teacher-Evaluation System Reportedly Put Teachers' Jobs at Risk Through Unverifiable Scores
- Dutch City Court Defended Home Value Generated by Black-Box Algorithm
How other frameworks describe this risk
- Intelligibility
- Attributing the responsibility for AI's failures
- General Evaluations (Difficulty of identification and measurement of capabilities)
- Model outputs inconsistent with chain-of-thought reasoning
- Lack of understanding of in-context learning in language models
- Model Evaluations (Interpretability/Explainability)
- Lack of transparency and interpretability
- Opacity (the black box problem)