MIT AI Risk Repository · domain 7: AI system safety, failures, & limitations

7.4 Lack of transparency or interpretability

Challenges in understanding or explaining the decision-making processes of AI systems, which can lead to mistrust, difficulty in enforcing compliance standards or holding relevant actors accountable for harms, and the inability to identify and correct errors.

Risk entries
42
Frameworks citing it
12
Recorded incidents
5
Incidents since 2020
2
Causal entity (risk entries)
Causal entity (risk entries) 21 0 AI: 21 AI 21 Other: 12 Other 12 Human: 8 Human 8 Not coded: 1 Not coded 1
Causal entity (risk entries)
LabelValue
AI21
Other12
Human8
Not coded1
Intent (risk entries)
Intent (risk entries) 23 0 Unintentional: 23 Unintentional 23 Other: 18 Other 18 Not coded: 1 Not coded 1
Intent (risk entries)
LabelValue
Unintentional23
Other18
Not coded1
Timing (risk entries)
Timing (risk entries) 21 0 Post-deployment: 21 Post-deployment 21 Other: 14 Other 14 Pre-deployment: 6 Pre-deployment 6 Not coded: 1 Not coded 1
Timing (risk entries)
LabelValue
Post-deployment21
Other14
Pre-deployment6
Not coded1
Recorded incidents per yearIncident date; current year partial
Recorded incidents per year 1 0 2016: 1 2016 1 2017: 1 2017 1 2018: 1 2018 1 2021: 1 2021 1 2022: 1 2022 1
Recorded incidents per year
LabelValue
20161
20171
20181
20211
20221
Entries by levelRisk categories, subcategories and additional evidence coded to this subdomain
Entries by level 23 0 Risk Category: 23 Risk Category 23 Risk Sub-Category: 19 Risk Sub-Category 19
Entries by level
LabelValue
Risk Category23
Risk Sub-Category19
  • Intelligibility

    "How can we build agent’s whose decisions we can understand? Con- nects explainable decisions (Berkeley) and informed oversight (MIRI)."

    AGI Safety Literature Review (Everitt2018 ) · Human · Unintentional · Pre-deployment

  • Attributing the responsibility for AI's failures

    "This section, constituting almost 8% of the articles, addresses the implications arising from AI acting and learning without direct human supervision, encompassing two main issues: a responsibility g...

    What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024) · Other · Other · Other

  • General Evaluations (Difficulty of identification and measurement of capabilities)

    "The capabilities of general-purpose AI systems can be difficult to measure, compared to the capabilities of more limited and fixed-purpose AI systems. This is in part due to a broader distribution of...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Other

  • Model Evaluations (Interpretability/Explainability)

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Human · Other · Pre-deployment

  • Model outputs inconsistent with chain-of-thought reasoning

    "Chain-of-thought reasoning is sometimes employed to get a better understanding of the model’s output, where it encourages transparent reasoning in text form. However, in some cases, this reasoning is...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · AI · Unintentional · Post-deployment

  • Lack of understanding of in-context learning in language models

    "In-context learning allows the model to learn a new task or improve its perfor- mance by providing examples in the prompt, without changing its weights [101]. Even though this technique is highly eff...

    Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024) · Other · Other · Other

  • Lack of transparency and interpretability

    "Today's Frontier AI is difficult to interpret and lacks transparency. Contextual understanding of the training data is not explicitly embedded within these models. They can fail to capture perspectiv...

    Future Risks of Frontier AI (GOS2023) · AI · Unintentional · Pre-deployment

  • Opacity (the black box problem)

    "Opacity surrounding the technical, internal decision-making processes of generative AI models is popularly known as the “black box problem.”277 Generative AI models, most ubiquitously built on deep n...

    Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024) · Other · Unintentional · Other

  • Transparency - Explainability

    Being a multifaceted concept, the term 'transparency' is both used to refer to technical explainability as well as organizational openness. Regarding the former, papers underscore the need for mechani...

    Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024) · Not coded · Not coded · Not coded

  • Lack of transparency

    "The idea of a "black box" making decisions without any explanation, without offering insight in the process, has a couple of disadvantages: it may fail to gain the trust of its users and it may fail...

    A framework for ethical Ai at the United Nations (Hogenhout2021) · AI · Unintentional · Other

  • Non-disclosure

    "Content might not be clearly disclosed as AI generated."

    AI Risk Atlas (IBM2025) · Human · Other · Post-deployment

  • Inaccessible training data

    "Without access to the training data, the types of explanations a model can provide are limited and more likely to be incorrect."

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Untraceable attribution

    "The content of the training data used for generating the model’s output is not accessible."

    AI Risk Atlas (IBM2025) · Other · Other · Post-deployment

  • Unexplainable output

    "Explanations for model output decisions might be difficult, imprecise, or not possible to obtain."

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Unreliable source attribution

    "Source attribution is the AI system's ability to describe from what training data it generated a portion or all its output. Since current techniques are based on approximations, these attributions mi...

    AI Risk Atlas (IBM2025) · AI · Unintentional · Post-deployment

  • Lack of model transparency

    "Lack of model transparency is due to insufficient documentation of the model design, development, and evaluation process and the absence of insights into the inner workings of the model."

    AI Risk Atlas (IBM2025) · Human · Unintentional · Other

  • Transparency and explainability

    "A recurring complaint among participants was a lack of knowledge about how AI systems made judgements. They emphasized the significance of making AI systems more visible and explainable so that peopl...

    Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks (Kumar2023) · AI · Unintentional · Post-deployment

  • Trust and reliability

    "The participants of the study emphasized the importance of trustworthiness and reliability in AI systems. The authors emphasized the importance of preserving precision and objectivity in the outcomes...

    Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks (Kumar2023) · AI · Other · Post-deployment

  • Explainability & Reasoning

    The ability to explain the outputs to users and reason correctly

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Unintentional · Post-deployment

  • Lack of Interpretability

    Due to the black box nature of most machine learning models, users typically are not able to understand the reasoning behind the model decisions

    Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024) · AI · Unintentional · Post-deployment

  • Decision making transparency

    "We face significant challenges bringing transparency to artificial network decisionmaking processes. Will we have transparency in AI decision making?"

    Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016) · AI · Other · Post-deployment

  • Explainability

    "A recurrent concern about AI algorithms is the lack of explainability for the model, which means information about how the algorithm arrives at its results is deficient (Deeks, 2019). Specifically, f...

    Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023) · Other · Unintentional · Post-deployment

  • Prompt engineering

    "With the wide application of generative AI, the ability to interact with AI efficiently and effectively has become one of the most important media literacies. Hence, it is imperative for generative A...

    Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023) · Human · Other · Post-deployment

  • Value Chain and Component Integration

    "Non-transparent or untraceable integration of upstream third-party components, including data that has been improperly obtained or not processed and cleaned due to increased automation from GAI; im...

    Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024) · Human · Unintentional · Pre-deployment

  • Lack of transparency

    "In situations in which the development and use of AI are not explained to the user, or in which the decision processes do not provide the criteria or steps that constitute the decision, the use of AI...

    Social Impacts of Artificial Intelligence and Mitigation Recommendations: An Exploratory Study (Paes2023) · AI · Unintentional · Post-deployment