MIT AI Risk Repository

Browse AI risks

554 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

554 entries · page 11 of 12

  1. 30.05.01 · Risk Sub-Category

    Explainability & Reasoning

    Lack of Interpretability

    Due to the black box nature of most machine learning models, users typically are not able to understand the reasoning behind the model decisions

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  2. 33.02.03 · Risk Sub-Category

    Technology concerns

    Explainability

    "A recurrent concern about AI algorithms is the lack of explainability for the model, which means information about how the algorithm arrives at its results is deficient (Deeks, 2019). Specifically, for generative AI models, there is no transparency to the reasoning of how the model arrives at the results (Dwivedi et al., 2023). The lack of transparency raises several issues. First, it might be difficult for users to interpret and understand the output (Dwivedi et al., 2023). It would also be difficult for users to discover potential mistakes in the output (Rudin, 2019). Further, when the inte

    From Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023)

  3. "A recurring complaint among participants was a lack of knowledge about how AI systems made judgements. They emphasized the significance of making AI systems more visible and explainable so that people may have confidence in their outputs and hold them accountable for their activities. Because AI systems are typically opaque, making it difficult for users to understand the rationale behind their judgements, ethical concerns about AI, as well as issues of transparency and explainability, arise. This lack of understanding can generate suspicion and reluctance to adopt AI technology, as well as m

    From Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks (Kumar2023)

  4. 39.19.00 · Risk Category

    Accountability

    An essential feature of decision-making in humans, AI, and also HLI-based agents is accountability. Implementing this feature in machines is a difficult task because many challenges should be considered to organize an AI-based model that is accountable. It should be noted that this issue in human decision-making is not ideal, and many factors such as bias, diversity, fairness, paradox, and ambiguity may affect it. In addition, the human decision-making process is based on personal flexibility, context-sensitive paradigms, empathy, and complex moral judgments. Therefore, all of these challenges

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  5. 39.21.00 · Risk Category

    Reproducibility

    How a learning model can be reproduced when it is obtained based on various sets of data and a large space of parameters. This problem becomes more challenging in data-driven learning procedures without transparent instructions

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  6. 39.25.00 · Risk Category

    Verifiability

    In many applications of AI-based systems such as medical healthcare and military services, the lack of verification of code may not be tolerable... due to some characteristics such as the non-linear and complex structure of AI-based solutions, existing solutions have been generally considered “black boxes”, not providing any information about what exactly makes them appear in their predictions and decision-making processes.

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  7. 42.06.00 · Risk Category

    Opacity

    "Stems from the mismatch between mathematical optimization in high-dimensionality characteristic of machine learning and the demands of human-scale reasoning and styles of semantic interpretation."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  8. 45.01.01 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of explainability)

    "AI algorithms, represented by deep learning, have complex internal workings. Their black-box or grey-box inference process results in unpredictable and untraceable outputs, making it challenging to quickly rectify them or trace their origins for accountability should any anomalies arise."

    From AI Safety Governance Framework (TC2602024)

  9. 47.01.05 · Risk Sub-Category

    Technical and operational risks

    Opacity (the black box problem)

    "Opacity surrounding the technical, internal decision-making processes of generative AI models is popularly known as the “black box problem.”277 Generative AI models, most ubiquitously built on deep neural networks with hundreds of billions of internal connections,278 have become so complex that their internal decision-making processes are no longer traceable or interpretable to even the most advanced expert observers. This means that, while the inputs and outputs of a system can be observed, developers cannot explain in detail why specific inputs correspond to specific outputs."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  10. "Non-transparent or untraceable integration of upstream third-party components, including data that has been improperly obtained or not processed and cleaned due to increased automation from GAI; improper supplier vetting across the AI lifecycle; or other issues that diminish transparency or accountability for downstream users."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  11. 51.06.00 · Risk Category

    Intelligibility

    "How can we build agent’s whose decisions we can understand? Con- nects explainable decisions (Berkeley) and informed oversight (MIRI)."

    From AGI Safety Literature Review (Everitt2018 )

  12. "Today's Frontier AI is difficult to interpret and lacks transparency. Contextual understanding of the training data is not explicitly embedded within these models. They can fail to capture perspectives of underrepresented groups or the limitations within which they are expected to perform without fine tuning or reinforcement learning with human feedback (RLHF)."

    From Future Risks of Frontier AI (GOS2023)

  13. 61.02.15 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Complexity-induced knowledge gap

    "The complexity of AI models and systems makes it challenging to demonstrate harm or establish a clear causal link between AI actions and their consequences."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  14. 62.18.05 · Risk Sub-Category

    Model Evaluations (Interpretability/Explainability)

    Model outputs inconsistent with chain-of-thought reasoning

    "Chain-of-thought reasoning is sometimes employed to get a better understanding of the model’s output, where it encourages transparent reasoning in text form. However, in some cases, this reasoning is not consistent with the final answer given by the AI model, and as such does not give sufficient transparency [113]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  15. 65.17.01 · Risk Sub-Category

    Output risks (Explainability)

    Inaccessible training data

    "Without access to the training data, the types of explanations a model can provide are limited and more likely to be incorrect."

    From AI Risk Atlas (IBM2025)

  16. 65.17.03 · Risk Sub-Category

    Output risks (Explainability)

    Unexplainable output

    "Explanations for model output decisions might be difficult, imprecise, or not possible to obtain."

    From AI Risk Atlas (IBM2025)

  17. 65.17.04 · Risk Sub-Category

    Output risks (Explainability)

    Unreliable source attribution

    "Source attribution is the AI system's ability to describe from what training data it generated a portion or all its output. Since current techniques are based on approximations, these attributions might be incorrect."

    From AI Risk Atlas (IBM2025)

  18. 65.22.06 · Risk Sub-Category

    Non-technical risks (Governance)

    Lack of model transparency

    "Lack of model transparency is due to insufficient documentation of the model design, development, and evaluation process and the absence of insights into the inner workings of the model."

    From AI Risk Atlas (IBM2025)

  19. 70.04.03 · Risk Sub-Category

    Social Risks

    Lack of transparency, explainability, and trust

    "Understanding how AI reaches conclusions or why AI systems perform specific actions motivates an entire branch of interpretability research [111], but physical embodiment raises the stakes for understanding these systems. For example, transparency of planned actions and explainability of decision-making is crucial when an AV suddenly changes lanes. A lack of transparency and explainability could lead to a lack of trust, which could become a critical and socially destabilizing issue with the widespread deployment of EAI [112–114]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  20. 61.01.08 · Risk Sub-Category

    Types of systemic risks from general-purpose AI

    Harms to non-humans

    "Large-scale harms to animals and the development of AI capable of suffering."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  21. 61.02.38 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Pattern recognition capability

    "AI models and systems could exacerbate financial bubbles by reinforcing market trends."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  22. 61.02.45 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Trading capabilities

    "AI may contribute to increased market volatility by accelerating transactions and influencing financial trends in unpredictable ways."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  23. 63.01.00 · Risk Category

    Miscoordination

    "Miscoordination arises when agents, despite a mutual and clear objective, cannot align their behaviours to achieve this objective. Unlike the case of differing objectives, in common-interest settings there is a more easily well-defined notion of ‘optimal’ behaviour and we describe agents as miscoordinating to the extent that they fall short of this optimum. Note that for common-interest settings it is not sufficient for agents’ objectives to be the same in the sense of being symmetric (e.g., when two agents both want the same prize, but only one can win). Rather, agents must have identical pr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  24. 63.01.01 · Risk Sub-Category

    Miscoordination

    Incompatible strategies

    "Incompatible Strategies. Even if all agents can perform well in isolation, miscoordination can still occur due to the agents choosing incompatible strategies (Cooper et al., 1990). Competitive (i.e., two- player zero-sum) settings allow designers to produce agents that are maximally capable without taking other players into account. Crucially, this is possible because playing a strategy at equilibrium in the zero-sum setting guarantees a certain payoff, even if other players deviate from the equilibrium (Nash, 1951). On the other hand, common-interest (and mixed-motive) settings often allow a

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  25. 63.01.02 · Risk Sub-Category

    Miscoordination

    Credit Assignment

    "Credit Assignment. While agents can often learn to jointly solve tasks and thus avoid coordination failures, learning is made more challenging in the multi-agent setting due to the problem of credit assignment (Du et al., 2023; Li et al., 2025, see also Section 3.1 on information asymmetries and Section 3.4, which discusses distributional shift). That is, in the presence of other learning agents, it can be unclear which agents’ actions caused a positive or negative outcome to obtain, especially if the environment is complex. Moreover, in multi-principal settings, agents may not have been trai

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  26. 63.01.03 · Risk Sub-Category

    Miscoordination

    Limited Interactions

    "Limited Interactions. Sometimes learning from historical interactions with the relevant agents may not be possible, or may be possible using only limited interactions. In such cases, some other form of information exchange is required for agents to be able to reliably coordinate their actions, such as via communication (Crawford & Sobel, 1982; Farrell & Rabin, 1996a) or a correlation device (Aumann, 1974, 1987). While advances in language modelling mean that there are likely to be fewer settings in which the inability of advanced AI systems to communicate leads to miscoordination, situations

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  27. 63.04.02 · Risk Sub-Category

    Information Asymmetries

    Bargaining

    "Bargaining. As a classic example of these strategic considerations is that when agents attempt to come to an agreement despite diverging interests, information asymmetries can lead to bargaining inef- ficiencies (Myerson & Satterthwaite, 1983). Relevant uncertainties about other agents can include how much they value possible agreements, their outside options, or their beliefs about others. The essential reason for such inefficiencies is that, under uncertainty about their counterparties, agents must make a trade-off between the rewards of making more favourable demands and the risk of other

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  28. 63.05.01 · Risk Sub-Category

    Network Effects

    Error propagation

    "Error Propagation. One well-known issue with communication networks is that information can be corrupted as it propagates through the network.24 As AI systems become capable of generating and processing more and more kinds of information, AI agents could end up ‘polluting the epistemic commons’ (Huang & Siddarth, 2023; Kay et al., 2024) of both other agents (Ju et al., 2024) and humans (see Case Study 7 and Section 3.1) Another increasingly important framework is the use of individual AI agents as part of teams and scaffolded chains of delegation, which transmit not only information but instr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  29. 63.06.00 · Risk Category

    Selection Pressures

    "Selection pressures (Section 3.3): some aspects of training and selection by those deploying and using AI agents can lead to undesirable behaviour;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  30. 63.06.01 · Risk Sub-Category

    Selection Pressures

    Undesirable Dispositions from Competition

    "Undesirable Dispositions from Competition. It is plausible that evolution selected for certain conflict-prone dispostions in humans, such as vengefulness, aggression, risk-seeking, selfishness, dishon- esty, deception, and spitefulness towards out-groups (Grafen, 1990; Han, 2022; Konrad & Morath, 2012; McNally & Jackson, 2013; Nowak, 2006; Rusch, 2014). Such traits could also be selected for in ML systems that are trained in more competitive multi-agent settings. For example, this might happen if systems are selected based on their performance relative to other agents (and so one agent’s loss

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  31. 63.06.02 · Risk Sub-Category

    Selection Pressures

    Undesirable Dispositions from Human Data

    "Undesirable Dispositions from Human Data. It is well-understood that models trained on human data – such as being pre-trained on human-written text or fine-tuned on human feedback – can exhibit human biases. For these reasons, there has already been considerable attention to measuring biases related to protected characteristics such as sex and ethnicity (e.g., Ferrara, 2023; Liang et al., 2021; Nadeem et al., 2020; Nangia et al., 2020), which can be amplified in multi-agent settings (Acerbi & Stubbersfield, 2023, see also Case Study 7). More recently, there has been increasing attention paid

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  32. "Destabilising dynamics (Section 3.4): systems that adapt in response to one another can produce dangerous feedback loops and unpredictability;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  33. 63.07.01 · Risk Sub-Category

    Destabilising Dynamics

    Feedback Loops

    "Feedback Loops. One of the best-known historical examples to illustrate destabilising dynamics in the context of autonomous agents is the 2010 flash crash, in which algorithmic trading agents entered into an unexpected feedback loop (Commission & Commission, 2010, see also Case Study 10).37 More generally, a feedback loop occurs when the output of a system is used as part of its input, creating a cycle that can either amplify or dampen the system’s behaviour. In multi-agent settings, feedback loops often arise from the interactions between agents, as each agent’s actions affect the environmen

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  34. 63.07.02 · Risk Sub-Category

    Destabilising Dynamics

    Cyclic Behaviour

    "Cyclic Behaviour. The dynamics described above are highly non-linear (small changes to the system’s state can result in large changes to its trajectory). Similar non-linear dynamics can emerge in multi- agent learning and lead to a variety of phenomena that do not occur in single-agent learning (Barfuss et al., 2019; Barfuss & Mann, 2022; Galla & Farmer, 2013; Leonardos et al., 2020; Nagarajan et al., 2020). One of the simplest examples of this phenomenon is Q-learning (Watkins & Dayan, 1992): in the case of a single agent, convergence to an optimal policy is guaranteed under modest condition

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  35. 63.07.04 · Risk Sub-Category

    Destabilising Dynamics

    Phase Transitions

    "Phase Transitions. Finally, small external changes to the system – such as the introduction of new agents or a distributional shift – can cause phase transitions, where the system undergoes an abrupt qualitative shift in overall behaviour (Barfuss et al., 2024). Formally, this corresponds to bifurcations in the system’s parameter space, which lead to the creation or destruction of dynamical attractors, resulting in complex and unpredictable dynamics (Crawford, 1991; Zeeman, 1976). For example, Leonardos & Piliouras (2022) show that changes to the exploration hyperparameter of RL agents can le

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  36. 63.07.05 · Risk Sub-Category

    Destabilising Dynamics

    Distributional Shift

    "Distributional Shift. Individual ML systems can perform poorly in contexts different from those in which they were trained. A key source of these distributional shifts is the actions and adaptations of other agents (Narang et al., 2023; Papoudakis et al., 2019; Piliouras & Yu, 2022), which in single-agent approaches are often simply or ignored or at best modelled exogenously. Indeed, the sheer number and variance of behaviours that can be exhibited other agents means that multi-agent systems pose an especially challenging generalisation problem for individual learners (Agapiou et al., 2022; L

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  37. 63.08.01 · Risk Sub-Category

    Commitment and Trust

    Inefficient Outcomes

    "Inefficient Outcomes. Without careful planning and the appropriate safeguards, we may soon be entering a world overrun by increasingly competent and autonomous software agents, able to act with little restriction. The abilities of these agents to persuade, deceive, and obfuscate their activities, as well as the fact they can be deployed remotely and easily created or destroyed by their deployer, means that by default they may garner little trust (from humans or from other agents). Such a world may end up being rife with economic inefficiencies (Krier, 2023; Schmitz, 2001), political problems

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  38. 63.08.03 · Risk Sub-Category

    Commitment and Trust

    Rigidity and Mistaken Commitments

    "Rigidity and Mistaken Commitments. Even when it is desirable to be able to make threats in order to deter socially harmful behaviour, doing so using AI agents effectively removes the human from the loop, which could prove disastrous in high-stakes contexts (e.g., a false positive in a nuclear sub- marine’s warning system; see also Case Study 11), or when irresponsible actors are enabled in making disproportionate or mistaken commitments."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  39. 63.09.00 · Risk Category

    Emergent Agency

    "Emergent agency (Section 3.6): qualitatively different goals or capabilities can emerge from the composition of innocuous independent systems or behaviours;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  40. 63.09.01 · Risk Sub-Category

    Emergent Agency

    Emergent Capabilities

    "Emergent Capabilities. Dangerous emergent capabilities could arise when a multi-agent system over- comes the safety-enhancing limitations of the individual systems, such as individual models’ narrow domains of application or myopia caused by a lack of long-term planning and long-term memory. For example, narrow systems for research planning, predicting the properties of molecules, and synthesising new chemicals could, when combined, lead to a complex ‘test and iterate’ automated workflow capable of designing dangerous new chemical compounds far beyond the scope of the initial systems’ capabil

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  41. 63.09.02 · Risk Sub-Category

    Emergent Agency

    Emergent Goals

    "Emergent Goals. Ascribing goals to a system is not always straightforward. For our present purposes, it will suffice to adopt a Dennetian perspective (Dennett, 1971), ascribing goals and intentions only when it is useful (i.e., predictive) to do so.51 While it might not be helpful to describe individual narrow AI tools as having goals, their combination may act as a (seemingly) goal-directed collective. For example, a group of moderation bots on a major social networking site could subtly but systematically manipulate the overall political perspectives of the user population, even though, ind

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  42. 72.03.02 · Risk Sub-Category

    Accident Risks

    Impact on Financial Stability

    "The integration of general-purpose AI into high-frequency trading, market-making, or systemic risk management could exacerbate systemic risk by exhibiting unexpected behavioral patterns during market stress. Moreover, the concentration of a few homogeneous foundation models across financial institutions may foster correlated decision-making and herd-following behaviors. The widespread adoption of AI agents could also amplify volatility through emergent phenomena from multi-agent interactions.23 All of these could precipitate a cascading global-scale financial system instability, with potentia

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  43. 72.06.05 · Risk Sub-Category

    Model Propensities

    Multi-agent collusion propensity:

    "Multiple agents tend to coordinate actions through covert means to maximize common interests (possibly harming third-party interests or evading regulation), even if individual agents are designed with safety constraints, their collusive behavior may still trigger systemic risks such as market manipulation or cascading failures that are difficult to detect and mitigate, and may develop specialized communication protocols to avoid monitoring."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  44. 74.01.00 · Risk Category

    Inherent Risk

    "In terms of inherent risk, LLMs could potentially reveal sensitive information from their utilized corpora for pre-training or fine-tuning, thereby raising issues of privacy leakage [37, 145, 226]. Meanwhile, it is well-known that LLMs may experi- ence hallucinations, resulting in the production of texts that are inaccurate and misleading [194]. Finally, since the values embedded in LLM-generated texts usually directly reflect the distribution of their training data, often sourced from the Internet, there exists a substantial risk that LLMs will overfit to a narrow set of human values or even

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  45. "Risks from Unreliability stem from general purpose AI models that lack reliability, robustness, transparency, corrigibility, and interpretability, making it challenging to predict and control their behaviour fully. This includes Discrimination and Stereotype Reproduction, Misinformation and Privacy Violations, and Accidents."

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  46. "Work focused at understanding indirect ways in which AI could contribute to existential threats, such as by shaping societal “turbulence”193 and other existential risk factors.194 This covers various long-term impacts on societal parameters such as science, cooperation, power, epistemics, and values:"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  47. From Future Risks of Frontier AI (GOS2023)

  48. From Future Risks of Frontier AI (GOS2023)

  49. 61.02.16 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Conflicting objectives in design

    "Designers and operators of AI may face conflicting objectives that compromise safety."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  50. 62.01.02 · Risk Sub-Category

    Dimension - Intent

    Unintentional

    "Risks can be realized by intentional or unintentional actions, and in some cases the intent is difficult to establish. To manage these risks, rigorous evaluations and red teaming can be performed, guardrails can be put in place, and model release can be gradual, such that AI model malfunctions have either low likeli- hood or low probability of occurrence. To prevent intentional misuse, acceptable use policies can be in place, and for riskier models Know Your Customer (KYC) measures can also be implemented by model providers."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.