MIT AI Risk Repository

Browse AI risks

422 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Gipiškis2024 ×

422 entries · page 8 of 9

  1. 59.18.00 · Risk Category

    Lack of explainability

    "The explainability of AI systems based on so-called black-box models is often limited. This opaqueness of AI systems can prevent developers from detecting shortcomings in the data or the model itself and decrease the performance and safety levels of the AI system."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  2. 59.24.00 · Risk Category

    Concept drift

    "Concept drift refers to a change in the rela- tionship between input variables and model output. If not treated appropriately, concept drift can reduce the reliability of AI systems."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  3. 61.02.15 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Complexity-induced knowledge gap

    "The complexity of AI models and systems makes it challenging to demonstrate harm or establish a clear causal link between AI actions and their consequences."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  4. 61.02.37 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Opaque AI networks

    "The complexity and opacity of AI models and systems make it difficult to predict and manage their behavior."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  5. 62.16.03 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Difficulty of identification and measurement of capabilities)

    "The capabilities of general-purpose AI systems can be difficult to measure, compared to the capabilities of more limited and fixed-purpose AI systems. This is in part due to a broader distribution of potential risks, a lack of well-defined metrics to evaluate these risks, and risks from unpredictable (or emergent) AI model properties."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  6. 62.18.05 · Risk Sub-Category

    Model Evaluations (Interpretability/Explainability)

    Model outputs inconsistent with chain-of-thought reasoning

    "Chain-of-thought reasoning is sometimes employed to get a better understanding of the model’s output, where it encourages transparent reasoning in text form. However, in some cases, this reasoning is not consistent with the final answer given by the AI model, and as such does not give sufficient transparency [113]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  7. 62.19.10 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Lack of understanding of in-context learning in language models

    "In-context learning allows the model to learn a new task or improve its perfor- mance by providing examples in the prompt, without changing its weights [101]. Even though this technique is highly effective, its working mechanism is not well understood. Since many potential misuses are directly related to prompting, it becomes difficult to guarantee safety when the exact mechanism of in-context learning is not fully investigated [13]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  8. 65.14.01 · Risk Sub-Category

    Output risks (misuse)

    Non-disclosure

    "Content might not be clearly disclosed as AI generated."

    From AI Risk Atlas (IBM2025)

  9. 65.17.01 · Risk Sub-Category

    Output risks (Explainability)

    Inaccessible training data

    "Without access to the training data, the types of explanations a model can provide are limited and more likely to be incorrect."

    From AI Risk Atlas (IBM2025)

  10. 65.17.02 · Risk Sub-Category

    Output risks (Explainability)

    Untraceable attribution

    "The content of the training data used for generating the model’s output is not accessible."

    From AI Risk Atlas (IBM2025)

  11. 65.17.03 · Risk Sub-Category

    Output risks (Explainability)

    Unexplainable output

    "Explanations for model output decisions might be difficult, imprecise, or not possible to obtain."

    From AI Risk Atlas (IBM2025)

  12. 65.17.04 · Risk Sub-Category

    Output risks (Explainability)

    Unreliable source attribution

    "Source attribution is the AI system's ability to describe from what training data it generated a portion or all its output. Since current techniques are based on approximations, these attributions might be incorrect."

    From AI Risk Atlas (IBM2025)

  13. 65.22.06 · Risk Sub-Category

    Non-technical risks (Governance)

    Lack of model transparency

    "Lack of model transparency is due to insufficient documentation of the model design, development, and evaluation process and the absence of insights into the inner workings of the model."

    From AI Risk Atlas (IBM2025)

  14. 70.04.03 · Risk Sub-Category

    Social Risks

    Lack of transparency, explainability, and trust

    "Understanding how AI reaches conclusions or why AI systems perform specific actions motivates an entire branch of interpretability research [111], but physical embodiment raises the stakes for understanding these systems. For example, transparency of planned actions and explainability of decision-making is crucial when an AV suddenly changes lanes. A lack of transparency and explainability could lead to a lack of trust, which could become a critical and socially destabilizing issue with the widespread deployment of EAI [112–114]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  15. 09.06.01 · Risk Sub-Category

    AI rights and responsibilities

    AI rights and responsibilities

    "We note literature—which gives us the domain termed Robot Rights—addressing the rights of the AI itself as we develop and implement it. We find arguments against [38] the affordance of rights for artificial agents: that they should be equals in ability but not in rights, that they should be inferior by design and expendable when needed, and that since they can be designed not to feel pain (or anything) they do not have the same rights as humans. On a more theoretical level, we find literature asking more fundamental questions, such as: at what point is a simulation of life (e.g. artificial in

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  16. 09.06.03 · Risk Sub-Category

    AI death

    AI death

    "The literature suggests that throughout the development of an AI we may go through several generations of agents which do not perform as expected [37] [43]. In this case, such agents may be placed into a suspended state, terminated, or deleted. Further, we could propose scenarios where research funding for a facility running such agents is exhausted, resulting in the inadvertent termination of a project. In these cases, is deletion or termination of AI programs (the moral patient) by a moral agent an act of murder? This, an example of Robot Ethics, raises issues of personhood which parallel r

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  17. 61.01.08 · Risk Sub-Category

    Types of systemic risks from general-purpose AI

    Harms to non-humans

    "Large-scale harms to animals and the development of AI capable of suffering."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  18. 61.02.38 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Pattern recognition capability

    "AI models and systems could exacerbate financial bubbles by reinforcing market trends."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  19. 61.02.45 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Trading capabilities

    "AI may contribute to increased market volatility by accelerating transactions and influencing financial trends in unpredictable ways."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  20. 62.31.02#2 · Risk Sub-Category

    Impacts of AI (Financial Impacts)

    Financial instability due to model homogeneity

    "The widespread use of similar models or algorithms across the financial sec- tor can lead to synchronized reactions to market signals, increasing volatility, triggering flash crashes, or market illiquidity [4]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  21. 63.01.00 · Risk Category

    Miscoordination

    "Miscoordination arises when agents, despite a mutual and clear objective, cannot align their behaviours to achieve this objective. Unlike the case of differing objectives, in common-interest settings there is a more easily well-defined notion of ‘optimal’ behaviour and we describe agents as miscoordinating to the extent that they fall short of this optimum. Note that for common-interest settings it is not sufficient for agents’ objectives to be the same in the sense of being symmetric (e.g., when two agents both want the same prize, but only one can win). Rather, agents must have identical pr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  22. 63.01.01 · Risk Sub-Category

    Miscoordination

    Incompatible strategies

    "Incompatible Strategies. Even if all agents can perform well in isolation, miscoordination can still occur due to the agents choosing incompatible strategies (Cooper et al., 1990). Competitive (i.e., two- player zero-sum) settings allow designers to produce agents that are maximally capable without taking other players into account. Crucially, this is possible because playing a strategy at equilibrium in the zero-sum setting guarantees a certain payoff, even if other players deviate from the equilibrium (Nash, 1951). On the other hand, common-interest (and mixed-motive) settings often allow a

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  23. 63.01.02 · Risk Sub-Category

    Miscoordination

    Credit Assignment

    "Credit Assignment. While agents can often learn to jointly solve tasks and thus avoid coordination failures, learning is made more challenging in the multi-agent setting due to the problem of credit assignment (Du et al., 2023; Li et al., 2025, see also Section 3.1 on information asymmetries and Section 3.4, which discusses distributional shift). That is, in the presence of other learning agents, it can be unclear which agents’ actions caused a positive or negative outcome to obtain, especially if the environment is complex. Moreover, in multi-principal settings, agents may not have been trai

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  24. 63.01.03 · Risk Sub-Category

    Miscoordination

    Limited Interactions

    "Limited Interactions. Sometimes learning from historical interactions with the relevant agents may not be possible, or may be possible using only limited interactions. In such cases, some other form of information exchange is required for agents to be able to reliably coordinate their actions, such as via communication (Crawford & Sobel, 1982; Farrell & Rabin, 1996a) or a correlation device (Aumann, 1974, 1987). While advances in language modelling mean that there are likely to be fewer settings in which the inability of advanced AI systems to communicate leads to miscoordination, situations

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  25. 63.02.00 · Risk Category

    Conflict

    "In the vast majority of real-world strategic interactions, agents’ objectives are neither identical nor completely opposed. Indeed, if AI agents are sufficiently aligned to their users or deployers, we should expect some degree of both cooperation and competition, mirroring human society. These mixed-motive settings include the possibility of mutual gains, but also the risk of conflict due to selfish incentives. In what follows, we examine the extent to which advanced AI might precipitate or exacerbate such risks."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  26. 63.02.01 · Risk Sub-Category

    Conflict

    Social Dilemmas

    "Social Dilemmas. As noted in our definition, conflict can arise in any situation in which selfish incentives diverge from the collective good, known as a social dilemma (Dawes & Messick, 2000; Hardin, 1968; Kollock, 1998; Ostrom, 1990). While this is by no means a modern problem, advances in AI could further enable actors to pursue their selfish incentives by overcoming the technical, legal, or social barriers that standardly help to prevent this. To take a plausible, near-term (if very low-stakes) example, an automated AI assistant could easily reserve a table at every restaurant in town in

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  27. 63.02.02 · Risk Sub-Category

    Conflict

    Military Domains

    "Perhaps the most obvious and worrying instances of AI conflict are those in which human conflict is already a major concern, such as military domains (although other, less salient forms of conflict such as international trade wars are also cause for concern). For example, beyond applications of more narrow AI tools in lethal autonomous weapons systems (Horowitz, 2021), future AI systems might serve as advisors or negotiators in high-stakes military decisions (Black et al., 2024; Manson, 2024). Indeed, companies such as Palantir have already developed LLM-powered tools for military planning (P

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  28. 63.02.03 · Risk Sub-Category

    Conflict

    Coercion and Extortion

    "Advanced AI systems might also lead to various forms of coercion and extortion in less extreme settings (Ellsberg, 1968; Harrenstein et al., 2007). These threats might target humans directly (such as the revelation of private information extracted by advanced AI surveillance tools), or other AI systems that are deployed on behalf of humans (such as by hacking a system to limit its resources or operational capacity; see also Section 3.7). Increasing AI cyber-offensive capabilities – including those that target other AI systems via adversarial attacks and jailbreaking (Gleave et al., 2020; Yami

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  29. 63.03.00 · Risk Category

    Collusion

    "Collusion has long been a topic of intense study in economics, law, and politics, among other disciplines. While there is no universal definition of collusion, it generally refers to secretive cooperation between two or more parties at the expense of one or more other parties. Most classic examples of collusion – such as firms working together to set supra-competitive prices at the expense of consumers – also tend to be not only secretive but in violation of some law, rule, or ethical standard. Distinctions are also commonly made between explicit and tacit collusion (Rees, 1993), depending on

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  30. 63.03.01 · Risk Sub-Category

    Collusion

    Markets

    "Markets. The quintessential case of collusion in mixed-motive settings is markets, in which efficiency results from competition, not cooperation. While this is not a new problem, collusion between AI systems is especially concerning since they may operate inscrutably due to the speed, scale, complexity, or subtlety of their actions.17 Warnings of this possibility have come from technologists, economists, and legal scholars (Beneke & Mackenrodt, 2019; Brown & MacKay, 2023; Ezrachi & Stucke, 2017; Harrington, 2019; Mehra, 2016). Importantly, AI systems can collude even when collusion is not int

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  31. 63.03.02 · Risk Sub-Category

    Collusion

    Steganography

    "Steganography. In the near future we will likely see LLMs communicating with each other to jointly accomplish tasks. To try to prevent collusion, we could monitor and constrain their communication (e.g., to be in natural language). However, models might secretly learn to communicate by concealing messages within other, non-secret text. Recent work on steganography using ML has demonstrated that this concern is well-founded (Hu et al., 2018; Mathew et al., 2024; Roger & Greenblatt, 2023; Schroeder de Witt et al., 2023b; Yang et al., 2019, see also Case Study 5). Secret communication could also

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  32. "Information asymmetries (Section 3.1): private information can lead to miscoordination, deception, and conflict;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  33. 63.04.01 · Risk Sub-Category

    Information Asymmetries

    Communication constraints

    "Communication Constraints. A fundamental source of information asymmetries is that constraints on information exchange can exist, even when agents share a common goal (see Section 2.1). These might be constraints on space (i.e., the amount of information that can be communicated) if the information that needs to be communicated is especially complex, time if a snap decision is required before all information can be communicated, or both."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  34. 63.04.02 · Risk Sub-Category

    Information Asymmetries

    Bargaining

    "Bargaining. As a classic example of these strategic considerations is that when agents attempt to come to an agreement despite diverging interests, information asymmetries can lead to bargaining inef- ficiencies (Myerson & Satterthwaite, 1983). Relevant uncertainties about other agents can include how much they value possible agreements, their outside options, or their beliefs about others. The essential reason for such inefficiencies is that, under uncertainty about their counterparties, agents must make a trade-off between the rewards of making more favourable demands and the risk of other

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  35. 63.04.03 · Risk Sub-Category

    Information Asymmetries

    Deception

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  36. 63.05.00 · Risk Category

    Network Effects

    "Network effects (Section 3.2): minor changes in properties or connection patterns of agents in a network can lead to dramatic changes in the behaviour of the whole group;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  37. 63.05.01 · Risk Sub-Category

    Network Effects

    Error propagation

    "Error Propagation. One well-known issue with communication networks is that information can be corrupted as it propagates through the network.24 As AI systems become capable of generating and processing more and more kinds of information, AI agents could end up ‘polluting the epistemic commons’ (Huang & Siddarth, 2023; Kay et al., 2024) of both other agents (Ju et al., 2024) and humans (see Case Study 7 and Section 3.1) Another increasingly important framework is the use of individual AI agents as part of teams and scaffolded chains of delegation, which transmit not only information but instr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  38. 63.05.02 · Risk Sub-Category

    Network Effects

    Network rewiring

    "Network Rewiring. A different class of problems concerns not changes in the content transmitted through the network but changes in the network structure itself (Albert et al., 2000)."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  39. 63.05.03 · Risk Sub-Category

    Network Effects

    Homogeneity and correlated failures

    "Homogeneity and Correlated Failures. The current paradigm driving the state of the art in AI is the ‘foundation model’ (Bommasani et al., 2021): large-scale ML models pre-trained on broad data, which can be repurposed for a wide range of downstream applications. The costs required to create such models (and continuing returns to scale) means that only well-resourced actors can create cutting- edge models (Epoch, 2023; Hoffmann et al., 2022; Kaplan et al., 2020), making them relatively few in number. If current trends continue, it is likely that many AI agents will be powered by a small number

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  40. 63.06.00 · Risk Category

    Selection Pressures

    "Selection pressures (Section 3.3): some aspects of training and selection by those deploying and using AI agents can lead to undesirable behaviour;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  41. 63.06.01 · Risk Sub-Category

    Selection Pressures

    Undesirable Dispositions from Competition

    "Undesirable Dispositions from Competition. It is plausible that evolution selected for certain conflict-prone dispostions in humans, such as vengefulness, aggression, risk-seeking, selfishness, dishon- esty, deception, and spitefulness towards out-groups (Grafen, 1990; Han, 2022; Konrad & Morath, 2012; McNally & Jackson, 2013; Nowak, 2006; Rusch, 2014). Such traits could also be selected for in ML systems that are trained in more competitive multi-agent settings. For example, this might happen if systems are selected based on their performance relative to other agents (and so one agent’s loss

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  42. 63.06.02 · Risk Sub-Category

    Selection Pressures

    Undesirable Dispositions from Human Data

    "Undesirable Dispositions from Human Data. It is well-understood that models trained on human data – such as being pre-trained on human-written text or fine-tuned on human feedback – can exhibit human biases. For these reasons, there has already been considerable attention to measuring biases related to protected characteristics such as sex and ethnicity (e.g., Ferrara, 2023; Liang et al., 2021; Nadeem et al., 2020; Nangia et al., 2020), which can be amplified in multi-agent settings (Acerbi & Stubbersfield, 2023, see also Case Study 7). More recently, there has been increasing attention paid

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  43. 63.06.03 · Risk Sub-Category

    Selection Pressures

    Undesirable Capabilities

    "Undesirable Capabilities. As agents interact, they iteratively exploit each other’s weaknesses, forc- ing them to address these weaknesses and gain new capabilities. This co-adaptation between agents can quickly lead to emergent self-supervised autocurricula (where agents create their own challenges, driving open-ended skill acquisition through interaction), generating agents with ever-more sophisticated strate- gies in order to out-compete each other (Leibo et al., 2019). This effect is so powerful that harnessing it has been critical to the success of superhuman systems, such as the use of

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  44. "Destabilising dynamics (Section 3.4): systems that adapt in response to one another can produce dangerous feedback loops and unpredictability;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  45. 63.07.01 · Risk Sub-Category

    Destabilising Dynamics

    Feedback Loops

    "Feedback Loops. One of the best-known historical examples to illustrate destabilising dynamics in the context of autonomous agents is the 2010 flash crash, in which algorithmic trading agents entered into an unexpected feedback loop (Commission & Commission, 2010, see also Case Study 10).37 More generally, a feedback loop occurs when the output of a system is used as part of its input, creating a cycle that can either amplify or dampen the system’s behaviour. In multi-agent settings, feedback loops often arise from the interactions between agents, as each agent’s actions affect the environmen

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  46. 63.07.02 · Risk Sub-Category

    Destabilising Dynamics

    Cyclic Behaviour

    "Cyclic Behaviour. The dynamics described above are highly non-linear (small changes to the system’s state can result in large changes to its trajectory). Similar non-linear dynamics can emerge in multi- agent learning and lead to a variety of phenomena that do not occur in single-agent learning (Barfuss et al., 2019; Barfuss & Mann, 2022; Galla & Farmer, 2013; Leonardos et al., 2020; Nagarajan et al., 2020). One of the simplest examples of this phenomenon is Q-learning (Watkins & Dayan, 1992): in the case of a single agent, convergence to an optimal policy is guaranteed under modest condition

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  47. 63.07.03 · Risk Sub-Category

    Destabilising Dynamics

    Chaos

    "Chaos. Unlike the systems that tend towards fixed points or cycles described above, chaotic systems are inherently unpredictable and highly sensitive to initial conditions. While it might seem easy to dismiss such notions as mathematical exoticisms, recent work has shown that, in fact, chaotic dynamics are not only possible in a wide range of multi-agent learning setups (Andrade et al., 2021; Galla & Farmer, 2013; Palaiopanos et al., 2017; Sato et al., 2002; Vlatakis-Gkaragkounis et al., 2023), but can become the norm as the number of agents increases (Bielawski et al., 2021; Cheung & Piliour

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  48. 63.07.04 · Risk Sub-Category

    Destabilising Dynamics

    Phase Transitions

    "Phase Transitions. Finally, small external changes to the system – such as the introduction of new agents or a distributional shift – can cause phase transitions, where the system undergoes an abrupt qualitative shift in overall behaviour (Barfuss et al., 2024). Formally, this corresponds to bifurcations in the system’s parameter space, which lead to the creation or destruction of dynamical attractors, resulting in complex and unpredictable dynamics (Crawford, 1991; Zeeman, 1976). For example, Leonardos & Piliouras (2022) show that changes to the exploration hyperparameter of RL agents can le

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  49. 63.07.05 · Risk Sub-Category

    Destabilising Dynamics

    Distributional Shift

    "Distributional Shift. Individual ML systems can perform poorly in contexts different from those in which they were trained. A key source of these distributional shifts is the actions and adaptations of other agents (Narang et al., 2023; Papoudakis et al., 2019; Piliouras & Yu, 2022), which in single-agent approaches are often simply or ignored or at best modelled exogenously. Indeed, the sheer number and variance of behaviours that can be exhibited other agents means that multi-agent systems pose an especially challenging generalisation problem for individual learners (Agapiou et al., 2022; L

    From Multi-Agent Risks from Advanced AI (Hammond2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.