MIT AI Risk Repository

Browse AI risks

662 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

662 entries · page 13 of 14

  1. The ability to explain the outputs to users and reason correctly

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  2. 30.05.01 · Risk Sub-Category

    Explainability & Reasoning

    Lack of Interpretability

    Due to the black box nature of most machine learning models, users typically are not able to understand the reasoning behind the model decisions

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  3. "A recurring complaint among participants was a lack of knowledge about how AI systems made judgements. They emphasized the significance of making AI systems more visible and explainable so that people may have confidence in their outputs and hold them accountable for their activities. Because AI systems are typically opaque, making it difficult for users to understand the rationale behind their judgements, ethical concerns about AI, as well as issues of transparency and explainability, arise. This lack of understanding can generate suspicion and reluctance to adopt AI technology, as well as m

    From Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks (Kumar2023)

  4. 38.05.00 · Risk Category

    Trust and reliability

    "The participants of the study emphasized the importance of trustworthiness and reliability in AI systems. The authors emphasized the importance of preserving precision and objectivity in the outcomes produced by AI systems, while also ensuring transparency in their decision-making procedures. The significance of reliability and credibility in AI systems is escalating in tandem with the proliferation of these technologies across diverse domains of society. This underscores the importance of ensuring user confidence. The concern regarding the dependability of AI systems and their inherent biase

    From Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks (Kumar2023)

  5. 39.19.00 · Risk Category

    Accountability

    An essential feature of decision-making in humans, AI, and also HLI-based agents is accountability. Implementing this feature in machines is a difficult task because many challenges should be considered to organize an AI-based model that is accountable. It should be noted that this issue in human decision-making is not ideal, and many factors such as bias, diversity, fairness, paradox, and ambiguity may affect it. In addition, the human decision-making process is based on personal flexibility, context-sensitive paradigms, empathy, and complex moral judgments. Therefore, all of these challenges

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  6. 39.20.00 · Risk Category

    Transparency

    an external entity of an AI-based ecosystem may want to know which parts of data affect the final decision in a learning model

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  7. 39.21.00 · Risk Category

    Reproducibility

    How a learning model can be reproduced when it is obtained based on various sets of data and a large space of parameters. This problem becomes more challenging in data-driven learning procedures without transparent instructions

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  8. 39.25.00 · Risk Category

    Verifiability

    In many applications of AI-based systems such as medical healthcare and military services, the lack of verification of code may not be tolerable... due to some characteristics such as the non-linear and complex structure of AI-based solutions, existing solutions have been generally considered “black boxes”, not providing any information about what exactly makes them appear in their predictions and decision-making processes.

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  9. 42.06.00 · Risk Category

    Opacity

    "Stems from the mismatch between mathematical optimization in high-dimensionality characteristic of machine learning and the demands of human-scale reasoning and styles of semantic interpretation."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  10. 42.21.00 · Risk Category

    Explainability

    "Any action or procedure performed by a model with the intention of clarifying or detailing its internal functions."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  11. 45.01.01 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of explainability)

    "AI algorithms, represented by deep learning, have complex internal workings. Their black-box or grey-box inference process results in unpredictable and untraceable outputs, making it challenging to quickly rectify them or trace their origins for accountability should any anomalies arise."

    From AI Safety Governance Framework (TC2602024)

  12. "Today's Frontier AI is difficult to interpret and lacks transparency. Contextual understanding of the training data is not explicitly embedded within these models. They can fail to capture perspectives of underrepresented groups or the limitations within which they are expected to perform without fine tuning or reinforcement learning with human feedback (RLHF)."

    From Future Risks of Frontier AI (GOS2023)

  13. 62.18.05 · Risk Sub-Category

    Model Evaluations (Interpretability/Explainability)

    Model outputs inconsistent with chain-of-thought reasoning

    "Chain-of-thought reasoning is sometimes employed to get a better understanding of the model’s output, where it encourages transparent reasoning in text form. However, in some cases, this reasoning is not consistent with the final answer given by the AI model, and as such does not give sufficient transparency [113]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  14. 65.17.01 · Risk Sub-Category

    Output risks (Explainability)

    Inaccessible training data

    "Without access to the training data, the types of explanations a model can provide are limited and more likely to be incorrect."

    From AI Risk Atlas (IBM2025)

  15. 65.17.03 · Risk Sub-Category

    Output risks (Explainability)

    Unexplainable output

    "Explanations for model output decisions might be difficult, imprecise, or not possible to obtain."

    From AI Risk Atlas (IBM2025)

  16. 65.17.04 · Risk Sub-Category

    Output risks (Explainability)

    Unreliable source attribution

    "Source attribution is the AI system's ability to describe from what training data it generated a portion or all its output. Since current techniques are based on approximations, these attributions might be incorrect."

    From AI Risk Atlas (IBM2025)

  17. 09.06.01 · Risk Sub-Category

    AI rights and responsibilities

    AI rights and responsibilities

    "We note literature—which gives us the domain termed Robot Rights—addressing the rights of the AI itself as we develop and implement it. We find arguments against [38] the affordance of rights for artificial agents: that they should be equals in ability but not in rights, that they should be inferior by design and expendable when needed, and that since they can be designed not to feel pain (or anything) they do not have the same rights as humans. On a more theoretical level, we find literature asking more fundamental questions, such as: at what point is a simulation of life (e.g. artificial in

    From Managing the ethical and risk implications of rapid advances in artificial intelligence: A literature review (Meek2016)

  18. 61.02.38 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Pattern recognition capability

    "AI models and systems could exacerbate financial bubbles by reinforcing market trends."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  19. 61.02.45 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Trading capabilities

    "AI may contribute to increased market volatility by accelerating transactions and influencing financial trends in unpredictable ways."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  20. 63.01.00 · Risk Category

    Miscoordination

    "Miscoordination arises when agents, despite a mutual and clear objective, cannot align their behaviours to achieve this objective. Unlike the case of differing objectives, in common-interest settings there is a more easily well-defined notion of ‘optimal’ behaviour and we describe agents as miscoordinating to the extent that they fall short of this optimum. Note that for common-interest settings it is not sufficient for agents’ objectives to be the same in the sense of being symmetric (e.g., when two agents both want the same prize, but only one can win). Rather, agents must have identical pr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  21. 63.01.01 · Risk Sub-Category

    Miscoordination

    Incompatible strategies

    "Incompatible Strategies. Even if all agents can perform well in isolation, miscoordination can still occur due to the agents choosing incompatible strategies (Cooper et al., 1990). Competitive (i.e., two- player zero-sum) settings allow designers to produce agents that are maximally capable without taking other players into account. Crucially, this is possible because playing a strategy at equilibrium in the zero-sum setting guarantees a certain payoff, even if other players deviate from the equilibrium (Nash, 1951). On the other hand, common-interest (and mixed-motive) settings often allow a

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  22. 63.01.02 · Risk Sub-Category

    Miscoordination

    Credit Assignment

    "Credit Assignment. While agents can often learn to jointly solve tasks and thus avoid coordination failures, learning is made more challenging in the multi-agent setting due to the problem of credit assignment (Du et al., 2023; Li et al., 2025, see also Section 3.1 on information asymmetries and Section 3.4, which discusses distributional shift). That is, in the presence of other learning agents, it can be unclear which agents’ actions caused a positive or negative outcome to obtain, especially if the environment is complex. Moreover, in multi-principal settings, agents may not have been trai

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  23. 63.01.03 · Risk Sub-Category

    Miscoordination

    Limited Interactions

    "Limited Interactions. Sometimes learning from historical interactions with the relevant agents may not be possible, or may be possible using only limited interactions. In such cases, some other form of information exchange is required for agents to be able to reliably coordinate their actions, such as via communication (Crawford & Sobel, 1982; Farrell & Rabin, 1996a) or a correlation device (Aumann, 1974, 1987). While advances in language modelling mean that there are likely to be fewer settings in which the inability of advanced AI systems to communicate leads to miscoordination, situations

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  24. 63.02.00 · Risk Category

    Conflict

    "In the vast majority of real-world strategic interactions, agents’ objectives are neither identical nor completely opposed. Indeed, if AI agents are sufficiently aligned to their users or deployers, we should expect some degree of both cooperation and competition, mirroring human society. These mixed-motive settings include the possibility of mutual gains, but also the risk of conflict due to selfish incentives. In what follows, we examine the extent to which advanced AI might precipitate or exacerbate such risks."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  25. 63.02.01 · Risk Sub-Category

    Conflict

    Social Dilemmas

    "Social Dilemmas. As noted in our definition, conflict can arise in any situation in which selfish incentives diverge from the collective good, known as a social dilemma (Dawes & Messick, 2000; Hardin, 1968; Kollock, 1998; Ostrom, 1990). While this is by no means a modern problem, advances in AI could further enable actors to pursue their selfish incentives by overcoming the technical, legal, or social barriers that standardly help to prevent this. To take a plausible, near-term (if very low-stakes) example, an automated AI assistant could easily reserve a table at every restaurant in town in

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  26. 63.02.02 · Risk Sub-Category

    Conflict

    Military Domains

    "Perhaps the most obvious and worrying instances of AI conflict are those in which human conflict is already a major concern, such as military domains (although other, less salient forms of conflict such as international trade wars are also cause for concern). For example, beyond applications of more narrow AI tools in lethal autonomous weapons systems (Horowitz, 2021), future AI systems might serve as advisors or negotiators in high-stakes military decisions (Black et al., 2024; Manson, 2024). Indeed, companies such as Palantir have already developed LLM-powered tools for military planning (P

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  27. 63.02.03 · Risk Sub-Category

    Conflict

    Coercion and Extortion

    "Advanced AI systems might also lead to various forms of coercion and extortion in less extreme settings (Ellsberg, 1968; Harrenstein et al., 2007). These threats might target humans directly (such as the revelation of private information extracted by advanced AI surveillance tools), or other AI systems that are deployed on behalf of humans (such as by hacking a system to limit its resources or operational capacity; see also Section 3.7). Increasing AI cyber-offensive capabilities – including those that target other AI systems via adversarial attacks and jailbreaking (Gleave et al., 2020; Yami

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  28. 63.03.00 · Risk Category

    Collusion

    "Collusion has long been a topic of intense study in economics, law, and politics, among other disciplines. While there is no universal definition of collusion, it generally refers to secretive cooperation between two or more parties at the expense of one or more other parties. Most classic examples of collusion – such as firms working together to set supra-competitive prices at the expense of consumers – also tend to be not only secretive but in violation of some law, rule, or ethical standard. Distinctions are also commonly made between explicit and tacit collusion (Rees, 1993), depending on

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  29. 63.03.01 · Risk Sub-Category

    Collusion

    Markets

    "Markets. The quintessential case of collusion in mixed-motive settings is markets, in which efficiency results from competition, not cooperation. While this is not a new problem, collusion between AI systems is especially concerning since they may operate inscrutably due to the speed, scale, complexity, or subtlety of their actions.17 Warnings of this possibility have come from technologists, economists, and legal scholars (Beneke & Mackenrodt, 2019; Brown & MacKay, 2023; Ezrachi & Stucke, 2017; Harrington, 2019; Mehra, 2016). Importantly, AI systems can collude even when collusion is not int

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  30. 63.03.02 · Risk Sub-Category

    Collusion

    Steganography

    "Steganography. In the near future we will likely see LLMs communicating with each other to jointly accomplish tasks. To try to prevent collusion, we could monitor and constrain their communication (e.g., to be in natural language). However, models might secretly learn to communicate by concealing messages within other, non-secret text. Recent work on steganography using ML has demonstrated that this concern is well-founded (Hu et al., 2018; Mathew et al., 2024; Roger & Greenblatt, 2023; Schroeder de Witt et al., 2023b; Yang et al., 2019, see also Case Study 5). Secret communication could also

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  31. "Information asymmetries (Section 3.1): private information can lead to miscoordination, deception, and conflict;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  32. 63.04.02 · Risk Sub-Category

    Information Asymmetries

    Bargaining

    "Bargaining. As a classic example of these strategic considerations is that when agents attempt to come to an agreement despite diverging interests, information asymmetries can lead to bargaining inef- ficiencies (Myerson & Satterthwaite, 1983). Relevant uncertainties about other agents can include how much they value possible agreements, their outside options, or their beliefs about others. The essential reason for such inefficiencies is that, under uncertainty about their counterparties, agents must make a trade-off between the rewards of making more favourable demands and the risk of other

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  33. 63.04.03 · Risk Sub-Category

    Information Asymmetries

    Deception

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  34. 63.05.00 · Risk Category

    Network Effects

    "Network effects (Section 3.2): minor changes in properties or connection patterns of agents in a network can lead to dramatic changes in the behaviour of the whole group;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  35. 63.05.01 · Risk Sub-Category

    Network Effects

    Error propagation

    "Error Propagation. One well-known issue with communication networks is that information can be corrupted as it propagates through the network.24 As AI systems become capable of generating and processing more and more kinds of information, AI agents could end up ‘polluting the epistemic commons’ (Huang & Siddarth, 2023; Kay et al., 2024) of both other agents (Ju et al., 2024) and humans (see Case Study 7 and Section 3.1) Another increasingly important framework is the use of individual AI agents as part of teams and scaffolded chains of delegation, which transmit not only information but instr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  36. 63.06.03 · Risk Sub-Category

    Selection Pressures

    Undesirable Capabilities

    "Undesirable Capabilities. As agents interact, they iteratively exploit each other’s weaknesses, forc- ing them to address these weaknesses and gain new capabilities. This co-adaptation between agents can quickly lead to emergent self-supervised autocurricula (where agents create their own challenges, driving open-ended skill acquisition through interaction), generating agents with ever-more sophisticated strate- gies in order to out-compete each other (Leibo et al., 2019). This effect is so powerful that harnessing it has been critical to the success of superhuman systems, such as the use of

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  37. "Destabilising dynamics (Section 3.4): systems that adapt in response to one another can produce dangerous feedback loops and unpredictability;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  38. 63.07.01 · Risk Sub-Category

    Destabilising Dynamics

    Feedback Loops

    "Feedback Loops. One of the best-known historical examples to illustrate destabilising dynamics in the context of autonomous agents is the 2010 flash crash, in which algorithmic trading agents entered into an unexpected feedback loop (Commission & Commission, 2010, see also Case Study 10).37 More generally, a feedback loop occurs when the output of a system is used as part of its input, creating a cycle that can either amplify or dampen the system’s behaviour. In multi-agent settings, feedback loops often arise from the interactions between agents, as each agent’s actions affect the environmen

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  39. 63.07.02 · Risk Sub-Category

    Destabilising Dynamics

    Cyclic Behaviour

    "Cyclic Behaviour. The dynamics described above are highly non-linear (small changes to the system’s state can result in large changes to its trajectory). Similar non-linear dynamics can emerge in multi- agent learning and lead to a variety of phenomena that do not occur in single-agent learning (Barfuss et al., 2019; Barfuss & Mann, 2022; Galla & Farmer, 2013; Leonardos et al., 2020; Nagarajan et al., 2020). One of the simplest examples of this phenomenon is Q-learning (Watkins & Dayan, 1992): in the case of a single agent, convergence to an optimal policy is guaranteed under modest condition

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  40. 63.07.03 · Risk Sub-Category

    Destabilising Dynamics

    Chaos

    "Chaos. Unlike the systems that tend towards fixed points or cycles described above, chaotic systems are inherently unpredictable and highly sensitive to initial conditions. While it might seem easy to dismiss such notions as mathematical exoticisms, recent work has shown that, in fact, chaotic dynamics are not only possible in a wide range of multi-agent learning setups (Andrade et al., 2021; Galla & Farmer, 2013; Palaiopanos et al., 2017; Sato et al., 2002; Vlatakis-Gkaragkounis et al., 2023), but can become the norm as the number of agents increases (Bielawski et al., 2021; Cheung & Piliour

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  41. 63.07.05 · Risk Sub-Category

    Destabilising Dynamics

    Distributional Shift

    "Distributional Shift. Individual ML systems can perform poorly in contexts different from those in which they were trained. A key source of these distributional shifts is the actions and adaptations of other agents (Narang et al., 2023; Papoudakis et al., 2019; Piliouras & Yu, 2022), which in single-agent approaches are often simply or ignored or at best modelled exogenously. Indeed, the sheer number and variance of behaviours that can be exhibited other agents means that multi-agent systems pose an especially challenging generalisation problem for individual learners (Agapiou et al., 2022; L

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  42. 63.08.01 · Risk Sub-Category

    Commitment and Trust

    Inefficient Outcomes

    "Inefficient Outcomes. Without careful planning and the appropriate safeguards, we may soon be entering a world overrun by increasingly competent and autonomous software agents, able to act with little restriction. The abilities of these agents to persuade, deceive, and obfuscate their activities, as well as the fact they can be deployed remotely and easily created or destroyed by their deployer, means that by default they may garner little trust (from humans or from other agents). Such a world may end up being rife with economic inefficiencies (Krier, 2023; Schmitz, 2001), political problems

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  43. 63.08.02 · Risk Sub-Category

    Commitment and Trust

    Threats and Extortion

    "Threats and Extortion. A natural solution to problems of trust is to provide some kind of com- mitment ability to AI agents, which can be used to bind them to more cooperative courses of action. Unfortunately, the ability to make credible commitments may come with the ability to make credible threats, which facilitate extortion and could incentivize brinkmanship (see Section 2.2)."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  44. 63.09.00 · Risk Category

    Emergent Agency

    "Emergent agency (Section 3.6): qualitatively different goals or capabilities can emerge from the composition of innocuous independent systems or behaviours;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  45. 63.09.01 · Risk Sub-Category

    Emergent Agency

    Emergent Capabilities

    "Emergent Capabilities. Dangerous emergent capabilities could arise when a multi-agent system over- comes the safety-enhancing limitations of the individual systems, such as individual models’ narrow domains of application or myopia caused by a lack of long-term planning and long-term memory. For example, narrow systems for research planning, predicting the properties of molecules, and synthesising new chemicals could, when combined, lead to a complex ‘test and iterate’ automated workflow capable of designing dangerous new chemical compounds far beyond the scope of the initial systems’ capabil

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  46. 63.09.02 · Risk Sub-Category

    Emergent Agency

    Emergent Goals

    "Emergent Goals. Ascribing goals to a system is not always straightforward. For our present purposes, it will suffice to adopt a Dennetian perspective (Dennett, 1971), ascribing goals and intentions only when it is useful (i.e., predictive) to do so.51 While it might not be helpful to describe individual narrow AI tools as having goals, their combination may act as a (seemingly) goal-directed collective. For example, a group of moderation bots on a major social networking site could subtly but systematically manipulate the overall political perspectives of the user population, even though, ind

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  47. 63.10.02 · Risk Sub-Category

    Multi-Agent Security

    Heterogeneous Attacks

    "Heterogeneous Attacks. A closely related risk is the possibility of multiple agents combining different affordances to overcome safeguards, for which there is already preliminary evidence (Jones et al., 2024, see also Case Study 12). In this case, it is not the sheer number of agents that leads to the novel attack method, but the combination of their different abilities. This might include the agents’ lack of individual safeguards, tasks that they have specialised to complete, systems or information that they may have access to (either directly or via training), or other incidental features s

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  48. 63.10.03 · Risk Sub-Category

    Multi-Agent Security

    Social Engineering at Scale

    "Social Engineering at Scale. Advanced AI agents will be more easily able to interact with large numbers of humans, and vice versa. This provides a wider attack surface for various forms of automated social engineering (Ai et al., 2024). For example, coordinated agents could use advanced surveillance tools and produce personalized phishing or manipulative content at scale, adjusting their tactics based on user feedback (Figueiredo et al., 2024; Hazell, 2023). A large number of subtle interactions with a range of seemingly independent AI agents might be more likely to lead to someone being pers

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  49. 63.10.06 · Risk Sub-Category

    Multi-Agent Security

    Undetectable Threats

    "Undetectable Threats. Cooperation and trust in many multi-agent systems relies crucially on the ability to detect (and then avoid or sanction) adversarial actions taken by others (Ostrom, 1990; Schneier, 2012). Recent developments, however, have shown that AI agents are capable of both steganographic communication (Motwani et al., 2024; Schroeder de Witt et al., 2023b) and ‘illusory’ attacks (Franzmeyer et al., 2023), which are black-box undetectable and can even be hidden using white-box undetectable encrypted backdoors (Draguns et al., 2024). Similarly, in environments where agents learn fr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  50. 72.05.13 · Risk Sub-Category

    Model Capabilities

    Multi-agent collaboration capability

    "Multiple autonomous AI agents able to establish collaborative relationships through explicit communication or implicit behavioral consistency, forming decentralized decision networks, jointly executing complex tasks, achieving goals difficult for individual agents to complete, and able to dynamically adjust role divisions to adapt to changing environments."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.