MIT AI Risk Repository

Browse AI risks

336 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

336 entries · page 6 of 7

  1. 20.02.00 · Risk Category

    AI Ethics

    "Ethical challenges are widely discussed in the literature and are at the heart of the debate on how to govern and regulate AI technology in the future (Bostrom & Yudkowsky, 2014; IEEE, 2017; Wirtz et al., 2019). Lin et al. (2008, p. 25) formulate the problem as follows: “there is no clear task specification for general moral behavior, nor is there a single answer to the question of whose morality or what morality should be implemented in AI”. Ethical behavior mostly depends on an underlying value system. When AI systems interact in a public environment and influence citizens, they are expecte

    From The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration (Wirtz2020)

  2. 20.02.02 · Risk Sub-Category

    AI Ethics

    Compatibility of AI vs. human value judgement

    "Compatibility of machine and human value judgment refers to the challenge whether human values can be globally implemented into learning AI systems without the risk of developing an own or even divergent value system to govern their behavior and possibly become harmful to humans."

    From The Dark Sides of Artificial Intelligence: An Integrated AI Governance Framework for Public Administration (Wirtz2020)

  3. 37.01.02 · Risk Sub-Category

    Design of AI

    Balancing AI's risks

    "This category constitutes more than 16% of the articles and focuses on addressing the potential risks associated with AI systems. Given the ubiquity of AI technologies, these articles explore the implications of AI risks across various contexts linked to design and unpredictability, military purposes, emergency procedures, and AI takeover."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  4. 42.04.00 · Risk Category

    Moral

    "Less moral responsibility humans will feel regarding their life-or-death decisions with the increase of machines autonomy."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  5. 45.02.06 · Risk Sub-Category

    Safety risks in AI Applications

    Real-world risks (inducing traditional economic and social security risks)

    "Hallucinations and erroneous decisions of models and algorithms, along with issues such as system performance degradation, interruption, and loss of control caused by improper use or external attacks, will pose security threats to users' personal safety, property, and socioeconomic security and stability."

    From AI Safety Governance Framework (TC2602024)

  6. "Christiano (2016) argues that the universal distribution M (Hutter, 2005; Solomonoff, 1964a,b, 1978) is malign. The argument is somewhat intricate, and is based on the idea that a hypothesis about the world often includes simulations of other agents, and that these agents may have an incentive to influence anyone making decisions based on the distribution. While it is unclear to what extent this type of problem would affect any practical agent, it bears some semblance to aggressive memes, which do cause problems for human reasoning (Dennett, 1990)."

    From AGI Safety Literature Review (Everitt2018 )

  7. 51.12.00 · Risk Category

    Meta-cognition

    "Agents that reason about their own computational resources and logically uncertain events can encounter strange paradoxes due to Godelian limitations (Fallenstein and Soares, 2015; Soares and Fallenstein, 2014, 2017) and shortcomings of probability theory (Soares and Fallenstein, 2014, 2015, 2017). They may also be reflectively unstable, preferring to change the principles by which they select actions (Arbital, 2018)."

    From AGI Safety Literature Review (Everitt2018 )

  8. "The distribution of the data used for training a model should match the operational data ́s distribution while consisting of sufficiently many samples. An important aspect of matching distributions between training and operational data is that also data which is rarely confronting the AI system in operation is represented in the training data."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  9. "In the case of sparse data quantity, the simulation or generation of data is a valid alternative. However, it is essential to make sure that the simulated data is sufficiently similar to real data, especially in the way the AI system perceives them. Otherwise, generalization to operational data and reliable operational behavior can not be guaranteed."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  10. 59.17.00 · Risk Category

    Over- and underfitting

    "Over- and underfitting describe the over or insufficient adaption of a model to training data. Both phenomena can cause an AI system to behave unreliably if confronted with operational data."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  11. 59.23.00 · Risk Category

    Data drift

    "Data drift is a phenomenon in that distribution of operational input data departs from those used during training. This can cause a degradation in performance."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  12. 62.15.00 · Risk Sub-Category

    Model Development

    Training-related (Robust overfitting in adversarial training)

    "Adversarial training can be affected by robust overfitting, where the model’s robustness on test data decreases during further training, particularly after the learning rate decay. This issue has been consistently observed across various datasets and algorithms in adversarial training settings [163, 230]. Robust over- fitting can affect the model’s ability to generalize effectively and reduce its resilience to adversarial attacks."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  13. 62.15.02 · Risk Sub-Category

    Model Development

    Training-related (Poor model confidence calibration)

    "Models can be affected by poor confidence calibration [85], where the predicted probabilities do not accurately reflect the true likelihood of ground truth cor- rectness. This miscalibration makes it difficult to interpret the model’s predic- tions reliably, as high accuracy does not guarantee that the confidence levels are meaningful. This can cause overconfidence in incorrect predictions or un- derconfidence in correct ones."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  14. 62.15.10 · Risk Sub-Category

    Model Development

    Fine-tuning related (Catastrophic forgetting due to continual instruction fine-tuning)

    "Catastrophic forgetting occurs when a model loses its ability to retain previously learned tasks (or factual information) after being trained on new ones. In language models, this can occur due to continual instruction tuning. This tendency may become more pronounced as the model’s size increases [127]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  15. 62.34.01 · Risk Sub-Category

    Impacts of AI (Bias)

    Homogenization or correlated failures in model derivatives

    "Homogenization refers to common methodologies and models used across down- stream GPAI systems, which may lead to uniform failures and amplification of biases [176, 30]. This risk arises when numerous downstream AI systems are built upon a few large-scale foundation models."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  16. 65.06.02 · Risk Sub-Category

    Training Data Risks (Accuracy)

    Unrepresentative data

    "Unrepresentative data occurs when the training or fine-tuning data is not sufficiently representative of the underlying population or does not measure the phenomenon of interest."

    From AI Risk Atlas (IBM2025)

  17. 70.01.02 · Risk Sub-Category

    Physical Risks

    Accidental harm

    "Automation in sectors ranging from manufacturing to healthcare has and will increasingly put humans into close contact with EAI systems [7]. This interaction increases the risk of accidental physical harm. Though accidental harm has been a longstanding issue in industrial robotics, increased AI capabilities could exacerbate this risk; several recent reports document an increase in industrial injuries following the introduction of AI-controlled robots [66–68]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  18. 71.01.03 · Risk Sub-Category

    Scientific Domain of Agents

    Radiological Risks

    "Radiological risks involve both immediate operational hazards, such as exposure incidents or containment failures during the automated handling of radioactive materials, and broader security concerns regarding the potential misuse of AI systems in nuclear research."

    From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025)

  19. 71.01.04 · Risk Sub-Category

    Scientific Domain of Agents

    Physical (Mechanical ) Risks

    "Physical (mechanical) risks are associated with robotics and automated systems, which could lead to equipment malfunctions or physical harm in laboratory settings."

    From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025)

  20. 33.02.03 · Risk Sub-Category

    Technology concerns

    Explainability

    "A recurrent concern about AI algorithms is the lack of explainability for the model, which means information about how the algorithm arrives at its results is deficient (Deeks, 2019). Specifically, for generative AI models, there is no transparency to the reasoning of how the model arrives at the results (Dwivedi et al., 2023). The lack of transparency raises several issues. First, it might be difficult for users to interpret and understand the output (Dwivedi et al., 2023). It would also be difficult for users to discover potential mistakes in the output (Rudin, 2019). Further, when the inte

    From Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023)

  21. 37.02.04 · Risk Sub-Category

    Human-AI interaction

    Attributing the responsibility for AI's failures

    "This section, constituting almost 8% of the articles, addresses the implications arising from AI acting and learning without direct human supervision, encompassing two main issues: a responsibility gap and AI's moral status."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  22. 42.01.00 · Risk Category

    Accountability

    "The ability to determine whether a decision was made in accordance with procedural and substantive standards and to hold someone responsible if those standards are not met."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  23. 47.01.05 · Risk Sub-Category

    Technical and operational risks

    Opacity (the black box problem)

    "Opacity surrounding the technical, internal decision-making processes of generative AI models is popularly known as the “black box problem.”277 Generative AI models, most ubiquitously built on deep neural networks with hundreds of billions of internal connections,278 have become so complex that their internal decision-making processes are no longer traceable or interpretable to even the most advanced expert observers. This means that, while the inputs and outputs of a system can be observed, developers cannot explain in detail why specific inputs correspond to specific outputs."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  24. 59.18.00 · Risk Category

    Lack of explainability

    "The explainability of AI systems based on so-called black-box models is often limited. This opaqueness of AI systems can prevent developers from detecting shortcomings in the data or the model itself and decrease the performance and safety levels of the AI system."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  25. 59.24.00 · Risk Category

    Concept drift

    "Concept drift refers to a change in the rela- tionship between input variables and model output. If not treated appropriately, concept drift can reduce the reliability of AI systems."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  26. 61.02.15 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Complexity-induced knowledge gap

    "The complexity of AI models and systems makes it challenging to demonstrate harm or establish a clear causal link between AI actions and their consequences."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  27. 61.02.37 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Opaque AI networks

    "The complexity and opacity of AI models and systems make it difficult to predict and manage their behavior."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  28. 62.16.03 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Difficulty of identification and measurement of capabilities)

    "The capabilities of general-purpose AI systems can be difficult to measure, compared to the capabilities of more limited and fixed-purpose AI systems. This is in part due to a broader distribution of potential risks, a lack of well-defined metrics to evaluate these risks, and risks from unpredictable (or emergent) AI model properties."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  29. 62.19.10 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Lack of understanding of in-context learning in language models

    "In-context learning allows the model to learn a new task or improve its perfor- mance by providing examples in the prompt, without changing its weights [101]. Even though this technique is highly effective, its working mechanism is not well understood. Since many potential misuses are directly related to prompting, it becomes difficult to guarantee safety when the exact mechanism of in-context learning is not fully investigated [13]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  30. 65.17.02 · Risk Sub-Category

    Output risks (Explainability)

    Untraceable attribution

    "The content of the training data used for generating the model’s output is not accessible."

    From AI Risk Atlas (IBM2025)

  31. 70.04.03 · Risk Sub-Category

    Social Risks

    Lack of transparency, explainability, and trust

    "Understanding how AI reaches conclusions or why AI systems perform specific actions motivates an entire branch of interpretability research [111], but physical embodiment raises the stakes for understanding these systems. For example, transparency of planned actions and explainability of decision-making is crucial when an AV suddenly changes lanes. A lack of transparency and explainability could lead to a lack of trust, which could become a critical and socially destabilizing issue with the widespread deployment of EAI [112–114]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  32. 62.31.02#2 · Risk Sub-Category

    Impacts of AI (Financial Impacts)

    Financial instability due to model homogeneity

    "The widespread use of similar models or algorithms across the financial sec- tor can lead to synchronized reactions to market signals, increasing volatility, triggering flash crashes, or market illiquidity [4]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  33. 63.04.01 · Risk Sub-Category

    Information Asymmetries

    Communication constraints

    "Communication Constraints. A fundamental source of information asymmetries is that constraints on information exchange can exist, even when agents share a common goal (see Section 2.1). These might be constraints on space (i.e., the amount of information that can be communicated) if the information that needs to be communicated is especially complex, time if a snap decision is required before all information can be communicated, or both."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  34. 63.05.02 · Risk Sub-Category

    Network Effects

    Network rewiring

    "Network Rewiring. A different class of problems concerns not changes in the content transmitted through the network but changes in the network structure itself (Albert et al., 2000)."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  35. 63.05.03 · Risk Sub-Category

    Network Effects

    Homogeneity and correlated failures

    "Homogeneity and Correlated Failures. The current paradigm driving the state of the art in AI is the ‘foundation model’ (Bommasani et al., 2021): large-scale ML models pre-trained on broad data, which can be repurposed for a wide range of downstream applications. The costs required to create such models (and continuing returns to scale) means that only well-resourced actors can create cutting- edge models (Epoch, 2023; Hoffmann et al., 2022; Kaplan et al., 2020), making them relatively few in number. If current trends continue, it is likely that many AI agents will be powered by a small number

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  36. 63.06.01 · Risk Sub-Category

    Selection Pressures

    Undesirable Dispositions from Competition

    "Undesirable Dispositions from Competition. It is plausible that evolution selected for certain conflict-prone dispostions in humans, such as vengefulness, aggression, risk-seeking, selfishness, dishon- esty, deception, and spitefulness towards out-groups (Grafen, 1990; Han, 2022; Konrad & Morath, 2012; McNally & Jackson, 2013; Nowak, 2006; Rusch, 2014). Such traits could also be selected for in ML systems that are trained in more competitive multi-agent settings. For example, this might happen if systems are selected based on their performance relative to other agents (and so one agent’s loss

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  37. 63.06.02 · Risk Sub-Category

    Selection Pressures

    Undesirable Dispositions from Human Data

    "Undesirable Dispositions from Human Data. It is well-understood that models trained on human data – such as being pre-trained on human-written text or fine-tuned on human feedback – can exhibit human biases. For these reasons, there has already been considerable attention to measuring biases related to protected characteristics such as sex and ethnicity (e.g., Ferrara, 2023; Liang et al., 2021; Nadeem et al., 2020; Nangia et al., 2020), which can be amplified in multi-agent settings (Acerbi & Stubbersfield, 2023, see also Case Study 7). More recently, there has been increasing attention paid

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  38. 63.07.04 · Risk Sub-Category

    Destabilising Dynamics

    Phase Transitions

    "Phase Transitions. Finally, small external changes to the system – such as the introduction of new agents or a distributional shift – can cause phase transitions, where the system undergoes an abrupt qualitative shift in overall behaviour (Barfuss et al., 2024). Formally, this corresponds to bifurcations in the system’s parameter space, which lead to the creation or destruction of dynamical attractors, resulting in complex and unpredictable dynamics (Crawford, 1991; Zeeman, 1976). For example, Leonardos & Piliouras (2022) show that changes to the exploration hyperparameter of RL agents can le

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  39. "Commitment and trust (Section 3.5): difficulties in forming credible commitments, trust, or reputation can prevent mutual gains in AI-AI and human-AI interactions;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  40. "Multi-agent security (Section 3.7): multi-agent systems give rise to new kinds of security threats and vulnerabilities."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  41. 63.10.04 · Risk Sub-Category

    Multi-Agent Security

    Vulnerable AI Agents

    "Vulnerable AI Agents. The use of AI agents as delegates or representatives of humans or organisa- tions also introduces the possibility of attacks on AI agents themselves. In other words, agents can be considered vulnerable extensions of their principals, introducing a novel attack surface (SecureWorks, 2023). Attacks on an AI agent could be used to extract private information about their principal (Wei & Liu, 2024; Wu et al., 2024a), or to manipulate the agent to take actions that the principal would find undesirable (Zhang et al., 2024a). This includes attacks that have direct relevance for

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  42. 63.10.05 · Risk Sub-Category

    Multi-Agent Security

    Cascading Security Failures

    "Cascading Security Failures. Localised attacks in multi-agent systems can result in catastrophic macroscopic outcomes (Motter & Lai, 2002, see also Sections 3.2 and 3.4). These cascades can be hard to mitigate or recover from because component failure may be difficult to detect or localise in multi-agent systems (Lamport et al., 1982), and authentication challenges can facilitate false flag attacks (Skopik & Pahi, 2020). Computer worms represent a classic example of a cybersecurity threat that relies inherently on networked systems. Recent work has provided preliminary evidence that similar a

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  43. 72.03.02 · Risk Sub-Category

    Accident Risks

    Impact on Financial Stability

    "The integration of general-purpose AI into high-frequency trading, market-making, or systemic risk management could exacerbate systemic risk by exhibiting unexpected behavioral patterns during market stress. Moreover, the concentration of a few homogeneous foundation models across financial institutions may foster correlated decision-making and herd-following behaviors. The widespread adoption of AI agents could also amplify volatility through emergent phenomena from multi-agent interactions.23 All of these could precipitate a cascading global-scale financial system instability, with potentia

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  44. "A foremost lesson of game theory is that optimal decision-making within a single-agent setting (i.e. selfishly optimizing for an agent’s own utility) can produce sub-optimal outcomes in the presence of other strategic agents. Failing to account for the strategic nature of other agents can cause an agent to adopt strategies under which potentially everyone, including the agent itself, ends up worse off (Schelling, 1981; Harsanyi, 1995; Roughgarden, 2005; Nisan, 2007). Examples include collective action problems (or ‘social dilemmas’) such as arms races or the depletion of common resources, as

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  45. 73.02.01 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Foundationality May Cause Correlated Failures

    "Another important characteristic of LLM development is foundationality — due to the expense of large- scale pretraining, many deployed instances share similar or identical learned components. Foundation- ality may both be a blessing and a curse. On the one hand, it may be possible to exploit the similarity in the design of LLM-agents to facilitate cooperation (Critch et al., 2022; Conitzer and Oesterheld, 2023; Oesterheld et al., 2023). On the other hand, foundationality may leave LLM-agents vulnerable to correlated failures both in terms of safety and capabilities due to increased output hom

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  46. 73.02.02 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Groups of LLM-Agents May Show Emergent Functionality

    "Multi-agent learning, either through explicit finetuning or implicit in-context learning, may enable LLM-agents to influence each other during their interactions (Foerster et al., 2018). Under some environmental settings, this can create feedback loops that result in novel and emergent behaviors that would not manifest in the absence of multi-agent interactions (Hammond et al., 2024, Section 3.6). Emergent functionality is a safety risk in two ways. Firstly, it may itself be dangerous (Shevlane et al., 2023). Secondly, it makes assurance harder as such emergent behaviors are difficult to pre

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  47. 46.04.01 · Risk Sub-Category

    Socio-technical and Infrastructural

    Deception - Systemic abberations

  48. From Future Risks of Frontier AI (GOS2023)

  49. 58.02.01 · Risk Sub-Category

    Physical

    Bodily Injury

    "Bodily injury - Physical pain, injury, illness, or disease suffered by an individual or group due to the malfunction, use or misuse of a technology system."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  50. 58.02.03 · Risk Sub-Category

    Physical

    Personal Health Deterioration

    "Personal health deterioration - Physical deterioration of an individual or animal over time, increasing their risk of disease, organ failure, prolonged hospital stay or death, etc."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.