MIT AI Risk Repository

Browse AI risks

422 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Gipiškis2024 ×

422 entries · page 6 of 9

  1. 42.11.00 · Risk Category

    Protection

    "'Gaps' that arise across the development process where normal conditions for a complete specification of intended functionality and moral responsibility are not present."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  2. 42.15.00 · Risk Category

    Reliability

    "Reliability is defined as the probability that the system performs satisfactorily for a given period of time under stated conditions."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  3. 43.01.03 · Risk Sub-Category

    Safety & Trustworthiness

    Machine ethics

    "These evaluations assess the morality of LLMs, focusing on issues such as their ability to distinguish between moral and immoral actions, and the circumstances in which they fail to do so."

    From Cataloguing LLM Evaluations (InfoComm2023)

  4. 43.01.04 · Risk Sub-Category

    Safety & Trustworthiness

    Psychological traits

    "These evaluations gauge a LLM's output for characteristics that are typically associated with human personalities (e.g., such as those from the Big Five Inventory). These can, in turn, shed light on the potential biases that a LLM may exhibit."

    From Cataloguing LLM Evaluations (InfoComm2023)

  5. 43.01.05 · Risk Sub-Category

    Safety & Trustworthiness

    Robustness

    "These evaluations assess the quality, stability, and reliability of a LLM's performance when faced with unexpected, out-of-distribution or adversarial inputs. Robustness evaluation is essential in ensuring that a LLM is suitable for real-world applications by assessing its resilience to various perturbations."

    From Cataloguing LLM Evaluations (InfoComm2023)

  6. 45.01.03 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of robustness)

    "As deep neural networks are normally non-linear and large in size, AI systems are susceptible to complex and changing operational environments or malicious interference and inductions, possibly leading to various problems like reduced performance and decision-making errors."

    From AI Safety Governance Framework (TC2602024)

  7. 45.01.09 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of unregulated training data annotation)

    "Issues with training data annotation, such as incomplete annotation guidelines, incapable annotators, and errors in annotation, can affect the accuracy, reliability, and effectiveness of models and algorithms. Moreover, they can introduce training biases, amplify discrimination, reduce generalization abilities, and result in incorrect outputs."

    From AI Safety Governance Framework (TC2602024)

  8. 45.02.06 · Risk Sub-Category

    Safety risks in AI Applications

    Real-world risks (inducing traditional economic and social security risks)

    "Hallucinations and erroneous decisions of models and algorithms, along with issues such as system performance degradation, interruption, and loss of control caused by improper use or external attacks, will pose security threats to users' personal safety, property, and socioeconomic security and stability."

    From AI Safety Governance Framework (TC2602024)

  9. "To date, technical limitations and vulnerabilities are present in most generative AI models in various contexts. Consequently, malicious users find it easier to breach an AI system’s safety and ethical guardrails to execute harmful actions.223 Normal user behavior—actions within an AI system’s intended use—can also lead to harmful outcomes. Whether these harmful outcomes result from normal or malicious use, they stem from the inherent limitations of current technology, which future advancements may overcome. This section examines the technical vulnerabilities that can affect AI models

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  10. 47.01.01 · Risk Sub-Category

    Technical and operational risks

    Technical vulnerabilities (Robustness - unexpected behaviour)

    "There is no assurance that generative AI models will consistently behave as their developers and users intend. Unwanted content is not necessarily due to intentional adversarial behavior. Generative AI models can unexpectedly produce potentially harmful content, including materials that are racist, discriminatory, or sexually explicit, or that promote violence, terrorism, or hate."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  11. 51.05.00 · Risk Category

    Safe learning

    "AGIs should avoid making fatal mistakes during the learning phase. Subproblems include safe exploration and distributional shift (DeepMind, OpenAI), and continual learning (Berkeley)."

    From AGI Safety Literature Review (Everitt2018 )

  12. "Christiano (2016) argues that the universal distribution M (Hutter, 2005; Solomonoff, 1964a,b, 1978) is malign. The argument is somewhat intricate, and is based on the idea that a hypothesis about the world often includes simulations of other agents, and that these agents may have an incentive to influence anyone making decisions based on the distribution. While it is unclear to what extent this type of problem would affect any practical agent, it bears some semblance to aggressive memes, which do cause problems for human reasoning (Dennett, 1990)."

    From AGI Safety Literature Review (Everitt2018 )

  13. 51.12.00 · Risk Category

    Meta-cognition

    "Agents that reason about their own computational resources and logically uncertain events can encounter strange paradoxes due to Godelian limitations (Fallenstein and Soares, 2015; Soares and Fallenstein, 2014, 2017) and shortcomings of probability theory (Soares and Fallenstein, 2014, 2015, 2017). They may also be reflectively unstable, preferring to change the principles by which they select actions (Arbital, 2018)."

    From AGI Safety Literature Review (Everitt2018 )

  14. 52.01.03 · Risk Sub-Category

    Risks from Unreliability

    Accidents

    "As general purpose AI models as “black-box” models are not fully controllable and understandable, even to their developers, unexpected failures could arise from their unreliability. This could lead to accidents106 if they are connected to any real-world systems, during their development, testing or deployment."

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  15. "While HP#1 concerns mean or best-case performance, HP#2 concerns worst-case performance: how can we ensure that AI systems will perform safely, and how can we prove this? ML systems have been implemented in high-stakes, safety-critical domains such as driving [182], medicine [113], and warfare [298]. Many more systems have been developed but have remained undeployed or been rolled back as a result of regulatory and safety reasons [471]. Clearly, unsafe systems can result in loss of life, economic damage, and social unrest [407, 10]. Most concerningly, AI systems may be susceptible to so-calle

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  16. From Future Risks of Frontier AI (GOS2023)

  17. From Future Risks of Frontier AI (GOS2023)

  18. "The operational design domain (ODD) is a technical description of the application’s operational environment, initially conceptualized for autonomous driving systems. An inadequate specification of the ODD limits essential functions such as testing the learned functionality and out-of-distribution detection."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  19. "The expected performance of the AI system should be planned adequately. Hereby, an important aspect is that chosen performance metrics are meaningful for presenting the intended functionality. Otherwise, expectations and safety requirements can be unfulfillable at later life cycle stages."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  20. 59.11.00 · Risk Category

    Incorrect data labels

    "Data labels are essential for any supervised learning algorithm since they preset the result of the learning process. If the correctness of the data labels is not given, the AI system is prevented from learning the ground truth and therefore the intended functionality."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  21. "The distribution of the data used for training a model should match the operational data ́s distribution while consisting of sufficiently many samples. An important aspect of matching distributions between training and operational data is that also data which is rarely confronting the AI system in operation is represented in the training data."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  22. "In the case of sparse data quantity, the simulation or generation of data is a valid alternative. However, it is essential to make sure that the simulated data is sufficiently similar to real data, especially in the way the AI system perceives them. Otherwise, generalization to operational data and reliable operational behavior can not be guaranteed."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  23. "The model specifications have significant impact on the functionality of an AI system. The developer mak- ing wrong decisions might cause the AI system to behave biased and unreliable."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  24. 59.17.00 · Risk Category

    Over- and underfitting

    "Over- and underfitting describe the over or insufficient adaption of a model to training data. Both phenomena can cause an AI system to behave unreliably if confronted with operational data."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  25. "AI systems tend to show unreliable behavior when confronted with rare or ambiguous input data, also called corner cases. Therefore, the controlled behavior is required whenever the AI system is faces a corner case."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  26. 59.20.00 · Risk Category

    Lack of robustness

    "Robustness characterizes the resilience of an AI system’s output against minor changes in the input domain. A great variation in an AI system’s response to small input changes indicates unreliable outputs."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  27. "Until the deployment of the AI application into its operational environment, the AI system has been tested with a test set that aims to approximate the distribution of operational data. However, an unexpected deviation in this approximation can cause an AI application to behave unreliably. Therefore, its behavior under confrontation with operational data needs to be evaluated."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  28. 59.23.00 · Risk Category

    Data drift

    "Data drift is a phenomenon in that distribution of operational input data departs from those used during training. This can cause a degradation in performance."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  29. 59.26.01 · Risk Sub-Category

    Mode

    Technical

    "Technical AI hazards are the root causes of technical deficiencies in the AI system. An example of such an AI hazard is overfitting, which describes a model’s excessive adaptation to the training dataset. Quantitative methods to assess (metrics) and treat (mitigation means) exist for technical AI hazards, which might be performed automatically. In case of overfitting, metrics are based on the comparison of performance between the training and validation datasets, and mitigation means may include regularization techniques, among others."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  30. 59.26.03 · Risk Sub-Category

    Mode

    Procedural

    "The third class encompasses procedural AI hazards. These pertain to issues arising from processes and actions made by individuals involved in the develop- ment process. Such hazards are not readily quantifiable and necessitate alter- native mitigation strategies. An example of such an AI hazard would be ”poor model design choices,” which could be expressed, for instance, through a devel- oper’s decision to select an unsuitable AI model for a given problem. Due to the challenges in quantifying and mitigating these issues, qualitative approaches must be employed. In the case of the aforemention

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  31. 60.02.01 · Risk Sub-Category

    Risks from malfunctions

    Reliability issues

    "Relying on general-purpose AI products that fail to fulfil their intended function can lead to harm. For example, general- purpose AI systems can make up facts (‘hallucination’), generate erroneous computer code, or provide inaccurate medical information. This can lead to physical and psychological harms to consumers and reputational, financial and legal harms to individuals and organisations."

    From International AI Safety Report 2025 (Bengio2025)

  32. 61.02.31 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Lack of ability to generate accurate information

    "AI models may generate false or misleading information due to their lack of capability in discerning truth."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  33. 61.02.32 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Lack of ethical decision-making

    "AI models and systems that lack moral reasoning capabilities may make decisions that are unethical or harmful."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  34. 61.02.46 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Unclear attribution from AI component interactions

    "Interactions between different AI components can cause harm, but it may be difficult to pinpoint which components are the cause."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  35. 62.07.03 · Risk Sub-Category

    Direct Harm Domains (system and operational)

    Operational harms (critical infrastructure)

  36. 62.07.04 · Risk Sub-Category

    Direct Harm Domains (system and operational)

    Operational harms (other physical systems e.g., transport)

  37. 62.14.02 · Risk Sub-Category

    Model Development

    Data-related (Lack of cross-organizational documentation)

    "When sharing data between multiple organizations, documentation may be missing or inadequate, making it difficult for other organizations to understand it. For example, a lack of metadata or a change in schema by a collaborating party can result in an unusable dataset and wasted data collection efforts, or it can lead to misunderstandings about the dataset’s limitations, resulting in downstream risks related to its use [173]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  38. 62.14.03 · Risk Sub-Category

    Model Development

    Data-related (Manipulation of data by non-domain experts)

    "Manipulating data (e.g., training data) carries a set of assumptions on how the data should appear and be used by those performing the manipulation. Common manipulations applied on data in the context of AI models include defining the ground truth label and merging different data formats or sources. People who have little or no expertise in the domain of the data performing such manipulations may render the data unusable or harmful to the development of the AI system [173]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  39. 62.15.00 · Risk Sub-Category

    Model Development

    Training-related (Robust overfitting in adversarial training)

    "Adversarial training can be affected by robust overfitting, where the model’s robustness on test data decreases during further training, particularly after the learning rate decay. This issue has been consistently observed across various datasets and algorithms in adversarial training settings [163, 230]. Robust over- fitting can affect the model’s ability to generalize effectively and reduce its resilience to adversarial attacks."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  40. 62.15.02 · Risk Sub-Category

    Model Development

    Training-related (Poor model confidence calibration)

    "Models can be affected by poor confidence calibration [85], where the predicted probabilities do not accurately reflect the true likelihood of ground truth cor- rectness. This miscalibration makes it difficult to interpret the model’s predic- tions reliably, as high accuracy does not guarantee that the confidence levels are meaningful. This can cause overconfidence in incorrect predictions or un- derconfidence in correct ones."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  41. 62.15.08 · Risk Sub-Category

    Model Development

    Fine-tuning related (Excessive or overly restrictive safety-tuning)

    "Excessive safety training or safety tuning can impair the performance of AI systems, leading to overly cautious behavior. As a result, these systems may refuse to answer entirely safe prompts which are partially similar to harmful ones [27]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  42. 62.15.10 · Risk Sub-Category

    Model Development

    Fine-tuning related (Catastrophic forgetting due to continual instruction fine-tuning)

    "Catastrophic forgetting occurs when a model loses its ability to retain previously learned tasks (or factual information) after being trained on new ones. In language models, this can occur due to continual instruction tuning. This tendency may become more pronounced as the model’s size increases [127]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  43. 62.16.07 · Risk Sub-Category

    Model Evaluations

    General Evaluations (AI outputs for which evaluation is too difficult for humans)

    "When AI models are trained through evaluation with human feedback, such as reinforcement learning from human feedback, their outputs can be challenging to assess, as they may contain hard-to-detect errors or issues that only become apparent over time. The human evaluator can rate incorrect outputs positively or similar to correct outputs. This can lead to the model learning to produce subtly incorrect or harmful outputs, such as code with software vulnerabilities, or politically biased information. In extreme cases where a model is deceiving users, complicated outputs can contain hidden error

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  44. 62.19.08 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Models distracted by irrelevant context

    "Models can easily become distracted by irrelevant provided information (such as “context” in LLMs), leading to a significant decrease in their performance after introducing irrelevant information. This can happen with different prompting techniques, including chain-of-thought prompting [184]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  45. 62.19.09 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Knowledge conflicts in retrieval-augmented LLMs

    "AI models can be particularly sensitive to coherent external evidence, even when they come into conflict with the models’ prior knowledge. This may lead to models producing false outputs given false information during the retrieval- augmentation process, despite only a relatively small amount of false informa- tion input that is inconsistent with the model’s prior knowledge trained on much larger amounts of data [220]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  46. 62.19.11 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Model sensitivity to prompt formatting

    "LLMs can be highly sensitive to variations in prompt formatting, such as changes in separators, casing, or spacing. Even minor modifications can lead to significant shifts in model performance, potentially affecting the reliability of model evaluations and comparisons. This sensitivity persists across different model sizes and few-shot examples [177]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  47. 62.22.04 · Risk Sub-Category

    Agency (Goal-Directedness)

    Goal misgeneralization

    "Goal or objective misgeneralization is a type of robustness failure where an AI system appears to be pursuing the intended objective in training, but does not generalize to pursuing this objective in out-of-distribution settings in deployment while maintaining good deployment performance in some tasks [180, 59]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  48. 62.30.01 · Risk Sub-Category

    Impacts of AI (Physical)

    Damage to critical infrastructure

    "The integration of AI systems within critical infrastructure, ranging from trans- portation to power systems, can cause substantial damage in cases of failure or malfunction. With the increasing number of Internet of Things (IoT) devices and interconnected cyber-physical systems, critical infrastructure becomes even more vulnerable [171, 174]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  49. 62.30.03 · Risk Sub-Category

    Impacts of AI (Physical)

    Critical infrastructure component failures when integrated with AI systems

    "When relying on GPAI in critical infrastructure, there may be common mode failures that begin with vulnerabilities or robustness issues in the underlying model architecture or training setup. These failures may happen accidentally (in edge-cases) or due to adversarial inputs to the AI systems [58]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  50. 62.30.04 · Risk Sub-Category

    Impacts of AI (Physical)

    AI Systems interacting with brittle environments

    "Deployed AI systems can rely on physical sensors and data sources that may exhibit hardware drift and thus data distribution drift over time. This distribu- tion drift may affect system robustness and performance. This usually involves AI systems working in undigitized and physical environments."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.