MIT AI Risk Repository

Browse AI risks

190 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

190 entries · page 4 of 4

  1. 45.01.09 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of unregulated training data annotation)

    "Issues with training data annotation, such as incomplete annotation guidelines, incapable annotators, and errors in annotation, can affect the accuracy, reliability, and effectiveness of models and algorithms. Moreover, they can introduce training biases, amplify discrimination, reduce generalization abilities, and result in incorrect outputs."

    From AI Safety Governance Framework (TC2602024)

  2. 51.05.00 · Risk Category

    Safe learning

    "AGIs should avoid making fatal mistakes during the learning phase. Subproblems include safe exploration and distributional shift (DeepMind, OpenAI), and continual learning (Berkeley)."

    From AGI Safety Literature Review (Everitt2018 )

  3. "Christiano (2016) argues that the universal distribution M (Hutter, 2005; Solomonoff, 1964a,b, 1978) is malign. The argument is somewhat intricate, and is based on the idea that a hypothesis about the world often includes simulations of other agents, and that these agents may have an incentive to influence anyone making decisions based on the distribution. While it is unclear to what extent this type of problem would affect any practical agent, it bears some semblance to aggressive memes, which do cause problems for human reasoning (Dennett, 1990)."

    From AGI Safety Literature Review (Everitt2018 )

  4. "The operational design domain (ODD) is a technical description of the application’s operational environment, initially conceptualized for autonomous driving systems. An inadequate specification of the ODD limits essential functions such as testing the learned functionality and out-of-distribution detection."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  5. "The expected performance of the AI system should be planned adequately. Hereby, an important aspect is that chosen performance metrics are meaningful for presenting the intended functionality. Otherwise, expectations and safety requirements can be unfulfillable at later life cycle stages."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  6. 59.11.00 · Risk Category

    Incorrect data labels

    "Data labels are essential for any supervised learning algorithm since they preset the result of the learning process. If the correctness of the data labels is not given, the AI system is prevented from learning the ground truth and therefore the intended functionality."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  7. "The distribution of the data used for training a model should match the operational data ́s distribution while consisting of sufficiently many samples. An important aspect of matching distributions between training and operational data is that also data which is rarely confronting the AI system in operation is represented in the training data."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  8. "In the case of sparse data quantity, the simulation or generation of data is a valid alternative. However, it is essential to make sure that the simulated data is sufficiently similar to real data, especially in the way the AI system perceives them. Otherwise, generalization to operational data and reliable operational behavior can not be guaranteed."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  9. "The model specifications have significant impact on the functionality of an AI system. The developer mak- ing wrong decisions might cause the AI system to behave biased and unreliable."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  10. 59.26.01 · Risk Sub-Category

    Mode

    Technical

    "Technical AI hazards are the root causes of technical deficiencies in the AI system. An example of such an AI hazard is overfitting, which describes a model’s excessive adaptation to the training dataset. Quantitative methods to assess (metrics) and treat (mitigation means) exist for technical AI hazards, which might be performed automatically. In case of overfitting, metrics are based on the comparison of performance between the training and validation datasets, and mitigation means may include regularization techniques, among others."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  11. 59.26.03 · Risk Sub-Category

    Mode

    Procedural

    "The third class encompasses procedural AI hazards. These pertain to issues arising from processes and actions made by individuals involved in the develop- ment process. Such hazards are not readily quantifiable and necessitate alter- native mitigation strategies. An example of such an AI hazard would be ”poor model design choices,” which could be expressed, for instance, through a devel- oper’s decision to select an unsuitable AI model for a given problem. Due to the challenges in quantifying and mitigating these issues, qualitative approaches must be employed. In the case of the aforemention

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  12. 62.14.02 · Risk Sub-Category

    Model Development

    Data-related (Lack of cross-organizational documentation)

    "When sharing data between multiple organizations, documentation may be missing or inadequate, making it difficult for other organizations to understand it. For example, a lack of metadata or a change in schema by a collaborating party can result in an unusable dataset and wasted data collection efforts, or it can lead to misunderstandings about the dataset’s limitations, resulting in downstream risks related to its use [173]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  13. 62.14.03 · Risk Sub-Category

    Model Development

    Data-related (Manipulation of data by non-domain experts)

    "Manipulating data (e.g., training data) carries a set of assumptions on how the data should appear and be used by those performing the manipulation. Common manipulations applied on data in the context of AI models include defining the ground truth label and merging different data formats or sources. People who have little or no expertise in the domain of the data performing such manipulations may render the data unusable or harmful to the development of the AI system [173]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  14. 62.15.00 · Risk Sub-Category

    Model Development

    Training-related (Robust overfitting in adversarial training)

    "Adversarial training can be affected by robust overfitting, where the model’s robustness on test data decreases during further training, particularly after the learning rate decay. This issue has been consistently observed across various datasets and algorithms in adversarial training settings [163, 230]. Robust over- fitting can affect the model’s ability to generalize effectively and reduce its resilience to adversarial attacks."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  15. 65.02.01 · Risk Sub-Category

    Training Data Risks (Data laws)

    Data usage restrictions

    "Laws and other restrictions can limit or prohibit the use of some data for specific AI use cases."

    From AI Risk Atlas (IBM2025)

  16. 65.02.02 · Risk Sub-Category

    Training Data Risks (Data laws)

    Data acquisition restrictions

    "Laws and other regulations might limit the collection of certain types of data for specific AI use cases."

    From AI Risk Atlas (IBM2025)

  17. 65.02.03 · Risk Sub-Category

    Training Data Risks (Data laws)

    Data transfer restrictions

    "Laws and other restrictions can limit or prohibit transferring data."

    From AI Risk Atlas (IBM2025)

  18. 65.06.01 · Risk Sub-Category

    Training Data Risks (Accuracy)

    Data contamination

    "Data contamination occurs when incorrect data is used for training. For example, data that is not aligned with model’s purpose or data that is already set aside for other development tasks such as testing and evaluation."

    From AI Risk Atlas (IBM2025)

  19. 65.06.02 · Risk Sub-Category

    Training Data Risks (Accuracy)

    Unrepresentative data

    "Unrepresentative data occurs when the training or fine-tuning data is not sufficiently representative of the underlying population or does not measure the phenomenon of interest."

    From AI Risk Atlas (IBM2025)

  20. 65.07.02 · Risk Sub-Category

    Training Data Risks (Value alignment)

    Improper data curation

    "Improper collection and preparation of training or tuning data includes data label errors and by using data with conflicting information or misinformation."

    From AI Risk Atlas (IBM2025)

  21. "Non-transparent or untraceable integration of upstream third-party components, including data that has been improperly obtained or not processed and cleaned due to increased automation from GAI; improper supplier vetting across the AI lifecycle; or other issues that diminish transparency or accountability for downstream users."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  22. 51.06.00 · Risk Category

    Intelligibility

    "How can we build agent’s whose decisions we can understand? Con- nects explainable decisions (Berkeley) and informed oversight (MIRI)."

    From AGI Safety Literature Review (Everitt2018 )

  23. "Today's Frontier AI is difficult to interpret and lacks transparency. Contextual understanding of the training data is not explicitly embedded within these models. They can fail to capture perspectives of underrepresented groups or the limitations within which they are expected to perform without fine tuning or reinforcement learning with human feedback (RLHF)."

    From Future Risks of Frontier AI (GOS2023)

  24. "Throughout the development of an AI system, it is vital to document every decision and action taken. This is not only essential to optimize the development process itself but also required for the auditability of the AI system."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  25. "The transparency to end users of the AI system increases the user’s trust in the AI application. If not adequately integrated into the design, this might prevent the proper operation and cause potential misuse of the AI application."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  26. 63.06.00 · Risk Category

    Selection Pressures

    "Selection pressures (Section 3.3): some aspects of training and selection by those deploying and using AI agents can lead to undesirable behaviour;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  27. 59.25.01 · Risk Sub-Category

    AI lifecycle stage

    (1) Scoping

    "A majority of them possess an initial stage devoted to the planning and scoping of the AI system."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  28. 59.25.02 · Risk Sub-Category

    AI lifecycle stage

    (2) Data collection and preparation

  29. 59.25.03 · Risk Sub-Category

    AI lifecycle stage

    (3) Modeling

  30. "As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  31. 62.04.01 · Risk Sub-Category

    Dimension - Technical Attributes (AI inadequacy - technical failure)

    Supervised/unsupervised AI (AI data quality related - biased training data)

    "As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  32. 62.04.02 · Risk Sub-Category

    Dimension - Technical Attributes (AI inadequacy - technical failure)

    Supervised/unsupervised AI (AI training performance related - Robustness)

    "As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  33. 62.04.03 · Risk Sub-Category

    Dimension - Technical Attributes (AI inadequacy - technical failure)

    Supervised/unsupervised AI (AI training performance related - Accuracy)

    "As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  34. 62.04.04 · Risk Sub-Category

    Dimension - Technical Attributes (AI inadequacy - technical failure)

    Supervised/unsupervised AI (AI training performance related - Reliability)

    "As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  35. 62.04.05 · Risk Sub-Category

    Dimension - Technical Attributes (AI inadequacy - technical failure)

    Reinforcement learning AI (Training design related)

    "As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  36. 62.04.06 · Risk Sub-Category

    Dimension - Technical Attributes (AI inadequacy - technical failure)

    Reinforcement learning AI (Training performance related)

    "As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  37. 62.06.01 · Risk Sub-Category

    Dimension - Stage of Risk Emergence

    Pre-deployment

    "For GPAIs or foundation models, risks emerge during training, prior to being repurposed and deployed in more specific AI systems or applications. Risk assessments can be conducted before deployment, and monitoring of AI models can occur as required throughout the deployment phase. In certain cases, version updates or model recalls may be warranted post-deployment."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  38. 62.14.00 · Risk Category

    Model Development

  39. 62.16.00 · Risk Category

    Model Evaluations

    "This section catalogs the risk sources and risk management measures related to model evaluations (often called evals). We categorize them into the fol- lowing groups: general evaluations, benchmarking, red teaming, auditing, and interpretability/explainability. The subsection on general evaluations consists of items that are common to various evaluation techniques, while the other subsections are specific to their respective evaluation types."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.