MIT AI Risk Repository

Browse AI risks

190 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

190 entries · page 2 of 4

  1. 65.08.01 · Risk Sub-Category

    Training Data Risks (Robustness)

    Data poisoning

    "A type of adversarial attack where an adversary or malicious insider injects intentionally corrupted, false, misleading, or incorrect samples into the training or fine-tuning datasets."

    From AI Risk Atlas (IBM2025)

  2. "The previous section explored jailbreaks and other forms of adversarial prompts as ways to elicit harmful capabilities acquired during pretraining. These methods make no assumptions about the training data. On the other hand, poisoning attacks (Biggio et al., 2012) perturb training data to introduce specific vulnerabilities, called backdoors, that can then be exploited at inference time by the adversary. This is a challenging problem in current large language models because they are trained on data gathered from untrusted sources (e.g. internet), which can easily be poisoned by an adversary (

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  3. 74.02.02 · Risk Sub-Category

    Malicious Use

    Jailbreak in LLM Malicious Use - Poisoning Training Data

    "In the data collecting and pre-training phase, malicious adversaries can Jailbreak LLMs through poisoning their training data to make the model to output harmful content."

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  4. 74.02.03 · Risk Sub-Category

    Malicious Use

    Jailbreak in LLM Malicious Use - Backdoor Attack

    "However, there are still ones who can leave holes in the training dataset, making LLMs appear safe on average, but generate harmful content under other specific conditions. This kind of attack can be categorized as "backdoor attack". Evan et al. developed a backdoor model that behaves as expected when trained, but exhibits different and potentially harmful behavior when deployed [81]. The results show that these backdoor behaviors persist even after multiple security training techniques are applied."

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  5. 74.02.04 · Risk Sub-Category

    Malicious Use

    Jailbreak in LLM Malicious Use - White & Black Box Attacks

    "In the fine-tuning and alignment phase, elaborately- designed instruction datasets can be utilized to fine-tune LLMs to drive them to perform undesirable behaviors, such as generating harmful information or content that violates ethical norms, and thus achieve a jailbreak. Based on the accessibility to the model parameters, we can categorize them into white-box and black-box attacks. For white-box attacks, we can jailbreak the model by modifying its parameter weights. In [107], Lermen et al. used LoRA to fine-tune the Llama2’s 7B, 13B, and 70B as well as Mixtral on AdvBench and RefusalBench d

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  6. 02.09.02 · Risk Sub-Category

    Hallucinations

    Noisy Training Data

    "Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation. When training on large-scale corpora, this issue becomes more serious because it is difficult to eliminate all the noise from the massive pre-training data."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  7. 02.09.03 · Risk Sub-Category

    Hallucinations

    Defective Decoding Process

    In general, LLMs employ the Transformer architecture [32] and generate content in an autoregressive manner, where the prediction of the next token is conditioned on the previously generated token sequence. Such a scheme could accumulate errors [105]. Besides, during the decoding process, top-p sampling [28] and top-k sampling [27] are widely adopted to enhance the diversity of the generated content. Nevertheless, these sampling strategies can introduce “randomness” [113], [136], thereby increasing the potential of hallucinations"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  8. 22.01.02 · Risk Sub-Category

    Malicious Use (Intentional)

    Unleashing AI Agents

    "people could build AIs that pursue dangerous goals’"

    From An Overview of Catastrophic AI Risks (Hendrycks2023)

  9. 53.02.02 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Acquisition of a goal to harm society

    "cases of AI systems being given the outright goal of harming humanity (ChaosGPT);"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  10. 37.01.00 · Risk Category

    Design of AI

    "ethical concerns regarding how AI is designed and who designs it"

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  11. 54.04.02 · Risk Sub-Category

    Within-country issues: domestic inequality

    Privatization of AI

    "Researchers in deep learning and those with greater research impact are more likely to migrate to industry, raising concerns about the “privatization of AI knowledge” [278]. Specically, if the most sophisticated AI approaches become proprietary and are used only within private research labs, then it will be impossible for universities to teach them, let alone contribute to leading research."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  12. 13.01.07 · Risk Sub-Category

    Impacts: The Technical Base System

    Data and Content Moderation Labor

    "Two key ethical concerns in the use of crowdwork for generative AI systems are: crowdworkers are frequently subject to working conditions that are taxing and debilitative to both physical and mental health, and there is a widespread deficit in documenting the role crowdworkers play in AI development. This contributes to a lack of transparency and explainability in resulting model outputs. Manual review is necessary to limit the harmful outputs of AI systems, including generative AI systems. A common harmful practice is to intentionally employ crowdworkers with few labor protections, often tak

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  13. 18.06.05 · Risk Sub-Category

    Socioeconomic and environmental harms

    Exploitative data sourcing and enrichment

    "Perpetuating exploitative labour practices to build AI systems (sourcing, user testing)"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  14. 54.01.01 · Risk Sub-Category

    Negative impacts of AI use

    Under-recognized work

    "Without training data, ML cannot take place. Much of this data comes from paid clickwork (also called “platform work” [170] or “microwork” [558]), unpaid crowdsourcing, and unpaid user behavior capture. Clickworkers, mainly in the global south, perform repetitive data-labeling tasks for use in the training of ML models [558]. The market value of such annotations “is projected to reach $13.7 billion by 2030” [228] and the annotation industry is widely reported to have little concern for workers’ rights. Besides welfare and rights, the invisibility of this contribution arguably contributes to a

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  15. 61.02.25 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Exploitation in AI development

    "Outsourcing tasks like data labeling to low-income countries can perpetuate inequality."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  16. 65.23.07 · Risk Sub-Category

    Non-technical risks (Societal impact)

    Human exploitation

    "When workers who train AI models such as ghost workers are not provided with adequate working conditions, fair compensation, and good health care benefits that also include mental health."

    From AI Risk Atlas (IBM2025)

  17. 66.04.07 · Risk Sub-Category

    Societal and Cultural

    Labor exploitation

    "Use/misuse of labour to help train, develop, manage or optimise a technology system or set of systems, including under-paid and/or offshore"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  18. 05.17.00 · Risk Category

    Copyright - Authorship

    The emergence of generative AI raises issues regarding disruptions to existing copyright norms. Frequently discussed in the literature are violations of copyright and intellectual property rights stemming from the unauthorized collection of text or image training data. Another concern relates to generative models memorizing or plagiarizing copyrighted content. Additionally, there are open questions and debates around the copyright or ownership of model outputs, the protection of creative prompts, and the general blurring of traditional concepts of authorship.

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  19. "The extent and effectiveness of legal protections for intellectual property have been thrown into question with the rise of generative AI. Generative AI trains itself on vast pools of data that often include IP-protected works.

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  20. 47.03.03 · Risk Sub-Category

    Legal challenges

    Copyright challenges (training models using copyrighted output)

    "Generative AI companies are regularly accused of violating copyright law by training AI models on copyrighted works without gaining permission or paying compensation to the copyright owners. In fact, a substantial number of copyrighted documents and books have been incorporated into the training datasets of generative AI models."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  21. "There are also issues around intellectual property rights for content in training datasets"

    From Future Risks of Frontier AI (GOS2023)

  22. "The risks associated with the race to develop the first AGI, including the development of poor quality and unsafe AGI, and heightened political and control issues."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  23. 45.01.13 · Risk Sub-Category

    AI's inherent safety risks

    Risks from AI systems (Risks of supply chain security)

    "The AI industry relies on a highly globalized supply chain. However, certain countries may use unilateral coercive measures, such as technology barriers and export restrictions, to create development obstacles and maliciously disrupt the global AI supply chain. This can lead to significant risks of supply disruptions for chips, software, and tools."

    From AI Safety Governance Framework (TC2602024)

  24. 55.02.04 · Risk Sub-Category

    Worsened conflict

    Resource conflicts driven by AI development

    "AI development may itself become a new flash point for conflicts—causing more conflict to occur— especially conflicts over AI-relevant resources (such as data centres, semiconductor manufacturing facilities and raw materials)."

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  25. 61.02.17 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Dangerous development races

    "Competitive pressures could lead to the neglect of safety measures in AI development."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  26. 62.29.03a · Additional evidence

    Impacts of AI (General)

    Competitive pressures in GPAI product release

  27. "The capabilities of current risk management and legal processes in the context of the development of an AGI."

    From The risks associated with Artificial General Intelligence: A systematic review (McLean2023)

  28. 24.01.02 · Risk Sub-Category

    Capability failures

    Difficult to develop metrics for evaluating benefits or harms caused by AI assistants

    "Another difficulty facing AI assistant systems is that it is challenging to develop metrics for evaluating particular aspects of benefits or harms caused by the assistant – especially in a sufficiently expansive sense, which could involve much of society (see Chapter 19). Having these metrics is useful both for assessing the risk of harm from the system and for using the metric as a training signal."

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  29. 62.16.02 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Limited coverage of capabilities evaluations)

    "GPAI model developers might run capabilities evaluations to determine whether it has dangerous or dual-use capabilities, and then decide whether it is safe to deploy. Such capabilities evaluations can fail to demonstrate all the capabilities of a model. For example, evaluations may miss certain capabilities that are difficult to assess, prohibitively costly to verify, or obscured by the model’s tendency to refuse responses due to safety training, even if it possesses some of these capabilities."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  30. 62.16.06 · Risk Sub-Category

    Model Evaluations

    General Evaluations (Biased evaluations of encoded human values)

    "Encoded human values in AI models that are easier to evaluate might be preferred for inclusion in evaluations over those that are more difficult to measure [13]. This might come at the expense of more desirable but harder-to-quantify values. This bias can lead to an imbalance, where easier-to-measure values dominate the evaluation process, while other important values are underrepresented."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  31. 62.16.08 · Risk Sub-Category

    Model Evaluations

    Benchmarking (Benchmark leakage or data contamination)

    "Benchmark leakage [235, 224, 221, 161] can happen when an AI model is trained or fine-tuned with evaluation-related data. This can lead to an unreliable model evaluation, especially if the data contains question-answer pairs from bench- marks."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  32. 62.16.09 · Risk Sub-Category

    Model Evaluations

    Benchmarking (Raw data contamination)

    "This type of contamination [170] occurs when the raw and unlabeled data of a benchmark is used as part of the training set. Such data may not be properly formatted and may contain noise, especially if the contamination happens before the data is pre-processed into the benchmark. If this contamination occurs, it could cast doubt on the few-shot and zero-shot performance of the model on that benchmark."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  33. 62.16.10 · Risk Sub-Category

    Model Evaluations

    Benchmarking (Cross-lingual data contamination)

    "Models that have been trained on data encoded in multiple languages, such as LLMs trained on web-crawled data, may contain contamination that is obscured by translation [226]. The most basic form of this is when a benchmark is trans- lated to another language and then fed to the model as training data. The fact that the benchmark is translated before becoming training data can obscure the contamination from detection methods, giving false assurance that the model has generalized on the capabilities that the benchmark tests for."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  34. 62.16.11 · Risk Sub-Category

    Model Evaluations

    Benchmarking (Guideline contamination)

    "Guideline contamination refers to scenarios where instructions for the collec- tion, annotation, or use of the dataset are exposed to the model [170]. These instructions may contain explicit data-label pairs that can improve the model’s capabilities for the task."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  35. 62.16.12 · Risk Sub-Category

    Model Evaluations

    Benchmarking (Annotation contamination)

    "Annotation contamination refers to scenarios where the model is exposed to the benchmark labels during training [170]. This type of contamination can make the model learn the acceptable distribution of outputs. Combining this with raw data contamination of the test split, any evaluation made with the benchmark is invalidated because the entire test split is essentially leaked to the model."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  36. 62.16.14 · Risk Sub-Category

    Model Evaluations

    Benchmark Inaccuracy (Benchmarks may not accurately evaluate capabilities)

    "Benchmarks of AI systems can both underestimate and overestimate the capa- bilities of those AI systems. Underestimates can happen if an evaluation is not comprehensive enough, if the benchmark is saturated by existing models, or if the capabilities in question depend on a complicated setup, such as realistic computer programming tasks. Overestimates of capabilities can occur if an AI system is trained or fine-tuned on the contents of the benchmark, leading to overfitting."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  37. 62.16.15 · Risk Sub-Category

    Model Evaluations

    Benchmark Inaccuracy (Benchmark saturation)

    "Benchmark saturation refers to benchmarks reaching their evaluation ceiling. The tendency towards benchmark saturation has been demonstrated in various benchmarks [19]. When benchmarks reach or are close to saturation, they stop being effective measures for new models, as more nuanced capability gains might not be detected."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  38. 62.16.16 · Risk Sub-Category

    Model Evaluations

    Benchmark Limitations (Insufficient benchmarks for AI safety evaluation)

    "Benchmarks dedicated to measuring the performance of AI systems (e.g., on programming or math tasks) are more well-developed than those for assessing safety and harms in AI systems [234]. This gap can lead to AI systems excelling in specific tasks while exhibiting harmful behaviors that go undetected. More safety-related evaluation datasets can help in identifying previously overlooked undesirable model behaviors."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  39. 62.16.17 · Risk Sub-Category

    Model Evaluations

    Benchmark Limitations (Underestimating capabilities that are not covered by benchmarks)

    "A lack of test coverage by benchmarks on specific abilities of a model can obscure the model’s capabilities from both the developer and the user [160]. This can lead to a false sense of safety and trust due to a lack of understanding of the model’s limitations."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  40. 62.17.01 · Risk Sub-Category

    Model Evaluations (Auditing)

    Conflicts of interest in auditor selection

    "Conflicts of interest can arise if there is no independence in the auditor selection process or if the auditors are closely associated with the developer [123, 157]. In such cases, the conflict of interest can appear even if third-party evaluators are involved. In the case of external auditing, the potential candidates might be selected from a narrow group of auditors, or have conflicting financial incentives for whether to report model shortcomings publicly."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  41. 62.17.02 · Risk Sub-Category

    Model Evaluations (Auditing)

    Auditor capacity mismatch

    "Auditors may not be able to address all of the specific safety, performance, or validation needs. Reports of passing audits may be more inclusive than can be justified due to a lack of knowledge of specific risks and how they can be tested, or a lack of capacity to perform sufficiently rigorous testing."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  42. 62.17.03 · Risk Sub-Category

    Model Evaluations (Auditing)

    Auditor failure

    "Auditors may not publicly disclose risks they find, may be required to not pub- licize shortcomings, or may not receive sufficient cooperation from the relevant internal parties."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  43. 65.01.01 · Risk Sub-Category

    Training Data Risks (Transparency)

    Lack of training data transparency

    "Without accurate documentation on how a model's data was collected, curated, and used to train a model, it might be harder to satisfactorily explain the behavior of the model with respect to the data."

    From AI Risk Atlas (IBM2025)

  44. 65.01.02 · Risk Sub-Category

    Training Data Risks (Transparency)

    Uncertain data provenance

    "Data provenance refers to tracing history of data, which includes its ownership, origin, and transformations. Without standardized and established methods for verifying where the data came from, there are no guarantees that the data is the same as the original source and has the correct usage terms."

    From AI Risk Atlas (IBM2025)

  45. 65.22.02 · Risk Sub-Category

    Non-technical risks (Governance)

    Unrepresentative risk testing

    "Testing is unrepresentative when the test inputs are mismatched with the inputs that are expected during deployment."

    From AI Risk Atlas (IBM2025)

  46. 65.22.03 · Risk Sub-Category

    Non-technical risks (Governance)

    Incomplete usage definition

    "Since foundation models can be used for many purposes, a model’s intended use is important for defining the relevant risks of that model. As the use changes, the relevant risks might correspondingly change."

    From AI Risk Atlas (IBM2025)

  47. 65.22.04 · Risk Sub-Category

    Non-technical risks (Governance)

    Lack of data transparency

    "Lack of data transparency is due to insufficient documentation of training or tuning dataset details. "

    From AI Risk Atlas (IBM2025)

  48. 65.22.07 · Risk Sub-Category

    Non-technical risks (Governance)

    Lack of testing diversity

    "AI model risks are socio-technical, so their testing needs input from a broad set of disciplines and diverse testing practices."

    From AI Risk Atlas (IBM2025)

  49. 39.02.00 · Risk Category

    Energy Consumption

    Some learning algorithms, including deep learning, utilize iterative learning processes [23]. This approach results in high energy consumption.

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.