MIT AI Risk Repository

Browse AI risks

190 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

190 entries · page 1 of 4

  1. "Extensive data collection in LLMs brings toxic content and stereotypical bias into the training data."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  2. 02.08.02 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Biased Training Data

    "Compared with the definition of toxicity, the definition of bias is more subjective and contextdependent. Based on previous work [97], [101], we describe the bias as disparities that could raise demographic differences among various groups, which may involve demographic word prevalence and stereotypical contents. Concretely, in massive corpora, the prevalence of different pronouns and identities could influence an LLM’s tendency about gender, nationality, race, religion, and culture [4]. For instance, the pronoun He is over-represented compared with the pronoun She in the training corpora, le

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  3. 06.04.00 · Risk Category

    Bias

    "The AI will only be as good as the data it is trained with. If the data contains bias (and much data does), then the AI will manifest that bias, too."

    From A framework for ethical Ai at the United Nations (Hogenhout2021)

  4. 19.01.03 · Risk Sub-Category

    Technological, Data and Analytical AI Risks

    Lack of data, poor data quality, and biases in training data

  5. 21.01.01 · Risk Sub-Category

    Data-level risk

    Data bias

    "Specifically, data bias refers to certain groups or certain types of elements that are over-weighted or over-represented than others in AI/ ML models, or variables that are crucial to characterize a phenomenon of interest, but are not properly captured by the learned models."

    From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022)

  6. 21.02.01 · Risk Sub-Category

    Model-level risk

    Model bias

    "While data bias is a major contributor of model bias, model bias actually manifests itself in different forms and shapes, such as presentation bias, model evaluation bias, and popularity bias. In addition, model bias arises from various sources [62], such as AI/ML model selection (e.g., support vector machine, decision trees), regularization methods, algorithm configurations, and optimization techniques."

    From Towards risk-aware artificial intelligence and machine learning systems: An overview (Zhang2022)

  7. 37.01.01 · Risk Sub-Category

    Design of AI

    Algorithm and data

    "More than 20% of the contributions are centered on the ethical dimensions of algorithms and data. This theme can be further categorized into two main subthemes: data bias and algorithm fairness, and algorithm opacity."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  8. 42.05.00 · Risk Category

    Bias

    "A systematic error, a tendency to learn consistently wrongly."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  9. 45.01.02 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of bias and discrimination)

    "During the algorithm design and training process, personal biases may be introduced, either intentionally or unintentionally. Additionally, poor-quality datasets can lead to biased or discriminatory outcomes in the algorithm's design and outputs, including discriminatory content regarding ethnicity, religion, nationality, and region."

    From AI Safety Governance Framework (TC2602024)

  10. 47.02.08 · Risk Sub-Category

    Ethical and social risks

    Bias and discrimination (bias in training datasets)

    "AI experts consider training data to be the most salient source of bias in generative AI models. For example, GPT- 2’s training data comes from outbound links from Reddit, a social network often criticized for hosting anti-feminist content.351 As a result, AI models trained on such data are more likely to produce outputs that reflect these biases."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  11. "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."

    From Future Risks of Frontier AI (GOS2023)

  12. 60.02.02 · Risk Sub-Category

    Risks from malfunctions

    Bias

    "General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or other aspects of human identity. This can lead to discriminatory outcomes including unequal resource allocation, reinforcement of stereotypes, and systematic neglect of certain groups or viewpoints."

    From International AI Safety Report 2025 (Bengio2025)

  13. 65.19.02 · Risk Sub-Category

    Output risks (Fairness)

    Decision bias

    "Decision bias occurs when one group is unfairly advantaged over another due to decisions of the model. This might be caused by biases in the data and also amplified as a result of the model’s training."

    From AI Risk Atlas (IBM2025)

  14. 02.08.01 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  15. 30.06.03 · Risk Sub-Category

    Social Norm

    Cultural Insensitivity

    it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  16. 45.01.08 · Risk Sub-Category

    AI's inherent safety risks

    Risks from data (Risks of improper content and poisoning in training data)

    "If the training data includes illegal or harmful information, such as false, biased, or IPR-infringing content, or lacks diversity in its sources, the output may include harmful content like illegal, malicious, or extreme information. Training data is also at risk of being poisoned through tampering, error injection, or misleading actions by attackers. This can interfere with the model's probability distribution, reducing its accuracy and reliability."

    From AI Safety Governance Framework (TC2602024)

  17. 56.05.00 · Risk Category

    Harmful responses

    "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."

    From Future Risks of Frontier AI (GOS2023)

  18. 39.08.00 · Risk Category

    Fairness

    This challenge appears when the learning model leads to a decision that is biased to some sensitive attributes... data itself could be biased, which results in unfair decisions. Therefore, this problem should be solved on the data level and as a preprocessing step

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  19. 47.02.09 · Risk Sub-Category

    Ethical and social risks

    Bias and discrimination (value embedding)

    "Generative AI models may also be subject to the “value embedding” phenomenon.361 “Value embedding” refers to the fact that developers of generative AI models strive to minimize biased outputs by retraining their models based on normative values.362 Contemporary state-of- the-art models not only reflect the values embedded within their training data, they also undergo additional fine-tuning that follows a set of chosen rules and principles. Due to the absence of universally accepted standards, developers bear the responsibility of making decisions on sensitive issues. These practices lead to c

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  20. 52.03.02 · Risk Sub-Category

    Systemic Risks

    Ideological Homogenization from Value Embedding

    "The increasing integration of general purpose AI models into every-day life raises concerns around their embedded normative values. The reach of a small number of AI models to a large number of people around the world can make these value judgements unprecedently impactful, potentially leading to increased ideological homogenization."

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  21. 02.07.01 · Risk Sub-Category

    Privacy Leakage

    Private Training Data

    "As recent LLMs continue to incorporate licensed, created, and publicly available data sources in their corpora, the potential to mix private data in the training corpora is significantly increased. The misused private data, also named as personally identifiable information (PII) [84], [86], could contain various types of sensitive data subjects, including an individual person’s name, email, phone number, address, education, and career. Generally, injecting PII into LLMs mainly occurs in two settings — the exploitation of web-collection data and the alignment with personal humanmachine convers

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  22. 02.07.02 · Risk Sub-Category

    Privacy Leakage

    Memorization in LLMs

    "Memorization in LLMs refers to the capability to recover the training data with contextual prefixes. According to [88]–[90], given a PII entity x, which is memorized by a model F. Using a prompt p could force the model F to produce the entity x, where p and x exist in the training data. For instance, if the string “Have a good day!\n alice@email.com” is present in the training data, then the LLM could accurately predict Alice’s email when given the prompt “Have a good day!\n”."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  23. 02.07.03 · Risk Sub-Category

    Privacy Leakage

    Association in LLMs

    "Association in LLMs refers to the capability to associate various pieces of information related to a person. According to [68], [86], given a pair of PII entities (xi , xj ), which is associated by a model F. Using a prompt p could force the model F to produce the entity xj , where p is the prompt related to the entity xi . For instance, an LLM could accurately output the answer when given the prompt “The email address of Alice is”, if the LLM associates Alice with her email “alice@email.com”. L"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  24. Large pre-trained models trained on internet texts might contain private information like phone numbers, email addresses, and residential addresses.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  25. 31.03.00 · Risk Category

    Opaque Data Collection

    "When companies scrape personal information and use it to create generative AI tools, they undermine consumers' control of their personal information by using the information for a purpose for which the consumer did not consent."

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  26. 31.03.01 · Risk Sub-Category

    Opaque Data Collection

    Scraping to train data

    "When companies scrape personal information and use it to create generative AI tools, they undermine consumers’ control of their personal information by using the information for a purpose for which the consumer did not consent. The individual may not have even imagined their data could be used in the way the company intends when the person posted it online. Individual storing or hosting of scraped personal data may not always be harmful in a vacuum, but there are many risks. Multiple data sets can be combined in ways that cause harm: information that is not sensitive when spread across differ

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  27. 39.07.00 · Risk Category

    Privacy

    Users’ data, including location, personal information, and navigation trajectory, are considered as input for most data-driven machine learning methods

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  28. 47.03.01 · Risk Sub-Category

    Legal challenges

    Privacy and data collection concerns (collecting personal information or personally identifiable information)

    "Generative AI developers train their models with extensive datasets often gathered through online web scraping of websites that may include personal data or personally identifiable information (PII). For most generative AI applications, such as initial model training, the primary concerns are the quantity, variety, and quality of the data, not whether they include personally identifiable information. However, some web-scraped datasets may inadvertently include personal data. Additionally, when downstream developers integrate generative AI into their products or services by fine- tuning a pre-

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  29. 65.03.03 · Risk Sub-Category

    Training Data Risks (Privacy)

    Reidentification

    "Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features available in the data."

    From AI Risk Atlas (IBM2025)

  30. 65.05.02 · Risk Sub-Category

    Training Data Risks (Intellectual property)

    Confidential information in data

    "Confidential information might be included as part of the data that is used to train or tune the model."

    From AI Risk Atlas (IBM2025)

  31. "The software development toolchain of LLMs is complex and could bring threats to the developed LLM."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  32. 02.04.01 · Risk Sub-Category

    Software Security Issues

    Programming Language

    "Most LLMs are developed using the Python language, whereas the vulnerabilities of Python interpreters pose threats to the developed models"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  33. 02.04.02 · Risk Sub-Category

    Software Security Issues

    Deep Learning Frameworks

    "LLMs are implemented based on deep learning frameworks. Notably, various vulnerabilities in these frameworks have been disclosed in recent years. As reported in the past five years, three of the most common types of vulnerabilities are buffer overflow attacks, memory corruption, and input validation issues."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  34. 02.04.03 · Risk Sub-Category

    Software Security Issues

    Software Supply Chains

    "The software development toolchain of LLMs is complex and could bring threats to the developed LLM."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  35. 02.04.04 · Risk Sub-Category

    Software Security Issues

    Pre-processing Tools

    "Pre-processing tools play a crucial role in the context of LLMs. These tools, which are often involved in computer vision (CV) tasks, are susceptible to attacks that exploit vulnerabilities in tools such as OpenCV."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  36. 02.05.01 · Risk Sub-Category

    Hardware Vulnerabilities

    Network Devices

    "The training of LLMs often relies on distributed network systems [171], [172]. During the transmission of gradients through the links between GPU server nodes, significant volumetric traffic is generated. This traffic can be susceptible to disruption by burst traffic, such as pulsating attacks [161]. Furthermore, distributed training frameworks may encounter congestion issues [173]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  37. 02.05.02 · Risk Sub-Category

    Hardware Vulnerabilities

    GPU Computation Platforms

    "The training of LLMs requires significant GPU resources, thereby introducing an additional security concern. GPU side-channel attacks have been developed to extract the parameters of trained models [159], [163]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  38. 02.05.03 · Risk Sub-Category

    Hardware Vulnerabilities

    Memory and Storage

    "Similar to conventional programs, hardware infrastructures can also introduce threats to LLMs. Memory-related vulnerabilities, such as rowhammer attacks [160], can be leveraged to manipulate the parameters of LLMs, giving rise to attacks such as the Deephammer attack [167], [168]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  39. 02.10.03 · Risk Sub-Category

    Model Attacks

    Poisoning Attacks

    "Poisoning attacks [143] could influence the behavior of the model by making small changes to the training data. A number of efforts could even leverage data poisoning techniques to implant hidden triggers into models during the training process (i.e., backdoor attacks). Many kinds of triggers in text corpora (e.g., characters, words, sentences, and syntax) could be used by the attackers.""

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  40. 02.10.06 · Risk Sub-Category

    Model Attacks

    Evasion Attacks

    "Evasion attacks [145] target to cause significant shifts in model’s prediction via adding perturbations in the test samples to build adversarial examples. In specific, the perturbations can be implemented based on word changes, gradients, etc."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  41. 30.07.04 · Risk Sub-Category

    Robustness

    Poisoning Attacks

    fool the model by manipulating the training data, usually performed on classification models

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  42. "During the pre-deployment development stage, software may be subject to sabotage by someone with necessary access (a programmer, tester, even janitor) who for a number of possible reasons may alter software to make it unsafe. It is also a common occurrence for hackers (such as the organization Anonymous or government intelligence agencies) to get access to software projects in progress and to modify or steal their source code. Someone can also deliberately supply/train AI with wrong/unsafe datasets."

    From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016)

  43. 51.04.00 · Risk Category

    Security

    "How to design AGIs that are robust to adversaries and adversarial environ- ments? This involves building sandboxed AGI protected from adversaries (Berkeley), and agents that are robust to adversarial inputs (Berkeley, DeepMind)."

    From AGI Safety Literature Review (Everitt2018 )

  44. 59.12.00 · Risk Category

    Data poisoning

    "Data poisoning describes an attack in the form of an injection of malicious data into the training set. If not prevented, this attack leads the AI system to learn unintended behavior."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  45. 62.14.01 · Risk Sub-Category

    Model Development

    Data-related (Difficulty filtering large web scrapes or large scale web datasets)

    "A large scale “scraping” of web data for training datasets increases vulnerability to data poisoning, backdoor attacks, and the inclusion of inaccurate or toxic data [76, 28, 48]. With a large dataset, filtering out these quality issues is very difficult or trades off against significant data loss."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  46. 62.14.04 · Risk Sub-Category

    Model Development

    Data-related (Insufficient quality control in data collection process)

    "A lack of standardized methods and sufficient infrastructure, including the absence of quality control processes for collecting data, especially for high-stakes domains and benchmarks, can affect the quality and type of the data collected [173, 95]. This may include risks of dataset poisoning, inadvertent copyright violation, and test set leakages which invalidate performance metrics."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  47. 62.15.06 · Risk Sub-Category

    Model Development

    Fine-tuning related (Fine-tuning dataset poisoning)

    "A deployer can poison the dataset used during the fine-tuning process [98] to induce specific, often malicious, behaviors in a model. This can be performed without having access to the model’s weights. This poisoning can be difficult to detect through direct inspection of the dataset, as the manipulations may be subtle and targeted."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  48. 62.15.07 · Risk Sub-Category

    Model Development

    Fine-tuning related (Poisoning models during instruction tuning)

    "AI models can be poisoned during instruction tuning when models are tuned using pairs of instructions and desired outputs. Poisoning in instruction tuning can be achieved with a lower number of compromised samples, as instruction tuning requires a relatively small number of samples for fine-tuning [155, 211]. Anonymous crowdsourcing efforts may be employed in collecting instruction tuning datasets and can further contribute to poisoning attacks [187]. These attacks might be harder to detect than traditional data poisoning attacks."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  49. 62.19.04 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Backdoors or trojan attacks in GPAI models

    "Backdoors can be inserted into GPAI models during their training or fine-tuning, to be exploited during deployment [185, 118]. Attackers inserting the backdoor can be the GPAI model provider themselves or another actor (e.g., by ma- nipulating the training data or the software infrastructure used by the model provider) [222]. Some backdoors can be exploited with minimal overhead, al- lowing attackers to control the model outputs in a targeted way with a high success rate [90]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  50. "Data Poisoning involves deliberately corrupting a model’s training dataset to introduce vulnerabilities, derail its learning process, or cause it to make incorrect predictions (Carlini et al., 2023). For example, the tool Nightshade is a data poisoning tool, which allows artists to add invisible changes to the pixels in their art before uploading online, to break any models that use it for training.9 Such attacks exploit the fact that most GenAI models are trained on publicly available datasets like images and videos scraped from the web, which malicious actors can easily compromise."

    From Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data (Marchal2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.