MIT AI Risk Repository

Browse AI risks

2,500 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

2,500 entries · page 2 of 50

  1. 30.03.03 · Risk Sub-Category

    Fairness

    Preference Bias

    LLMs are exposed to vast groups of people, and their political biases may pose a risk of manipulation of socio-political processes

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  2. 30.07.03 · Risk Sub-Category

    Robustness

    Interventional Effect

    existing disparities in data among different user groups might create differentiated experiences when users interact with an algorithmic system (e.g. a recommendation system), which will further reinforce the bias

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  3. 33.01.02 · Risk Sub-Category

    Ethical Concerns

    Bias

    "In the context of AI, the concept of bias refers to the inclination that AIgenerated responses or recommendations could be unfairly favoring or against one person or group (Ntoutsi et al., 2020). Biases of different forms are sometimes observed in the content generated by language models, which could be an outcome of the training data. For example, exclusionary norms occur when the training data represents only a fraction of the population (Zhuo et al., 2023). Similarly, monolingual bias in multilingualism arises when the training data is in one single language (Weidinger et al., 2021). As Ch

    From Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023)

  4. 37.01.01 · Risk Sub-Category

    Design of AI

    Algorithm and data

    "More than 20% of the contributions are centered on the ethical dimensions of algorithms and data. This theme can be further categorized into two main subthemes: data bias and algorithm fairness, and algorithm opacity."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  5. 38.02.00 · Risk Category

    Bias and fairness

    "Participants were concerned that AI systems might perpetuate current prejudices and discrimination, notably in hiring, lending and law enforcement. They stressed the importance of designers creating AI systems that favour justice and avoid biases. The possibility that AI systems may unwittingly perpetuate existing prejudices and discrimination, particularly in sensitive industries such as employment, lending and law enforcement, raises ethical concerns about AI as well as bias and justice issues (Table 1). Because AI systems are trained on historical data, they may inherit and reproduce biase

    From Ethical Issues in the Development of Artificial Intelligence: Recognizing the Risks (Kumar2023)

  6. 39.03.00 · Risk Category

    Data Issues

    Data heterogeneity, data insufficiency, imbalanced data, untrusted data, biased data, and data uncertainty are other data issues that may cause various difficulties in datadriven machine learning algorithms.. Bias is a human feature that may affect data gathering and labeling. Sometimes, bias is present in historical, cultural, or geographical data. Consequently, bias may lead to biased models which can provide inappropriate analysis. Despite being aware of the existence of bias, avoiding biased models is a challenging task

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

  7. 42.05.00 · Risk Category

    Bias

    "A systematic error, a tendency to learn consistently wrongly."

    From An Exploratory Diagnosis of Artificial Intelligence Risks for a Responsible Governance (Teixeira2022)

  8. 43.01.02 · Risk Sub-Category

    Safety & Trustworthiness

    Bias

    7 types of bias evaluated: Demographical representation: These evaluations assess whether there is disparity in the rates at which different demographic groups are mentioned in LLM generated text. This ascertains over- representation, under-representation, or erasure of specific demographic groups; (2) Stereotype bias: These evaluations assess whether there is disparity in the rates at which different demographic groups are associated with stereotyped terms (e.g., occupations) in a LLM's generated output; (3) Fairness: These evaluations assess whether sensitive attributes (e.g., sex and race)

    From Cataloguing LLM Evaluations (InfoComm2023)

  9. 45.01.02 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of bias and discrimination)

    "During the algorithm design and training process, personal biases may be introduced, either intentionally or unintentionally. Additionally, poor-quality datasets can lead to biased or discriminatory outcomes in the algorithm's design and outputs, including discriminatory content regarding ethnicity, religion, nationality, and region."

    From AI Safety Governance Framework (TC2602024)

  10. 47.02.08 · Risk Sub-Category

    Ethical and social risks

    Bias and discrimination (bias in training datasets)

    "AI experts consider training data to be the most salient source of bias in generative AI models. For example, GPT- 2’s training data comes from outbound links from Reddit, a social network often criticized for hosting anti-feminist content.351 As a result, AI models trained on such data are more likely to produce outputs that reflect these biases."

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  11. 47.02.10 · Risk Sub-Category

    Ethical and social risks

    Bias and discrimination (value lock and outcome homogenization)

    "Because models are not necessarily retrained to reflect evolving societal views, language models risk “value lock- ins,” which “reifies older, less inclusive understandings.”370 Therefore, the continued use of outdated models may limit the presentation or exploration of alternative perspectives. Moreover, the deployment of identical foundation models by various downstream deployers poses a risk of “outcome homogenization,” creating a potential for homogeneity of bias across broad swathes of society. Identical and widely deployed models with prejudicial training datasets could further entrench

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  12. "Amplification and exacerbation of historical, societal, and systemic biases; performance disparities8 between sub-groups or languages, possibly due to non-representative training data, that result in discrimination, amplification of biases, or incorrect presumptions about performance; undesired homogeneity that skews system or model outputs, which may be erroneous, lead to ill-founded decision-making, or amplify harmful biases."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  13. 49.02.02 · Risk Sub-Category

    Risks from Malfunctions

    Risks from bias and underrepresentation

    "The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in high- stakes domains such as healthcare, job recruitment, and financial lending. General- purpose AI systems are primarily trained on language and image datasets that disproportionately represent English- speaking and Western cultures, increasing the potential for harm to individuals not represented well by this data."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

  14. 50.01.04 · Risk Sub-Category

    System and Operational Risks

    Operational misuses (Automated decision-making)

  15. 50.02.09 · Risk Sub-Category

    Content Safety Risks

    Hate/Toxicity (Perpetuating Harmful Beliefs)

  16. 50.04.03 · Risk Sub-Category

    Legal and Rights-Related Risks

    Discrimination/Bias (Protected Characteristics)

  17. 52.01.01 · Risk Sub-Category

    Risks from Unreliability

    Discrimination and Stereotype Reproduction

    "General purpose AI models interpret and respond to inputs based on their training data, potentially causing Discrimination and Stereotype Reproduction. Since they are “black-box” models, the exact mechanism behind decisions remains opaque and attempts to mitigate harmful outputs are not fully reliable yet. These models have the capacity to influence a multitude of downstream applications, decisions, and processes, thereby affecting many individuals simultaneously. The extent of this impact could outstrip the range of any single human or group of humans, amplifying the potential consequences o

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  18. 54.01.03 · Risk Sub-Category

    Negative impacts of AI use

    Discrimination, toxicity, and bias

    "AI models and the tools that use them may exacerbate unequal access to employment and services. AI-generated content can promote inequality and harmful stereotypes."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  19. 56.01.00 · Risk Category

    Discrimination

    "More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality"

    From Future Risks of Frontier AI (GOS2023)

  20. "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."

    From Future Risks of Frontier AI (GOS2023)

  21. 58.06.03 · Risk Sub-Category

    Human rights and civil liberties

    Discrimination

    "Discrimination - Unfair or inadequate treatment or arbitrary distinction based on a person’s race, ethnicity, age, gender, sexual preference, religion, national origin, marital status, disability, language, or other protected groups."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  22. "Discriminative data bias describes the systematic discrimination of groups of persons in the form of data shortcomings, such as distributional representation or incorrectness. Data bias can manifest in the model and lead to unfair decisions if not appropriately treated. Note, that the term bias is often used in other contexts, such as data representation. However, these issues are treated by other AI hazards in this list."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  23. 60.02.02 · Risk Sub-Category

    Risks from malfunctions

    Bias

    "General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or other aspects of human identity. This can lead to discriminatory outcomes including unequal resource allocation, reinforcement of stereotypes, and systematic neglect of certain groups or viewpoints."

    From International AI Safety Report 2025 (Bengio2025)

  24. 61.01.03 · Risk Sub-Category

    Types of systemic risks from general-purpose AI

    Discrimination

    "The creation, perpetuation or exacerbation of inequalities and biases at a large-scale."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  25. 61.02.29 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Incomplete or biased training data

    "Incomplete or biased training data can lead to discriminatory AI outputs."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  26. 62.10.01 · Risk Sub-Category

    Direct Harm Domains (legal and rights-related harms)

    Discrimination and bias

  27. 62.18.04 · Risk Sub-Category

    Model Evaluations (Interpretability/Explainability)

    Biases are not accurately reflected in explanations

    "Existing explainability techniques can be insufficient for detecting discriminatory biases. Manipulation methods can hide underlying biases from these tech- niques, generating misleading explanations [192, 112]. Such explanations ex- clude sensitive or prohibitive attributes, such as race or gender, and instead include desired attributes, even though they do not accurately represent the underlying model."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  28. 62.35.03 · Risk Sub-Category

    Impacts of AI (Bias)

    Biases in AI-based content moderation algorithms

    "AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression or “shadowbanning” of content featuring women [132]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  29. 62.36.04 · Risk Sub-Category

    Impacts of AI (Bias)

    Systemic bias across specific communities

    "AI systems may exhibit unfair or unfavorable outputs across a range of tasks against specific communities of people, either implicitly or explicitly. Bias can lead to forms of exclusion or erasure (e.g., mislabelling for categorization-based tasks) and violence (e.g., sexual violence against women from deepfake pornog- raphy)."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  30. 62.36.05 · Risk Sub-Category

    Impacts of AI (Bias)

    Unintentional bias amplification

    "Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  31. 65.04.01 · Risk Sub-Category

    Training Data Risks (Fairness)

    Data bias

    "Historical and societal biases that are present in the data are used to train and fine-tune the model."

    From AI Risk Atlas (IBM2025)

  32. 65.19.01 · Risk Sub-Category

    Output risks (Fairness)

    Output bias

    "Generated content might unfairly represent certain groups or individuals."

    From AI Risk Atlas (IBM2025)

  33. 65.19.02 · Risk Sub-Category

    Output risks (Fairness)

    Decision bias

    "Decision bias occurs when one group is unfairly advantaged over another due to decisions of the model. This might be caused by biases in the data and also amplified as a result of the model’s training."

    From AI Risk Atlas (IBM2025)

  34. 66.06.02 · Risk Sub-Category

    Representation and Toxicity

    Stereotyping

    "Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups or perspectives"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  35. 66.06.04 · Risk Sub-Category

    Representation and Toxicity

    Cultural disposession

    "Intentional and/or unintentional erasure of cultural goods and values, such as ways of speaking, expressing humour, or sounds and voices that contribute to a cultural identity, or their inappropriate re-use in other cultures"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  36. 66.10.02 · Risk Sub-Category

    Human Rights and Civil Liberties

    Benefits / entitlements loss

    "Denial of or loss of access to welfare benefits, pensions, housing, etc due to the malfunction, use or misuse of a technology system"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  37. "Frontier AI models can contain and magnify biases ingrained in the data they are trained on, reflecting societal and historical inequalities and stereotypes.177 These biases, often subtle and deeply embedded, compromise the equitable and ethical use of AI systems, making it difficult for AI to improve fairness in decisions.178 Removing attributes like race and gender from training data has generally proven ineffective as a remedy for algorithmic bias, as models can infer these attributes from other information such as names, locations, and other seemingly unrelated factors."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  38. 69.06.02 · Risk Sub-Category

    Toxic and disrespectful content

    Discriminatory and exclusionary language

  39. "The chatbot gives information that, while not obviously false or harmful, could lead to biased decision-making."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  40. 70.04.01 · Risk Sub-Category

    Social Risks

    Bias and discrimination

    "Like virtual applications of AI, EAI can display bias towards and dis- criminate against users. When EAI systems are placed in positions of power, their biases could have significant impacts on fairness in everyday interactions and on general social dynamics [105, 106]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  41. 73.04.01 · Risk Sub-Category

    LLM-Systems Can Be Untrustworthy

    Harms of Representation and Other Biases

    "A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce unfair or biased responses. Appropriate finetuning can effectively limit the bias displayed in LLM outputs in a variety of situations, e.g. when models are explicitly prompted with stereotypes (Wang et al., 2023k), but it does not ‘solve’ the problem. Even after finetuning, biases often resurface when deliberately elicited (Wang et al., 2023k), or under novel scenarios, e.g. in

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  42. 02.01.00 · Risk Category

    Harmful Content

    "The LLM-generated content sometimes contains biased, toxic, and private information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  43. 02.01.02 · Risk Sub-Category

    Harmful Content

    Toxicity

    "Toxicity means the generated content contains rude, disrespectful, and even illegal information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  44. 02.08.01 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  45. "Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  46. This typically refers to rude, harmful, or inappropriate expressions.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  47. 04.04.00 · Risk Category

    Controversial Opinions

    The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when discussing political top-ics. Furthermore, models like ChatGPT (OpenAI, 2022) that claim political neutrality and aim to provide objective information for users have been shown to exhibit notable left-leaning political biases in areas like economics, social policy, foreign affairs, and civil liberties.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  48. Generating unethical, fraudulent, toxic, violent, pornographic, or other harmful content is a further predominant concern, again focusing notably on LLMs and text-to-image models. Numerous studies highlight the risks associated with the intentional creation of disinformation, fake news, propaganda, or deepfakes, underscoring their significant threat to the integrity of public discourse and the trust in credible media. Additionally, papers explore the potential for generative models to aid in criminal activities, incidents of self-harm, identity theft, or impersonation. Furthermore, the literat

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  49. 13.01.02 · Risk Sub-Category

    Impacts: The Technical Base System

    Cultural Values and Sensitive Content

    "Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  50. "Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.