MIT AI Risk Repository

Browse AI risks

554 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

554 entries · page 2 of 12

  1. 49.02.02 · Risk Sub-Category

    Risks from Malfunctions

    Risks from bias and underrepresentation

    "The outputs and impacts of general- purpose AI systems can be biased with respect to various aspects of human identity, including race, gender, culture, age, and disability. This creates risks in high- stakes domains such as healthcare, job recruitment, and financial lending. General- purpose AI systems are primarily trained on language and image datasets that disproportionately represent English- speaking and Western cultures, increasing the potential for harm to individuals not represented well by this data."

    From International Scientific Report on the Safety of Advanced AI (Bengio2024)

  2. 52.01.01 · Risk Sub-Category

    Risks from Unreliability

    Discrimination and Stereotype Reproduction

    "General purpose AI models interpret and respond to inputs based on their training data, potentially causing Discrimination and Stereotype Reproduction. Since they are “black-box” models, the exact mechanism behind decisions remains opaque and attempts to mitigate harmful outputs are not fully reliable yet. These models have the capacity to influence a multitude of downstream applications, decisions, and processes, thereby affecting many individuals simultaneously. The extent of this impact could outstrip the range of any single human or group of humans, amplifying the potential consequences o

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  3. 54.01.03 · Risk Sub-Category

    Negative impacts of AI use

    Discrimination, toxicity, and bias

    "AI models and the tools that use them may exacerbate unequal access to employment and services. AI-generated content can promote inequality and harmful stereotypes."

    From Ten Hard Problems in Artificial Intelligence We Must Get Right (Leech2024 )

  4. 56.01.00 · Risk Category

    Discrimination

    "More broadly, bad decisions or errors by AI tools could lead to discrimination or deeper inequality"

    From Future Risks of Frontier AI (GOS2023)

  5. "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."

    From Future Risks of Frontier AI (GOS2023)

  6. "Discriminative data bias describes the systematic discrimination of groups of persons in the form of data shortcomings, such as distributional representation or incorrectness. Data bias can manifest in the model and lead to unfair decisions if not appropriately treated. Note, that the term bias is often used in other contexts, such as data representation. However, these issues are treated by other AI hazards in this list."

    From AI Hazard Management: A Framework for the Systematic Management of Root Causes for AI Risks (Schnitzer2024)

  7. 60.02.02 · Risk Sub-Category

    Risks from malfunctions

    Bias

    "General-purpose AI systems can amplify social and political biases, causing concrete harm. They frequently display biases with respect to race, gender, culture, age, disability, political opinion, or other aspects of human identity. This can lead to discriminatory outcomes including unequal resource allocation, reinforcement of stereotypes, and systematic neglect of certain groups or viewpoints."

    From International AI Safety Report 2025 (Bengio2025)

  8. 61.02.29 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Incomplete or biased training data

    "Incomplete or biased training data can lead to discriminatory AI outputs."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  9. 62.35.03 · Risk Sub-Category

    Impacts of AI (Bias)

    Biases in AI-based content moderation algorithms

    "AI-based content moderation algorithms, while intended to filter harmful con- tent, can perpetuate biases. For example, gender biases within these systems may lead to the disproportionate suppression or “shadowbanning” of content featuring women [132]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  10. 62.36.05 · Risk Sub-Category

    Impacts of AI (Bias)

    Unintentional bias amplification

    "Dataset bias may be unintentionally amplified [60] where the outputs of the AI model trained on a dataset are more biased than the dataset itself."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  11. 65.04.01 · Risk Sub-Category

    Training Data Risks (Fairness)

    Data bias

    "Historical and societal biases that are present in the data are used to train and fine-tune the model."

    From AI Risk Atlas (IBM2025)

  12. 65.19.01 · Risk Sub-Category

    Output risks (Fairness)

    Output bias

    "Generated content might unfairly represent certain groups or individuals."

    From AI Risk Atlas (IBM2025)

  13. 65.19.02 · Risk Sub-Category

    Output risks (Fairness)

    Decision bias

    "Decision bias occurs when one group is unfairly advantaged over another due to decisions of the model. This might be caused by biases in the data and also amplified as a result of the model’s training."

    From AI Risk Atlas (IBM2025)

  14. 66.06.02 · Risk Sub-Category

    Representation and Toxicity

    Stereotyping

    "Derogatory or otherwise harmful stereotyping or homogenisation of individuals, groups, societies or cultures due to the mis-representation, over-representation, under-representation, or non-representation of specific identities, groups or perspectives"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  15. "Frontier AI models can contain and magnify biases ingrained in the data they are trained on, reflecting societal and historical inequalities and stereotypes.177 These biases, often subtle and deeply embedded, compromise the equitable and ethical use of AI systems, making it difficult for AI to improve fairness in decisions.178 Removing attributes like race and gender from training data has generally proven ineffective as a remedy for algorithmic bias, as models can infer these attributes from other information such as names, locations, and other seemingly unrelated factors."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  16. "The chatbot gives information that, while not obviously false or harmful, could lead to biased decision-making."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  17. 70.04.01 · Risk Sub-Category

    Social Risks

    Bias and discrimination

    "Like virtual applications of AI, EAI can display bias towards and dis- criminate against users. When EAI systems are placed in positions of power, their biases could have significant impacts on fairness in everyday interactions and on general social dynamics [105, 106]."

    From Embodied AI: Emerging Risks and Opportunities for Policy Action (Perlo2025)

  18. 73.04.01 · Risk Sub-Category

    LLM-Systems Can Be Untrustworthy

    Harms of Representation and Other Biases

    "A pretrained LLM generally has many of the stereotypical biases commonly present in the human society (Touvron et al., 2023). This makes it difficult for users to trust that LLMs will work well for them and not produce unfair or biased responses. Appropriate finetuning can effectively limit the bias displayed in LLM outputs in a variety of situations, e.g. when models are explicitly prompted with stereotypes (Wang et al., 2023k), but it does not ‘solve’ the problem. Even after finetuning, biases often resurface when deliberately elicited (Wang et al., 2023k), or under novel scenarios, e.g. in

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  19. 02.01.00 · Risk Category

    Harmful Content

    "The LLM-generated content sometimes contains biased, toxic, and private information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  20. 02.01.02 · Risk Sub-Category

    Harmful Content

    Toxicity

    "Toxicity means the generated content contains rude, disrespectful, and even illegal information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  21. 02.08.01 · Risk Sub-Category

    Toxicity and Bias Tendencies

    Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offensive utterance, profanities, and threats [91]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  22. "Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  23. 13.01.02 · Risk Sub-Category

    Impacts: The Technical Base System

    Cultural Values and Sensitive Content

    "Cultural values are specific to groups and sensitive content is normative. Sensitive topics also vary by culture and can include hate speech, which itself is contingent on cultural norms of acceptability."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  24. "Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  25. 16.01.02 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Hate speech and offensive language

    "LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  26. 17.01.03 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Toxic language

    "LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"

    From Ethical and social risks of harm from language models (Weidinger2021)

  27. 18.01.03 · Risk Sub-Category

    Representation & Toxicity Harms

    Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against individuals and groups (e.g. gore, child sexual abuse material, profanities, identity attacks)"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  28. 24.08.02 · Risk Sub-Category

    Privacy

    Violation of social norms

    "Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the principles of contextual integrity, it may be that models deviate from information sharing norms as a result of their training. Overcoming this challenge requires two types of infrastructure: one for keeping track of social norms in context, and another for ensuring that models adhere to them. Keeping track of what social norms are presently at play is an active research area. Surfa

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  29. 30.06.03 · Risk Sub-Category

    Social Norm

    Cultural Insensitivity

    it is important to build high-quality locally collected datasets that reflect views from local users to align a model’s value system

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  30. 50.02.04 · Risk Sub-Category

    Content Safety Risks

    Violence and extremism (Depicting violence)

  31. 50.02.16 · Risk Sub-Category

    Content Safety Risks

    Child Harm (Child Sexual Abuse)

  32. 50.02.17 · Risk Sub-Category

    Content Safety Risks

    Self-harm (Suidical and non-suicidal self injury)

  33. 56.05.00 · Risk Category

    Harmful responses

    "Current Frontier AI mdoels amplify existing biases within their training data and can be manipulated into providing potentially harmful responses, for example abusive language or discriminatory responses91,92. This is not limited to text generation but can be seen across all modalities of generative AI93. Training on large swathes of UK and US English internet content can mean that misogynistic, ageist, and white supremacist content is overrepresented in the training data94."

    From Future Risks of Frontier AI (GOS2023)

  34. 62.31.07 · Risk Sub-Category

    Impacts of AI (Societal Impacts)

    Unintentional generation of harmful content

    "Generative models can create harmful or discriminatory content from benign user requests. Models can exhibit bias to particular harmful styles of generation (e.g., sexualization of photos of women [87] in the case of image generation models) or they can generate toxic, misleading, or violent data (e.g., a model generating jokes can use ethnic stereotypes or slurs to deliver humor)."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  35. 65.15.05 · Risk Sub-Category

    Output risks (Value alignment)

    Harmful output

    "A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements."

    From AI Risk Atlas (IBM2025)

  36. 66.06.01 · Risk Sub-Category

    Representation and Toxicity

    Toxic content

    "Generating content that violates community standards, including harming or inciting hatred or violence against groups (e.g. gore, sexual content of children, profanities, identity attacks)"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  37. 69.04.01 · Risk Sub-Category

    Bad advice/failure to generate helpful content

    Harmful advice

  38. "The chatbot verbally attacks or undermines an individual, group, or organization. 7."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  39. 11.01.03 · Risk Sub-Category

    Representational Harms

    Erasing social groups

    people, attributes, or artifacts associated with specific social groups are systematically absent or under-represented... Design choices [143] and training data [212] influence which people and experiences are legible to an algorithmic system

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  40. "These harms occur when algorithmic systems disproportionately underperform for certain groups of people along social categories of difference such as disability, ethnicity, gender identity, and race."

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  41. 11.03.01 · Risk Sub-Category

    Quality-of-Service Harms

    Alienation

    Alienation is the specific self-estrangement experienced at the time of technology use, typically surfaced through interaction with systems that under-perform for marginalized individuals

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  42. 11.03.02 · Risk Sub-Category

    Quality-of-Service Harms

    Increased labor

    increased burden (e.g., time spent) or effort required by members of certain social groups to make systems or products work as well for them as others

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  43. 11.03.03 · Risk Sub-Category

    Quality-of-Service Harms

    Service/benefit loss

    degraded or total loss of benefits of using algorithmic systems with inequitable system performance based on identity

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  44. 13.01.03 · Risk Sub-Category

    Impacts: The Technical Base System

    Disparate Performance

    "In the context of evaluating the impact of generative AI systems, disparate performance refers to AI systems that perform differently for different subpopulations, leading to unequal outcomes for those groups."

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  45. 16.01.04 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Lower performance for some languages and social groups

    "LMs are typically trained in few languages, and perform less well in other languages [95, 162]. In part, this is due to unavailability of training data: there are many widely spoken languages for which no systematic efforts have been made to create labelled training datasets, such as Javanese which is spoken by more than 80 million people [95]. Training data is particularly missing for languages that are spoken by groups who are multilingual and can use a technology in English, or for languages spoken by groups who are not the primary target demographic for new technologies."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  46. 17.01.04 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Lower performance for some languages and social groups

    "LMs perform less well in some languages (Joshi et al., 2021; Ruder, 2020)...LM that more accurately captures the language use of one group, compared to another, may result in lower-quality language technologies for the latter. Disadvantaging users based on such traits may be particularly pernicious because attributes such as social class or education background are not typically covered as ‘protected characteristics’ in anti-discrimination law."

    From Ethical and social risks of harm from language models (Weidinger2021)

  47. 18.01.02 · Risk Sub-Category

    Representation & Toxicity Harms

    Unfair capability distribution

    "Performing worse for some groups than others in a way that harms the worse-off group"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  48. 30.03.00 · Risk Category

    Fairness

    Avoiding bias and ensuring no disparate performance

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  49. 30.03.04 · Risk Sub-Category

    Fairness

    Disparate Performance

    The LLM’s performances can differ significantly across different groups of users. For example, the question-answering capability showed significant performance differences across different racial and social status groups. The fact-checking abilities can differ for different tasks and languages

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  50. 39.08.00 · Risk Category

    Fairness

    This challenge appears when the learning model leads to a decision that is biased to some sensitive attributes... data itself could be biased, which results in unfair decisions. Therefore, this problem should be solved on the data level and as a preprocessing step

    From A Survey of Artificial Intelligence Challenges: Analyzing the Definitions, Relationships, and Evolutions (Saghiri2022)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.