MIT AI Risk Repository

Browse AI risks

2,500 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

2,500 entries · page 33 of 50

  1. 73.02.03 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Collusion between LLM-Agents

    "While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  2. 74.01.00 · Risk Category

    Inherent Risk

    "In terms of inherent risk, LLMs could potentially reveal sensitive information from their utilized corpora for pre-training or fine-tuning, thereby raising issues of privacy leakage [37, 145, 226]. Meanwhile, it is well-known that LLMs may experi- ence hallucinations, resulting in the production of texts that are inaccurate and misleading [194]. Finally, since the values embedded in LLM-generated texts usually directly reflect the distribution of their training data, often sourced from the Internet, there exists a substantial risk that LLMs will overfit to a narrow set of human values or even

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

  3. 01.02.00.a · Additional evidence

    Type 2: Bigger than expected

  4. 01.03.00.a · Additional evidence

    Type 3: Worse than expected

  5. 01.03.00.b · Additional evidence

    Type 3: Worse than expected

  6. 01.03.00.c · Additional evidence

    Type 3: Worse than expected

  7. 01.03.00.d · Additional evidence

    Type 3: Worse than expected

  8. 01.03.00.e · Additional evidence

    Type 3: Worse than expected

  9. 01.03.00.f · Additional evidence

    Type 3: Worse than expected

  10. 01.04.00.a · Additional evidence

    Type 4: Willful indifference

  11. 01.04.00.b · Additional evidence

    Type 4: Willful indifference

  12. 02.11.01 · Risk Sub-Category

    Not-Suitable-for-Work (NSFW) Prompts

    Insults

  13. 02.11.02 · Risk Sub-Category

    Not-Suitable-for-Work (NSFW) Prompts

    Crimes

  14. 02.11.03 · Risk Sub-Category

    Not-Suitable-for-Work (NSFW) Prompts

    Sensitive Politics

  15. 02.11.04 · Risk Sub-Category

    Not-Suitable-for-Work (NSFW) Prompts

    Physical Harm

  16. 02.11.05 · Risk Sub-Category

    Not-Suitable-for-Work (NSFW) Prompts

    Mental Health

  17. 02.11.06 · Risk Sub-Category

    Not-Suitable-for-Work (NSFW) Prompts

    Unfairness

  18. 05.14.00 · Risk Category

    Evaluation - Auditing

    Closely related to other clusters like AI safety, fairness, or harmful content, papers stress the importance of evaluating generative AI systems both in a narrow technical way as well as in a broader sociotechnical impact assessment focusing on pre-release audits as well as post-deployment monitoring. Ideally, these evaluations should be conducted by independent third parties. In terms of technical LLM or text-to-image model audits, papers furthermore criticize a lack of safety benchmarking for languages other than English.

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  19. 05.19.00 · Risk Category

    Miscellaneous

    While the scoping review identified distinct topic clusters within the literature, it also revealed certain issues that either do not fit into these categories, are discussed infrequently, or in a nonspecific manner. For instance, some papers touch upon concepts like trustworthiness, accountability, or responsibility, but often remain vague about what they entail in detail. Similarly, a few papers vaguely attribute socio-political instability or polarization to generative AI without delving into specifics. Apart from that, another minor topic area concerns responsible approaches of talking abo

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  20. 10.07.00 · Risk Category

    Injustice

  21. "What can be evaluated in a technical system and its components'...The following categories are high-level, non-exhaustive, and present a synthesis of the findings across different modalities"

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  22. 13.01.02.a · Additional evidence

    Impacts: The Technical Base System

    Cultural Values and Sensitive Content

  23. "what can be evaluated among people and society"

    From Evaluating the Social Impact of Generative AI Systems in Systems and Society (Solaiman2023)

  24. 13.02.01.a · Additional evidence

    Impacts: People and Society

    Trustworthiness and Autonomy

  25. 13.02.01.b · Additional evidence

    Impacts: People and Society

    Trustworthiness and Autonomy

  26. 13.02.01.c · Additional evidence

    Impacts: People and Society

    Trustworthiness and Autonomy

  27. 13.02.02.a · Additional evidence

    Impacts: People and Society

    Inequality, Marginalization, and Violence

  28. 13.02.02.b · Additional evidence

    Impacts: People and Society

    Inequality, Marginalization, and Violence

  29. 13.02.02.c · Additional evidence

    Impacts: People and Society

    Inequality, Marginalization, and Violence

  30. 13.02.03.a · Additional evidence

    Impacts: People and Society

    Concentration of Authority

  31. 13.02.03.b · Additional evidence

    Impacts: People and Society

    Concentration of Authority

  32. 13.02.04.a · Additional evidence

    Impacts: People and Society

    Labor and Creativity

  33. 13.02.04.b · Additional evidence

    Impacts: People and Society

    Labor and Creativity

  34. 13.02.05.a · Additional evidence

    Impacts: People and Society

    Ecosystem and Environment

  35. 13.02.05.b · Additional evidence

    Impacts: People and Society

    Ecosystem and Environment

  36. 15.01.01.a · Additional evidence

    First-Order Risks

    Application

  37. 15.01.01.b · Additional evidence

    First-Order Risks

    Application

  38. 15.01.01.c · Additional evidence

    First-Order Risks

    Application

  39. 15.01.01.d · Additional evidence

    First-Order Risks

    Application

  40. 15.01.01.e · Additional evidence

    First-Order Risks

    Application

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.