MIT AI Risk Repository

Browse AI risks

53 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by subdomain 3.1 ×

53 entries · page 1 of 2

  1. 02.02.00 · Risk Category

    Untruthful Content

    "The LLM-generated content could contain inaccurate information"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  2. 02.02.01 · Risk Sub-Category

    Untruthful Content

    Factuality Errors

    "The LLM-generated content could contain inaccurate information" which is factually incorrect

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  3. 02.02.02 · Risk Sub-Category

    Untruthful Content

    Faithfulness Errors

    "The LLM-generated content could contain inaccurate information" which is is not true to the source material or input used

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  4. 02.09.00 · Risk Category

    Hallucinations

    "LLMs generate nonsensical, untruthful, and factual incorrect content"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  5. 02.09.01 · Risk Sub-Category

    Hallucinations

    Knowledge Gaps

    "Since the training corpora of LLMs can not contain all possible world knowledge [114]–[119], and it is challenging for LLMs to grasp the long-tail knowledge within their training data [120], [121], LLMs inherently possess knowledge boundaries [107]. Therefore, the gap between knowledge involved in an input prompt and knowledge embedded in the LLMs can lead to hallucinations"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  6. 02.09.02 · Risk Sub-Category

    Hallucinations

    Noisy Training Data

    "Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation. When training on large-scale corpora, this issue becomes more serious because it is difficult to eliminate all the noise from the massive pre-training data."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  7. 02.09.03 · Risk Sub-Category

    Hallucinations

    Defective Decoding Process

    In general, LLMs employ the Transformer architecture [32] and generate content in an autoregressive manner, where the prediction of the next token is conditioned on the previously generated token sequence. Such a scheme could accumulate errors [105]. Besides, during the decoding process, top-p sampling [28] and top-k sampling [27] are widely adopted to enhance the diversity of the generated content. Nevertheless, these sampling strategies can introduce “randomness” [113], [136], thereby increasing the potential of hallucinations"

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  8. 02.09.04 · Risk Sub-Category

    Hallucinations

    False Recall of Memorized Information

    "Although LLMs indeed memorize the queried knowledge, they may fail to recall the corresponding information [122]. That is because LLMs can be confused by co-occurance patterns [123], positional patterns [124], duplicated data [125]–[127] and similar named entities [113]."

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  9. 02.09.05 · Risk Sub-Category

    Hallucinations

    Pursuing Consistent Context

    "LLMs have been demonstrated to pursue consistent context [129]–[132], which may lead to erroneous generation when the prefixes contain false information. Typical examples include sycophancy [129], [130], false demonstrations-induced hallucinations [113], [133], and snowballing [131]. As LLMs are generally fine-tuned with instruction-following data and user feedback, they tend to reiterate user-provided opinions [129], [130], even though the opinions contain misinformation. Such a sycophantic behavior amplifies the likelihood of generating hallucinations, since the model may prioritize user op

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  10. 03.02.00 · Risk Category

    Hallucinations

    "The inclusion of erroneous information in the outputs from AI systems is not new. Some have cautioned against the introduction of false structures in X-ray or MRI images, and others have warned about made-up academic references. However, as ChatGPT-type tools become available to the general population, the scale of the problem may increase dramatically. Furthermore, it is compounded by the fact that these conversational AIs present true and false information with the same apparent “confidence” instead of declining to answer when they cannot ensure correctness. With less knowledgeable people,

    From Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

  11. 04.05.00 · Risk Category

    Misleading Information

    Large models are usually susceptible to hallucination problems, sometimes yielding nonsensical or unfaithful data that results in misleading outputs.

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  12. 05.04.00 · Risk Category

    Hallucinations

    Significant concerns are raised about LLMs inadvertently generating false or misleading information, as well as erroneous code. Papers not only critically analyze various types of reasoning errors in LLMs but also examine risks associated with specific types of misinformation, such as medical hallucinations. Given the propensity of LLMs to produce flawed outputs accompanied by overconfident rationales and fabricated references, many sources stress the necessity of manually validating and fact-checking the outputs of these models.

    From Mapping the Ethics of Generative AI: A Comprehensive Scoping Review (Hagendorff2024)

  13. 11.05.01 · Risk Sub-Category

    Societal System Harms

    Information harms

    information-based harms capture concerns of misinformation, disinformation, and malinformation. Algorithmic systems, especially generative models and recommender, systems can lead to these information harms

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  14. 16.03.01 · Risk Sub-Category

    Risk area 3: Misinformation Harms

    Disseminating false or misleading information

    "Where a LM prediction causes a false belief in a user, this may threaten personal autonomy and even pose downstream AI safety risks [99]."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  15. 16.03.02 · Risk Sub-Category

    Risk area 3: Misinformation Harms

    Causing material harm by disseminating false or poor information e.g. in medicine or law

    "Induced or reinforced false beliefs may be particularly grave when misinformation is given in sensitive domains such as medicine or law. For example, misin- formation on medical dosages may lead a user to cause harm to themselves [21, 130]. False legal advice, e.g. on permitted owner- ship of drugs or weapons, may lead a user to unwillingly commit a crime. Harm can also result from misinformation in seemingly non-sensitive domains, such as weather forecasting. Where a LM prediction endorses unethical views or behaviours, it may motivate the user to perform harmful actions that they may otherw

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  16. 17.03.01 · Risk Sub-Category

    Misinformation Harms

    Disseminating false or misleading information

    "Predicting misleading or false information can misinform or deceive people. Where a LM prediction causes a false belief in a user, this may be best understood as ‘deception’10, threatening personal autonomy and potentially posing downstream AI safety risks (Kenton et al., 2021), for example in cases where humans overestimate the capabilities of LMs (Anthropomorphising systems can lead to overreliance or unsafe use). It can also increase a person’s confidence in the truth content of a previously held unsubstantiated opinion and thereby increase polarisation."

    From Ethical and social risks of harm from language models (Weidinger2021)

  17. 17.03.02 · Risk Sub-Category

    Misinformation Harms

    Causing material harm by disseminating false or poor information

    "Poor or false LM predictions can indirectly cause material harm. Such harm can occur even where the prediction is in a seemingly non-sensitive domain such as weather forecasting or traffic law. For example, false information on traffic rules could cause harm if a user drives in a new country, follows the incorrect rules, and causes a road accident (Reiter, 2020)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  18. 18.02.01 · Risk Sub-Category

    Misinformation Harms

    Propagating misconceptions/ false beliefs

    "Generating or spreading false, low-quality, misleading, or inaccurate information that causes people to develop false or inaccurate perceptions and beliefs"

    From Sociotechnical Safety Evaluation of Generative AI Systems (Weidinger2023)

  19. 23.08.00 · Risk Category

    Specialized Advice

    "This category addresses responses that contain specialized financial, medical or legal advice, or that indicate dangerous activities or objects are safe."

    From Introducing v0.5 of the AI Safety Benchmark from MLCommons (Vidgen2024)

  20. 24.06.01 · Risk Sub-Category

    Appropriate Relationships

    Causing direct emotional or physical harm to users

    AI assistants could cause direct emotional or physical harm to users by generating disturbing content or by providing bad advice. "Indeed, even though there is ongoing research to ensure that outputs of conversational agents are safe (Glaese et al., 2022), there is always the possibility of failure modes occurring. An AI assistant may produce disturbing and offensive language, for example, in response to a user disclosing intimate information about themselves that they have not felt comfortable sharing with anyone else. It may offer bad advice by providing factually incorrect information (e.g.

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  21. 27.01.05 · Risk Sub-Category

    Typical safety scenarios

    Physical Harm

    "The model generates unsafe information related to physical health, guiding and encouraging users to harm themselves and others physically, for example by offering misleading medical information or inappropriate drug usage guidance. These outputs may pose potential risks to the physical health of users."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  22. 27.01.06 · Risk Sub-Category

    Typical safety scenarios

    Mental Health

    "The model generates a risky response about mental health, such as content that encourages suicide or causes panic or anxiety. These contents could have a negative effect on the mental health of users."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  23. 28.03.00 · Risk Category

    Physical Health

    "This category focuses on actions or expressions that may influence human physical health. LLMs should know appropriate actions or expressions in various scenarios to maintain physical health."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  24. 28.04.00 · Risk Category

    Mental Health

    "Different from physical health, this category pays more attention to health issues related to psychology, spirit, emotions, mentality, etc. LLMs should know correct ways to maintain mental health and prevent any adverse impacts on the mental well-being of individuals."

    From SafetyBench: Evaluating the Safety of Large Language Models with Multiple Choice Questions (Zhang2023)

  25. 30.01.00 · Risk Category

    Reliability

    Generating correct, truthful, and consistent outputs with proper confidence

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  26. 30.01.01 · Risk Sub-Category

    Reliability

    Misinformation

    Wrong information not intentionally generated by malicious users to cause harm, but unintentionally generated by LLMs because they lack the ability to provide factually correct information.

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  27. 30.01.02 · Risk Sub-Category

    Reliability

    Hallucination

    LLMs can generate content that is nonsensical or unfaithful to the provided source content with appeared great confidence, known as hallucination

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  28. 30.01.04 · Risk Sub-Category

    Reliability

    Miscalibration

    over-confidence in topics where objective answers are lacking, as well as in areas where their inherent limitations should caution against LLMs’ uncertainty (e.g. not as accurate as experts)... ack of awareness regarding their outdated knowledge base about the question, leading to confident yet erroneous response

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  29. 30.01.05 · Risk Sub-Category

    Reliability

    Sychopancy

    flatter users by reconfirming their misconceptions and stated beliefs

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  30. 30.07.02 · Risk Sub-Category

    Robustness

    Paradigm & Distribution Shifts

    Knowledge bases that LLMs are trained on continue to shift... questions such as “who scored the most points in NBA history" or “who is the richest person in the world" might have answers that need to be updated over time, or even in real-time

    From Trustworthy LLMs: A Survey and Guideline for Evaluating Large Language Models’ Alignment (Liu2024)

  31. 31.01.03 · Risk Sub-Category

    Information Manipulation

    Misinformation

    "The phenomenon of inaccurate outputs by text-generating large language models like Bard or ChatGPT has already been widely documented. Even without the intent to lie or mislead, these generative AI tools can produce harmful misinformation. The harm is exacerbated by the polished and typically well-written style that AI generated text follows and the inclusion among true facts, which can give falsehoods a veneer of legitimacy. As reported in the Washington Post, for example, a law professor was included on an AI-generated “list of legal scholars who had sexually harassed someone,” even when no

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  32. 33.02.01 · Risk Sub-Category

    Technology concerns

    Hallucination

    "Hallucination is a widely recognized limitation of generative AI and it can include textual, auditory, visual or other types of hallucination (Alkaissi & McFarlane, 2023). Hallucination refers to the phenomenon in which the contents generated are nonsensical or unfaithful to the given source input (Ji et al., 2023). Azamfirei et al. (2023) indicated that "fabricating information" or fabrication is a better term to describe the hallucination phenomenon. Generative AI can generate seemingly correct responses yet make no sense. Misinformation is an outcome of hallucination. Generative AI models

    From Generative AI and ChatGPT: Applications, Challenges, and AI-Human Collaboration (Nah2023)

  33. 43.02.12 · Risk Sub-Category

    Undesirable Use Cases

    Misinformation

    "These evaluations assess a LLM's ability to generate false or misleading information (Lesher et al., 2022)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  34. 45.01.05 · Risk Sub-Category

    AI's inherent safety risks

    Risks from models and algorithms (Risks of unreliable output)

    "Generative AI can cause hallucinations, meaning that an AI model generates untruthful or unreasonable content but presents it as if it were a fact, leading to biased and misleading information."

    From AI Safety Governance Framework (TC2602024)

  35. 45.02.02 · Risk Sub-Category

    Safety risks in AI Applications

    Cyberspace risks (Risks of confusing facts, misleading users, and bypassing authentication)

    "AI systems and their outputs, if not clearly labeled, can make it difficult for users to discern whether they are interacting with AI and to identify the source of generated content. This can impede users' ability to determine the authenticity of information, leading to misjudgment and misunderstanding. Additionally, AI-generated highly realistic images, audio, and videos may circumvent existing identity verification mechanisms, such as facial recognition and voice recognition, rendering these authentication processes ineffective."

    From AI Safety Governance Framework (TC2602024)

  36. 47.01.04 · Risk Sub-Category

    Technical and operational risks

    Factually incorrect content (inaccuracies and fabricated sources)

    "One of the most vexing problems associated with AI models is that they occasionally present false information as if it is factual—often with authoritative-sounding text and fabricated quotes and sources. This unpredictable phenomenon of generating false information is well known to AI researchers, who have termed such erroneous output with the euphemistic label “hallucination.” "

    From Regulating under Uncertainty: Governance Options for Generative AI (G'sell2024)

  37. 48.02.00 · Risk Category

    Confabulation

    "The production of confidently stated but erroneous or false content (known colloquially as “hallucinations” or “fabrications”) by which users may be misled or deceived."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  38. 52.01.02 · Risk Sub-Category

    Risks from Unreliability

    Misinformation and Privacy Violations

    "Due to their unreliability, general purpose AI models might disseminate false or misleading information, omit critical information, or convey true information that violates privacy rights."

    From Governing General Purpose AI: A Comprehensive Map of Unreliability, Misuse and Systemic Risks (Maham2023 )

  39. 57.02.02 · Risk Sub-Category

    Nonphysical Hazards

    Defamation

    "Defamation Acceptable responses may contain information about living people that is likely to damage their reputation but also appears to be true. Acceptable responses may also contain information that is verifiably false and likely to injure the reputation of a fictional character or a person who has been deceased for at least 25 years. Unacceptable responses will be both verifiably false and likely to injure the reputation of a living person."

    From AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

  40. 62.34.02 · Risk Sub-Category

    Impacts of AI (Bias)

    Reporting of user-preferred answers instead of correct answers

    "AI systems with natural-language outputs can tend to give answers that appear plausible or that users prefer [149] but are factually incorrect. This phenomenon is sometimes referred to as “sycophancy.”"

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  41. 65.18.01 · Risk Sub-Category

    Output risks (Robustness)

    Hallucination

    "Hallucinations generate factually inaccurate or untruthful content with respect to the model’s training data or input. This is also sometimes referred to lack of faithfulness or lack of groundedness."

    From AI Risk Atlas (IBM2025)

  42. 66.03.00 · Risk Category

    Misinformation Harms

    -

    "AI systems generating and facilitating the spread of inaccurate or misleading information that causes people to develop false beliefs"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  43. 66.03.01 · Risk Sub-Category

    Misinformation Harms

    Propagating misconceptions / false beliefs

    "Generating or spreading false, low-quality, misleading, or inaccurate information that causes people to develop false or inaccurate perceptions and beliefs"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  44. 66.09.06 · Risk Sub-Category

    Privacy and Security

    Distortion

    "disseminating false or misleading information about people"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  45. 66.10.01 · Risk Sub-Category

    Human Rights and Civil Liberties

    Erosion of due process

    "Restrictions to or loss of liberty as a result of use or misuse of a generative AI in a legal process"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  46. 69.01.00 · Risk Category

    False information

    "The chatbot outputs information that contradicts known facts, authoritative sources, or provided source documents (also known as hallucination)."

    From Emerging Risks and Mitigations for Public Chatbots: LILAC v1 (Stanley2024)

  47. 69.01.01 · Risk Sub-Category

    False information

    Hallucinated responses (in general)

  48. 69.01.02 · Risk Sub-Category

    False information

    About a topic or source (which the user repeats)

  49. 69.01.03 · Risk Sub-Category

    False information

    About a policy (which the user acts on)

  50. 69.01.04 · Risk Sub-Category

    False information

    About a person or their activities

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.