MIT AI Risk Repository

Browse AI risks

336 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset

336 entries · page 2 of 7

  1. 61.02.42 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Risks from network interconnectivity

    "The interconnectedness of AI networks can create vulnerabilities, where issues in one part of the network can have cascading effects across the system."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  2. 62.19.06 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Vulnerabilities arising from additional modalities in multimodal models

    "Additional modalities can introduce new attack vectors in multimodal models as well as expand the scope of the previous attacks, ranging from jailbreaking to poisoning [13]. Typically, different modalities have different robustness levels, allowing malicious actors to choose the most vulnerable part of the model to attack [119, 181]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  3. 62.19.07 · Risk Sub-Category

    Attacks on GPAIs/GPAI Failure Modes

    Vulnerabilities to jailbreaks exploiting long context windows (many- shot jailbreaking)

    "Language models with long context windows are vulnerable to new types of ex- ploitations that are ineffective on models with shorter context windows. While few-shot jailbreaking, which involves providing few examples of the desired harmful output, might not trigger a harmful response, many-shot jailbreak- ing, which involves a higher number of such examples, increases the likelihood of eliciting an undesirable output. These vulnerabilities become more significant as context windows expand with newer model releases [7]."

    From Risk Sources and Risk Management Measures in Support of Standards for General-Purpose AI Systems (Gipiškis2024)

  4. "LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardized evaluation makes it difficult to compare them. We also do not have efficient white-box methods to evaluate adver- sarial robustness. Multi-modal LLMs may further allow novel types of jailbreaks via additional modalities. Finally, the lack of robust privilege levels within the LLM input means that jailbreaking and prompt-injection attacks may be particularly hard to eliminate

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  5. 73.07.01 · Risk Sub-Category

    Jailbreaks and Prompt Injections Threaten Security of LLMs

    Exploiting Limited Generalization of Safety Finetuning

    "Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. using encoded text (Wei et al., 2023c) or low-resource languages (Deng et al., 2023a; Yong et al., 2023) (see also Section 3.2)."

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  6. 24.11.00 · Risk Category

    Misinformation risks

    "The rapid integration of AI systems with advanced capabilities, such as greater autonomy, content generation, memorisation and planning skills (see Chapter 4) into personalised assistants also raises new and more specific challenges related to misinformation, disinformation and the broader integrity of our information environment. "

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  7. 58.07.06 · Risk Sub-Category

    Societal and Cultural

    Historical revisionism

    "Historical revisionism - Deliberate or unintentional reinterpretation of established/orthodox historical events or accounts held by societies, communities, academics."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  8. 02.09.05 · Risk Sub-Category

    Hallucinations

    Pursuing Consistent Context

    "LLMs have been demonstrated to pursue consistent context [129]–[132], which may lead to erroneous generation when the prefixes contain false information. Typical examples include sycophancy [129], [130], false demonstrations-induced hallucinations [113], [133], and snowballing [131]. As LLMs are generally fine-tuned with instruction-following data and user feedback, they tend to reiterate user-provided opinions [129], [130], even though the opinions contain misinformation. Such a sycophantic behavior amplifies the likelihood of generating hallucinations, since the model may prioritize user op

    From Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  9. 11.05.01 · Risk Sub-Category

    Societal System Harms

    Information harms

    information-based harms capture concerns of misinformation, disinformation, and malinformation. Algorithmic systems, especially generative models and recommender, systems can lead to these information harms

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  10. 45.02.02 · Risk Sub-Category

    Safety risks in AI Applications

    Cyberspace risks (Risks of confusing facts, misleading users, and bypassing authentication)

    "AI systems and their outputs, if not clearly labeled, can make it difficult for users to discern whether they are interacting with AI and to identify the source of generated content. This can impede users' ability to determine the authenticity of information, leading to misjudgment and misunderstanding. Additionally, AI-generated highly realistic images, audio, and videos may circumvent existing identity verification mechanisms, such as facial recognition and voice recognition, rendering these authentication processes ineffective."

    From AI Safety Governance Framework (TC2602024)

  11. 66.09.06 · Risk Sub-Category

    Privacy and Security

    Distortion

    "disseminating false or misleading information about people"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  12. 69.01.05 · Risk Sub-Category

    False information

    Spreads and self-perpetuates mis/disinformation

  13. 31.01.05 · Risk Sub-Category

    Information Manipulation

    Clickbait and feeding the surveillance advertising ecosystem

    "Beyond misinformation and disinformation, generative AI can be used to create clickbait headlines and articles, which manipulate how users navigate the internet and applications. For example, generative AI is being used to create full articles, regardless of their veracity, grammar, or lack of common sense, to drive search engine optimization and create more webpages that users will click on. These mechanisms attempt to maximize clicks and engagement at the truth’s expense, degrading users’ experiences in the process. Generative AI continues to feed this harmful cycle by spreading misinformat

    From Generating Harms - Generative AI's impact and paths forwards (EPIC2023)

  14. 55.04.05 · Risk Sub-Category

    Worsened epistemic processes for society

    Reduced decision-making capacity as a result of decreased trust in information

    "In addition, the increased awareness of these trends in information production and distribution could make it harder for anyone to evaluate the trustworthiness of any information source, reducing overall trust in information. In all of these scenarios, it would be much harder for humanity to make good decisions on important issues, particularly due to declining trust in credible multipartisan sources, which could hamper attempts at cooperation and collective action. The vaccine and mask hesitancy that exacerbated Covid-19, for example, were likely the result of insufficient trust in public he

    From A Survey of the Potential Long-term Impacts of AI: How AI Could Lead to Long-term Changes in Science, Cooperation, Power, Epistemics and Values (Clarke2023)

  15. 58.03.08 · Risk Sub-Category

    Psychological

    Radicalisation

    "Radicalisation - Adoption of extreme political, social, or religious ideals and aspirations due to the nature or misuse of an algorithmic system, potentially resulting in abuse, violence, or terrorism."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  16. 58.08.05 · Risk Sub-Category

    Political and Economic

    Institutional trust loss

    "Institutional trust loss - Erosion of trust in public institutions and weakened checks and balances due to mis/disinformation, influence operations, over-dependence on technology, etc."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  17. 61.02.20 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Detection challenges in content

    "The difficulty in distinguishing synthetic content from authentic material adds to information risks."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  18. 66.03.02 · Risk Sub-Category

    Misinformation Harms

    Pollution of information ecosystems

    "Contaminating publicly available information with false or inaccurate information (i.e., the generative tool's output is disseminated beyond the end user)"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  19. 66.03.03 · Risk Sub-Category

    Misinformation Harms

    Erosion of trust in public information

  20. 66.04.01 · Risk Sub-Category

    Societal and Cultural

    Overburdening ecosystems

    "Pollution of a space/ecosystem that is expected to be free of AI involvement/influence (e.g., creative material submission portals, job applications)"

    From A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents (Li2025)

  21. 67.01.01 · Risk Sub-Category

    Societal harms

    Degradation of the information environment

    "Frontier AI can cheaply generate realistic content which can falsely portray people and events. There is potential risk of compromised decision-making by individuals and institutions who rely on inaccurate or misleading publicly available information, as well as lower overall trust in true information."

    From Capabilities and Risks from Frontier AI (DSIT2023)

  22. "Some uses of AI have been deeply concerning, namely voice cloning [58] and the generation of deep fake videos [59]. For example, in March 2022, in the early days of the Russian invasion of Ukraine, hackers broadcast via the Ukrainian news website Ukraine 24 a deep fake video of President Volodymyr Zelensky capitulating and calling on his soldiers to lay down their weapons [60]. The necessary software to create these fakes is readily available on the Internet, and the hardware requirements are modest by today’s standards [61]. Other nefarious uses of AI include accelerating password cracking [

    From Navigating the Landscape of AI Ethics and Responsibility (Cunha2023)

  23. LMs, due to their remarkable capabilities, carry the same potential for malice as other technological products. For instance, they may be used in information warfare to generate deceptive information or unlawful content, thereby having a significant impact on individuals and society. As current LMs are increasingly built as agents to accomplish user objectives, they may disregard the moral and safety guidelines if operating without adequate supervision. Instead, they may execute user commands mechanically without considering the potential damage. They might interact unpredictably with humans a

    From Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  24. 61.02.22 · Risk Sub-Category

    Sources of systemic risks from general-purpose AI

    Dual-use nature

    "AI’s potential for both beneficial and harmful applications complicates efforts to manage its societal impacts effectively."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  25. 71.02.02 · Risk Sub-Category

    User Intent

    Malicious and Indirect

    "Benign intermediate for harmful end objective"

    From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025)

  26. 11.05.03 · Risk Sub-Category

    Societal System Harms

    Civic and political harms

    Political harms emerge when “people are disenfranchised and deprived of appropriate political power and influence” [186, p. 162]. These harms focus on the domain of government, and focus on how algorithmic systems govern through individualized nudges or micro-directives [187], that may destabilize governance systems, erode human rights, be used as weapons of war [188], and enact surveillant regimes that disproportionately target and harm people of color

    From Sociotechnical Harms of Algorithmic Systems: Scoping a Taxonomy for Harm Reduction (Shelby2023)

  27. "Informational and communicational AI risks refer particularly to informational manipulation through AI systems that influence the provision of information (Rahwan, 2018; Wirtz & Müller, 2019), AIbased disinformation and computational propaganda, as well as targeted censorship through AI systems that use respectively modified algorithms, and thus restrict freedom of speech."

    From Governance of artificial intelligence: A risk and guideline-based integrative framework (Wirtz2022)

  28. 19.02.01 · Risk Sub-Category

    Informational and Communicational AI Risks

    Manipulation and control of information provision (e.g., personalised adds, filtered news)

  29. 24.03.14 · Risk Sub-Category

    Malicious Uses

    Authoritarian Surveillance, Censorship, and Use: Delegation of Decision-Making Authority to Malicious Actors

    "Finally, the principal value proposition of AI assistants is that they can either enhance or automate decision-making capabilities of people in society, thus lowering the cost and increasing the accuracy of decision-making for its user. However, benefiting from this enhancement necessarily means delegating some degree of agency away from a human and towards an automated decision-making system—motivating research fields such as value alignment. This introduces a whole new form of malicious use which does not break the tripwire of what one might call an ‘attack’ (social engineering, cyber offen

    From The Ethics of Advanced AI Assistants (Gabriel2024)

  30. 46.02.02 · Risk Sub-Category

    Financial and Economic Damage

    Propaganda - Extremist schemes

  31. "The distortion of the information ecosystem, including the spread of misinformation, fake news, and other forms of deceptive content [28], is categorized as “Information Manipulation.”"

    From GenAI against humanity: nefarious applications of generative artificial intelligence and large language models (Ferrara2023)

  32. 46.03.01 · Risk Sub-Category

    Information Manipulation

    Deception - Information control

  33. "Lastly, broader harms that can impact communities, societal structures, and critical infrastructures, including threats to democratic processes, social cohesion, and technological systems, are captured under “Societal, Socio-technical, and Infrastructural Damage.”"

    From GenAI against humanity: nefarious applications of generative artificial intelligence and large language models (Ferrara2023)

  34. 46.04.02 · Risk Sub-Category

    Socio-technical and Infrastructural

    Propaganda - Synthetic realities

  35. 46.04.03 · Risk Sub-Category

    Socio-technical and Infrastructural

    Dishonesty - Targeted surveillance

  36. 50.03.12 · Risk Sub-Category

    Societal Risks

    Manipulation (Sowing Division)

  37. 50.03.14 · Risk Sub-Category

    Societal Risks

    Defamation

  38. "Political and Economic - Manipulation of political beliefs, damage to political institutions and the effective delivery of government services."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  39. "Large-scale influence on communication and information systems, and epistemic processes more generally."

    From A Taxonomy of Systemic Risks from General-Purpose AI (Uuk2025)

  40. 37.01.03 · Risk Sub-Category

    Design of AI

    Threats to human institutions and life

    "This group comprises 11% of the articles and centers on risks stemming from AI systems designed with malicious intent or that can end up in a threat to human life. It can be divided into two key themes: threats to law and democracy, and transhumanism."

    From What Ethics Can Say on Artificial Intelligence: Insights from a Systematic Literature Review (Giarmoleo2024)

  41. "Eased access to or synthesis of materially nefarious information or design capabilities related to chemical, biological, radiological, or nuclear (CBRN) weapons or other dangerous materials or agents."

    From Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST2024)

  42. 50.01.03 · Risk Sub-Category

    System and Operational Risks

    Security risks (availability)

  43. 50.04.08 · Risk Sub-Category

    Legal and Rights-Related Risks

    Criminal Activities (Other Unlawful/Criminal Activities)

  44. 53.03.06 · Risk Sub-Category

    Direct catastrophe from AI

    Failures in or misuse of intermediary (non-AGI) AI systems, resulting in catastrophe

    "Deployment of “prepotent” AI systems that are non-general but capable of outperforming human collective efforts on various key dimensions;170 → Militarization of AI enabling mass attacks using swarms of lethal autonomous weapons systems;171 → Military use of AI leading to (intentional or unintentional) nuclear escalation, either because machine learning systems are directly integrated in nuclear command and control systems in ways that result in escalation172 or because conventional AI-enabled systems (e.g., autonomous ships) are deployed in ways that result in provocation and escalation;173

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  45. 58.05.01 · Risk Sub-Category

    Financial and business

    Business operations/infrastructure damage

    "Business operations/infrastructure damage - Damage, disruption, or destruction of a business system and/or its components due to malfunction, cyberattacks, etc."

    From A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  46. 68.02.00 · Risk Category

    Cyber offense

    "Cyber risks, especially in the context of cyber offense, are an existing threat that may be exacerbated by AI. [108] demonstrated that teams of LLM agents can exploit zero-day vulnerabilities when given a description of the vulnerability and toy capture-the-flag problems. While cyber risks are not typically regarded as catastrophic, [3] argues that cyberwarfare is an underappreciated risk that poses a credible threat of catastrophic harm."

    From Dimensional Characterization and Pathway Modeling for Catastrophic AI Risks (Chin2025)

  47. 71.01.02 · Risk Sub-Category

    Scientific Domain of Agents

    Biological Risks

    "Biological risks encompass the dangerous modification of pathogens and unethical manipulation of genetic material, potentially leading to unforeseen biohazardous outcomes."

    From Risks of AI Scientists: Prioritizing Safeguarding Over Autonomy (Tang2025)

  48. 46.01.02 · Risk Sub-Category

    Personal Loss and Identity Theft

    Propaganda - Digital impersonations

    "AI-generated impersonation for identity theft might be found at the intersection of “Harm to the Person” and “Deception.”"

    From GenAI against humanity: nefarious applications of generative artificial intelligence and large language models (Ferrara2023)

  49. "Then, we have the potential for financial loss, fraud, market manipulation, and other economic harms, which fall under “Financial and Economic Damage.”

    From GenAI against humanity: nefarious applications of generative artificial intelligence and large language models (Ferrara2023)

  50. 46.02.01 · Risk Sub-Category

    Financial and Economic Damage

    Deception - Bespoke ransom

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.