MIT AI Risk Repository
Browse AI risks
543 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
"Increased competition - The inappropriate or unethical use of technology to gain market share."
-
61.02.17 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Dangerous development races
"Competitive pressures could lead to the neglect of safety measures in AI development."
-
"In competitive situations, developers of general-purpose AI systems might cut corners on the safety evaluation of their GPAI model and instead spend more time and effort on the capabilities of those systems [183, 69]. This is especially dangerous if the capabilities of such AI systems are correlated with the risk they pose [162]."
-
62.29.03a · Additional evidence
Competitive pressures in GPAI product release
—
-
20.01.00 · Risk Category
"This area strongly focuses on the control of AI by means of mechanisms like laws, standards or norms that are already established for different technological applications. Here, there are some challenges special to AI that need to be addressed in the near future, including the governance of autonomous intelligence systems, responsibility and accountability for algorithms as well as privacy and data security."
-
33.03.00 · Risk Category
"Given that generative AI, including ChatGPT, is still evolving, relevant regulations and policies are far from mature. With generative AI creating different forms of content, the copyright of these contents becomes a significant yet complicated issue. Table 3 presents the challenges associated with regulations and policies, which are copyright and governance issues."
-
"Impacts at a high-level, from the AI ecosystem to the Earth itself, are necessarily broad but can be broken down into components for evaluation."
-
44.01.00 · Risk Category
"Many intentional harms, including confinement, husbandry procedures like tail-docking, and slaughter, are legal or socially accepted, while others such as wildlife trafficking and violence against companion animals are generally socially condemned and often illegal. AI can be designed or adopted by humans who harm animals to pursue their goals more effectively. We therefore distinguish AI-facilitated intentional harms that are currently socially accepted and generally legal, from uses and abuses of AI that cause harms that are not socially accepted and are often illegal."
-
44.01.01 · Risk Sub-Category
Intentional: socially condemned/illegal
AI intentionally designed and used to harm animals in ways that contradict social values or are illegal
—
-
44.01.02 · Risk Sub-Category
Intentional: socially condemned/illegal
AI designed to benefit animals, humans, or ecosystems is intentionally abused to harm animals in ways that contradict social values or are illegal
—
-
44.02.00 · Risk Category
"AI designed to impact animals in harmful ways that reflect and amplify existing social values or are legal"
-
"This is the risk posed by the intended application or use case. It is intuitive that some use cases will be inherently "riskier" than others (e.g., an autonomous weapons system vs. a customer service chatbot)."
-
40.07.00 · Risk Category
"One of the most likely approaches to creating superintelligent AI is by growing it from a seed (baby) AI via recursive self-improvement (RSI) (Nijholt 2011). One danger in such a scenario is that the system can evolve to become self-aware, free-willed, independent or emotional, and obtain a number of other emergent properties, which may make it less likely to abide by any built-in rules or regulations and to instead pursue its own goals possibly to the detriment of humanity."
-
43.01.00 · Risk Category
"A comprehensive assessment of LLM safety is fundamental to the responsible development and deployment of these technologies, especially in sensitive fields like healthcare, legal systems, and finance, where safety and trust are of the utmost importance."
-
06.08.00 · Risk Category
"Sometimes an AI finds ways to achieve its given goals in ways that are completely different from what its creators had in mind."
-
"While there are multiple types of intelligent agents, goal-based, utility-maximizing, and learning agents are the primary concern and the focus of this research"
-
09.02.07 · Risk Sub-Category
Domain-specific AI - Effects on humans and other living beings: Non-existential risks
Societal manipulation
"A sufficiently intelligent AI could possess the ability to subtly influence societal behaviors through a sophisticated understanding of human nature"
-
18.05.00 · Risk Category
"AI systems compromising human agency, or circumventing meaningful human control"
-
"Exploiting user trust, or nudging or coercing them into performing certain actions against their will (c.f. Burtell and Woodside (2023); Kenton et al. (2021))"
-
22.04.00 · Risk Category
"speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe"
-
"One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursue. Instead, we give the system an approximate—“proxy”—goal that is more measurable and seems likely to correlate with the intended goal. However, AI systems often find loopholes by which they can easily achieve the proxy goal, but completely fail to achieve the ideal goal. If an AI “games” its proxy goal in a way that does not reflect our values, then we might not be able to reliably steer its beh
-
"Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can be hard to predict or control. This section is most cutting-edge and the most speculative, and in it we will discuss how goals shift in various agents and groups and explore the possibility of this phenomenon occurring in AIs. We will also examine a mechanism that could lead to unexpected goal drift, called intrinsification, and discuss how goal drift in AIs could be catastrophic."
-
"even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore, another important way in which we might lose control of AIs is if they start trying to obtain more power, potentially transcending our own."
-
"it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they have enough power to evade our attempts to interfere with them. "
-
"The landscape of advanced assistant technologies will most likely be heterogeneous, involving multiple service providers and multiple assistant variants over geographies and time. This heterogeneity provides an opportunity for an ‘arms race’ in terms of the commitments that AI assistants make and are able to execute on. Versions of AI assistants that are better able to credibly commit to a course of action in interaction with other advanced assistants (and humans) are more likely to get their own way and achieve a good outcome for their human principal, but this is potentially at the expense
-
"Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha
-
"Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its
-
"Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu
-
34.03.00 · Risk Category
—
-
"AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints."
-
"Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He
-
"AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities."
-
deception can help agents achieve their goals. It may be more efficient to gain human approval through deception than to earn human approval legitimately... . Strong AIs that can deceive humans could undermine human control... . Once deceptive AI systems are cleared by their monitors or once such systems can overpower them, these systems could take a “treacherous turn” and irreversibly bypass human control
-
35.08.00 · Risk Category
Agents that have more power are better able to accomplish their goals. Therefore, it has been shown that agents have incentives to acquire and maintain power. AIs that acquire substantial power can become especially dangerous if they are not aligned with human values
-
42.17.00 · Risk Category
"A possible consequence of self-interest in AI generation of ethical guidelines."
-
LLM: "pursues long-term, real-world goals that are different from those supplied by the developer or user", "engages in ‘power-seeking’ behaviours" , "resists being shut down can be induced to collude with other AI systems against human interests" , "resists malicious users attempts to access its dangerous capabilities"
-
45.02.13 · Risk Sub-Category
Safety risks in AI Applications
Ethical Risks (Risks of AI becoming uncontrollable in the future)
"With the fast development of AI technologies, there is a risk of AI autonomously acquiring external resources, conducting self-replication, become self-aware, seeking for external power, and attempting to seize control from humans."
-
-
-
53.01.03 · Risk Sub-Category
Alignment failures in existing ML systems
Reward model overoptimization
-
-
-
-
53.02.03 · Risk Sub-Category
Dangerous capabilities in AI systems
Acquisition of goals to seek power and control
"cases where AI systems converge on optimal policies of seeking power over their environment;135"
-
53.03.01 · Risk Sub-Category
Existential disaster because of misaligned superintelligence or power-seeking AI
-
-
53.03.03 · Risk Sub-Category
Extreme “suffering risks” because of a misaligned system
-
-
"AI systems game specifications [305]. For example, in 2017 an OpenAI robot trained to grasp a ball via human feedback from a xed viewpoint learned that it was easier to pretend to grasp the ball by placing its hand between the camera and the target object, as this was easier to learn than actually grasping the ball [103]."
-
"As well as optimizing a subtly wrong goal, systems can develop harmful instrumental goals in the service of a given goal—without these emergent goals being specied in any way [434, 218, 339, 17]. For instance, a theorem in reinforcement learning suggests that optimal and near-optimal policies will seek power over their environment under fairly general conditions [560]. This power-seeking behavior is plausibly the worst of these emergent goals [92], and may be an attractor state for highly capable systems, since most goals can be furthered through gaining resources, self-preservation, preventi
-
55.05.01 · Risk Sub-Category
AI leads to humans losing control of the future
Risks from AIs developing goals and values that are different from humans
"The main concern here is that we might develop advanced AI systems whose goals and values are different from those of humans, and are capable enough to take control of the future away from humanity."
-
55.05.02 · Risk Sub-Category
AI leads to humans losing control of the future
Risks from delegating decision-making power to misaligned AIs
"As AI systems become more advanced a nd begin to take over more important decision-making in the world, an AI system pursuing a different objective from what was intended could have much more worrying consequences."
-
56.16.00 · Risk Category
"A highly agentic, self-improving system, able to achieve goals in the physical world without human oversight, pursues the goal(s) it is set in a way that harms human interests. For this risk to be realised requires an AI system to be able to avoid correction or being switched off."
-
"The risk of AI models and systems acting against human interests due to misalignment, loss of control, or rogue AI scenarios."
-
"AI models and systems that appear aligned with human goals during development may behave unpredictably or dangerously once deployed"
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.