MIT AI Risk Repository
Browse AI risks
336 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
55.01.02 · Risk Sub-Category
Risks from accelerating scientific progress
Faster scientific progress makes it harder for governance to keep pace with development
"Exacerbating these problems is that faster scientific progress would make it even harder for governance to keep pace with the deployment of new technologies. When these technologies are especially powerful or dangerous, such as those discussed above, insufficient governance can magnify their harms.8 This is known as the pacing problem, and it is an issue that technology governance already faces [47], for a variety of reasons"
-
61.02.12 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Challenges in perceiving, measuring, and recognizing harm
"Harm from AI often manifests subtly or over the long term, making it difficult to identify, measure, and address effectively."
-
61.02.40 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Rapid development outpacing regulation
"The fast pace of AI development may outstrip regulatory and legal frameworks."
-
61.02.41 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Resistance to international law
"AI models and systems may prove difficult to regulate or control under international law."
-
61.02.47 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Unpredictability of AI development trajectory
"The unpredictable trajectory of AI development complicates governance and risk management."
-
"Once a model is deployed, it can be exposed to benchmark data provided by the users [95, 170]. The model may then be further trained by these user inputs containing benchmark data."
-
"Benchmark saturation refers to benchmarks reaching their evaluation ceiling. The tendency towards benchmark saturation has been demonstrated in various benchmarks [19]. When benchmarks reach or are close to saturation, they stop being effective measures for new models, as more nuanced capability gains might not be detected."
-
62.16.16 · Risk Sub-Category
Benchmark Limitations (Insufficient benchmarks for AI safety evaluation)
"Benchmarks dedicated to measuring the performance of AI systems (e.g., on programming or math tasks) are more well-developed than those for assessing safety and harms in AI systems [234]. This gap can lead to AI systems excelling in specific tasks while exhibiting harmful behaviors that go undetected. More safety-related evaluation datasets can help in identifying previously overlooked undesirable model behaviors."
-
62.16.17 · Risk Sub-Category
Benchmark Limitations (Underestimating capabilities that are not covered by benchmarks)
"A lack of test coverage by benchmarks on specific abilities of a model can obscure the model’s capabilities from both the developer and the user [160]. This can lead to a false sense of safety and trust due to a lack of understanding of the model’s limitations."
-
"Determining who is responsible for an AI model is challenging without good documentation and governance processes."
-
"Determining responsibility when EAI causes harm requires new accountability and liability frameworks that address the complexities of highly autonomous physical systems. Human users may disagree with decisions taken by expert EAI systems, raising significant questions of delegation and responsibility [108]. Lack of EAI accountability could lead to confusion for users and breakdowns in traditional justice systems [109]."
-
"EAI deployment could fundamentally reshape society, particularly if the speed of technological development outpaces society’s ability to adapt [103, 120]. For example, EAI systems could provide physical threats of violence and mass surveillance capabilities to back up AI-enabled authoritarianism [121]."
-
44.03.00 · Risk Category
"AI designed to benefit animals, humans, or ecosystems has unintended harmful impact on animals"
-
47.04.05 · Risk Sub-Category
Environmental, economical, and societal challenges
Environmental cost (energy consumption)
"Training large AI models requires a substantial amount of computing power to handle vast datasets, which translates into high energy consumption."
-
47.04.06 · Risk Sub-Category
Environmental, economical, and societal challenges
Environmental cost (water consumption)
"Data centers use water for cooling to prevent servers from overheating. The water consumption associated with AI training and inference processes can be substantial, impacting local water resources."
-
48.05.00 · Risk Category
"Impacts due to high compute resource utilization in training or operating GAI models, and related outcomes that may adversely impact ecosystems."
-
"Biodiversity loss - Over-expansion of technology infrastructure, or inadequate alignment of technology with sustainable practices, leading to deforestation, habitat destruction, and fragmentation and loss of biodiversity."
-
"Carbon emissions - Release of carbon dioxide, nitric oxide and other gases, increasing carbon emissions, exacerbating climate change, and negatively impacting local communities."
-
"Electronic waste - Electrical or electronic equipment that is waste, including all components, sub-assemblies and consumables that are part of the equipment at the time the equipment becomes waste"
-
61.02.23 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Energy-intensive processes
"AI data collection, storage, and model training are energy-intensive, contributing to environmental risks."
-
"Training and deploying large models require substantial energy expenditure. The trend toward developing larger models exacerbates this issue. This can lead to excessive energy usage and have a negative environmental impact."
-
"Action(s) that lead directly or indirectly to the damage or destruction of tangible property eg. buildings, possessions, vehicles, robots"
-
"Excessive energy use resulting in energy bottlenecks and shortages for communities, organisations and businesses"
-
68.05.00 · Risk Category
"AI models are often trained using large amounts of computation. This process is very energy intensive, potentially leading to significant greenhouse emissions depending on the energy sources [132]. Experts believe drastically increasing carbon emissions could accelerate climate change, which may constitute a catastrophic risk [133]."
-
"Short-term or long-term Negative effects on the natural environment"
-
15.01.00 · Risk Category
"First-order risks can be generally broken down into risks arising from intended and unintended use, system design and implementation choices, and properties of the chosen dataset and learning components."
-
40.05.00 · Risk Category
"While it is most likely that any advanced intelligent software will be directly designed or evolved, it is also possible that we will obtain it as a complete package from some unknown source. For example, an AI could be extracted from a signal obtained in SETI (Search for Extraterrestrial Intelligence) research, which is not guaranteed to be human friendly (Carrigan Jr 2004, Turchin March 15, 2013)."
-
40.06.00 · Risk Category
"While highly rare, it is known, that occasionally individual bits may be flipped in different hardware devices due to manufacturing defects or cosmic rays hitting just the right spot (Simonite March 7, 2008). This is similar to mutations observed in living organisms and may result in a modification of an intelligent system."
-
49.02.00 · Risk Category
None provided.
-
The general tenet of AI alignment involves training generative AI systems to be harmless, helpful, and honest, ensuring their behavior aligns with and respects human values. However, a central debate in this area concerns the methodological challenges in selecting appropriate values. While AI systems can acquire human values through feedback, observation, or debate, there remains ambiguity over which individuals are qualified or legitimized to provide these guiding signals. Another prominent issue pertains to deceptive alignment, which might cause generative AI systems to tamper evaluations. A
-
08.02.00 · Risk Category
"The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement."
-
08.06.00 · Risk Category
"The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race."
-
09.03.02 · Risk Sub-Category
AGI - Effects on humans and other living beings: Existential risks
Unpredictable outcomes
"Our culture, lifestyle, and even probability of survival may change drastically. Because the intentions programmed into an artificial agent cannot be guaranteed to lead to a positive outcome, Machine Ethics becomes a topic that may not produce guaranteed results, and Safety Engineering may correspondingly degrade our ability to utilize the technology fully."
-
12.06.00 · Risk Category
"The speculative potential for future advanced AI systems to harm human civilization, either through misuse or due to challenges in aligning AI objectives with human values."
-
This is the difficulty of controlling the ML system
-
19.01.01 · Risk Sub-Category
Technological, Data and Analytical AI Risks
Loss of control of autonomous systems and unforeseen behaviour due to lack of transparency and self-programming/ reprogramming
—
-
The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these feedback loops arise from the interaction of multiple agents in a population... Within highly complex systems, the emergence of runaway processes may be hard to predict, because the conditions under which positive feedback loops occur may be non-obvious. The system of interacting AI assistants, their human principals, other humans and other algorithms will certainly be highly complex. Therefore, ther
-
34.01.00 · Risk Category
we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and finally shift our emphasis towards an examination of misaligned behaviors and dangerous capabilities
-
"Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b)."
-
35.04.00 · Risk Category
AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by default our systems may pursue simplified proxies of human values. The result could be suboptimal or even catastrophic if a sufficiently powerful AI successfully optimizes its flawed objective to an extreme degree
-
"This category encompasses nearly 17% of the articles and addresses the overall imperative of establishing a harmonious coexistence between humans and machines, and the key concerns that gives rise to this need."
-
"'Loss of control’ scenarios are potential future scenarios in which society can no longer meaningfully constrain some advanced general- purpose AI agents, even if it becomes clear they are causing harm. These scenarios are hypothesised to arise through a combination of social and technical factors, such as pressures to delegate decisions to general- purpose AI systems, and limitations of existing techniques used to influence the behaviours of general- purpose AI systems."
-
51.03.00 · Risk Category
"If we get something wrong in the design or construction of an agent, will the agent cooperate in us trying to fix it? This is called error-tolerant design by MIRI-AF and corrigibility by Soares, Fallenstein, et al. (2015). The problem is connected to safe interruptibility as considered by DeepMind."
-
54.03.00 · Risk Category
"How do we ensure AI acts according to our values? Equivalently, how do we prevent poorly-understood AI systems from advancing goals we do not endorse? Whereas HP#2 concerns the prevention of harm caused by incompetent systems, HP#3 seeks to align competent AIs with humans, through methods which ensure their behavior is compatible with the user’s intentions."
-
62.16.05 · Risk Sub-Category
General Evaluations (Inaccurate measurement of model encoded human values)
"There is a lack of robust frameworks for understanding and evaluating if the output of AI systems robustly conforms to human values, as opposed to if the systems have learned to produce outputs that are only partially correlated with them (i.e., mimicking) [13]. Additionally, outputs by AI models often do not perfectly reflect the representation of human values learned by the model, and it is not known how these values evolve and transition across different stages of model training and deployment. Such evaluations may be especially challenging with LLMs that adopt different personas with diff
-
"For LLM-agents, both the goal and environment observations are typically specified in the prompt through natural language. While natural language may provide a richer and more natural means of specifying goals than alternatives such as hand-engineering objective functions, natural language still suffers from underspecification (Grice, 1975; Piantadosi et al., 2012). Furthermore, in practice, users may neglect fully specifying their goals, especially the information pertaining to elements of the environment that ought not to be changed (the classic frame problem (Shanahan, 2016)). Such undersp
-
"As the general capabilities of LLM-empowered systems improve, the negative consequences and risks induced by these systems also get increasingly alarming accordingly, especially in high-stakes areas [28, 146]. Although they may not be intentionally introduced, severe problematic issues related to human values can be raised. Specifically, even before language models become extremely large, pre-trained language models have already exhibited a certain degree of value judgments. For example, Schramowski et al. [171] reveal the existence of the moral direction with the sentence embeddings of moral
-
"Risk to the existence of humanity."
-
67.04.03 · Risk Sub-Category
Capabilities that could be used to reduce human control - Manipulation
"There is evidence that language models tend to respond as though they share the user’s stated views, and larger models do this more than smaller ones.276 The ability to predict people’s views and generate text that they will endorse could be useful for manipulation."
-
"Accidents include unintended failure modes that, in principle, could be considered the fault of the system or the developer"
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.