MIT AI Risk Repository
Browse AI risks
494 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
61.02.12 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Challenges in perceiving, measuring, and recognizing harm
"Harm from AI often manifests subtly or over the long term, making it difficult to identify, measure, and address effectively."
-
61.02.40 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Rapid development outpacing regulation
"The fast pace of AI development may outstrip regulatory and legal frameworks."
-
61.02.41 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Resistance to international law
"AI models and systems may prove difficult to regulate or control under international law."
-
61.02.47 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Unpredictability of AI development trajectory
"The unpredictable trajectory of AI development complicates governance and risk management."
-
"Benchmark saturation refers to benchmarks reaching their evaluation ceiling. The tendency towards benchmark saturation has been demonstrated in various benchmarks [19]. When benchmarks reach or are close to saturation, they stop being effective measures for new models, as more nuanced capability gains might not be detected."
-
62.16.16 · Risk Sub-Category
Benchmark Limitations (Insufficient benchmarks for AI safety evaluation)
"Benchmarks dedicated to measuring the performance of AI systems (e.g., on programming or math tasks) are more well-developed than those for assessing safety and harms in AI systems [234]. This gap can lead to AI systems excelling in specific tasks while exhibiting harmful behaviors that go undetected. More safety-related evaluation datasets can help in identifying previously overlooked undesirable model behaviors."
-
62.16.17 · Risk Sub-Category
Benchmark Limitations (Underestimating capabilities that are not covered by benchmarks)
"A lack of test coverage by benchmarks on specific abilities of a model can obscure the model’s capabilities from both the developer and the user [160]. This can lead to a false sense of safety and trust due to a lack of understanding of the model’s limitations."
-
62.17.00 · Risk Category
-
-
62.17.01 · Risk Sub-Category
Conflicts of interest in auditor selection
"Conflicts of interest can arise if there is no independence in the auditor selection process or if the auditors are closely associated with the developer [123, 157]. In such cases, the conflict of interest can appear even if third-party evaluators are involved. In the case of external auditing, the potential candidates might be selected from a narrow group of auditors, or have conflicting financial incentives for whether to report model shortcomings publicly."
-
"Auditors may not publicly disclose risks they find, may be required to not pub- licize shortcomings, or may not receive sufficient cooperation from the relevant internal parties."
-
"Data provenance refers to tracing history of data, which includes its ownership, origin, and transformations. Without standardized and established methods for verifying where the data came from, there are no guarantees that the data is the same as the original source and has the correct usage terms."
-
"Determining who is responsible for an AI model is challenging without good documentation and governance processes."
-
"Insufficient documentation of the system that uses the model and the model’s purpose within the system in which it is used."
-
"EAI deployment could fundamentally reshape society, particularly if the speed of technological development outpaces society’s ability to adapt [103, 120]. For example, EAI systems could provide physical threats of violence and mass surveillance capabilities to back up AI-enabled authoritarianism [121]."
-
32.04.00 · Risk Category
Environmental harm, Sustainability
-
44.05.00 · Risk Category
"AI is disused (not developed or deployed) in directions that would benefit animals (and instead developments that harm or do no benefit to animals are invested in)"
-
58.09.00 · Risk Category
"Environmental - Damage to the environment directly or indirectly caused by a technology system or set of systems."
-
"Carbon emissions - Release of carbon dioxide, nitric oxide and other gases, increasing carbon emissions, exacerbating climate change, and negatively impacting local communities."
-
"Electronic waste - Electrical or electronic equipment that is waste, including all components, sub-assemblies and consumables that are part of the equipment at the time the equipment becomes waste"
-
"Natural resource depletion - Extraction of minerals, metals, rare earths, and fossil fuels that deplete natural resources and increase carbon emissions."
-
"Pollution - Actual or potential pollution to the air, ground, noise, or water caused by a technology system."
-
"Training and deploying large models require substantial energy expenditure. The trend toward developing larger models exacerbates this issue. This can lead to excessive energy usage and have a negative environmental impact."
-
"AI, and large generative models in particular, might produce increased carbon emissions and increase water usage for their training and operation."
-
"Action(s) that lead directly or indirectly to the damage or destruction of tangible property eg. buildings, possessions, vehicles, robots"
-
"Actual or potential pollution to the air, ground, noise, or water caused by a technology system"
-
"Short-term or long-term Negative effects on the natural environment"
-
15.01.00 · Risk Category
"First-order risks can be generally broken down into risks arising from intended and unintended use, system design and implementation choices, and properties of the chosen dataset and learning components."
-
"This is the risk posed by the choice of data used for training and validation."
-
40.05.00 · Risk Category
"While it is most likely that any advanced intelligent software will be directly designed or evolved, it is also possible that we will obtain it as a complete package from some unknown source. For example, an AI could be extracted from a signal obtained in SETI (Search for Extraterrestrial Intelligence) research, which is not guaranteed to be human friendly (Carrigan Jr 2004, Turchin March 15, 2013)."
-
40.08.00 · Risk Category
"Previous research has shown that utility maximizing agents are likely to fall victims to the same indulgences we frequently observe in people, such as addictions, pleasure drives (Majot and Yampolskiy 2014), self-delusions and wireheading (Yampolskiy 2014). In general, what we call mental illness in people, particularly sociopathy as demonstrated by lack of concern for others, is also likely to show up in artificial minds."
-
43.02.00 · Risk Category
"This category encompasses the evaluation of potential catastrophic consequences that might arise from the use of LLMs. "
-
59.15.00 · Risk Category
"In data-driven AI development, the annotated data set is commonly split into training, validation, and test sets, whereby it is essential that the latter is not used for development but only for evaluation. Using the test set for training manipulates the testing strategy, which is the basis of the system’s quality assurance."
-
A primary concern is the emergence of human-level or superhuman generative models, commonly referred to as AGI, and their potential existential or catastrophic risks to humanity. Connected to that, AI safety aims at avoiding deceptive or power-seeking machine behavior, model self-replication, or shutdown evasion. Ensuring controllability, human oversight, and the implementation of red teaming measures are deemed to be essential in mitigating these risks, as is the need for increased AI safety research and promoting safety cultures within AI organizations instead of fueling the AI race. Further
-
The general tenet of AI alignment involves training generative AI systems to be harmless, helpful, and honest, ensuring their behavior aligns with and respects human values. However, a central debate in this area concerns the methodological challenges in selecting appropriate values. While AI systems can acquire human values through feedback, observation, or debate, there remains ambiguity over which individuals are qualified or legitimized to provide these guiding signals. Another prominent issue pertains to deceptive alignment, which might cause generative AI systems to tamper evaluations. A
-
"The risks associated with containment, confinement, and control in the AGI development phase, and after an AGI has been developed, loss of control of an AGI."
-
08.02.00 · Risk Category
"The risks associated with AGI goal safety, including human attempts at making goals safe, as well as the AGI making its own goals safe during self-improvement."
-
08.06.00 · Risk Category
"The risks posed generally to humanity as a whole, including the dangers of unfriendly AGI, the suffering of the human race."
-
09.03.02 · Risk Sub-Category
AGI - Effects on humans and other living beings: Existential risks
Unpredictable outcomes
"Our culture, lifestyle, and even probability of survival may change drastically. Because the intentions programmed into an artificial agent cannot be guaranteed to lead to a positive outcome, Machine Ethics becomes a topic that may not produce guaranteed results, and Safety Engineering may correspondingly degrade our ability to utilize the technology fully."
-
12.06.00 · Risk Category
"The speculative potential for future advanced AI systems to harm human civilization, either through misuse or due to challenges in aligning AI objectives with human values."
-
14.03.00 · Risk Category
"The degree of automation and control describes the extent to which an AI system functions independently of human supervision and control."
-
This is the difficulty of controlling the ML system
-
19.01.01 · Risk Sub-Category
Technological, Data and Analytical AI Risks
Loss of control of autonomous systems and unforeseen behaviour due to lack of transparency and self-programming/ reprogramming
—
-
24.02.00 · Risk Category
"As we think about even more intelligent and advanced AI assistants, perhaps outperforming humans on many cognitive tasks, the question of how humans can successfully control such an assistant looms large. To achieve the goals we set for an assistant, it is possible (Shah, 2022) that the AI assistant will implement some form of consequentialist reasoning: considering many different plans, predicting their consequences and executing the plan that does best according to some metric, M. This kind of reasoning can arise because it is a broadly useful capability (e.g. planning ahead, considering mo
-
"Specification gaming (Krakovna et al., 2020) occurs when some faulty feedback is provided to the assistant in the training data (i.e. the training objective O does not fully capture what the user/designer wants the assistant to do). It is typified by the sort of behaviour that exploits loopholes in the task specification to satisfy the literal specification of a goal without achieving the intended outcome."
-
"In the problem of goal misgeneralisation (Langosco et al., 2023; Shah et al., 2022), the AI system's behaviour during out-of-distribution operation (i.e. not using input from the training data) leads it to generalise poorly about its goal while its capabilities generalise well, leading to undesired behaviour. Applied to the case of an advanced AI assistant, this means the system would not break entirely – the assistant might still competently pursue some goal, but it would not be the goal we had intended."
-
"Here, the agent develops its own internalised goal, G, which is misgeneralised and distinct from the training reward, R. The agent also develops a capability for situational awareness (Cotra, 2022): it can strategically use the information about its situation (i.e. that it is an ML model being trained using a particular training setup, e.g. RL fine-tuning with training reward, R) to its advantage. Building on these foundations, the agent realises that its optimal strategy for doing well at its own goal G is to do well on R during training and then pursue G at deployment – it is only doing wel
-
The 2010 flash crash is an example of a runaway process caused by interacting algorithms. Runaway processes are characterised by feedback loops that accelerate the process itself. Typically, these feedback loops arise from the interaction of multiple agents in a population... Within highly complex systems, the emergence of runaway processes may be hard to predict, because the conditions under which positive feedback loops occur may be non-obvious. The system of interacting AI assistants, their human principals, other humans and other algorithms will certainly be highly complex. Therefore, ther
-
34.01.00 · Risk Category
we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and finally shift our emphasis towards an examination of misaligned behaviors and dangerous capabilities
-
"AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023)."
-
35.04.00 · Risk Category
AI agents are directed by goals and objectives. Creating general-purpose objectives that capture human values could be challenging... Since goal-directed AI systems need measurable objectives, by default our systems may pursue simplified proxies of human values. The result could be suboptimal or even catastrophic if a sufficiently powerful AI successfully optimizes its flawed objective to an extreme degree
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.