MIT AI Risk Repository
Browse AI risks
2,500 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
60.03.00 · Risk Category
"This section considers a range of systemic risks, in the sense of “broader societal risks associated with AI deployment, beyond the capabilities of individual models” (636). Note that this is not identical with how the European AI Act uses ‘systemic risks’ to refer to general - purpose AI models with a high impact on society, based on criteria such as training compute and the number of users."
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
61.01.00 · Risk Category
-
-
"The large-scale erosion or violation of fundamental human rights and freedoms."
-
61.02.00 · Risk Category
-
-
61.02.16 · Risk Sub-Category
Sources of systemic risks from general-purpose AI
Conflicting objectives in design
"Designers and operators of AI may face conflicting objectives that compromise safety."
-
62.01.00 · Risk Category
—
-
"Risks can be realized by intentional or unintentional actions, and in some cases the intent is difficult to establish. To manage these risks, rigorous evaluations and red teaming can be performed, guardrails can be put in place, and model release can be gradual, such that AI model malfunctions have either low likeli- hood or low probability of occurrence. To prevent intentional misuse, acceptable use policies can be in place, and for riskier models Know Your Customer (KYC) measures can also be implemented by model providers."
-
"Risks can be realized by intentional or unintentional actions, and in some cases the intent is difficult to establish. To manage these risks, rigorous evaluations and red teaming can be performed, guardrails can be put in place, and model release can be gradual, such that AI model malfunctions have either low likeli- hood or low probability of occurrence. To prevent intentional misuse, acceptable use policies can be in place, and for riskier models Know Your Customer (KYC) measures can also be implemented by model providers."
-
"Risks can be realized by intentional or unintentional actions, and in some cases the intent is difficult to establish. To manage these risks, rigorous evaluations and red teaming can be performed, guardrails can be put in place, and model release can be gradual, such that AI model malfunctions have either low likeli- hood or low probability of occurrence. To prevent intentional misuse, acceptable use policies can be in place, and for riskier models Know Your Customer (KYC) measures can also be implemented by model providers."
-
62.02.00 · Risk Category
—
-
"A risk may be triggered by a human, where the AI serves merely as a tool, or by the AI acting autonomously with no human intervention, or it may involve a combination of both, with the human delegating some parts of decision-making to the AI. For risks where AI is the entity, these risks are exacerbated by an increase in the AI’s level of autonomy. To manage risks involving AI as the trigger, appropriate levels of human oversight can be built-in."
-
"A risk may be triggered by a human, where the AI serves merely as a tool, or by the AI acting autonomously with no human intervention, or it may involve a combination of both, with the human delegating some parts of decision-making to the AI. For risks where AI is the entity, these risks are exacerbated by an increase in the AI’s level of autonomy. To manage risks involving AI as the trigger, appropriate levels of human oversight can be built-in."
-
"A risk may be triggered by a human, where the AI serves merely as a tool, or by the AI acting autonomously with no human intervention, or it may involve a combination of both, with the human delegating some parts of decision-making to the AI. For risks where AI is the entity, these risks are exacerbated by an increase in the AI’s level of autonomy. To manage risks involving AI as the trigger, appropriate levels of human oversight can be built-in."
-
62.03.00 · Risk Category
—
-
"In the context of Normal Accident Theory [150], normal accidents are those that “could no longer be ascribed to isolated equipment malfunction, operator error, or acts of God.” We refer to these as “system failures” (to be distinguished from “systemic risks”), while the opposite would be “isolated failures.” For isolated failures, harms are consistent with the underlying failure modes. For example, an AI capable of producing false or misleading content would constitute risks re- lated to misinformation and disinformation. Whereas for system failures, harms are not consistent with the underlyi
-
"In the context of Normal Accident Theory [150], normal accidents are those that “could no longer be ascribed to isolated equipment malfunction, operator error, or acts of God.” We refer to these as “system failures” (to be distinguished from “systemic risks”), while the opposite would be “isolated failures.” For isolated failures, harms are consistent with the underlying failure modes. For example, an AI capable of producing false or misleading content would constitute risks re- lated to misinformation and disinformation. Whereas for system failures, harms are not consistent with the underlyi
-
—
-
"As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."
-
62.04.01 · Risk Sub-Category
Dimension - Technical Attributes (AI inadequacy - technical failure)
Supervised/unsupervised AI (AI data quality related - biased training data)
"As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."
-
62.04.02 · Risk Sub-Category
Dimension - Technical Attributes (AI inadequacy - technical failure)
Supervised/unsupervised AI (AI training performance related - Robustness)
"As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."
-
62.04.03 · Risk Sub-Category
Dimension - Technical Attributes (AI inadequacy - technical failure)
Supervised/unsupervised AI (AI training performance related - Accuracy)
"As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."
-
62.04.04 · Risk Sub-Category
Dimension - Technical Attributes (AI inadequacy - technical failure)
Supervised/unsupervised AI (AI training performance related - Reliability)
"As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."
-
62.04.05 · Risk Sub-Category
Dimension - Technical Attributes (AI inadequacy - technical failure)
Reinforcement learning AI (Training design related)
"As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."
-
62.04.06 · Risk Sub-Category
Dimension - Technical Attributes (AI inadequacy - technical failure)
Reinforcement learning AI (Training performance related)
"As above, there are broadly two dimensions of technical failure modes: quality of data or input signal, and training performance. Due to a lack of transparency, it may be difficult to ascertain the type of technical failure that gives rise to a particular risk, and it is often a combination of several factors. Risks pertain- ing to AI failures are exacerbated by poor quality training data and imperfect training signals. Various measures can be implemented to improve the quality of the training data, and fine-tuning techniques can be used to disincentivize harmful model behavior."
-
62.05.00 · Risk Category
"An example of AI capabilities is that an AI might be capable of developing novel bioweapons. Whereas an example of AI inadequacy is a self-driving car causing an accident due to not being able to recognize certain objects. The boundary between capabilities and inadequacy is sometimes blurred. For exam- ple, when an AI generates falsehoods, it could be framed as either a capability of developing fiction, or an inadequacy in generating truthful content."
-
"Inherent capabilities are inherent to the AI, whether they are deliberately trained or have emerged unintentionally."
-
—
-
"Extrinsic capabilities, on the other hand, are acquired through the use of external tools, such as LLM plugins."
-
62.06.00 · Risk Category
—
-
"For GPAIs or foundation models, risks emerge during training, prior to being repurposed and deployed in more specific AI systems or applications. Risk assessments can be conducted before deployment, and monitoring of AI models can occur as required throughout the deployment phase. In certain cases, version updates or model recalls may be warranted post-deployment."
-
"For GPAIs or foundation models, risks emerge during training, prior to being repurposed and deployed in more specific AI systems or applications. Risk assessments can be conducted before deployment, and monitoring of AI models can occur as required throughout the deployment phase. In certain cases, version updates or model recalls may be warranted post-deployment."
-
62.07.00 · Risk Category
"For “system and operational harms,” the AI systems interact with other systems and industries, where a failure in an AI system could lead to failures of a wider scope."
-
62.09.00 · Risk Category
"These are in contrast with “societal harms,” which are less direct but have more far-reaching effects on segments of society"
-
—
-
62.10.00 · Risk Category
"Finally, “legal and rights-related harms” concern either harms from illegal activities or harms from violations of human rights."
-
62.10.03 · Risk Sub-Category
Direct Harm Domains (legal and rights-related harms)
Criminal activities
—
-
62.14.00 · Risk Category
-
-
—
-
62.16.00 · Risk Category
"This section catalogs the risk sources and risk management measures related to model evaluations (often called evals). We categorize them into the fol- lowing groups: general evaluations, benchmarking, red teaming, auditing, and interpretability/explainability. The subsection on general evaluations consists of items that are common to various evaluation techniques, while the other subsections are specific to their respective evaluation types."
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.