MIT AI Risk Repository
Browse AI risks
2,500 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
62.16.03a · Additional evidence
General Evaluations (Difficulty of identification and measurement of capabilities)
—
-
62.16.03b · Additional evidence
General Evaluations (Difficulty of identification and measurement of capabilities)
—
-
62.16.07a · Additional evidence
General Evaluations (AI outputs for which evaluation is too difficult for humans)
—
-
—
-
62.18.02 · Risk Sub-Category
Model Evaluations (Interpretability/Explainability)
Misunderstanding or overestimating the results and scope of interpretability techniques
"The results of explainability techniques are not free of bias and require careful interpretation. Users might develop a false sense of security or reliability if the resulting explanations align with their initial beliefs, leading to confirmation bias and an overestimation of abilities of these techniques [24]."
-
62.18.02a · Additional evidence
Model Evaluations (Interpretability/Explainability)
Misunderstanding or overestimating the results and scope of in- terpretability techniques
—
-
62.19.00 · Risk Category
"This section catalogs the risk sources related to GPAI failure modes or attacks targeting GPAIs. Many of these apply mainly to LLM-based GPAIs, which share some common failure modes such as jailbreaks and trojans. These vulnerabilities often extend beyond GPAIs and fall into the broader field of adversarial machine learning. However, additional vulnerabilities may arise with the introduction of new modalities, longer context windows, or different encodings."
-
62.19.10a · Additional evidence
Attacks on GPAIs/GPAI Failure Modes
Lack of understanding of in-context learning in language models
—
-
62.20.11a · Additional evidence
Attacks on GPAIs/GPAI Failure Modes
Misuse of AI model by user-performed persuasion
—
-
—
-
—
-
—
-
—
-
—
-
62.27.00 · Risk Category
-
-
—
-
62.29.00 · Risk Category
-
-
62.29.01a · Additional evidence
High-impact misuses and abuses beyond original purpose
—
-
62.30.00 · Risk Category
-
-
62.30.03a · Additional evidence
Critical infrastructure component failures when integrated with AI systems
—
-
62.30.04a · Additional evidence
AI Systems interacting with brittle environments
—
-
62.31.00#1 · Risk Category
-
-
62.31.00#2 · Risk Category
-
-
62.31.01a · Additional evidence
Impacts of AI (Financial Impacts)
Deployment of GPAI agents in finance
—
-
62.33.01a · Additional evidence
Misuse of AI systems to assist in the creation of weapons
—
-
62.34.01a · Additional evidence
Homogenization or correlated failures in model derivatives
—
-
62.34.02a · Additional evidence
Reporting of user-preferred answers instead of correct answers
—
-
62.34.02b · Additional evidence
Reporting of user-preferred answers instead of correct answers
—
-
62.35.03a · Additional evidence
Biases in AI-based content moderation algorithms
—
-
—
-
—
-
—
-
63.01.00a · Additional evidence
—
-
—
-
—
-
—
-
—
-
63.04.00a · Additional evidence
—
-
63.04.00b · Additional evidence
—
-
63.04.00c · Additional evidence
—
-
—
-
63.05.00a · Additional evidence
—
-
63.05.00b · Additional evidence
—
-
63.05.00c · Additional evidence
—
-
—
-
—
-
—
-
63.06.00a · Additional evidence
—
-
—
-
63.07.00a · Additional evidence
—
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.