MIT AI Risk Repository
Browse AI risks
19 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
53.01.00 · Risk Category
-
-
53.01.01 · Risk Sub-Category
Alignment failures in existing ML systems
Faulty reward functions in the wild
-
-
-
-
53.01.03 · Risk Sub-Category
Alignment failures in existing ML systems
Reward model overoptimization
-
-
-
-
-
-
-
-
-
-
53.02.03 · Risk Sub-Category
Dangerous capabilities in AI systems
Acquisition of goals to seek power and control
"cases where AI systems converge on optimal policies of seeking power over their environment;135"
-
53.03.01 · Risk Sub-Category
Existential disaster because of misaligned superintelligence or power-seeking AI
-
-
53.03.03 · Risk Sub-Category
Extreme “suffering risks” because of a misaligned system
-
-
53.03.04 · Risk Sub-Category
Existential disaster because of conflict between AI systems and multi-system interactions
-
-
53.01.08 · Risk Sub-Category
Alignment failures in existing ML systems
Harms from increasingly agentic algorithmic systems
-
-
53.02.00 · Risk Category
-
-
"cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;"
-
"examples of cases where AI systems improve AI systems"
-
"the ability of simple software to autonomously spread around the internet in spite of countermeasures (various software worms and computer viruses)"
-
"The demonstrated ability of anonymous actors to accumulate resources online (e.g., Satoshi Nakamoto as an anonymous crypto billionaire)"
-
"Cases of AI systems deceiving humans to carry out tasks or meet goals.139"
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.