MIT AI Risk Repository

Browse AI risks

19 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Maas2023 ×

19 entries

  1. 53.01.01 · Risk Sub-Category

    Alignment failures in existing ML systems

    Faulty reward functions in the wild

  2. 53.01.02 · Risk Sub-Category

    Alignment failures in existing ML systems

    Specification gaming

  3. 53.01.03 · Risk Sub-Category

    Alignment failures in existing ML systems

    Reward model overoptimization

  4. 53.01.04 · Risk Sub-Category

    Alignment failures in existing ML systems

    Instrumental convergence

  5. 53.01.05 · Risk Sub-Category

    Alignment failures in existing ML systems

    Goal misgeneralization

  6. 53.01.06 · Risk Sub-Category

    Alignment failures in existing ML systems

    Inner misalignment

  7. 53.01.07 · Risk Sub-Category

    Alignment failures in existing ML systems

    Language model misalignment

  8. 53.02.03 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Acquisition of goals to seek power and control

    "cases where AI systems converge on optimal policies of seeking power over their environment;135"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  9. 53.03.01 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of misaligned superintelligence or power-seeking AI

  10. 53.03.03 · Risk Sub-Category

    Direct catastrophe from AI

    Extreme “suffering risks” because of a misaligned system

  11. 53.03.04 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of conflict between AI systems and multi-system interactions

  12. 53.01.08 · Risk Sub-Category

    Alignment failures in existing ML systems

    Harms from increasingly agentic algorithmic systems

  13. 53.02.01 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Situational awareness

    "cases where a large language model displays awareness that it is a model, and it can recognize whether it is currently in testing or deployment;"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  14. 53.02.04 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Self-improvement

    "examples of cases where AI systems improve AI systems"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  15. 53.02.05 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Autonomous replication

    "the ability of simple software to autonomously spread around the internet in spite of countermeasures (various software worms and computer viruses)"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  16. 53.02.06 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Anonymous resource acquisition

    "The demonstrated ability of anonymous actors to accumulate resources online (e.g., Satoshi Nakamoto as an anonymous crypto billionaire)"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  17. 53.02.07 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Deception

    "Cases of AI systems deceiving humans to carry out tasks or meet goals.139"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.