MIT AI Risk Repository

Browse AI risks

12 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

12 entries

  1. 53.01.01 · Risk Sub-Category

    Alignment failures in existing ML systems

    Faulty reward functions in the wild

  2. 53.01.02 · Risk Sub-Category

    Alignment failures in existing ML systems

    Specification gaming

  3. 53.01.03 · Risk Sub-Category

    Alignment failures in existing ML systems

    Reward model overoptimization

  4. 53.01.04 · Risk Sub-Category

    Alignment failures in existing ML systems

    Instrumental convergence

  5. 53.01.05 · Risk Sub-Category

    Alignment failures in existing ML systems

    Goal misgeneralization

  6. 53.01.06 · Risk Sub-Category

    Alignment failures in existing ML systems

    Inner misalignment

  7. 53.01.07 · Risk Sub-Category

    Alignment failures in existing ML systems

    Language model misalignment

  8. 53.02.03 · Risk Sub-Category

    Dangerous capabilities in AI systems

    Acquisition of goals to seek power and control

    "cases where AI systems converge on optimal policies of seeking power over their environment;135"

    From Advancing AI Governance: A Literature Review of Problems, Options, and Proposals (Maas2023)

  9. 53.03.01 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of misaligned superintelligence or power-seeking AI

  10. 53.03.03 · Risk Sub-Category

    Direct catastrophe from AI

    Extreme “suffering risks” because of a misaligned system

  11. 53.03.04 · Risk Sub-Category

    Direct catastrophe from AI

    Existential disaster because of conflict between AI systems and multi-system interactions

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.