{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"22.04.00","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":null,"description":"\"speculative technical mechanisms that might lead to rogue AIs and how a loss of control could bring about catastrophe\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.01","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Proxy Gaming","description":"\"One way we might lose control of an AI agent’s actions is if it engages in behavior known as “proxy gaming.” It is often difficult to specify and measure the exact goal that we want a system to pursue. Instead, we give the system an approximate—“proxy”—goal that is more measurable and seems likely to correlate with the intended goal. However, AI systems often find loopholes by which they can easily achieve the proxy goal, but completely fail to achieve the ideal goal. If an AI “games” its proxy goal in a way that does not reflect our values, then we might not be able to reliably steer its beh","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.02","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Goal Drift","description":"\"Even if we successfully control early AIs and direct them to promote human values, future AIs could end up with different goals that humans would not endorse. This process, termed “goal drift,” can be hard to predict or control. This section is most cutting-edge and the most speculative, and in it we will discuss how goals shift in various agents and groups and explore the possibility of this phenomenon occurring in AIs. We will also examine a mechanism that could lead to unexpected goal drift, called intrinsification, and discuss how goal drift in AIs could be catastrophic.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.03","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Power Seeking","description":"\"even if an agent started working to achieve an unintended goal, this would not necessarily be a problem, as long as we had enough power to prevent any harmful actions it wanted to attempt. Therefore, another important way in which we might lose control of AIs is if they start trying to obtain more power, potentially transcending our own.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"22.04.04","quick_ref":"Hendrycks2023","paper_title":"An Overview of Catastrophic AI Risks","level":"Risk Sub-Category","risk_category":"Rogue AIs (Internal)","risk_subcategory":"Deception","description":"\"it is plausible that AIs could learn to deceive us. They might, for example, pretend to be acting as we want them to, but then take a “treacherous turn” when we stop monitoring them, or when they have enough power to evade our attempts to interfere with them. \"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"}]}