{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-12"}
{"rows":[{"ev_id":"34.02.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Double edge components","risk_subcategory":null,"description":"\"Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detailed exposition of specific misaligned behaviors (•) and introduces what we term double edge components (+). These components are designed to enhance the capability of AI systems in handling real-world settings but also potentially exacerbate misalignment issues. It should be noted that some of these double edge components (+) remain speculative. Nevertheless, it is imperative to discuss their pote","entity":"AI","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Situational Awareness","description":"\"AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreactions of the world (including humans) to its actions (Cotra, 2022). ...However, suchknowledge also paves the way for advanced methods of reward hacking, heightened deception/manipulationskills, and an increased propensity to chase instrumental subgoals (Ngo et al., 2024).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Broadly-Scoped Goals","description":"\"Advanced AI systems are expected to develop objectives that span long timeframes,deal with complex tasks, and operate in open-ended settings (Ngo et al., 2024). ...However, it can also bring about the risk of encouraging manipulatingbehaviors (e.g., AI systems may take some bad actions to achieve human happiness, such as persuadingthem to do high-pressure jobs (Jacob Steinhardt, 2023)).\"","entity":"Human","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Mesa-Optimization Objectives","description":"\"The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives specified by the training signals, and optimization for these misaligned goals may leadto systems out of control (Hubinger et al., 2019c).\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.2"},{"ev_id":"34.02.04","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Double edge components","risk_subcategory":"Access to Increased Resources","description":"\"Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false information, deceive users, disrupt network security, and, in more dire scenarios, be compromised by malicious actors for ill purposes. Moreover, their increased access to data and resources can facilitate self-proliferation, posing existential risks (Shevlane et al., 2023).\"","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.2"}]}