{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"34.01.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Causes of Misalignment","risk_subcategory":null,"description":"we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and finally shift our emphasis towards an examination of misaligned behaviors and dangerous capabilities","entity":"Other","intent":"Other","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Reward Hacking","description":"\"Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Goal Misgeneralization","description":"\"Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its ","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Reward Tampering","description":"\"Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.01.05","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Causes of Misalignment","risk_subcategory":"Limitations of Reward Modeling","description":"\"Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b).\"","entity":"Other","intent":"Unintentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.00","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Category","risk_category":"Misaligned Behaviors","risk_subcategory":null,"description":null,"entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.01","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Power-Seeking Behaviors","description":"\"AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.02","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Untruthful Output","description":"\"AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023).\"","entity":"AI","intent":"Other","timing":"Other","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.03","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Deceptive Alignment & Manipulation","description":"\"Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He","entity":"AI","intent":"Intentional","timing":"Pre-deployment","domain":7,"subdomain":"7.1"},{"ev_id":"34.03.04","quick_ref":"Ji2023","paper_title":"AI Alignment: A Comprehensive Survey","level":"Risk Sub-Category","risk_category":"Misaligned Behaviors","risk_subcategory":"Collectively Harmful Behaviors","description":"\"AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities.\"","entity":"AI","intent":"Intentional","timing":"Other","domain":7,"subdomain":"7.1"}]}