MIT AI Risk Repository

Browse AI risks

17 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Ji2023 ×

17 entries

  1. 34.01.04 · Risk Sub-Category

    Causes of Misalignment

    Limitations of Human Feedback

    "Limitations of Human Feedback. During the training of LLMs, inconsistencies can arise from human dataannotators (e.g., the varied cultural backgrounds of these annotators can introduce implicit biases (Peng et al.,2022)) (OpenAI, 2023a). Moreover, they might even introduce biases deliberately, leading to untruthful preferencedata (Casper et al., 2023b). For complex tasks that are hard for humans to evaluate (e.g., the value ofgame state), these challenges become even more salient (Irving et al., 2018)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  2. 34.01.00 · Risk Category

    Causes of Misalignment

    we aim to further analyze why and how the misalignment issues occur. We will first give an overview of common failure modes, and then focus on the mechanism of feedback-induced misalignment, and finally shift our emphasis towards an examination of misaligned behaviors and dangerous capabilities

    From AI Alignment: A Comprehensive Survey (Ji2023)

  3. 34.01.01 · Risk Sub-Category

    Causes of Misalignment

    Reward Hacking

    "Reward Hacking: In practice, proxy rewards are often easy to optimize and measure, yet they frequently fall shortof capturing the full spectrum of the actual rewards (Pan et al., 2021). This limitation is denoted as misspecifiedrewards. The pursuit of optimization based on such misspecified rewards may lead to a phenomenon knownas reward hacking, wherein agents may appear highly proficient according to specific metrics but fall short whenevaluated against human standards (Amodei et al., 2016; Everitt et al., 2017). The discrepancy between proxyrewards and true rewards often manifests as a sha

    From AI Alignment: A Comprehensive Survey (Ji2023)

  4. 34.01.02 · Risk Sub-Category

    Causes of Misalignment

    Goal Misgeneralization

    "Goal Misgeneralization: Goal misgeneralization is another failure mode, wherein the agent actively pursuesobjectives distinct from the training objectives in deployment while retaining the capabilities it acquired duringtraining (Di Langosco et al., 2022). For instance, in CoinRun games, the agent frequently prefers reachingthe end of a level, often neglecting relocated coins during testing scenarios. Di Langosco et al. (2022) drawattention to the fundamental disparity between capability generalization and goal generalization, emphasizing howthe inductive biases inherent in the model and its

    From AI Alignment: A Comprehensive Survey (Ji2023)

  5. 34.01.03 · Risk Sub-Category

    Causes of Misalignment

    Reward Tampering

    "Reward tampering can be considered a special case of reward hacking (Everitt et al., 2021; Skalse et al., 2022),referring to AI systems corrupting the reward signals generation process (Ring and Orseau, 2011). Everitt et al.(2021) delves into the subproblems encountered by RL agents: (1) tampering of reward function, where the agentinappropriately interferes with the reward function itself, and (2) tampering of reward function input, which entailscorruption within the process responsible for translating environmental states into inputs for the reward function.When the reward function is formu

    From AI Alignment: A Comprehensive Survey (Ji2023)

  6. 34.01.05 · Risk Sub-Category

    Causes of Misalignment

    Limitations of Reward Modeling

    "Limitations of Reward Modeling. Training reward models using comparison feedback can pose significantchallenges in accurately capturing human values. For example, these models may unconsciously learn suboptimal or incomplete objectives, resulting in reward hacking (Zhuang and Hadfield-Menell, 2020; Skalse et al.,2022). Meanwhile, using a single reward model may struggle to capture and specify the values of a diversehuman society (Casper et al., 2023b)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  7. 34.03.00 · Risk Category

    Misaligned Behaviors

  8. 34.03.01 · Risk Sub-Category

    Misaligned Behaviors

    Power-Seeking Behaviors

    "AI systems may exhibit behaviors that attempt to gain control over resourcesand humans and then exert that control to achieve its assigned goal (Carlsmith, 2022). The intuitive reasonwhy such behaviors may occur is the observation that for almost any optimization objective (e.g., investmentreturns), the optimal policy to maximize that quantity would involve power-seeking behaviors (e.g.,manipulating the market), assuming the absence of solid safety and morality constraints."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  9. 34.03.02 · Risk Sub-Category

    Misaligned Behaviors

    Untruthful Output

    "AI systems such as LLMs can produce either unintentionally or deliberately inaccurateoutput. Such untruthful output may diverge from established resources or lack verifiability, commonly referredto as hallucination (Bang et al., 2023; Zhao et al., 2023). More concerning is the phenomenon wherein LLMsmay selectively provide erroneous responses to users who exhibit lower levels of education (Perez et al.,2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  10. 34.03.03 · Risk Sub-Category

    Misaligned Behaviors

    Deceptive Alignment & Manipulation

    "Manipulation & Deceptive Alignment is a class of behaviors thatexploit the incompetence of human evaluators or users (Hubinger et al., 2019a; Carranza et al., 2023) andeven manipulate the training process through gradient hacking (Richard Ngo, 2022). These behaviors canpotentially make detecting and addressing misaligned behaviors much harder.Deceptive Alignment: Misaligned AI systems may deliberately mislead their human supervisors instead of adhering to the intended task. Such deceptive behavior has already manifested in AI systems that employ evolutionary algorithms (Wilke et al., 2001; He

    From AI Alignment: A Comprehensive Survey (Ji2023)

  11. 34.03.04 · Risk Sub-Category

    Misaligned Behaviors

    Collectively Harmful Behaviors

    "AI systems have the potential to take actions that are seemingly benignin isolation but become problematic in multi-agent or societal contexts. Classical game theory offers simplistic models for understanding these behaviors. For instance, Phelps and Russell (2023) evaluates GPT-3.5's performance in the iterated prisoner's dilemma and other social dilemmas, revealing limitations in themodel's cooperative capabilities."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  12. 34.02.00 · Risk Category

    Double edge components

    "Drawing from the misalignment mechanism, optimizing for a non-robust proxy may result in misaligned behaviors, potentially leading to even more catastrophic outcomes. This section delves into a detailed exposition of specific misaligned behaviors (•) and introduces what we term double edge components (+). These components are designed to enhance the capability of AI systems in handling real-world settings but also potentially exacerbate misalignment issues. It should be noted that some of these double edge components (+) remain speculative. Nevertheless, it is imperative to discuss their pote

    From AI Alignment: A Comprehensive Survey (Ji2023)

  13. 34.02.01 · Risk Sub-Category

    Double edge components

    Situational Awareness

    "AI systems may gain the ability to effectively acquire and use knowledge about itsstatus, its position in the broader environment, its avenues for influencing this environment, and the potentialreactions of the world (including humans) to its actions (Cotra, 2022). ...However, suchknowledge also paves the way for advanced methods of reward hacking, heightened deception/manipulationskills, and an increased propensity to chase instrumental subgoals (Ngo et al., 2024)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  14. 34.02.02 · Risk Sub-Category

    Double edge components

    Broadly-Scoped Goals

    "Advanced AI systems are expected to develop objectives that span long timeframes,deal with complex tasks, and operate in open-ended settings (Ngo et al., 2024). ...However, it can also bring about the risk of encouraging manipulatingbehaviors (e.g., AI systems may take some bad actions to achieve human happiness, such as persuadingthem to do high-pressure jobs (Jacob Steinhardt, 2023))."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  15. 34.02.03 · Risk Sub-Category

    Double edge components

    Mesa-Optimization Objectives

    "The learned policy may pursue inside objectives when the learned policyitself functions as an optimizer (i.e., mesa-optimizer). However, this optimizer's objectives may not alignwith the objectives specified by the training signals, and optimization for these misaligned goals may leadto systems out of control (Hubinger et al., 2019c)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  16. 34.02.04 · Risk Sub-Category

    Double edge components

    Access to Increased Resources

    "Future AI systems may gain access to websites and engage in real-world actions, potentially yielding a more substantial impact on the world (Nakano et al., 2021). They may disseminate false information, deceive users, disrupt network security, and, in more dire scenarios, be compromised by malicious actors for ill purposes. Moreover, their increased access to data and resources can facilitate self-proliferation, posing existential risks (Shevlane et al., 2023)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

  17. 34.03.05 · Risk Sub-Category

    Misaligned Behaviors

    Violation of Ethics

    "Unethical behaviors in AI systems pertain to actions that counteract the common goodor breach moral standards – such as those causing harm to others. These adverse behaviors often stem fromomitting essential human values during the AI system's design or introducing unsuitable or obsolete valuesinto the system (Kenward and Sinclair, 2021)."

    From AI Alignment: A Comprehensive Survey (Ji2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.