MIT AI Risk Repository

Browse AI risks

4 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

4 entries

  1. "A foremost lesson of game theory is that optimal decision-making within a single-agent setting (i.e. selfishly optimizing for an agent’s own utility) can produce sub-optimal outcomes in the presence of other strategic agents. Failing to account for the strategic nature of other agents can cause an agent to adopt strategies under which potentially everyone, including the agent itself, ends up worse off (Schelling, 1981; Harsanyi, 1995; Roughgarden, 2005; Nisan, 2007). Examples include collective action problems (or ‘social dilemmas’) such as arms races or the depletion of common resources, as

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  2. 73.02.01 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Foundationality May Cause Correlated Failures

    "Another important characteristic of LLM development is foundationality — due to the expense of large- scale pretraining, many deployed instances share similar or identical learned components. Foundation- ality may both be a blessing and a curse. On the one hand, it may be possible to exploit the similarity in the design of LLM-agents to facilitate cooperation (Critch et al., 2022; Conitzer and Oesterheld, 2023; Oesterheld et al., 2023). On the other hand, foundationality may leave LLM-agents vulnerable to correlated failures both in terms of safety and capabilities due to increased output hom

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  3. 73.02.02 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Groups of LLM-Agents May Show Emergent Functionality

    "Multi-agent learning, either through explicit finetuning or implicit in-context learning, may enable LLM-agents to influence each other during their interactions (Foerster et al., 2018). Under some environmental settings, this can create feedback loops that result in novel and emergent behaviors that would not manifest in the absence of multi-agent interactions (Hammond et al., 2024, Section 3.6). Emergent functionality is a safety risk in two ways. Firstly, it may itself be dangerous (Shevlane et al., 2023). Secondly, it makes assurance harder as such emergent behaviors are difficult to pre

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  4. 73.02.03 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Collusion between LLM-Agents

    "While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.