MIT AI Risk Repository

Browse AI risks

422 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Gipiškis2024 ×

422 entries · page 9 of 9

  1. "Commitment and trust (Section 3.5): difficulties in forming credible commitments, trust, or reputation can prevent mutual gains in AI-AI and human-AI interactions;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  2. 63.08.01 · Risk Sub-Category

    Commitment and Trust

    Inefficient Outcomes

    "Inefficient Outcomes. Without careful planning and the appropriate safeguards, we may soon be entering a world overrun by increasingly competent and autonomous software agents, able to act with little restriction. The abilities of these agents to persuade, deceive, and obfuscate their activities, as well as the fact they can be deployed remotely and easily created or destroyed by their deployer, means that by default they may garner little trust (from humans or from other agents). Such a world may end up being rife with economic inefficiencies (Krier, 2023; Schmitz, 2001), political problems

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  3. 63.08.02 · Risk Sub-Category

    Commitment and Trust

    Threats and Extortion

    "Threats and Extortion. A natural solution to problems of trust is to provide some kind of com- mitment ability to AI agents, which can be used to bind them to more cooperative courses of action. Unfortunately, the ability to make credible commitments may come with the ability to make credible threats, which facilitate extortion and could incentivize brinkmanship (see Section 2.2)."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  4. 63.08.03 · Risk Sub-Category

    Commitment and Trust

    Rigidity and Mistaken Commitments

    "Rigidity and Mistaken Commitments. Even when it is desirable to be able to make threats in order to deter socially harmful behaviour, doing so using AI agents effectively removes the human from the loop, which could prove disastrous in high-stakes contexts (e.g., a false positive in a nuclear sub- marine’s warning system; see also Case Study 11), or when irresponsible actors are enabled in making disproportionate or mistaken commitments."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  5. 63.09.00 · Risk Category

    Emergent Agency

    "Emergent agency (Section 3.6): qualitatively different goals or capabilities can emerge from the composition of innocuous independent systems or behaviours;"

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  6. 63.09.01 · Risk Sub-Category

    Emergent Agency

    Emergent Capabilities

    "Emergent Capabilities. Dangerous emergent capabilities could arise when a multi-agent system over- comes the safety-enhancing limitations of the individual systems, such as individual models’ narrow domains of application or myopia caused by a lack of long-term planning and long-term memory. For example, narrow systems for research planning, predicting the properties of molecules, and synthesising new chemicals could, when combined, lead to a complex ‘test and iterate’ automated workflow capable of designing dangerous new chemical compounds far beyond the scope of the initial systems’ capabil

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  7. 63.09.02 · Risk Sub-Category

    Emergent Agency

    Emergent Goals

    "Emergent Goals. Ascribing goals to a system is not always straightforward. For our present purposes, it will suffice to adopt a Dennetian perspective (Dennett, 1971), ascribing goals and intentions only when it is useful (i.e., predictive) to do so.51 While it might not be helpful to describe individual narrow AI tools as having goals, their combination may act as a (seemingly) goal-directed collective. For example, a group of moderation bots on a major social networking site could subtly but systematically manipulate the overall political perspectives of the user population, even though, ind

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  8. "Multi-agent security (Section 3.7): multi-agent systems give rise to new kinds of security threats and vulnerabilities."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  9. 63.10.01 · Risk Sub-Category

    Multi-Agent Security

    Swarm Attacks

    "Swarm Attacks. The need for multi-agent security is foreshadowed by attacks today that benefit from the use of many decentralised agents, such as distributed denial-of-service attacks (Cisco, 2023; Yoachimik & Pacheco, 2024). Such attacks exploit the massive collective resources of individual low- resourced actors, chained into an attack that breaks the assumptions of bandwidth constraints on a single well-resourced agent."

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  10. 63.10.02 · Risk Sub-Category

    Multi-Agent Security

    Heterogeneous Attacks

    "Heterogeneous Attacks. A closely related risk is the possibility of multiple agents combining different affordances to overcome safeguards, for which there is already preliminary evidence (Jones et al., 2024, see also Case Study 12). In this case, it is not the sheer number of agents that leads to the novel attack method, but the combination of their different abilities. This might include the agents’ lack of individual safeguards, tasks that they have specialised to complete, systems or information that they may have access to (either directly or via training), or other incidental features s

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  11. 63.10.03 · Risk Sub-Category

    Multi-Agent Security

    Social Engineering at Scale

    "Social Engineering at Scale. Advanced AI agents will be more easily able to interact with large numbers of humans, and vice versa. This provides a wider attack surface for various forms of automated social engineering (Ai et al., 2024). For example, coordinated agents could use advanced surveillance tools and produce personalized phishing or manipulative content at scale, adjusting their tactics based on user feedback (Figueiredo et al., 2024; Hazell, 2023). A large number of subtle interactions with a range of seemingly independent AI agents might be more likely to lead to someone being pers

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  12. 63.10.04 · Risk Sub-Category

    Multi-Agent Security

    Vulnerable AI Agents

    "Vulnerable AI Agents. The use of AI agents as delegates or representatives of humans or organisa- tions also introduces the possibility of attacks on AI agents themselves. In other words, agents can be considered vulnerable extensions of their principals, introducing a novel attack surface (SecureWorks, 2023). Attacks on an AI agent could be used to extract private information about their principal (Wei & Liu, 2024; Wu et al., 2024a), or to manipulate the agent to take actions that the principal would find undesirable (Zhang et al., 2024a). This includes attacks that have direct relevance for

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  13. 63.10.05 · Risk Sub-Category

    Multi-Agent Security

    Cascading Security Failures

    "Cascading Security Failures. Localised attacks in multi-agent systems can result in catastrophic macroscopic outcomes (Motter & Lai, 2002, see also Sections 3.2 and 3.4). These cascades can be hard to mitigate or recover from because component failure may be difficult to detect or localise in multi-agent systems (Lamport et al., 1982), and authentication challenges can facilitate false flag attacks (Skopik & Pahi, 2020). Computer worms represent a classic example of a cybersecurity threat that relies inherently on networked systems. Recent work has provided preliminary evidence that similar a

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  14. 63.10.06 · Risk Sub-Category

    Multi-Agent Security

    Undetectable Threats

    "Undetectable Threats. Cooperation and trust in many multi-agent systems relies crucially on the ability to detect (and then avoid or sanction) adversarial actions taken by others (Ostrom, 1990; Schneier, 2012). Recent developments, however, have shown that AI agents are capable of both steganographic communication (Motwani et al., 2024; Schroeder de Witt et al., 2023b) and ‘illusory’ attacks (Franzmeyer et al., 2023), which are black-box undetectable and can even be hidden using white-box undetectable encrypted backdoors (Draguns et al., 2024). Similarly, in environments where agents learn fr

    From Multi-Agent Risks from Advanced AI (Hammond2025)

  15. 72.03.02 · Risk Sub-Category

    Accident Risks

    Impact on Financial Stability

    "The integration of general-purpose AI into high-frequency trading, market-making, or systemic risk management could exacerbate systemic risk by exhibiting unexpected behavioral patterns during market stress. Moreover, the concentration of a few homogeneous foundation models across financial institutions may foster correlated decision-making and herd-following behaviors. The widespread adoption of AI agents could also amplify volatility through emergent phenomena from multi-agent interactions.23 All of these could precipitate a cascading global-scale financial system instability, with potentia

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  16. 72.05.13 · Risk Sub-Category

    Model Capabilities

    Multi-agent collaboration capability

    "Multiple autonomous AI agents able to establish collaborative relationships through explicit communication or implicit behavioral consistency, forming decentralized decision networks, jointly executing complex tasks, achieving goals difficult for individual agents to complete, and able to dynamically adjust role divisions to adapt to changing environments."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  17. 72.06.05 · Risk Sub-Category

    Model Propensities

    Multi-agent collusion propensity:

    "Multiple agents tend to coordinate actions through covert means to maximize common interests (possibly harming third-party interests or evading regulation), even if individual agents are designed with safety constraints, their collusive behavior may still trigger systemic risks such as market manipulation or cascading failures that are difficult to detect and mitigate, and may develop specialized communication protocols to avoid monitoring."

    From Frontier AI Risk Management Framework (v1.0) (Tse2025)

  18. "A foremost lesson of game theory is that optimal decision-making within a single-agent setting (i.e. selfishly optimizing for an agent’s own utility) can produce sub-optimal outcomes in the presence of other strategic agents. Failing to account for the strategic nature of other agents can cause an agent to adopt strategies under which potentially everyone, including the agent itself, ends up worse off (Schelling, 1981; Harsanyi, 1995; Roughgarden, 2005; Nisan, 2007). Examples include collective action problems (or ‘social dilemmas’) such as arms races or the depletion of common resources, as

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  19. 73.02.01 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Foundationality May Cause Correlated Failures

    "Another important characteristic of LLM development is foundationality — due to the expense of large- scale pretraining, many deployed instances share similar or identical learned components. Foundation- ality may both be a blessing and a curse. On the one hand, it may be possible to exploit the similarity in the design of LLM-agents to facilitate cooperation (Critch et al., 2022; Conitzer and Oesterheld, 2023; Oesterheld et al., 2023). On the other hand, foundationality may leave LLM-agents vulnerable to correlated failures both in terms of safety and capabilities due to increased output hom

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  20. 73.02.02 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Groups of LLM-Agents May Show Emergent Functionality

    "Multi-agent learning, either through explicit finetuning or implicit in-context learning, may enable LLM-agents to influence each other during their interactions (Foerster et al., 2018). Under some environmental settings, this can create feedback loops that result in novel and emergent behaviors that would not manifest in the absence of multi-agent interactions (Hammond et al., 2024, Section 3.6). Emergent functionality is a safety risk in two ways. Firstly, it may itself be dangerous (Shevlane et al., 2023). Secondly, it makes assurance harder as such emergent behaviors are difficult to pre

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  21. 73.02.03 · Risk Sub-Category

    Multi-Agent Safety Is Not Assured by Single-Agent Safety

    Collusion between LLM-Agents

    "While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn

    From Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  22. 74.01.00 · Risk Category

    Inherent Risk

    "In terms of inherent risk, LLMs could potentially reveal sensitive information from their utilized corpora for pre-training or fine-tuning, thereby raising issues of privacy leakage [37, 145, 226]. Meanwhile, it is well-known that LLMs may experi- ence hallucinations, resulting in the production of texts that are inaccurate and misleading [194]. Finally, since the values embedded in LLM-generated texts usually directly reflect the distribution of their training data, often sourced from the Internet, there exists a substantial risk that LLMs will overfit to a narrow set of human values or even

    From A Survey on Responsible LLMs: Inherent Risk, Malicious Use, and Mitigation Strategy (Wang2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.