{"attribution":{"source":"MIT AI Risk Repository, Domain Taxonomy of AI Risks v1 (MIT AI Risk Initiative)","license":"CC BY 4.0","license_url":"https://creativecommons.org/licenses/by/4.0/","citation":"Slattery, P., Saeri, A. K., Grundy, E. A. C., Graham, J., Noetel, M., Uuk, R., Dao, J., Pour, S., Casper, S., & Thompson, N. (2025). The AI Risk Repository: A comprehensive meta-review, database, and taxonomy of risks from artificial intelligence. arXiv:2408.12622."},"exported_at":"2026-09-11"}
{"rows":[{"ev_id":"73.02.00","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":null,"description":"\"A foremost lesson of game theory is that optimal decision-making within a single-agent setting (i.e. selfishly optimizing for an agent’s own utility) can produce sub-optimal outcomes in the presence of other strategic agents. Failing to account for the strategic nature of other agents can cause an agent to adopt strategies under which potentially everyone, including the agent itself, ends up worse off (Schelling, 1981; Harsanyi, 1995; Roughgarden, 2005; Nisan, 2007). Examples include collective action problems (or ‘social dilemmas’) such as arms races or the depletion of common resources, as ","entity":"Other","intent":"Other","timing":"Other","domain":7,"subdomain":"7.6"},{"ev_id":"73.02.01","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":"Foundationality May Cause Correlated Failures","description":"\"Another important characteristic of LLM development is foundationality — due to the expense of large- scale pretraining, many deployed instances share similar or identical learned components. Foundation- ality may both be a blessing and a curse. On the one hand, it may be possible to exploit the similarity in the design of LLM-agents to facilitate cooperation (Critch et al., 2022; Conitzer and Oesterheld, 2023; Oesterheld et al., 2023). On the other hand, foundationality may leave LLM-agents vulnerable to correlated failures both in terms of safety and capabilities due to increased output hom","entity":"Other","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"73.02.02","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":"Groups of LLM-Agents May Show Emergent Functionality","description":"\"Multi-agent learning, either through explicit finetuning or implicit in-context learning, may enable LLM-agents to influence each other during their interactions (Foerster et al., 2018). Under some environmental settings, this can create feedback loops that result in novel and emergent behaviors that would not manifest in the absence of multi-agent interactions (Hammond et al., 2024, Section 3.6).  Emergent functionality is a safety risk in two ways. Firstly, it may itself be dangerous (Shevlane et al., 2023). Secondly, it makes assurance harder as such emergent behaviors are difficult to pre","entity":"Other","intent":"Other","timing":"Post-deployment","domain":7,"subdomain":"7.6"},{"ev_id":"73.02.03","quick_ref":"Anwar2024","paper_title":"Foundational Challenges in Assuring Alignment and Safety of Large Language Models","level":"Risk Sub-Category","risk_category":"Multi-Agent Safety Is Not Assured by Single-Agent Safety","risk_subcategory":"Collusion between LLM-Agents","description":"\"While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn","entity":"AI","intent":"Intentional","timing":"Post-deployment","domain":7,"subdomain":"7.6"}]}