MIT AI Risk Repository
Browse AI risks
2,500 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
73.02.03 · Risk Sub-Category
Multi-Agent Safety Is Not Assured by Single-Agent Safety
Collusion between LLM-Agents
"While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn
-
74.01.00 · Risk Category
"In terms of inherent risk, LLMs could potentially reveal sensitive information from their utilized corpora for pre-training or fine-tuning, thereby raising issues of privacy leakage [37, 145, 226]. Meanwhile, it is well-known that LLMs may experi- ence hallucinations, resulting in the production of texts that are inaccurate and misleading [194]. Finally, since the values embedded in LLM-generated texts usually directly reflect the distribution of their training data, often sourced from the Internet, there exists a substantial risk that LLMs will overfit to a narrow set of human values or even
-
01.01.00.a · Additional evidence
—
-
01.01.00.b · Additional evidence
—
-
01.01.00.c · Additional evidence
—
-
01.02.00.a · Additional evidence
—
-
01.03.00.a · Additional evidence
—
-
01.03.00.b · Additional evidence
—
-
01.03.00.c · Additional evidence
—
-
01.03.00.d · Additional evidence
—
-
01.03.00.e · Additional evidence
—
-
01.03.00.f · Additional evidence
—
-
01.04.00.a · Additional evidence
—
-
01.04.00.b · Additional evidence
—
-
N/A
-
N/A
-
N/A
-
N/A
-
N/A
-
N/A
-
05.14.00 · Risk Category
Closely related to other clusters like AI safety, fairness, or harmful content, papers stress the importance of evaluating generative AI systems both in a narrow technical way as well as in a broader sociotechnical impact assessment focusing on pre-release audits as well as post-deployment monitoring. Ideally, these evaluations should be conducted by independent third parties. In terms of technical LLM or text-to-image model audits, papers furthermore criticize a lack of safety benchmarking for languages other than English.
-
05.19.00 · Risk Category
While the scoping review identified distinct topic clusters within the literature, it also revealed certain issues that either do not fit into these categories, are discussed infrequently, or in a nonspecific manner. For instance, some papers touch upon concepts like trustworthiness, accountability, or responsibility, but often remain vague about what they entail in detail. Similarly, a few papers vaguely attribute socio-political instability or polarization to generative AI without delving into specifics. Apart from that, another minor topic area concerns responsible approaches of talking abo
-
09.01.00 · Risk Category
Domain-specific AI - Effects on humans and other living beings: Existential Risks
—
-
09.02.00 · Risk Category
Domain-specific AI - Effects on humans and other living beings: Non-existential risks
—
-
—
-
—
-
09.05.00 · Risk Category
—
-
09.06.00 · Risk Category
—
-
[not defined in text]
-
10.08.00 · Risk Category
[not defined in text]
-
13.01.00 · Risk Category
"What can be evaluated in a technical system and its components'...The following categories are high-level, non-exhaustive, and present a synthesis of the findings across different modalities"
-
13.01.02.a · Additional evidence
Impacts: The Technical Base System
Cultural Values and Sensitive Content
—
-
13.02.00 · Risk Category
"what can be evaluated among people and society"
-
—
-
—
-
—
-
13.02.02.a · Additional evidence
Inequality, Marginalization, and Violence
—
-
13.02.02.b · Additional evidence
Inequality, Marginalization, and Violence
—
-
13.02.02.c · Additional evidence
Inequality, Marginalization, and Violence
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.