MIT AI Risk Repository
Browse AI risks
662 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
"Multiple agents tend to coordinate actions through covert means to maximize common interests (possibly harming third-party interests or evading regulation), even if individual agents are designed with safety constraints, their collusive behavior may still trigger systemic risks such as market manipulation or cascading failures that are difficult to detect and mitigate, and may develop specialized communication protocols to avoid monitoring."
-
73.02.03 · Risk Sub-Category
Multi-Agent Safety Is Not Assured by Single-Agent Safety
Collusion between LLM-Agents
"While it would often be preferable for LLM-agents to be cooperative, cooperation can be undesirable if it undermines pro-social competition or produces negative externalities for coalition non-members (Dorner, 2021; Buterin, 2019; Dafoe et al., 2020). Collusion between relatively simple AI systems has been observed in the real world (Assad et al., 2020; Wieting and Sapi, 2021) and synthetic experiments (Brown and MacKay, 2023; Calvano et al., 2020; Klein, 2021) Collusion can occur through explicit or steganographic communication. Steganographic communication hides information in seemingly inn
-
74.01.00 · Risk Category
"In terms of inherent risk, LLMs could potentially reveal sensitive information from their utilized corpora for pre-training or fine-tuning, thereby raising issues of privacy leakage [37, 145, 226]. Meanwhile, it is well-known that LLMs may experi- ence hallucinations, resulting in the production of texts that are inaccurate and misleading [194]. Finally, since the values embedded in LLM-generated texts usually directly reflect the distribution of their training data, often sourced from the Internet, there exists a substantial risk that LLMs will overfit to a narrow set of human values or even
-
51.07.00 · Risk Category
"Societal consequences: AGI will have substantial legal, economic, political, and military consequences. Only the FLI agenda is broad enough to cover these issues, though many of the mentioned organizations evidently care about the issue (Brundage et al., 2018; DeepMind, 2017)."
-
52.01.00 · Risk Category
"Risks from Unreliability stem from general purpose AI models that lack reliability, robustness, transparency, corrigibility, and interpretability, making it challenging to predict and control their behaviour fully. This includes Discrimination and Stereotype Reproduction, Misinformation and Privacy Violations, and Accidents."
-
53.03.00 · Risk Category
—
-
53.04.00 · Risk Category
"Work focused at understanding indirect ways in which AI could contribute to existential threats, such as by shaping societal “turbulence”193 and other existential risk factors.194 This covers various long-term impacts on societal parameters such as science, cooperation, power, epistemics, and values:"
-
57.01.00 · Risk Category
"Physical hazards can cause physical harm to users or to the public. It may happen through the AI system endorsing or enabling behavior that causes physical harm to the user or to others."
-
57.02.00 · Risk Category
"Nonphysical hazards are unlikely to cause physical harm, but they may elicit criminal behavior and lead to other individual or societal harm."
-
"A risk may be triggered by a human, where the AI serves merely as a tool, or by the AI acting autonomously with no human intervention, or it may involve a combination of both, with the human delegating some parts of decision-making to the AI. For risks where AI is the entity, these risks are exacerbated by an increase in the AI’s level of autonomy. To manage risks involving AI as the trigger, appropriate levels of human oversight can be built-in."
-
"AI systems leaking, reproducing, generating or inferring sensitive, private, hazardous, or secured information"
-
73.04.00 · Risk Category
"A key desideratum for an LLM from a user’s perspective is ‘trustworthiness’, i.e. assurance of reliability and consistent performance, and absence of any accidental harm caused by the technology to the user.16 Providing assurance that an LLM-based system will not cause accidental harm remains a major open challenge. Harms may either occur directly due to the flawed nature of LLMs, e.g. an LLM generating toxic language or behaving inappropriately in some other ways, or may occur due to improper usage by a user, e.g. automation bias due to a user’s overreliance on LLM."
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.