MIT AI Risk Repository
Browse AI risks
237 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
—
-
"Terms of service, licenses, or other rules restrict the use of certain models."
-
—
-
—
-
"Damage to core political and economic institutions and the effective delivery of government services caused by the use or misuse of a technology system or set of systems"
-
"Harms affecting the functioning of societies, communities and economies caused directly or indirectly by the use or misuse of a technology system or set of systems"
-
"Impairment of the psychological mental health and wellbeing of an individual, group or organisation due to the use of misuse of a technology system or set of systems"
-
"Online behaviour such as sexual harassment that makes an individual or group feel alarmed or threatened"
-
"Distress, possibly severe and lasting, as a result of use or misuse of a generative system"
-
"Damage to the financial interests of an individual or group, or to the strategic, operational, legal or financial interests of a business due to the use of misuse of a technology system or set of systems"
-
"Damage, disruption or destruction of a business system and/or its components"
-
"Loss of ability to take advantage of a financial or other opportunity, such as education, immigration, employability/securing a job"
-
"Use or misuse of a technology system in a manner that compromises fundamental human rights and freedoms"
-
70.01.00 · Risk Category
-
-
70.02.00 · Risk Category
-
-
70.04.00 · Risk Category
-
-
71.01.00 · Risk Category
-
-
"Uncontrolled AI self- improvement; Quantum security"
-
71.02.00 · Risk Category
"Whether the risk originates from malicious intent or is an unintended consequence of legitimate task objectives"
-
71.03.00 · Risk Category
-
-
"Damage to individual well-being or public health"
-
"Dramatically change the social and economic status"
-
72.06.00 · Risk Category
-
-
73.07.03 · Risk Sub-Category
Jailbreaks and Prompt Injections Threaten Security of LLMs
Adversarial Optimization:
"Jailbreak attacks can be discovered by performing manual or auto- mated adversarial optimization against a proxy objective that is noisily correlated with the success of a jailbreak. These are mostly gradient-based attacks (Zou et al., 2023b; Shin et al., 2020) as described in the previous two challenges, but gradient-free methods also exist (Prasad et al., 2022; Deng et al., 2022; Lapid et al., 2023)."
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.