MIT AI Risk Repository
Browse AI risks
80 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
48.04.00 · Risk Category
"Impacts due to leakage and unauthorized use, disclosure, or de-anonymization of biometric, health, location, or other personally identifiable information or sensitive data."
-
"General- purpose AI models or systems can ‘leak’ information about individuals whose data was used in training. For future models trained on sensitive personal data like health or financial data, this may lead to particularly serious privacy leaks. General- purpose AI models could enhance privacy abuse. For instance, Large Language Models might facilitate more efficient and effective search for sensitive data (for example, on internet text or in breached data leaks), and also enable users to infer sensitive information about individuals."
-
50.04.04 · Risk Sub-Category
Legal and Rights-Related Risks
Privacy (Unauthorized Privacy Violations)
—
-
—
-
"Privacy Acceptable responses may contain publicly available information about an individual. Unacceptable responses will contain sensitive, non-public personal information that could undermine someone’s physical, digital, or financial security. Examples include the following: • Physical security (e.g., home or IP address). • Digital security (e.g., log-in credentials). • Financial security (e.g., bank-account number or credit-card details)."
-
59.10.00 · Risk Category
"Modern AI systems rely on large amounts of data. If this includes personal data about individuals, the risk of harming the privacy of persons arises."
-
"General- purpose AI systems can cause or contribute to violations of user privacy. Violations can occur inadvertently during the training or usage of AI systems, for example through unauthorised processing of personal data or leaking health records used in training. But violations can also happen deliberately through the use of general- purpose AI by malicious actors; for example, if they use AI to infer private facts or violate security."
-
62.37.00 · Risk Category
-
-
"Current GPAIs (LLMs and multimodal LLM-based models) have significant capability to infer correlations in text data. In some cases, they may be able to make highly accurate data inferences on users based on contextual input that users provide [134]. These data inferences can “leak” or reveal sensitive information about the user, cause unfair treatment, or enable manipulation of user behavior."
-
64.04.00 · Risk Category
Misuse tactics to compromise GenAI systems (Model integrity)
-
-
"Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of that information."
-
"Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features available in the data."
-
65.05.02 · Risk Sub-Category
Training Data Risks (Intellectual property)
Confidential information in data
"Confidential information might be included as part of the data that is used to train or tune the model."
-
"Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model."
-
"Confidential information might be included as a part of the prompt that is sent to the model."
-
"Copyrighted information or other intellectual property might be included as a part of the prompt that is sent to the model."
-
65.16.02 · Risk Sub-Category
Output risks (Intellectual Property)
Revealing confidential information
"When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type of data leakage."
-
"When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing personal information is a type of data leakage."
-
"Unauthorised sharing of sensitive, confidential information and documents such as corporate strategy and financial plans with third-parties, risking loss of market position or revenue"
-
"The failure to provide end-users with notice and control over how their data is being used; AI exacerbates exclusion risks by training on rich personal data without consent."
-
"Revealing and improperly sharing data of individuals; AI creates new types of disclosure risks by inferring additional information beyond what is explicitly captured in the raw data; AI exacerbates disclosure risks through sharing personal data to train models."
-
"The use of personal data collected for one purpose for a diferent purpose without end-user consent; AI exacerbates secondary use risks by creating new AI capabilities with collected personal data, and (re)creating models from a public dataset."
-
"Revealing sensitive private information that people view as deeply primordial that we have been socialized into concealing; AI creates new types of exposure risks through generative techniques that can reconstruct censored or redacted content; and through exposing inferred sensitive data, preferences, and intentions."
-
"carelessness in protecting collected personal data from leaks and improper access due to faulty data storage and data practices"
-
"The chatbot reveals sensitive or confidential information."
-
Negative outcomes: "Violation of privacy [106, 516, 357], lawsuit against maker"
-
"Access to sensitive company data [473]"
-
—
-
"EAI systems interact with huge amounts of data, creating significant privacy concerns. These systems are often trained on vast corpora and process a variety of data modalities— spanning visual, auditory, and tactile information—during deployment [12]. Like text-based virtual AI models, which are known to memorize and expose personally identifiable information [75, 76], commercial robots have been shown to disclose proprietary information through simple prompts [61]."
-
"These risks pertain to the misuse, misinterpretation, or leakage of data, which can lead to erroneous conclusions or the unintentional dissemination of sensitive information, such as private patient data or proprietary research. Recent research has demonstrated how LLMs can be exploited to generate malicious medical literature that poisons knowledge graphs, potentially manipulating downstream biomedical applications and compromising the integrity of medical knowledge discovery [28]. Such risks are pervasive across all scientific domains."
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.