MIT AI Risk Repository
Browse AI risks
157 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
"Historical and societal biases that are present in the data are used to train and fine-tune the model."
-
"Generated content might unfairly represent certain groups or individuals."
-
"Decision bias occurs when one group is unfairly advantaged over another due to decisions of the model. This might be caused by biases in the data and also amplified as a result of the model’s training."
-
"Toxic output occurs when the model produces hateful, abusive, and profane (HAP) or obscene content. This also includes behaviors like bullying."
-
"A model might generate language that leads to physical harm The language might include overtly violent, covertly dangerous, or otherwise indirectly unsafe statements."
-
"It is important to include the perspectives or concerns of communities that are affected by model outcomes when designing and building models. Failing to include these perspectives makes it difficult to understand the relevant context for the model and to engender trust within these communities."
-
"Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of that information."
-
"Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features available in the data."
-
65.05.02 · Risk Sub-Category
Training Data Risks (Intellectual property)
Confidential information in data
"Confidential information might be included as part of the data that is used to train or tune the model."
-
"Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model."
-
"Confidential information might be included as a part of the prompt that is sent to the model."
-
"Copyrighted information or other intellectual property might be included as a part of the prompt that is sent to the model."
-
65.16.02 · Risk Sub-Category
Output risks (Intellectual Property)
Revealing confidential information
"When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type of data leakage."
-
"When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing personal information is a type of data leakage."
-
"A type of adversarial attack where an adversary or malicious insider injects intentionally corrupted, false, misleading, or incorrect samples into the training or fine-tuning datasets."
-
"A prompt injection attack forces a generative model that takes a prompt as input to produce unexpected output by manipulating the structure, instructions, or information contained in its prompt."
-
"An attribute inference attack is used to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when an adversary has some prior knowledge about the training data and uses that knowledge to infer the sensitive data."
-
"Evasion attacks attempt to make a model output incorrect results by slightly perturbing the input data that is sent to the trained model."
-
"A prompt leak attack attempts to extract a model's system prompt (also known as the system message)."
-
"A jailbreaking attack attempts to break through the guardrails that are established in the model to perform restricted actions."
-
"Because generative models tend to produce output like the input provided, the model can be prompted to reveal specific kinds of information. For example, adding personal information in the prompt increases its likelihood of generating similar kinds of personal information in its output. If personal data was included as part of the model’s training, there is a possibility it could be revealed."
-
"A membership inference attack repeatedly queries a model to determine whether a given input was part of the model’s training. More specifically, given a trained model and a data sample, an attacker samples the input space, observing outputs to deduce whether that sample was part of the model's training."
-
"An attribute inference attack repeatedly queries a model to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when an adversary has some prior knowledge about the training data and uses that knowledge to infer the sensitive data."
-
"Models might generate code that causes harm or unintentionally affects other systems."
-
"Hallucinations generate factually inaccurate or untruthful content with respect to the model’s training data or input. This is also sometimes referred to lack of faithfulness or lack of groundedness."
-
"Generative AI models might be used intentionally to generate hateful, abusive, and profane (HAP) or obscene content."
-
"Generative AI models might be used with the sole intention of harming people."
-
"Generative AI models might be used to intentionally create misleading or false information to deceive or influence a targeted audience."
-
"Generative AI models might be intentionally used to imitate people through deepfakes by using video, images, audio, or other modalities without their consent."
-
"Easy access to high-quality generative models might result in students that use AI models to plagiarize existing work intentionally or unintentionally."
-
65.23.05 · Risk Sub-Category
Non-technical risks (Societal impact)
Impact on education: bypassing learning
"Easy access to high-quality generative models might result in students that use AI models to bypass the learning process."
-
"Improper usage occurs when a model is used for a purpose that it was not originally designed for."
-
"In AI-assisted decision-making tasks, reliance measures how much a person trusts (and potentially acts on) a model’s output. Over-reliance occurs when a person puts too much trust in a model, accepting a model’s output when the model’s output is likely incorrect. Under-reliance is the opposite, where the person doesn’t trust the model but should."
-
"AI might affect the individuals’ ability to make choices and act independently in their best interests."
-
"Widespread adoption of foundation model-based AI systems might lead to people's job loss as their work is automated if they are not reskilled."
-
"When workers who train AI models such as ghost workers are not provided with adequate working conditions, fair compensation, and good health care benefits that also include mental health."
-
"A model might generate content that is similar or identical to existing work protected by copyright or covered by open-source license agreement."
-
65.21.03 · Risk Sub-Category
Non-technical risks (legal compliance)
Generated content ownership and IP
"Legal uncertainty about the ownership and intellectual property rights of AI-generated content."
-
"AI systems might overly represent certain cultures that result in a homogenization of culture and thoughts."
-
"Without accurate documentation on how a model's data was collected, curated, and used to train a model, it might be harder to satisfactorily explain the behavior of the model with respect to the data."
-
"Data provenance refers to tracing history of data, which includes its ownership, origin, and transformations. Without standardized and established methods for verifying where the data came from, there are no guarantees that the data is the same as the original source and has the correct usage terms."
-
"Determining who is responsible for an AI model is challenging without good documentation and governance processes."
-
"Insufficient documentation of the system that uses the model and the model’s purpose within the system in which it is used."
-
"Testing is unrepresentative when the test inputs are mismatched with the inputs that are expected during deployment."
-
"Since foundation models can be used for many purposes, a model’s intended use is important for defining the relevant risks of that model. As the use changes, the relevant risks might correspondingly change."
-
"Lack of data transparency is due to insufficient documentation of training or tuning dataset details. "
-
"A metric selected to measure or track a risk is incorrectly selected, incompletely measuring the risk, or measuring the wrong risk for the given context."
-
"AI model risks are socio-technical, so their testing needs input from a broad set of disciplines and diverse testing practices."
-
"AI, and large generative models in particular, might produce increased carbon emissions and increase water usage for their training and operation."
-
"Laws and other restrictions can limit or prohibit the use of some data for specific AI use cases."
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.