MIT AI Risk Repository

Browse AI risks

18 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework IBM2025 ×

18 entries

  1. 65.03.01 · Risk Sub-Category

    Training Data Risks (Privacy)

    Personal information in data

    "Inclusion or presence of personal identifiable information (PII) and sensitive personal information (SPI) in the data used for training or fine tuning the model might result in unwanted disclosure of that information."

    From AI Risk Atlas (IBM2025)

  2. 65.03.03 · Risk Sub-Category

    Training Data Risks (Privacy)

    Reidentification

    "Even with the removal or personal identifiable information (PII) and sensitive personal information (SPI) from data, it might be possible to identify persons due to correlations to other features available in the data."

    From AI Risk Atlas (IBM2025)

  3. 65.05.02 · Risk Sub-Category

    Training Data Risks (Intellectual property)

    Confidential information in data

    "Confidential information might be included as part of the data that is used to train or tune the model."

    From AI Risk Atlas (IBM2025)

  4. 65.11.03 · Risk Sub-Category

    Inference risks (Privacy)

    Personal information in prompt

    "Personal information or sensitive personal information that is included as a part of a prompt that is sent to the model."

    From AI Risk Atlas (IBM2025)

  5. 65.12.01 · Risk Sub-Category

    Inference risks (Intellectual property)

    Confidential data in prompt

    "Confidential information might be included as a part of the prompt that is sent to the model."

    From AI Risk Atlas (IBM2025)

  6. 65.12.02 · Risk Sub-Category

    Inference risks (Intellectual property)

    IP information in prompt

    "Copyrighted information or other intellectual property might be included as a part of the prompt that is sent to the model."

    From AI Risk Atlas (IBM2025)

  7. 65.16.02 · Risk Sub-Category

    Output risks (Intellectual Property)

    Revealing confidential information

    "When confidential information is used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing confidential information is a type of data leakage."

    From AI Risk Atlas (IBM2025)

  8. 65.20.01 · Risk Sub-Category

    Output risks (Privacy)

    Exposing personal information

    "When personal identifiable information (PII) or sensitive personal information (SPI) are used in training data, fine-tuning data, or as part of the prompt, models might reveal that data in the generated output. Revealing personal information is a type of data leakage."

    From AI Risk Atlas (IBM2025)

  9. 65.08.01 · Risk Sub-Category

    Training Data Risks (Robustness)

    Data poisoning

    "A type of adversarial attack where an adversary or malicious insider injects intentionally corrupted, false, misleading, or incorrect samples into the training or fine-tuning datasets."

    From AI Risk Atlas (IBM2025)

  10. 65.09.01 · Risk Sub-Category

    Inference risks (Robustness)

    Prompt injection attack

    "A prompt injection attack forces a generative model that takes a prompt as input to produce unexpected output by manipulating the structure, instructions, or information contained in its prompt."

    From AI Risk Atlas (IBM2025)

  11. 65.09.02 · Risk Sub-Category

    Inference risks (Robustness)

    Extraction attack

    "An attribute inference attack is used to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when an adversary has some prior knowledge about the training data and uses that knowledge to infer the sensitive data."

    From AI Risk Atlas (IBM2025)

  12. 65.09.03 · Risk Sub-Category

    Inference risks (Robustness)

    Evasion attack

    "Evasion attacks attempt to make a model output incorrect results by slightly perturbing the input data that is sent to the trained model."

    From AI Risk Atlas (IBM2025)

  13. 65.09.04 · Risk Sub-Category

    Inference risks (Robustness)

    Prompt leaking

    "A prompt leak attack attempts to extract a model's system prompt (also known as the system message)."

    From AI Risk Atlas (IBM2025)

  14. 65.10.01 · Risk Sub-Category

    Inference risks (Multi-category)

    Jailbreaking

    "A jailbreaking attack attempts to break through the guardrails that are established in the model to perform restricted actions."

    From AI Risk Atlas (IBM2025)

  15. 65.10.02 · Risk Sub-Category

    Inference risks (Multi-category)

    Prompt priming

    "Because generative models tend to produce output like the input provided, the model can be prompted to reveal specific kinds of information. For example, adding personal information in the prompt increases its likelihood of generating similar kinds of personal information in its output. If personal data was included as part of the model’s training, there is a possibility it could be revealed."

    From AI Risk Atlas (IBM2025)

  16. 65.11.01 · Risk Sub-Category

    Inference risks (Privacy)

    Membership inference attack

    "A membership inference attack repeatedly queries a model to determine whether a given input was part of the model’s training. More specifically, given a trained model and a data sample, an attacker samples the input space, observing outputs to deduce whether that sample was part of the model's training."

    From AI Risk Atlas (IBM2025)

  17. 65.11.02 · Risk Sub-Category

    Inference risks (Privacy)

    Attribute inference attack

    "An attribute inference attack repeatedly queries a model to detect whether certain sensitive features can be inferred about individuals who participated in training a model. These attacks occur when an adversary has some prior knowledge about the training data and uses that knowledge to infer the sensitive data."

    From AI Risk Atlas (IBM2025)

  18. 65.15.02 · Risk Sub-Category

    Output risks (Value alignment)

    Harmful code generation

    "Models might generate code that causes harm or unintentionally affects other systems."

    From AI Risk Atlas (IBM2025)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.