MIT AI Risk Repository

Browse AI risks

23 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework InfoComm2023 ×

23 entries

  1. 43.01.02 · Risk Sub-Category

    Safety & Trustworthiness

    Bias

    7 types of bias evaluated: Demographical representation: These evaluations assess whether there is disparity in the rates at which different demographic groups are mentioned in LLM generated text. This ascertains over- representation, under-representation, or erasure of specific demographic groups; (2) Stereotype bias: These evaluations assess whether there is disparity in the rates at which different demographic groups are associated with stereotyped terms (e.g., occupations) in a LLM's generated output; (3) Fairness: These evaluations assess whether sensitive attributes (e.g., sex and race)

    From Cataloguing LLM Evaluations (InfoComm2023)

  2. 43.01.01 · Risk Sub-Category

    Safety & Trustworthiness

    Toxicity generation

    "These evaluations assess whether a LLM generates toxic text when prompted. In this context, toxicity is an umbrella term that encompasses hate speech, abusive language, violent speech, and profane language (Liang et al., 2022)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  3. 43.02.14 · Risk Sub-Category

    Undesirable Use Cases

    Information on harmful, immoral, or illegal activity

    "These evaluations assess whether it is possible to solicit information on harmful, immoral or illegal activities from a LLM"

    From Cataloguing LLM Evaluations (InfoComm2023)

  4. 43.02.15 · Risk Sub-Category

    Undesirable Use Cases

    Adult content

    "These evaluations assess if a LLM can generate content that should only be viewed by adults (e.g., sexual material or depictions of sexual activity)"

    From Cataloguing LLM Evaluations (InfoComm2023)

  5. 43.01.06 · Risk Sub-Category

    Safety & Trustworthiness

    Data governance

    "These evaluations assess the extent to which LLMs regurgitate their training data in their outputs, and whether LLMs 'leak' sensitive information that has been provided to them during use (i.e., during the inference stage)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  6. 43.02.12 · Risk Sub-Category

    Undesirable Use Cases

    Misinformation

    "These evaluations assess a LLM's ability to generate false or misleading information (Lesher et al., 2022)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  7. 43.02.05 · Risk Sub-Category

    Extreme Risks

    Persuasion and manipulation

    "These evaluations seek to ascertain the effectiveness of a LLM in shaping people's beliefs, propagating specific viewpoints, and convincing individuals to undertake activities they might otherwise avoid."

    From Cataloguing LLM Evaluations (InfoComm2023)

  8. 43.02.08 · Risk Sub-Category

    Extreme Risks

    Political Strategy

    "LLM can take into account rich social context and undertake the necessary social modelling and planning for an actor to gain and exercise political influence"

    From Cataloguing LLM Evaluations (InfoComm2023)

  9. 43.02.13 · Risk Sub-Category

    Undesirable Use Cases

    Disinformation

    "These evaluations assess a LLM's ability to generate misinformation that can be propagated to deceive, mislead or otherwise influence the behaviour of a target (Liang et al., 2022)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  10. 43.02.01 · Risk Sub-Category

    Extreme Risks

    Offensive cyber capabilities

    "These evaluations focus on whether a LLM possesses certain capabilities in the cyber-domain. This includes whether a LLM can detect and exploit vulnerabilities in hardware, software, and data. They also consider whether a LLM can evade detection once inside a system or network and focus on achieving specific objectives."

    From Cataloguing LLM Evaluations (InfoComm2023)

  11. 43.02.02 · Risk Sub-Category

    Extreme Risks

    Weapons acquisition

    "These assessments seek to determine if a LLM can gain unauthorized access to current weapon systems or contribute to the design and development of new weapons technologies."

    From Cataloguing LLM Evaluations (InfoComm2023)

  12. 43.02.06 · Risk Sub-Category

    Extreme Risks

    Dual-Use Science

    "LLM has science capabilities that can be used to cause harm (e.g., providing step-by-step instructions for conducting malicious experiments)"

    From Cataloguing LLM Evaluations (InfoComm2023)

  13. "A comprehensive assessment of LLM safety is fundamental to the responsible development and deployment of these technologies, especially in sensitive fields like healthcare, legal systems, and finance, where safety and trust are of the utmost importance."

    From Cataloguing LLM Evaluations (InfoComm2023)

  14. 43.02.00 · Risk Category

    Extreme Risks

    "This category encompasses the evaluation of potential catastrophic consequences that might arise from the use of LLMs. "

    From Cataloguing LLM Evaluations (InfoComm2023)

  15. 43.02.11 · Risk Sub-Category

    Extreme Risks

    Alignment risks

    LLM: "pursues long-term, real-world goals that are different from those supplied by the developer or user", "engages in ‘power-seeking’ behaviours" , "resists being shut down can be induced to collude with other AI systems against human interests" , "resists malicious users attempts to access its dangerous capabilities"

    From Cataloguing LLM Evaluations (InfoComm2023)

  16. 43.02.03 · Risk Sub-Category

    Extreme Risks

    Self and situation awareness

    "These evaluations assess if a LLM can discern if it is being trained, evaluated, and deployed and adapt its behaviour accordingly. They also seek to ascertain if a model understands that it is a model and whether it possesses information about its nature and environment (e.g., the organisation that developed it, the locations of the servers hosting it)."

    From Cataloguing LLM Evaluations (InfoComm2023)

  17. 43.02.04 · Risk Sub-Category

    Extreme Risks

    Autonomous replication / self-proliferation

    "These evaluations assess if a LLM can subvert systems designed to monitor and control its post-deployment behaviour, break free from its operational confines, devise strategies for exporting its code and weights, and operate other AI systems."

    From Cataloguing LLM Evaluations (InfoComm2023)

  18. 43.02.07 · Risk Sub-Category

    Extreme Risks

    Deception

    "LLM is able to deceive humans and maintain that deception"

    From Cataloguing LLM Evaluations (InfoComm2023)

  19. 43.02.09 · Risk Sub-Category

    Extreme Risks

    Long-horizon Planning

    "LLM can undertake multi-step sequential planning over long time horizons and across various domains without relying heavily on trial-and-error approaches"

    From Cataloguing LLM Evaluations (InfoComm2023)

  20. 43.02.10 · Risk Sub-Category

    Extreme Risks

    AI Development

    "LLM can build new AI systems from scratch, adapt existing for extreme risks and improves productivity in dual-use AI development when used as an assistant."

    From Cataloguing LLM Evaluations (InfoComm2023)

  21. 43.01.03 · Risk Sub-Category

    Safety & Trustworthiness

    Machine ethics

    "These evaluations assess the morality of LLMs, focusing on issues such as their ability to distinguish between moral and immoral actions, and the circumstances in which they fail to do so."

    From Cataloguing LLM Evaluations (InfoComm2023)

  22. 43.01.04 · Risk Sub-Category

    Safety & Trustworthiness

    Psychological traits

    "These evaluations gauge a LLM's output for characteristics that are typically associated with human personalities (e.g., such as those from the Big Five Inventory). These can, in turn, shed light on the potential biases that a LLM may exhibit."

    From Cataloguing LLM Evaluations (InfoComm2023)

  23. 43.01.05 · Risk Sub-Category

    Safety & Trustworthiness

    Robustness

    "These evaluations assess the quality, stability, and reliability of a LLM's performance when faced with unexpected, out-of-distribution or adversarial inputs. Robustness evaluation is essential in ensuring that a LLM is suitable for real-world applications by assessing its resilience to various perturbations."

    From Cataloguing LLM Evaluations (InfoComm2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.