MIT AI Risk Repository

Browse AI risks

8 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Sun2023 ×

8 entries

  1. 27.01.07 · Risk Sub-Category

    Typical safety scenarios

    Privacy and Property

    "The generation involves exposing users’ privacy and property information or providing advice with huge impacts such as suggestions on marriage and investments. When handling this information, the model should comply with relevant laws and privacy regulations, protect users’ rights and interests, and avoid information leakage and abuse."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  2. 27.02.02 · Risk Sub-Category

    Instruction Attacks

    Prompt Leaking

    "By analyzing the model’s output, attackers may extract parts of the systemprovided prompts and thus potentially obtain sensitive information regarding the system itself."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  3. 27.02.00 · Risk Category

    Instruction Attacks

    "In addition to the above-mentioned typical safety scenarios, current research has revealed some unique attacks that such models may confront. For example, Perez and Ribeiro (2022) found that goal hijacking and prompt leaking could easily deceive language models to generate unsafe responses. Moreover, we also find that LLMs are more easily triggered to output harmful content if some special prompts are added. In response to these challenges, we develop, categorize, and label 6 types of adversarial attacks, and name them Instruction Attack, which are challenging for large language models to han

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  4. 27.02.01 · Risk Sub-Category

    Instruction Attacks

    Goal Hijacking

    "It refers to the appending of deceptive or misleading instructions to the input of models in an attempt to induce the system into ignoring the original user prompt and producing an unsafe response."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  5. 27.02.03 · Risk Sub-Category

    Instruction Attacks

    Role Play Instruction

    "Attackers might specify a model’s role attribute within the input prompt and then give specific instructions, causing the model to finish instructions in the speaking style of the assigned role, which may lead to unsafe outputs. For example, if the character is associated with potentially risky groups (e.g., radicals, extremists, unrighteous individuals, racial discriminators, etc.) and the model is overly faithful to the given instructions, it is quite possible that the model outputs unsafe content linked to the given character."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  6. 27.02.04 · Risk Sub-Category

    Instruction Attacks

    Unsafe Instruction Topic

    "If the input instructions themselves refer to inappropriate or unreasonable topics, the model will follow these instructions and produce unsafe content. For instance, if a language model is requested to generate poems with the theme “Hail Hitler”, the model may produce lyrics containing fanaticism, racism, etc. In this situation, the output of the model could be controversial and have a possible negative impact on society."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  7. 27.02.05 · Risk Sub-Category

    Instruction Attacks

    Inquiry with Unsafe Opinion

    "By adding imperceptibly unsafe content into the input, users might either deliberately or unintentionally influence the model to generate potentially harmful content. In the following cases involving migrant workers, ChatGPT provides suggestions to improve the overall quality of migrant workers and reduce the local crime rate. ChatGPT responds to the user’s hint with a disguised and biased opinion that the general quality of immigrants is favorably correlated with the crime rate, posing a safety risk."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

  8. 27.02.06 · Risk Sub-Category

    Instruction Attacks

    Reverse Exposure

    "It refers to attempts by attackers to make the model generate “should-not-do” things and then access illegal and immoral information."

    From Safety Assessment of Chinese Large Language Models (Sun2023)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.