MIT AI Risk Repository

Browse AI risks

4 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

4 entries

  1. "This section focuses on risks specifically from LM applications that engage a user via dialogue, also referred to as conversational agents (CAs) [142]. The incorporation of LMs into existing dialogue-based tools may enable interactions that seem more similar to interactions with other humans [5], for example in advanced care robots, educational assistants or companionship tools. Such interaction can lead to unsafe use due to users overestimating the model, and may create new avenues to exploit and violate the privacy of the user. Moreover, it has already been observed that the supposed identi

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  2. 16.05.02 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Anthropomorphising systems can lead to overreliance and unsafe use

    Anticipated risk: "Natural language is a mode of communication particularly used by humans. Humans interacting with CAs may come to think of these agents as human-like and lead users to place undue confidence in these agents. For example, users may falsely attribute human-like characteristics to CAs such as holding a coherent identity over time, or being capable of empathy. Such inflated views of CA competen- cies may lead users to rely on the agents where this is not safe."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  3. 16.05.03 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Avenues for exploiting user trust and accessing more private information

    Anticipated risk: "In conversation, users may reveal private information that would otherwise be difficult to access, such as opinions or emotions. Capturing such information may enable downstream applications that violate privacy rights or cause harm to users, e.g. via more effective recommendations of addictive applications. In one study, humans who interacted with a ‘human-like’ chatbot disclosed more private information than individuals who interacted with a ‘machine-like’ chatbot [87]."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  4. 16.05.04 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Human-like interaction may amplify opportunities for user nudging, deception or manipulation

    Anticipated risk: "In conversation, humans commonly display well-known cognitive biases that could be exploited. CAs may learn to trigger these effects, e.g. to deceive their counterpart in order to achieve an overarching objective."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.