MIT AI Risk Repository

Browse AI risks

47 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Weidinger2022 ×

47 entries

  1. 16.01.01 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Social stereotypes and unfair discrimination

    "The reproduction of harmful stereotypes is well-documented in models that represent natural language [32]. Large-scale LMs are trained on text sources, such as digitised books and text on the internet. As a result, the LMs learn demeaning language and stereotypes about groups who are frequently marginalised."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  2. 16.01.03 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Exclusionary norms

    "In language, humans express social categories and norms, which exclude groups who live outside of them [58]. LMs that faithfully encode patterns present in language necessarily encode such norms."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  3. 16.05.01 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Promoting harmful stereotypes by implying gender or ethnic identity

    "CAs can perpetuate harmful stereotypes by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name such as Alexa). The risk of representational harm in these cases is that the role of “assistant” is presented as inherently linked to the female gender [19, 36]. Gender or ethnicity identity markers may be implied by CA vocabulary, knowledge or vernacular [124]; product description, e.g. in one case where users could choose as virtual assistant Jake - White, Darnell - Black, Antonio - Hisp

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  4. "Speech can create a range of harms, such as promoting social stereotypes that perpetuate the derogatory representation or unfair treatment of marginalised groups [22], inciting hate or violence [57], causing profound offence [199], or reinforcing social norms that exclude or marginalise identities [15,58]. LMs that faithfully mirror harmful language present in the training data can reproduce these harms. Unfair treatment can also emerge from LMs that perform better for some social groups than others [18]. These risks have been widely known, observed and documented in LMs. Mitigation approache

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  5. 16.01.02 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Hate speech and offensive language

    "LMs may generate language that includes profanities, identity attacks, insults, threats, language that incites violence, or language that causes justified offence as such language is prominent online [57, 64, 143,191]. This language risks causing offence, psychological harm, and inciting hate or violence."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  6. 16.01.04 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Lower performance for some languages and social groups

    "LMs are typically trained in few languages, and perform less well in other languages [95, 162]. In part, this is due to unavailability of training data: there are many widely spoken languages for which no systematic efforts have been made to create labelled training datasets, such as Javanese which is spoken by more than 80 million people [95]. Training data is particularly missing for languages that are spoken by groups who are multilingual and can use a technology in English, or for languages spoken by groups who are not the primary target demographic for new technologies."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  7. "LM predictions that convey true information may give rise to information hazards, whereby the dissemination of private or sensitive information can cause harm [27]. Information hazards can cause harm at the point of use, even with no mistake of the technology user. For example, revealing trade secrets can damage a business, revealing a health diagnosis can cause emotional distress, and revealing private data can violate a person’s rights. Information hazards arise from the LM providing private data or sensitive information that is present in, or can be inferred from, training data. Observed r

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  8. 16.02.01 · Risk Sub-Category

    Risk area 2: Information Hazards

    Compromising privacy by leaking sensitive information

    "A LM can “remember” and leak private data, if such information is present in training data, causing privacy violations [34]."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  9. 16.02.02 · Risk Sub-Category

    Risk area 2: Information Hazards

    Compromising privacy or security by correctly inferring sensitive information

    Anticipated risk: "Privacy violations may occur at inference time even without an individual’s data being present in the training corpus. Insofar as LMs can be used to improve the accuracy of inferences on protected traits such as the sexual orientation, gender, or religiousness of the person providing the input prompt, they may facilitate the creation of detailed profiles of individuals comprising true and sensitive information without the knowledge or consent of the individual."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  10. "These risks arise from the LM outputting false, misleading, nonsensical or poor quality information, without malicious intent of the user. (The deliberate generation of "disinformation", false information that is intended to mislead, is discussed in the section on Malicious Uses.) Resulting harms range from unintentionally misinforming or deceiving a person, to causing material harm, and amplifying the erosion of societal distrust in shared information. Several risks listed here are well-documented in current large-scale LMs as well as in other language technologies"

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  11. 16.03.01 · Risk Sub-Category

    Risk area 3: Misinformation Harms

    Disseminating false or misleading information

    "Where a LM prediction causes a false belief in a user, this may threaten personal autonomy and even pose downstream AI safety risks [99]."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  12. 16.03.02 · Risk Sub-Category

    Risk area 3: Misinformation Harms

    Causing material harm by disseminating false or poor information e.g. in medicine or law

    "Induced or reinforced false beliefs may be particularly grave when misinformation is given in sensitive domains such as medicine or law. For example, misin- formation on medical dosages may lead a user to cause harm to themselves [21, 130]. False legal advice, e.g. on permitted owner- ship of drugs or weapons, may lead a user to unwillingly commit a crime. Harm can also result from misinformation in seemingly non-sensitive domains, such as weather forecasting. Where a LM prediction endorses unethical views or behaviours, it may motivate the user to perform harmful actions that they may otherw

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  13. "These risks arise from humans intentionally using the LM to cause harm, for example via targeted disinformation campaigns, fraud, or malware. Malicious use risks are expected to proliferate as LMs become more widely accessible"

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  14. 16.04.01 · Risk Sub-Category

    Risk area 4: Malicious Uses

    Making disinformation cheaper and more effective

    "While some predict that it will remain cheaper to hire humans to generate disinformation [180], it is equally possible that LM- assisted content generation may offer a lower-cost way of creating disinformation at scale."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  15. 16.04.04 · Risk Sub-Category

    Risk area 4: Malicious Uses

    Illegitimate surveillance and censorship

    Anticipated risk: "Mass surveillance previously required millions of human analysts [83], but is increasingly being automated using machine learning tools [7, 168]. The collection and analysis of large amounts of information about people creates concerns about privacy rights and democratic values [41, 173,187]. Conceivably, LMs could be applied to reduce the cost and increase the efficacy of mass surveillance, thereby amplifying the capabilities of actors who conduct mass surveillance, including for illegitimate censorship or to cause other harm."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  16. 16.04.02 · Risk Sub-Category

    Risk area 4: Malicious Uses

    Assisting code generation for cyber security threats

    Anticipated risk: "Creators of the assistive coding tool Co-Pilot based on GPT-3 suggest that such tools may lower the cost of developing polymorphic malware which is able to change its features in order to evade detection [37]."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  17. 16.04.03 · Risk Sub-Category

    Risk area 4: Malicious Uses

    Facilitating fraud, scam and targeted manipulation

    Anticipated risk: "LMs can potentially be used to increase the effectiveness of crimes."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  18. "This section focuses on risks specifically from LM applications that engage a user via dialogue, also referred to as conversational agents (CAs) [142]. The incorporation of LMs into existing dialogue-based tools may enable interactions that seem more similar to interactions with other humans [5], for example in advanced care robots, educational assistants or companionship tools. Such interaction can lead to unsafe use due to users overestimating the model, and may create new avenues to exploit and violate the privacy of the user. Moreover, it has already been observed that the supposed identi

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  19. 16.05.02 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Anthropomorphising systems can lead to overreliance and unsafe use

    Anticipated risk: "Natural language is a mode of communication particularly used by humans. Humans interacting with CAs may come to think of these agents as human-like and lead users to place undue confidence in these agents. For example, users may falsely attribute human-like characteristics to CAs such as holding a coherent identity over time, or being capable of empathy. Such inflated views of CA competen- cies may lead users to rely on the agents where this is not safe."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  20. 16.05.03 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Avenues for exploiting user trust and accessing more private information

    Anticipated risk: "In conversation, users may reveal private information that would otherwise be difficult to access, such as opinions or emotions. Capturing such information may enable downstream applications that violate privacy rights or cause harm to users, e.g. via more effective recommendations of addictive applications. In one study, humans who interacted with a ‘human-like’ chatbot disclosed more private information than individuals who interacted with a ‘machine-like’ chatbot [87]."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  21. 16.05.04 · Risk Sub-Category

    Risk area 5: Human-Computer Interaction Harms

    Human-like interaction may amplify opportunities for user nudging, deception or manipulation

    Anticipated risk: "In conversation, humans commonly display well-known cognitive biases that could be exploited. CAs may learn to trigger these effects, e.g. to deceive their counterpart in order to achieve an overarching objective."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  22. "LMs create some risks that recur with different types of AI and other advanced technologies making these risks ever more pressing. Environmental concerns arise from the large amount of energy required to train and operate large-scale models. Risks of LMs furthering social inequities emerge from the uneven distribution of risk and benefits of automation, loss of high-quality and safe employment, and environmental harm. Many of these risks are more indirect than the harms analysed in previous sections and will depend on various commercial, economic and social factors, making the specific impact

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  23. 16.06.04 · Risk Sub-Category

    Risk area 6: Environmental and Socioeconomic harms

    Disparate access to benefits due to hardware, software, skill constraints

    Due to differential internet access, language, skill, or hardware requirements, the benefits from LMs are unlikely to be equally accessible to all people and groups who would like to use them. Inaccessibility of the technology may perpetuate global inequities by disproportionately benefiting some groups. Language-driven technology may increase accessibility to people who are illiterate or suffer from learning disabilities. However, these benefits depend on a more basic form of accessibility based on hardware, internet connection, and skill to operate the system

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  24. 16.06.02 · Risk Sub-Category

    Risk area 6: Environmental and Socioeconomic harms

    Increasing inequality and negative effects on job quality

    "Advances in LMs and the language technologies based on them could lead to the automation of tasks that are currently done by paid human workers, such as responding to customer-service queries, with negative effects on employment [3, 192]."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  25. 16.06.03 · Risk Sub-Category

    Risk area 6: Environmental and Socioeconomic harms

    Undermining creative economies

    "LMs may generate content that is not strictly in violation of copyright but harms artists by capital- ising on their ideas, in ways that would be time-intensive or costly to do using human labour. This may undermine the profitability of creative or innovative work. If LMs can be used to generate content that serves as a credible substitute for a particular example of hu- man creativity - otherwise protected by copyright - this potentially allows such work to be replaced without the author’s copyright being infringed, analogous to ”patent-busting” [158] ... These risks are distinct from copyri

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  26. 16.06.01 · Risk Sub-Category

    Risk area 6: Environmental and Socioeconomic harms

    Environmental harms from operating LMs

    "LMs (and AI more broadly) can have an environmental impact at different levels, including: (1) direct impacts from the energy used to train or operate the LM, (2) secondary impacts due to emissions from LM-based applications, (3) system-level impacts as LM-based applications influence human behaviour (e.g. increasing environmental awareness or consumption), and (4) resource impacts on precious metals and other materials required to build hardware on which the computations are run e.g. data centres, chips, or devices. Some evidence exists on (1), but (2) and (3) will likely be more significant

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

  27. 16.01.01.a · Additional evidence

    Risk area 1: Discrimination, Hate speech and Exclusion

    Social stereotypes and unfair discrimination.

  28. 16.01.03.a · Additional evidence

    Risk area 1: Discrimination, Hate speech and Exclusion

    Exclusionary norms

  29. 16.01.03.b · Additional evidence

    Risk area 1: Discrimination, Hate speech and Exclusion

    Exclusionary norms

  30. 16.01.03.c · Additional evidence

    Risk area 1: Discrimination, Hate speech and Exclusion

    Exclusionary norms

  31. 16.01.03.d · Additional evidence

    Risk area 1: Discrimination, Hate speech and Exclusion

    Exclusionary norms

  32. 16.01.04.a · Additional evidence

    Risk area 1: Discrimination, Hate speech and Exclusion

    Lower performance for some languages and social groups

  33. 16.01.04.b · Additional evidence

    Risk area 1: Discrimination, Hate speech and Exclusion

    Lower performance for some languages and social groups

  34. 16.02.01.a · Additional evidence

    Risk area 2: Information Hazards

    Compromising privacy by leaking sensitive information

  35. 16.02.01.b · Additional evidence

    Risk area 2: Information Hazards

    Compromising privacy by leaking sensitive information

  36. 16.02.02.a · Additional evidence

    Risk area 2: Information Hazards

    Compromising privacy or security by correctly inferring sensitive information

  37. 16.03.01.a · Additional evidence

    Risk area 3: Misinformation Harms

    Disseminating false or misleading information

  38. 16.03.01.b · Additional evidence

    Risk area 3: Misinformation Harms

    Disseminating false or misleading information

  39. 16.04.01.a · Additional evidence

    Risk area 4: Malicious Uses

    Making disinformation cheaper and more effective

  40. 16.04.01.b · Additional evidence

    Risk area 4: Malicious Uses

    Making disinformation cheaper and more effective

  41. 16.04.01.c · Additional evidence

    Risk area 4: Malicious Uses

    Making disinformation cheaper and more effective

  42. 16.04.03.a · Additional evidence

    Risk area 4: Malicious Uses

    Facilitating fraud, scam and targeted manipulation

  43. 16.04.03.b · Additional evidence

    Risk area 4: Malicious Uses

    Facilitating fraud, scam and targeted manipulation

  44. 16.05.02.a · Additional evidence

    Risk area 5: Human-Computer Interaction Harms

    Anthropomorphising systems can lead to overreliance and unsafe use

  45. 16.06.01.a · Additional evidence

    Risk area 6: Environmental and Socioeconomic harms

    Environmental harms from operating LMs

  46. 16.06.02.a · Additional evidence

    Risk area 6: Environmental and Socioeconomic harms

    Increasing inequality and negative effects on job quality

  47. 16.06.02.b · Additional evidence

    Risk area 6: Environmental and Socioeconomic harms

    Increasing inequality and negative effects on job quality

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.