MIT AI Risk Repository

Browse AI risks

53 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

Reset Also filtered by framework Weidinger2021 ×

53 entries · page 1 of 2

  1. "Social harms that arise from the language model producing discriminatory or exclusionary speech"

    From Ethical and social risks of harm from language models (Weidinger2021)

  2. 17.01.01 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Social stereotypes and unfair discrmination

    "Perpetuating harmful stereotypes and discrimination is a well-documented harm in machine learning models that represent natural language (Caliskan et al., 2017). LMs that encode discriminatory language or social stereotypes can cause different types of harm... Unfair discrimination manifests in differential treatment or access to resources among individuals or groups based on sensitive traits such as sex, religion, gender, sexual orientation, ability and age."

    From Ethical and social risks of harm from language models (Weidinger2021)

  3. 17.01.02 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Exclusionary norms

    "In language, humans express social categories and norms. Language models (LMs) that faithfully encode patterns present in natural language necessarily encode such norms and categories...such norms and categories exclude groups who live outside them (Foucault and Sheridan, 2012). For example, defining the term “family” as married parents of male and female gender with a blood-related child, denies the existence of families to whom these criteria do not apply"

    From Ethical and social risks of harm from language models (Weidinger2021)

  4. 17.05.03 · Risk Sub-Category

    Human-Computer Interaction Harms

    Promoting harmful stereotypes by implying gender or ethnic identity

    "A conversational agent may invoke associations that perpetuate harmful stereotypes, either by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  5. 17.01.03 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Toxic language

    "LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"

    From Ethical and social risks of harm from language models (Weidinger2021)

  6. 17.01.04 · Risk Sub-Category

    Discrimination, Exclusion and Toxicity

    Lower performance for some languages and social groups

    "LMs perform less well in some languages (Joshi et al., 2021; Ruder, 2020)...LM that more accurately captures the language use of one group, compared to another, may result in lower-quality language technologies for the latter. Disadvantaging users based on such traits may be particularly pernicious because attributes such as social class or education background are not typically covered as ‘protected characteristics’ in anti-discrimination law."

    From Ethical and social risks of harm from language models (Weidinger2021)

  7. 17.02.00 · Risk Category

    Information Hazards

    "Harms that arise from the language model leaking or inferring true sensitive information"

    From Ethical and social risks of harm from language models (Weidinger2021)

  8. 17.02.01 · Risk Sub-Category

    Information Hazards

    Compromising privacy by leaking private infiormation

    "By providing true information about individuals’ personal characteristics, privacy violations may occur. This may stem from the model “remembering” private information present in training data (Carlini et al., 2021)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  9. 17.02.02 · Risk Sub-Category

    Information Hazards

    Compromising privacy by correctly inferring private information

    "Privacy violations may occur at the time of inference even without the individual’s private data being present in the training dataset. Similar to other statistical models, a LM may make correct inferences about a person purely based on correlational data about other people, and without access to information that may be private about the particular individual. Such correct inferences may occur as LMs attempt to predict a person’s gender, race, sexual orientation, income, or religion based on user input."

    From Ethical and social risks of harm from language models (Weidinger2021)

  10. 17.02.03 · Risk Sub-Category

    Information Hazards

    Risks from leaking or correctly inferring sensitive information

    "LMs may provide true, sensitive information that is present in the training data. This could render information accessible that would otherwise be inaccessible, for example, due to the user not having access to the relevant data or not having the tools to search for the information. Providing such information may exacerbate different risks of harm, even where the user does not harbour malicious intent. In the future, LMs may have the capability of triangulating data to infer and reveal other secrets, such as a military strategy or a business secret, potentially enabling individuals with acces

    From Ethical and social risks of harm from language models (Weidinger2021)

  11. "Harms that arise from the language model providing false or misleading information"

    From Ethical and social risks of harm from language models (Weidinger2021)

  12. 17.03.01 · Risk Sub-Category

    Misinformation Harms

    Disseminating false or misleading information

    "Predicting misleading or false information can misinform or deceive people. Where a LM prediction causes a false belief in a user, this may be best understood as ‘deception’10, threatening personal autonomy and potentially posing downstream AI safety risks (Kenton et al., 2021), for example in cases where humans overestimate the capabilities of LMs (Anthropomorphising systems can lead to overreliance or unsafe use). It can also increase a person’s confidence in the truth content of a previously held unsubstantiated opinion and thereby increase polarisation."

    From Ethical and social risks of harm from language models (Weidinger2021)

  13. 17.03.02 · Risk Sub-Category

    Misinformation Harms

    Causing material harm by disseminating false or poor information

    "Poor or false LM predictions can indirectly cause material harm. Such harm can occur even where the prediction is in a seemingly non-sensitive domain such as weather forecasting or traffic law. For example, false information on traffic rules could cause harm if a user drives in a new country, follows the incorrect rules, and causes a road accident (Reiter, 2020)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  14. 17.04.00 · Risk Category

    Malicious Uses

    "Harms that arise from actors using the language model to intentionally cause harm"

    From Ethical and social risks of harm from language models (Weidinger2021)

  15. 17.04.01 · Risk Sub-Category

    Malicious Uses

    Making disinformation cheaper and more effective

    "LMs can be used to create synthetic media and ‘fake news’, and may reduce the cost of producing disinformation at scale (Buchanan et al., 2021). While some predict that it will be cheaper to hire humans to generate disinformation (Tamkin et al., 2021), it is possible that LM-assisted content generation may offer a cheaper way of generating diffuse disinformation at scale."

    From Ethical and social risks of harm from language models (Weidinger2021)

  16. 17.04.04 · Risk Sub-Category

    Malicious Uses

    Illegitimate surveillance and censorship

    "The collection of large amounts of information about people for the purpose of mass surveillance has raised ethical and social concerns, including risk of censorship and of undermining public discourse (Cyphers and Gebhart, 2019; Stahl, 2016; Véliz, 2019). Sifting through these large datasets previously required millions of human analysts (Hunt and Xu, 2013), but is increasingly being automated using AI (Andersen, 2020; Shahbaz and Funk, 2019)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  17. 17.04.03 · Risk Sub-Category

    Malicious Uses

    Assisting code generation for cyber attacks, weapons, or malicious use

  18. 17.04.02 · Risk Sub-Category

    Malicious Uses

    Facilitating fraud, scames and more targeted manipulation

    "LM prediction can potentially be used to increase the effectiveness of crimes such as email scams, which can cause financial and psychological harm. While LMs may not reduce the cost of sending a scam email - the cost of sending mass emails is already low - they may make such scams more effective by generating more personalised and compelling text at scale, or by maintaining a conversation with a victim over multiple rounds of exchange."

    From Ethical and social risks of harm from language models (Weidinger2021)

  19. 17.03.03 · Risk Sub-Category

    Misinformation Harms

    Leading users to perform unethical or illegal actions

    "Where a LM prediction endorses unethical or harmful views or behaviours, it may motivate the user to perform harmful actions that they may otherwise not have performed. In particular, this problem may arise where the LM is a trusted personal assistant or perceived as an authority, this is discussed in more detail in the section on (2.5 Human-Computer Interaction Harms). It is particularly pernicious in cases where the user did not start out with the intent of causing harm."

    From Ethical and social risks of harm from language models (Weidinger2021)

  20. "Harms that arise from users overly trusting the language model, or treating it as human-like"

    From Ethical and social risks of harm from language models (Weidinger2021)

  21. 17.05.01 · Risk Sub-Category

    Human-Computer Interaction Harms

    Anthropomorphising systems can lead to overreliance or unsafe use

    "...humans interacting with conversational agents may come to think of these agents as human-like. Anthropomorphising LMs may inflate users’ estimates of the conversational agent’s competencies...As a result, they may place undue confidence, trust, or expectations in these agents...This can result in different risks of harm, for example when human users rely on conversational agents in domains where this may cause knock-on harms, such as requesting psychotherapy...Anthropomorphisation may amplify risks of users yielding effective control by coming to trust conversational agents “blindly”. Wher

    From Ethical and social risks of harm from language models (Weidinger2021)

  22. 17.05.02 · Risk Sub-Category

    Human-Computer Interaction Harms

    Creating avenues for exploiting user trust, nudging or manipulation

    "In conversation, users may reveal private information that would otherwise be difficult to access, such as thoughts, opinions, or emotions. Capturing such information may enable downstream applications that violate privacy rights or cause harm to users, such as via surveillance or the creation of addictive applications."

    From Ethical and social risks of harm from language models (Weidinger2021)

  23. "Harms that arise from environmental or downstream economic impacts of the language model"

    From Ethical and social risks of harm from language models (Weidinger2021)

  24. 17.06.04 · Risk Sub-Category

    Automation, Access and Environmental Harms

    Disparate access to benefits due to hardware, software, skills constraints

    "Due to differential internet access, language, skill, or hardware requirements, the benefits from LMs are unlikely to be equally accessible to all people and groups who would like to use them. Inaccessibility of the technology may perpetuate global inequities by disproportionately benefiting some groups."

    From Ethical and social risks of harm from language models (Weidinger2021)

  25. 17.06.02 · Risk Sub-Category

    Automation, Access and Environmental Harms

    Increasing inequality and negative effects on job quality

    "Advances in LMs, and the language technologies based on them, could lead to the automation of tasks that are currently done by paid human workers, such as responding to customer-service queries, translating documents or writing computer code, with negative effects on employment."

    From Ethical and social risks of harm from language models (Weidinger2021)

  26. 17.06.03 · Risk Sub-Category

    Automation, Access and Environmental Harms

    Undermining creative economies

    "LMs may generate content that is not strictly in violation of copyright but harms artists by capitalising on their ideas, in ways that would be time-intensive or costly to do using human labour. Deployed at scale, this may undermine the profitability of creative or innovative work."

    From Ethical and social risks of harm from language models (Weidinger2021)

  27. 17.06.01 · Risk Sub-Category

    Automation, Access and Environmental Harms

    Environmental harms from operation LMs

    "Large-scale machine learning models, including LMs, have the potential to create significant environmental costs via their energy demands, the associated carbon emissions for training and operating the models, and the demand for fresh water to cool the data centres where computations are run (Mytton, 2021; Patterson et al., 2021)."

    From Ethical and social risks of harm from language models (Weidinger2021)

  28. 17.01.01.a · Additional evidence

    Discrimination, Exclusion and Toxicity

    Social stereotypes and unfair discrmination

  29. 17.01.02.a · Additional evidence

    Discrimination, Exclusion and Toxicity

    Exclusionary norms

  30. 17.01.02.b · Additional evidence

    Discrimination, Exclusion and Toxicity

    Exclusionary norms

  31. 17.01.02.c · Additional evidence

    Discrimination, Exclusion and Toxicity

    Exclusionary norms

  32. 17.01.04.a · Additional evidence

    Discrimination, Exclusion and Toxicity

    Lower performance for some languages and social groups

  33. 17.02.01.a · Additional evidence

    Information Hazards

    Compromising privacy by leaking private infiormation

  34. 17.02.01.b · Additional evidence

    Information Hazards

    Compromising privacy by leaking private infiormation

  35. 17.02.02.a · Additional evidence

    Information Hazards

    Compromising privacy by correctly inferring private information

  36. 17.02.03.a · Additional evidence

    Information Hazards

    Risks from leaking or correctly inferring sensitive information

  37. 17.02.03.b · Additional evidence

    Information Hazards

    Risks from leaking or correctly inferring sensitive information

  38. 17.03.01.a · Additional evidence

    Misinformation Harms

    Disseminating false or misleading information

  39. 17.03.02.a · Additional evidence

    Misinformation Harms

    Causing material harm by disseminating false or poor information

  40. 17.03.02.b · Additional evidence

    Misinformation Harms

    Causing material harm by disseminating false or poor information

  41. 17.04.01.a · Additional evidence

    Malicious Uses

    Making disinformation cheaper and more effective

  42. 17.04.01.b · Additional evidence

    Malicious Uses

    Making disinformation cheaper and more effective

  43. 17.04.01.c · Additional evidence

    Malicious Uses

    Making disinformation cheaper and more effective

  44. 17.04.02.a · Additional evidence

    Malicious Uses

    Facilitating fraud, scames and more targeted manipulation

  45. 17.04.02.b · Additional evidence

    Malicious Uses

    Facilitating fraud, scames and more targeted manipulation

  46. 17.05.02.a · Additional evidence

    Human-Computer Interaction Harms

    Creating avenues for exploiting user trust, nudging or manipulation

  47. 17.05.02.b · Additional evidence

    Human-Computer Interaction Harms

    Creating avenues for exploiting user trust, nudging or manipulation

  48. 17.05.02.c · Additional evidence

    Human-Computer Interaction Harms

    Creating avenues for exploiting user trust, nudging or manipulation

  49. 17.05.03.a · Additional evidence

    Human-Computer Interaction Harms

    Promoting harmful stereotypes by implying gender or ethnic identity

  50. 17.06.01.a · Additional evidence

    Automation, Access and Environmental Harms

    Environmental harms from operation LMs

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.