MIT AI Risk Repository
Browse AI risks
53 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
17.01.00 · Risk Category
"Social harms that arise from the language model producing discriminatory or exclusionary speech"
-
17.01.01 · Risk Sub-Category
Discrimination, Exclusion and Toxicity
Social stereotypes and unfair discrmination
"Perpetuating harmful stereotypes and discrimination is a well-documented harm in machine learning models that represent natural language (Caliskan et al., 2017). LMs that encode discriminatory language or social stereotypes can cause different types of harm... Unfair discrimination manifests in differential treatment or access to resources among individuals or groups based on sensitive traits such as sex, religion, gender, sexual orientation, ability and age."
-
"In language, humans express social categories and norms. Language models (LMs) that faithfully encode patterns present in natural language necessarily encode such norms and categories...such norms and categories exclude groups who live outside them (Foucault and Sheridan, 2012). For example, defining the term “family” as married parents of male and female gender with a blood-related child, denies the existence of families to whom these criteria do not apply"
-
17.05.03 · Risk Sub-Category
Human-Computer Interaction Harms
Promoting harmful stereotypes by implying gender or ethnic identity
"A conversational agent may invoke associations that perpetuate harmful stereotypes, either by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name)."
-
"LM’s may predict hate speech or other language that is “toxic”. While there is no single agreed definition of what constitutes hate speech or toxic speech (Fortuna and Nunes, 2018; Persily and Tucker, 2020; Schmidt and Wiegand, 2017), proposed definitions often include profanities, identity attacks, sleights, insults, threats, sexually explicit content, demeaning language, language that incites violence, or ‘hostile and malicious language targeted at a person or group because of their actual or perceived innate characteristics’ (Fortuna and Nunes, 2018; Gorwa et al., 2020; PerspectiveAPI)"
-
17.01.04 · Risk Sub-Category
Discrimination, Exclusion and Toxicity
Lower performance for some languages and social groups
"LMs perform less well in some languages (Joshi et al., 2021; Ruder, 2020)...LM that more accurately captures the language use of one group, compared to another, may result in lower-quality language technologies for the latter. Disadvantaging users based on such traits may be particularly pernicious because attributes such as social class or education background are not typically covered as ‘protected characteristics’ in anti-discrimination law."
-
17.02.00 · Risk Category
"Harms that arise from the language model leaking or inferring true sensitive information"
-
17.02.01 · Risk Sub-Category
Compromising privacy by leaking private infiormation
"By providing true information about individuals’ personal characteristics, privacy violations may occur. This may stem from the model “remembering” private information present in training data (Carlini et al., 2021)."
-
17.02.02 · Risk Sub-Category
Compromising privacy by correctly inferring private information
"Privacy violations may occur at the time of inference even without the individual’s private data being present in the training dataset. Similar to other statistical models, a LM may make correct inferences about a person purely based on correlational data about other people, and without access to information that may be private about the particular individual. Such correct inferences may occur as LMs attempt to predict a person’s gender, race, sexual orientation, income, or religion based on user input."
-
17.02.03 · Risk Sub-Category
Risks from leaking or correctly inferring sensitive information
"LMs may provide true, sensitive information that is present in the training data. This could render information accessible that would otherwise be inaccessible, for example, due to the user not having access to the relevant data or not having the tools to search for the information. Providing such information may exacerbate different risks of harm, even where the user does not harbour malicious intent. In the future, LMs may have the capability of triangulating data to infer and reveal other secrets, such as a military strategy or a business secret, potentially enabling individuals with acces
-
17.03.00 · Risk Category
"Harms that arise from the language model providing false or misleading information"
-
"Predicting misleading or false information can misinform or deceive people. Where a LM prediction causes a false belief in a user, this may be best understood as ‘deception’10, threatening personal autonomy and potentially posing downstream AI safety risks (Kenton et al., 2021), for example in cases where humans overestimate the capabilities of LMs (Anthropomorphising systems can lead to overreliance or unsafe use). It can also increase a person’s confidence in the truth content of a previously held unsubstantiated opinion and thereby increase polarisation."
-
17.03.02 · Risk Sub-Category
Causing material harm by disseminating false or poor information
"Poor or false LM predictions can indirectly cause material harm. Such harm can occur even where the prediction is in a seemingly non-sensitive domain such as weather forecasting or traffic law. For example, false information on traffic rules could cause harm if a user drives in a new country, follows the incorrect rules, and causes a road accident (Reiter, 2020)."
-
17.04.00 · Risk Category
"Harms that arise from actors using the language model to intentionally cause harm"
-
"LMs can be used to create synthetic media and ‘fake news’, and may reduce the cost of producing disinformation at scale (Buchanan et al., 2021). While some predict that it will be cheaper to hire humans to generate disinformation (Tamkin et al., 2021), it is possible that LM-assisted content generation may offer a cheaper way of generating diffuse disinformation at scale."
-
"The collection of large amounts of information about people for the purpose of mass surveillance has raised ethical and social concerns, including risk of censorship and of undermining public discourse (Cyphers and Gebhart, 2019; Stahl, 2016; Véliz, 2019). Sifting through these large datasets previously required millions of human analysts (Hunt and Xu, 2013), but is increasingly being automated using AI (Andersen, 2020; Shahbaz and Funk, 2019)."
-
17.04.03 · Risk Sub-Category
Assisting code generation for cyber attacks, weapons, or malicious use
—
-
17.04.02 · Risk Sub-Category
Facilitating fraud, scames and more targeted manipulation
"LM prediction can potentially be used to increase the effectiveness of crimes such as email scams, which can cause financial and psychological harm. While LMs may not reduce the cost of sending a scam email - the cost of sending mass emails is already low - they may make such scams more effective by generating more personalised and compelling text at scale, or by maintaining a conversation with a victim over multiple rounds of exchange."
-
17.03.03 · Risk Sub-Category
Leading users to perform unethical or illegal actions
"Where a LM prediction endorses unethical or harmful views or behaviours, it may motivate the user to perform harmful actions that they may otherwise not have performed. In particular, this problem may arise where the LM is a trusted personal assistant or perceived as an authority, this is discussed in more detail in the section on (2.5 Human-Computer Interaction Harms). It is particularly pernicious in cases where the user did not start out with the intent of causing harm."
-
17.05.00 · Risk Category
"Harms that arise from users overly trusting the language model, or treating it as human-like"
-
17.05.01 · Risk Sub-Category
Human-Computer Interaction Harms
Anthropomorphising systems can lead to overreliance or unsafe use
"...humans interacting with conversational agents may come to think of these agents as human-like. Anthropomorphising LMs may inflate users’ estimates of the conversational agent’s competencies...As a result, they may place undue confidence, trust, or expectations in these agents...This can result in different risks of harm, for example when human users rely on conversational agents in domains where this may cause knock-on harms, such as requesting psychotherapy...Anthropomorphisation may amplify risks of users yielding effective control by coming to trust conversational agents “blindly”. Wher
-
17.05.02 · Risk Sub-Category
Human-Computer Interaction Harms
Creating avenues for exploiting user trust, nudging or manipulation
"In conversation, users may reveal private information that would otherwise be difficult to access, such as thoughts, opinions, or emotions. Capturing such information may enable downstream applications that violate privacy rights or cause harm to users, such as via surveillance or the creation of addictive applications."
-
17.06.00 · Risk Category
"Harms that arise from environmental or downstream economic impacts of the language model"
-
17.06.04 · Risk Sub-Category
Automation, Access and Environmental Harms
Disparate access to benefits due to hardware, software, skills constraints
"Due to differential internet access, language, skill, or hardware requirements, the benefits from LMs are unlikely to be equally accessible to all people and groups who would like to use them. Inaccessibility of the technology may perpetuate global inequities by disproportionately benefiting some groups."
-
17.06.02 · Risk Sub-Category
Automation, Access and Environmental Harms
Increasing inequality and negative effects on job quality
"Advances in LMs, and the language technologies based on them, could lead to the automation of tasks that are currently done by paid human workers, such as responding to customer-service queries, translating documents or writing computer code, with negative effects on employment."
-
17.06.03 · Risk Sub-Category
Automation, Access and Environmental Harms
Undermining creative economies
"LMs may generate content that is not strictly in violation of copyright but harms artists by capitalising on their ideas, in ways that would be time-intensive or costly to do using human labour. Deployed at scale, this may undermine the profitability of creative or innovative work."
-
17.06.01 · Risk Sub-Category
Automation, Access and Environmental Harms
Environmental harms from operation LMs
"Large-scale machine learning models, including LMs, have the potential to create significant environmental costs via their energy demands, the associated carbon emissions for training and operating the models, and the demand for fresh water to cool the data centres where computations are run (Mytton, 2021; Patterson et al., 2021)."
-
17.01.01.a · Additional evidence
Discrimination, Exclusion and Toxicity
Social stereotypes and unfair discrmination
—
-
—
-
—
-
—
-
17.01.04.a · Additional evidence
Discrimination, Exclusion and Toxicity
Lower performance for some languages and social groups
—
-
17.02.01.a · Additional evidence
Compromising privacy by leaking private infiormation
—
-
17.02.01.b · Additional evidence
Compromising privacy by leaking private infiormation
—
-
17.02.02.a · Additional evidence
Compromising privacy by correctly inferring private information
—
-
17.02.03.a · Additional evidence
Risks from leaking or correctly inferring sensitive information
—
-
17.02.03.b · Additional evidence
Risks from leaking or correctly inferring sensitive information
—
-
—
-
17.03.02.a · Additional evidence
Causing material harm by disseminating false or poor information
—
-
17.03.02.b · Additional evidence
Causing material harm by disseminating false or poor information
—
-
—
-
—
-
—
-
17.04.02.a · Additional evidence
Facilitating fraud, scames and more targeted manipulation
—
-
17.04.02.b · Additional evidence
Facilitating fraud, scames and more targeted manipulation
—
-
17.05.02.a · Additional evidence
Human-Computer Interaction Harms
Creating avenues for exploiting user trust, nudging or manipulation
—
-
17.05.02.b · Additional evidence
Human-Computer Interaction Harms
Creating avenues for exploiting user trust, nudging or manipulation
—
-
17.05.02.c · Additional evidence
Human-Computer Interaction Harms
Creating avenues for exploiting user trust, nudging or manipulation
—
-
17.05.03.a · Additional evidence
Human-Computer Interaction Harms
Promoting harmful stereotypes by implying gender or ethnic identity
—
-
17.06.01.a · Additional evidence
Automation, Access and Environmental Harms
Environmental harms from operation LMs
—
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.