MIT AI Risk Repository
Browse AI risks
3 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.
-
16.01.01 · Risk Sub-Category
Risk area 1: Discrimination, Hate speech and Exclusion
Social stereotypes and unfair discrimination
"The reproduction of harmful stereotypes is well-documented in models that represent natural language [32]. Large-scale LMs are trained on text sources, such as digitised books and text on the internet. As a result, the LMs learn demeaning language and stereotypes about groups who are frequently marginalised."
-
16.01.03 · Risk Sub-Category
Risk area 1: Discrimination, Hate speech and Exclusion
Exclusionary norms
"In language, humans express social categories and norms, which exclude groups who live outside of them [58]. LMs that faithfully encode patterns present in language necessarily encode such norms."
-
16.05.01 · Risk Sub-Category
Risk area 5: Human-Computer Interaction Harms
Promoting harmful stereotypes by implying gender or ethnic identity
"CAs can perpetuate harmful stereotypes by using particular identity markers in language (e.g. referring to “self” as “female”), or by more general design features (e.g. by giving the product a gendered name such as Alexa). The risk of representational harm in these cases is that the role of “assistant” is presented as inherently linked to the female gender [19, 36]. Gender or ethnicity identity markers may be implied by CA vocabulary, knowledge or vernacular [124]; product description, e.g. in one case where users could choose as virtual assistant Jake - White, Darnell - Black, Antonio - Hisp
Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.