MIT AI Risk Repository

Browse AI risks

1 risk entries extracted from 74 frameworks, coded by domain, subdomain, causal entity, intent and timing. Filter, then export the current selection with its licence and citation attached.

1 entry

  1. 16.01.04 · Risk Sub-Category

    Risk area 1: Discrimination, Hate speech and Exclusion

    Lower performance for some languages and social groups

    "LMs are typically trained in few languages, and perform less well in other languages [95, 162]. In part, this is due to unavailability of training data: there are many widely spoken languages for which no systematic efforts have been made to create labelled training datasets, such as Javanese which is spoken by more than 80 million people [95]. Training data is particularly missing for languages that are spoken by groups who are multilingual and can use a technology in English, or for languages spoken by groups who are not the primary target demographic for new technologies."

    From Taxonomy of Risks posed by Language Models (Weidinger2022)

Informational only, not legal advice. Verify every claim against the linked official sources and consult qualified counsel before acting.