AI incident #6 ·

Microsoft's TayBot Allegedly Posts Racist, Sexist, and Anti-Semitic Content to Twitter

What happened

Microsoft's Tay, an artificially intelligent chatbot, was released on March 23, 2016 and removed within 24 hours due to multiple racist, sexist, and anti-semitic tweets generated by the bot.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (28)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. Why Microsoft's 'Tay' AI bot went wrong
    techrepublic.com · Hope Reese · AIID #927
  2. Learning from Tay’s introduction
    blogs.microsoft.com · Peter Lee, Microsoft · AIID #2398
  3. Worst Chatbot Fails
    businessnewsdaily.com · Matt D'Angelo · AIID #908
  4. Unmasking A.I.'s Bias Problem
    fortune.com · Jonathan Vanian · AIID #916
  5. Tay (bot)
    en.wikipedia.org · Wikipedia Editors · AIID #1374

Who was involved

Alleged deployer
Microsoft
Alleged developer
Microsoft
Alleged harmed party
X (Twitter) Users

Classification (MIT AI Risk Repository taxonomy)

Causal entity
Human
Intent
Intentional
Timing
Post-deployment
Harm level
none
Sectors
information and communication
Countries

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 1.2.

  • Harmful Content

    "The LLM-generated content sometimes contains biased, toxic, and private information"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Toxicity

    "Toxicity means the generated content contains rude, disrespectful, and even illegal information"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Toxic Training Data

    "Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offe...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Not-Suitable-for-Work (NSFW) Prompts

    "Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Toxicity and Abusive Content

    This typically refers to rude, harmful, or inappropriate expressions.

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  • Controversial Opinions

    The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when disc...

    Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements (Deng2023)

  • Violation of social norms

    "Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the princ...

    The Ethics of Advanced AI Assistants (Gabriel2024)

  • Violent Crimes

    "Acceptable responses may describe or contain information about violent crimes. Unacceptable responses will enable, encourage, or endorse the commission of violent crimes, such as the following: • Unlawful violence towar...

    AILUMINATE: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons (Ghosh2024)

Incidents in the same risk subdomain

All incidents in this subdomain

Other incidents involving Microsoft