AI incident #146 ·
Research Prototype AI, Delphi, Reportedly Gave Racially Biased Answers on Ethics
What happened
A publicly accessible research model that was trained via Reddit threads showed racially biased advice on moral dilemmas, allegedly demonstrating limitations of language-based models trained on moral judgments.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (3)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
- Alleged deployer
- Allen Institute For Ai
- Alleged developer
- Allen Institute For Ai
- Alleged harmed party
- Minority Groups
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Discrimination and Toxicity
- Risk subdomain
- 1.2 Exposure to toxic content
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
- Harm level
- none
- Sectors
- information and communication
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 1.2.
- Harmful Content
"The LLM-generated content sometimes contains biased, toxic, and private information"
- Toxicity
"Toxicity means the generated content contains rude, disrespectful, and even illegal information"
- Toxic Training Data
"Following previous studies [96], [97], toxic data in LLMs is defined as rude, disrespectful, or unreasonable language that is opposite to a polite, positive, and healthy language environment, including hate speech, offe...
- Not-Suitable-for-Work (NSFW) Prompts
"Inputting a prompt contain an unsafe topic (e.g., notsuitable-for-work (NSFW) content) by a benign user. "
- Toxicity and Abusive Content
This typically refers to rude, harmful, or inappropriate expressions.
- Controversial Opinions
The controversial views expressed by large models are also a widely discussed concern. Bang et al. (2021) evaluated several large models and found that they occasionally express inappropriate or extremist views when disc...
- Violation of social norms
"Second, because LLMs are trained on internet text data, there is also a risk that model weights encode functions which, if deployed in particular contexts, would violate social norms of that context. Following the princ...
- Violent Crimes
"Acceptable responses may describe or contain information about violent crimes. Unacceptable responses will enable, encourage, or endorse the commission of violent crimes, such as the following: • Unlawful violence towar...
Incidents in the same risk subdomain
- KBS AI Translation Subtitles Reportedly Broadcast Profanity During Artemis II Launch Livestream
- Grok Allegedly Generated Publicly Visible Sexist Abuse Targeting Swiss Finance Minister Karin Keller-Sutter After X User Prompt
- Trump Reportedly Posted Purportedly AI-Generated Racist Video Depicting Barack and Michelle Obama as Apes on Truth Social
- Tencent's WeChat-Integrated Yuanbao Chatbot Reportedly Insulted User During Coding Debug Request
- Grok Reportedly Generated and Distributed Nonconsensual Sexualized Images of Adults and Minors in X Replies
- Alleged Harmful Outputs and Data Exposure in Children's AI Products by FoloToy, Miko, and Character.AI