MIT AI Risk Repository · Risk Category · 40.02.00
On Purpose - Post Deployment
Description
"Just because developers might succeed in creating a safe AI, it doesn't mean that it will not become unsafe at some later point. In other words, a perfectly friendly AI could be switched to the "dark side" during the post-deployment stage. This can happen rather innocuously as a result of someone lying to the AI and purposefully supplying it with incorrect information or more explicitly as a result of someone giving the AI orders to perform illegal or dangerous actions against others."
From Taxonomy of Pathways to Dangerous Artificial Intelligence (Yampolskiy2016), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 4. Malicious actors
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment
Subdomain definition: Using AI systems to gain a personal advantage over others such as through cheating, fraud, scams, blackmail or targeted manipulation of beliefs or behavior. Examples include AI-facilitated plagiarism for research or education, impersonating a trusted or fake individual for illegitimate financial benefit, or creating humiliating or sexual imagery.
Real-world incidents in this subdomain
- Italian Mediaset Journalist Safiria Leccese's Image Was Reportedly Used in a Purportedly AI-Generated Fake Loan Scam
- Scammers Reportedly Used AI-Cloned Daughter's Voice to Defraud Bay Area Mother in Fake Kidnapping Call
- Texas Man Arturo Hernandez Allegedly Published AI-Generated Deepfake Pornography Depicting Women in TAKE IT DOWN Act Case
- Guelph, Ontario, Woman Reportedly Lost $14,000 in Purported Deepfake MrBeast Cryptocurrency Scam
- Purportedly AI-Recreated Clips from Beastie Boys' 'Sabotage' Video Reportedly Appeared in FBI Promotional Video Posted by Kash Patel
- Ahmedabad Aadhaar Fraud Racket Reportedly Used Purportedly AI-Generated Deepfakes to Change Businessman's Linked Mobile Number