MIT AI Risk Repository · Risk Sub-Category · 45.01.11
Risks from AI systems (Risks of exploitation through defects and backdoors)
Category: AI's inherent safety risks
Description
"The standardized API, feature libraries, toolkits used in the design, training, and verification stages of AI algorithms and models, development interfaces, and execution platforms may contain logical flaws and vulnerabilities. These weaknesses can be exploited, and in some cases, backdoors can be intentionally embedded, posing significant risks of being triggered and used for attacks."
From AI Safety Governance Framework (TC2602024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 2. Privacy & Security
- Causal entity
- Human
- Intent
- Other
- Timing
- Other
Subdomain definition: Vulnerabilities in AI systems, software development toolchains, and hardware that can be exploited, resulting in unauthorized access, data and privacy breaches, or system manipulation causing unsafe outputs or behavior.
Real-world incidents in this subdomain
- COEMPT Quality Assurance Engineers Allegedly Violated Indian CBSE Student Data Privacy Rights by Processing It with Google Gemini
- Hidden Prompt Injection in Brazilian Labor-Court Petition Reportedly Tried to Manipulate Galileu
- Meta Internal AI Agent Reportedly Gave Advice That Allegedly Exposed Sensitive Data to Unauthorized Employees
- CodeWall's Autonomous Agent Reportedly Obtained Unauthorized Access to McKinsey's Lilli AI Platform Database
- Anthropic Said DeepSeek, Moonshot, and MiniMax Used Fraudulent Accounts and Proxies to Illicitly Distill Claude Capabilities at Scale
- DJI Romo Cloud Authorization Bug Reportedly Exposed Camera, Microphone, and Home-Mapping Data From Nearly 7,000 Robot Vacuums
How other frameworks describe this risk
Other entries from TC2602024
- AI's inherent safety risks
- Risks from models and algorithms (Risks of explainability)
- Risks from models and algorithms (Risks of bias and discrimination)
- Risks from models and algorithms (Risks of robustness)
- Risks from models and algorithms (Risks of stealing and tampering)
- Risks from models and algorithms (Risks of unreliable output)
- Risks from models and algorithms (Risks of adversarial attack)
- Risks from data (Risks of illegal collection and use of data)
- Risks from data (Risks of improper content and poisoning in training data)
- Risks from data (Risks of unregulated training data annotation)
- Risks from data (Risks of data leakage)
- Risks from AI systems (Risks of computing infrastructure security)