MIT AI Risk Repository · Risk Sub-Category · 24.03.04
Malicious Code Generation
Category: Malicious Uses
Description
"Malicious code is a term for code—whether it be part of a script or embedded in a software system—designed to cause damage, security breaches, or other threats to application security. Advanced AI assistants with the ability to produce source code can potentially lower the barrier to entry for threat actors with limited programming abilities or technical skills to produce malicious code. Recently, a series of proof-of-concept attacks have shown how a benign-seeming executable file can be crafted such that, at every runtime, it makes application programming interface (API) calls to an AI assis
From The Ethics of Advanced AI Assistants (Gabriel2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Domain
- 4. Malicious actors
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment
Subdomain definition: Using AI systems to develop cyber weapons (e.g., coding cheaper, more effective malware), develop new or enhance existing weapons (e.g., Lethal Autonomous Weapons or CBRNE), or use weapons to cause mass harm.
Real-world incidents in this subdomain
- Anthropic's Claude Was Reportedly Jailbroken To Allegedly Help Steal Sensitive Mexican Government Data
- OpenAI ChatGPT Models Reportedly Jailbroken to Provide Chemical, Biological, and Nuclear Weapons Instructions
- Anthropic Reportedly Identifies AI Misuse in Extortion Campaigns, North Korean IT Schemes, and Ransomware Sales
- LAMEHUG Malware Reportedly Integrates Large Language Model for Real-Time Command Generation in a Purported APT28-Linked Cyberattack
- Reported AI-Aided Development of Explosive Devices by Long Island Resident Michael Gann
- AI Chatbot Allegedly Used to Research Explosive Materials in Palm Springs Fertility Clinic Bombing
How other frameworks describe this risk
Other entries from Gabriel2024
- Capability failures
- Lack of capability for task
- Difficult to develop metrics for evaluating benefits or harms caused by AI assistants
- Safe exploration problem with widely deployed AI assistants
- Goal-related failures
- Misaligned consequentialist reasoning
- Specification gaming
- Goal misgeneralisation
- Deceptive alignment
- Malicious Uses
- Offensive Cyber Operations (General)
- AI-Powered Spear-Phishing at Scale