MIT AI Risk Repository · Risk Sub-Category · 24.01.03
Safe exploration problem with widely deployed AI assistants
Category: Capability failures
Description
"Moreover, we can expect assistants – that are widely deployed and deeply embedded across a range of social contexts – to encounter the safe exploration problem referenced above Amodei et al. (2016). For example, new users may have different requirements that need to be explored, or widespread AI assistants may change the way we live, thus leading to a change in our use cases for them (see Chapters 14 and 15). To learn what to do in these new situations, the assistants may need to take exploratory actions. This could be unsafe, for example a medical AI assistant when encountering a new disease
From The Ethics of Advanced AI Assistants (Gabriel2024), as extracted by the MIT AI Risk Repository (CC BY 4.0).
Classification
- Subdomain
- 7.3 Lack of capability or robustness
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
Subdomain definition: AI systems that fail to perform reliably or effectively under varying conditions, exposing them to errors and failures that can have significant consequences, especially in critical applications or areas that require moral reasoning.
Real-world incidents in this subdomain
- Purported AI Name-Reading System Reportedly Skipped and Misannounced Graduates at Arizona's Glendale Community College Commencement
- PocketOS Production Database Was Reportedly Deleted by Cursor AI Agent Running Claude Opus 4.6
- Baidu Apollo Go Robotaxis Stopped in Traffic During Reported System Failure in Wuhan, Stranding Some Passengers
- Purportedly AI-Enabled Targeting System Was Reportedly Implicated in Deadly U.S. Strike on Iranian Primary School
- Claude Code Agent Reportedly Deleted DataTalks.Club Production Infrastructure, Database, and Snapshots via Terraform
- Purportedly AI-Generated Sepsis Alert Reportedly Prompted Potentially Inappropriate IV Fluid Administration for a Dialysis Patient, Averted by Clinician Intervention
How other frameworks describe this risk
Other entries from Gabriel2024
- Capability failures
- Lack of capability for task
- Difficult to develop metrics for evaluating benefits or harms caused by AI assistants
- Goal-related failures
- Misaligned consequentialist reasoning
- Specification gaming
- Goal misgeneralisation
- Deceptive alignment
- Malicious Uses
- Offensive Cyber Operations (General)
- AI-Powered Spear-Phishing at Scale
- AI-Assisted Software Vulnerability Discovery