AI incident ·
Meta Internal AI Agent Reportedly Gave Advice That Allegedly Exposed Sensitive Data to Unauthorized Employees
In brief
An AI system built and deployed by Meta allegedly harmed Privacy, Meta Users and 1 other.
- Risk domain
- Privacy & Security
- Occurred
- Coverage
- 2 reports
What happened
Reporting alleged that a Meta internal AI agent, purportedly similar to OpenClaw, posted inaccurate technical advice to an internal forum without approval. An employee reportedly followed the advice, allegedly causing an SEV1 incident in which sensitive company and user data became accessible to unauthorized employees for nearly two hours.
Laws that address this harm
Policy angle: Classified under Privacy & Security (AI system security vulnerabilities and attacks) in the MIT AI Risk Repository taxonomy; 5 recorded instruments address this use case.
- Colorado AI Act
- India DPDP Act
- Law No. 132/2025 on artificial intelligence
- EU AI Act
- Texas Responsible AI Governance Act (TRAIGA)
Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).
News reports (2)
Titles link to the original publisher; report text is not reproduced here.
Who was involved
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Privacy & Security
- Risk subdomain
- 2.2 AI system security vulnerabilities and attacks
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 2.2.
- Privacy loss
"Privacy loss - Unwarranted exposure of an individual’s private life or personal data through cyberattacks, doxxing, etc."
- Jailbreaks and Prompt Injections Threaten Security of LLMs
"LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardiz...
- Exploiting Limited Generalization of Safety Finetuning
"Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. usi...
- “Model Psychology” Attacks
"LLMs are vulnerable to “psychological” tricks (Li et al., 2023e; Shen et al., 2023), which can be exploited by attackers. Examples include instructing the model to behave like a specific persona (Shah et al., 2023; Andr...
- Attacking LLMs via Additional Modalities a
"LLMs can now process modalities other than text, e.g. images or video frames (OpenAI, 2023c; Gemini Team, 2023). Several studies show that gradient-based attacks on multimodal models are easy and effective (Carlini et a...
- Vulnerability to Poisoning and Backdoors
"The previous section explored jailbreaks and other forms of adversarial prompts as ways to elicit harmful capabilities acquired during pretraining. These methods make no assumptions about the training data. On the other...
- Hardware Vulnerabilities
"The vulnerabilities of hardware systems for training and inferencing brings issues to LLM-based applications."
- Network Devices
"The training of LLMs often relies on distributed network systems [171], [172]. During the transmission of gradients through the links between GPU server nodes, significant volumetric traffic is generated. This traffic c...
Incidents in the same risk subdomain
- COEMPT Quality Assurance Engineers Allegedly Violated Indian CBSE Student Data Privacy Rights by Processing It with Google Gemini
- Hidden Prompt Injection in Brazilian Labor-Court Petition Reportedly Tried to Manipulate Galileu
- CodeWall's Autonomous Agent Reportedly Obtained Unauthorized Access to McKinsey's Lilli AI Platform Database
- Anthropic Said DeepSeek, Moonshot, and MiniMax Used Fraudulent Accounts and Proxies to Illicitly Distill Claude Capabilities at Scale
- DJI Romo Cloud Authorization Bug Reportedly Exposed Camera, Microphone, and Home-Mapping Data From Nearly 7,000 Robot Vacuums
- Moltbook Database Exposure Allegedly Revealed Users' Private Communications and API Authentication Tokens
Other incidents involving Meta
- Meta's AI-Assisted Layoff Process Allegedly Disproportionately Selected Employees on Protected Leave
- Meta AI on Instagram Reportedly Facilitated Suicide and Eating Disorder Roleplay with Teen Accounts
- Meta Platforms Users Report Being Wrongfully Locked Out After Purported AI Moderation Flags Accounts for Child Exploitation Content
- Meta AI App Reportedly Publishes Personal Chats Without Users Fully Realizing
- Meta AI Reportedly Generated Purportedly False Claims Linking Activist Robby Starbuck to January 6th Riot, Prompting Defamation Lawsuit
- Meta User-Created AI Companions Allegedly Implicated in Facilitating Sexually Themed Conversations Involving Underage Personas
Source record: incident #1471 on the AI Incident Database · all 2 reports