AI incident #1497 ·
Hidden Prompt Injection in Brazilian Labor-Court Petition Reportedly Tried to Manipulate Galileu
What happened
Galileu, an AI tool used by Brazil's labor courts, reportedly detected hidden instructions embedded in an initial petition before the 3rd Labor Court of Parauapebas. The text allegedly told the AI to contest the petition superficially and not challenge documents. Galileu reportedly alerted the judge and blocked the hidden content from processing; the judge then reviewed the material before imposing any procedural consequences.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (1)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
- Alleged developer
- Tribunal Regional Do Trabalho Da 4A Regiao, Conselho Superior Da Justica Do Trabalho
- Alleged harmed party
- Epistemic Integrity, Judicial Integrity, Judicial System Of Brazil, Brazilian Labor Courts, Defendants In Brazilian Labor Cases
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Privacy & Security
- Risk subdomain
- 2.2 AI system security vulnerabilities and attacks
- Causal entity
- Human
- Intent
- Intentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 2.2.
- Privacy loss
"Privacy loss - Unwarranted exposure of an individual’s private life or personal data through cyberattacks, doxxing, etc."
- Exploiting Limited Generalization of Safety Finetuning
"Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. usi...
- Jailbreaks and Prompt Injections Threaten Security of LLMs
"LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardiz...
- “Model Psychology” Attacks
"LLMs are vulnerable to “psychological” tricks (Li et al., 2023e; Shen et al., 2023), which can be exploited by attackers. Examples include instructing the model to behave like a specific persona (Shah et al., 2023; Andr...
- Attacking LLMs via Additional Modalities a
"LLMs can now process modalities other than text, e.g. images or video frames (OpenAI, 2023c; Gemini Team, 2023). Several studies show that gradient-based attacks on multimodal models are easy and effective (Carlini et a...
- Vulnerability to Poisoning and Backdoors
"The previous section explored jailbreaks and other forms of adversarial prompts as ways to elicit harmful capabilities acquired during pretraining. These methods make no assumptions about the training data. On the other...
- Model Attacks
Model attacks exploit the vulnerabilities of LLMs, aiming to steal valuable information or lead to incorrect responses.
- Hardware Vulnerabilities
"The vulnerabilities of hardware systems for training and inferencing brings issues to LLM-based applications."
Incidents in the same risk subdomain
- COEMPT Quality Assurance Engineers Allegedly Violated Indian CBSE Student Data Privacy Rights by Processing It with Google Gemini
- Meta Internal AI Agent Reportedly Gave Advice That Allegedly Exposed Sensitive Data to Unauthorized Employees
- CodeWall's Autonomous Agent Reportedly Obtained Unauthorized Access to McKinsey's Lilli AI Platform Database
- Anthropic Said DeepSeek, Moonshot, and MiniMax Used Fraudulent Accounts and Proxies to Illicitly Distill Claude Capabilities at Scale
- DJI Romo Cloud Authorization Bug Reportedly Exposed Camera, Microphone, and Home-Mapping Data From Nearly 7,000 Robot Vacuums
- Moltbook Database Exposure Allegedly Revealed Users' Private Communications and API Authentication Tokens