AI incident #1497 ·

Hidden Prompt Injection in Brazilian Labor-Court Petition Reportedly Tried to Manipulate Galileu

What happened

Galileu, an AI tool used by Brazil's labor courts, reportedly detected hidden instructions embedded in an initial petition before the 3rd Labor Court of Parauapebas. The text allegedly told the AI to contest the petition superficially and not challenge documents. Galileu reportedly alerted the judge and blocked the hidden content from processing; the judge then reviewed the material before imposing any procedural consequences.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (1)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. Galileu: Sistema identifica tentativa de manipulação em petição e alerta magistrado
    trt4.jus.br · Secretaria-Geral de Tecnologia e Inovação, General Secretariat for Technology and Innovation · AIID #7314

Who was involved

Alleged harmed party
Epistemic Integrity, Judicial Integrity, Judicial System Of Brazil, Brazilian Labor Courts, Defendants In Brazilian Labor Cases

Classification (MIT AI Risk Repository taxonomy)

Causal entity
Human
Intent
Intentional
Timing
Post-deployment
Harm level
Sectors
Countries

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 2.2.

  • Privacy loss

    "Privacy loss - Unwarranted exposure of an individual’s private life or personal data through cyberattacks, doxxing, etc."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Exploiting Limited Generalization of Safety Finetuning

    "Safety tuning is performed over a much narrower distribution compared to the pretraining distribution. This leaves the model vulnerable to attacks that exploit gaps in the generalization of the safety training, e.g. usi...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Jailbreaks and Prompt Injections Threaten Security of LLMs

    "LLMs are not adversarially robust and are vulnerable to security failures such as jailbreaks and prompt-injection attacks. While a number of jailbreak attacks have been proposed in the literature, the lack of standardiz...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • “Model Psychology” Attacks

    "LLMs are vulnerable to “psychological” tricks (Li et al., 2023e; Shen et al., 2023), which can be exploited by attackers. Examples include instructing the model to behave like a specific persona (Shah et al., 2023; Andr...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Attacking LLMs via Additional Modalities a

    "LLMs can now process modalities other than text, e.g. images or video frames (OpenAI, 2023c; Gemini Team, 2023). Several studies show that gradient-based attacks on multimodal models are easy and effective (Carlini et a...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Vulnerability to Poisoning and Backdoors

    "The previous section explored jailbreaks and other forms of adversarial prompts as ways to elicit harmful capabilities acquired during pretraining. These methods make no assumptions about the training data. On the other...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Model Attacks

    Model attacks exploit the vulnerabilities of LLMs, aiming to steal valuable information or lead to incorrect responses.

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Hardware Vulnerabilities

    "The vulnerabilities of hardware systems for training and inferencing brings issues to LLM-based applications."

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

Incidents in the same risk subdomain

All incidents in this subdomain