AI incident #1135 ·

Preprints Reportedly from Researchers from Multiple Universities Allegedly Contain Covert AI Prompts

What happened

Hidden prompts reportedly were discovered in at least 17 academic preprints on arXiv that purportedly instructed AI tools to deliver only positive peer reviews. The lead authors are reportedly affiliated with 14 institutions in eight countries, including Waseda University, KAIST, Peking University, and the University of Washington. The alleged concealed instructions, some of which were reportedly embedded using white text or tiny fonts, were purportedly intended to influence any reviewers who re

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (1)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. 'Positive review only': Researchers hide AI prompts in papers
    asia.nikkei.com · Shogo Sugiyama, Ryosuke Eguchi · AIID #5475

Who was involved

Alleged harmed party
Peer Review Process, Academic Integrity, Academic Conferences, Research Community

Classification (MIT AI Risk Repository taxonomy)

Causal entity
Human
Intent
Intentional
Timing
Pre-deployment
Harm level
Sectors
Countries

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 4.3.

  • Impersonation/identity theft

    "Impersonation/identity theft - Theft of an individual, group or organisation’s identity by a third-party in order to defraud, mock or otherwise harm them."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • IP/copyright loss

    "IP/copyright loss - Misuse or abuse of an individual or organisation’s intellectual property, including copyright, trademarks, and patents."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Dehumanisation/objectification

    "Dehumanisation/objectification - Use or misuse of a technology system to depict and/or treat people as not human, less than human, or as objects."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Defamation/libel/slander

    "Defamation/libel/slander - Use of a technology system to create, facilitate or amplify false perception(s) about an individual, group, or organisation."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Financial and business

    "Financial and Business - Use or misuse of a technology system in a manner that damages the financial interests of an individual or group, or which causes strategic, operational, legal or financial harm to a business or...

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Cheating/plagiarism

    "Cheating/plagiarism - Use of another person’s or group’s words or ideas without consent and/or acknowledgement."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Cybersecurity

    "LLMs may exacerbate cybersecurity risks in various ways (Newman, 2024). Firstly, LLMs may significantly amplify the effectiveness of deceptive operations aimed at tricking people into disclosing sensitive information or...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Domain-Specific Misuses

    "Improvements in LLMs may exert greater pressure to apply LLMs to various domains, such as health and education (Eloundou et al., 2023). Crude efforts to use LLMs in such domains, however, may incur harm and should be di...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

Incidents in the same risk subdomain

All incidents in this subdomain