AI incident #1044 ·

Reported Emergence of 'Vegetative Electron Microscopy' in Scientific Papers Traced to Purported AI Training Data Contamination

What happened

Researchers reportedly traced the appearance of the nonsensical phrase "vegetative electron microscopy" in scientific papers to contamination in AI training data. Testing indicated that large language models such as GPT-3, GPT-4, and Claude 3.5 may reproduce the term. The error allegedly originated from a digitization mistake that merged unrelated words during scanning, and a later translation error between Farsi and English.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (2)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

  1. A weird phrase is plaguing scientific papers – and we traced it back to a glitch in AI training data
    theconversation.com · Aaron J. Snoswell, Kevin Witzenberger, Rayane El Masri · AIID #5095

Who was involved

Alleged developer
Openai, Anthropic
Alleged harmed party
Researchers, Scientific Authors, Scientific Publishers, Peer Reviewers, Scholars, Readers Of Scientific Publications, Scientific Record, Academic Integrity

Classification (MIT AI Risk Repository taxonomy)

Risk domain
Misinformation
Causal entity
AI
Intent
Unintentional
Timing
Post-deployment
Harm level
Sectors
Countries

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 3.1.

  • Pursuing Consistent Context

    "LLMs have been demonstrated to pursue consistent context [129]–[132], which may lead to erroneous generation when the prefixes contain false information. Typical examples include sycophancy [129], [130], false demonstra...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Knowledge Gaps

    "Since the training corpora of LLMs can not contain all possible world knowledge [114]–[119], and it is challenging for LLMs to grasp the long-tail knowledge within their training data [120], [121], LLMs inherently posse...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Hallucinations

    "LLMs generate nonsensical, untruthful, and factual incorrect content"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Faithfulness Errors

    "The LLM-generated content could contain inaccurate information" which is is not true to the source material or input used

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Defective Decoding Process

    In general, LLMs employ the Transformer architecture [32] and generate content in an autoregressive manner, where the prediction of the next token is conditioned on the previously generated token sequence. Such a scheme...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Untruthful Content

    "The LLM-generated content could contain inaccurate information"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Factuality Errors

    "The LLM-generated content could contain inaccurate information" which is factually incorrect

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Noisy Training Data

    "Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation....

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

Incidents in the same risk subdomain

All incidents in this subdomain