AIPolicyTracker

AI incident ·

Study Highlights Persistent Hallucinations in Legal AI Systems

2 news reports Snapshot 7 Sep 2026

In brief

An AI system built by Thomson Reuters and Lexisnexis and deployed by Legal Professionals, Law Firms and 1 other allegedly harmed Legal Professionals, Clients Of Lawyers and 1 other.

Risk domain
Misinformation False or misleading information
Occurred
Coverage
2 reportsMay 2024

What happened

Stanford University’s Human-Centered AI Institute (HAI) conducted a study in which they designed a "pre-registered dataset of over 200 open-ended legal queries" to test AI products by LexisNexis (creator of Lexis+ AI) and Thomson Reuters (creator of Westlaw AI-Assisted Research and Ask Practical Law AI). The researchers found that these legal models hallucinate in 1 out of 6 (or more) benchmarking queries.

Laws that address this harm

Policy angle: Classified under Misinformation (False or misleading information) in the MIT AI Risk Repository taxonomy; 5 recorded instruments address this use case.

Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).

News reports (2)

Titles link to the original publisher; report text is not reproduced here.

  1. AI on Trial: Legal Models Hallucinate in 1 out of 6 (or More) Benchmarking Queries
    hai.stanford.edu · Varun Magesh, Faiz Surani, Matthew Dahl
  2. We asked ChatGPT for legal advice—here are five reasons why you shouldn't
    theconversation.com · Francine Ryan, Elizabeth Hardie

Who was involved

Alleged developer
Thomson Reuters, Lexisnexis
Alleged harmed party
Legal Professionals, Clients Of Lawyers, Legal System

Classification (MIT AI Risk Repository taxonomy)

Risk domain
Misinformation
Causal entity
AI
Intent
Unintentional
Timing
Post-deployment
Harm level
—
Sectors
—
Countries
—

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 3.1.

  • Pursuing Consistent Context

    "LLMs have been demonstrated to pursue consistent context [129]–[132], which may lead to erroneous generation when the prefixes contain false information. Typical examples include sycophancy [129], [130], false demonstra...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Defective Decoding Process

    In general, LLMs employ the Transformer architecture [32] and generate content in an autoregressive manner, where the prediction of the next token is conditioned on the previously generated token sequence. Such a scheme...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Noisy Training Data

    "Another important source of hallucinations is the noise in training data, which introduces errors in the knowledge stored in model parameters [111]–[113]. Generally, the training data inherently harbors misinformation....

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Hallucinations

    "LLMs generate nonsensical, untruthful, and factual incorrect content"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Faithfulness Errors

    "The LLM-generated content could contain inaccurate information" which is is not true to the source material or input used

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Factuality Errors

    "The LLM-generated content could contain inaccurate information" which is factually incorrect

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Untruthful Content

    "The LLM-generated content could contain inaccurate information"

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

  • Knowledge Gaps

    "Since the training corpora of LLMs can not contain all possible world knowledge [114]–[119], and it is challenging for LLMs to grasp the long-tail knowledge within their training data [120], [121], LLMs inherently posse...

    Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems (Cui2024)

Incidents in the same risk subdomain

All incidents in this subdomain

Source record: incident #704 on the AI Incident Database · all 2 reports