AI incident ·
The New York Times Reportedly Sues OpenAI and Microsoft Over Alleged Unauthorized AI Training on Its Content
In brief
An AI system built and deployed by Openai and Microsoft allegedly harmed Writers, The New York Times and 5 others.
- Risk domain
- Privacy & Security
- Occurred
- Coverage
- 2 reports
What happened
The New York Times alleges that OpenAI and Microsoft used millions of its articles without permission to train AI models, including ChatGPT. The lawsuit claims the companies scraped and reproduced copyrighted content without compensation, in turn undermining the Times’s business and competing with its journalism. Some AI outputs allegedly regurgitate Times articles verbatim. The lawsuit seeks damages and demands the destruction of AI models trained on its content.
Laws that address this harm
Policy angle: Classified under Privacy & Security (Compromise of privacy by obtaining, leaking or correctly inferring sensitive information) in the MIT AI Risk Repository taxonomy; 5 recorded instruments address this use case.
- Colorado AI Act
- India DPDP Act
- Law No. 132/2025 on artificial intelligence
- EU AI Act
- Texas Responsible AI Governance Act (TRAIGA)
Matched from the record's risk domain and country to the instruments recorded here. A reviewer can correct the match in the repository (data/external/incident_overrides.yaml).
News reports (2)
Titles link to the original publisher; report text is not reproduced here.
Who was involved
- Alleged deployer
- Openai, Microsoft
- Alleged developer
- Openai, Microsoft
- Alleged harmed party
- Writers, The New York Times, Publishers, Media Organizations, Journalists, Journalistic Integrity, Epistemic Integrity
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Privacy & Security
- Risk subdomain
- 2.1 Compromise of privacy by obtaining, leaking or correctly inferring sensitive information
- Causal entity
- AI
- Intent
- Unintentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 2.1.
- Risks to privacy
"General- purpose AI models or systems can ‘leak’ information about individuals whose data was used in training. For future models trained on sensitive personal data like health or financial data, this may lead to partic...
- Risks to privacy
"General- purpose AI systems can cause or contribute to violations of user privacy. Violations can occur inadvertently during the training or usage of AI systems, for example through unauthorised processing of personal d...
- Private Training Data
"As recent LLMs continue to incorporate licensed, created, and publicly available data sources in their corpora, the potential to mix private data in the training corpora is significantly increased. The misused private d...
- Association in LLMs
"Association in LLMs refers to the capability to associate various pieces of information related to a person. According to [68], [86], given a pair of PII entities (xi , xj ), which is associated by a model F. Using a pr...
- Privacy Leakage
"Privacy Leakage means the generated content includes sensitive personal information"
- Memorization in LLMs
"Memorization in LLMs refers to the capability to recover the training data with contextual prefixes. According to [88]–[90], given a PII entity x, which is memorized by a model F. Using a prompt p could force the model...
- Privacy Leakage
"The model is trained with personal data in the corpus and unintentionally exposing them during the conversation."
- Privacy and regulation violations
"Some of the broken systems discussed above are also very invasive of people’s privacy, controlling, for instance, the length of someone’s last romantic relationship [51]. More recently, ChatGPT was banned in Italy over...
Incidents in the same risk subdomain
- Grok Reportedly Disclosed Adult Performer Siri Dahl's Legal Name and Birthdate, Allegedly Contributing to Doxxing and Harassment
- NPR Host David Greene Alleged Google's NotebookLM Replicated His Voice Without Consent, Prompting Lawsuit
- Border Patrol Agent Allegedly Claimed Facial Recognition Identified Minneapolis ICE Observer and Global Entry Was Reportedly Revoked Three Days Later
- Perplexity AI Reportedly Accused in Federal Lawsuit of Purported Copyright Infringement and False Attribution of Chicago Tribune Content
- Secret Desires AI Platform Reportedly Exposed Nearly Two Million Sensitive Images in Cloud Storage Leak
- ChatGPT Reportedly Found to Reproduce Protected German Lyrics in Copyright Case
Other incidents involving Openai, Microsoft
Source record: incident #995 on the AI Incident Database · all 2 reports