AI incident #1001 ·
LLM Scrapers Allegedly Target Multiple Open Source Projects Disrupting the FOSS Ecosystem
What happened
In mid-March 2025, KDE's GitLab infrastructure was reportedly disrupted by aggressive AI web scrapers originating from Alibaba IP ranges. These bots allegedly ignored robots.txt and spoofed browser headers, which in turn purportedly overwhelmed the site and caused outages for developers. Similar incidents reportedly affected other FOSS projects like GNOME, SourceHut, and Fedora. The scraping is allegedly tied to large language model training, and reportedly imposes real costs and delays.
Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.
News reports (2)
Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.
Who was involved
- Alleged deployer
- Alibaba, Generative Ai Developers
- Alleged developer
- Alibaba, Generative Ai Developers
- Alleged harmed party
- Sysadmins, Sourcehut, Read The Docs, Linux Weekly News, Kde, Inkscape, Gnome, Foss Projects And Communities, Fedora, Diaspora, Curl
Classification (MIT AI Risk Repository taxonomy)
- Risk domain
- Socioeconomic & Environmental Harms
- Causal entity
- AI
- Intent
- Intentional
- Timing
- Post-deployment
- Harm level
- —
- Sectors
- —
- Countries
- —
Risk entries describing this failure mode
Entries from the MIT AI Risk Repository coded to subdomain 6.1.
- Monopolisation
"Monopolisation - Abuse of market power through the control of prices, thereby limiting competition and creating unfair barriers to entry."
- Power concentration
"Power concentration - Amplification of concentration of economic and/or political wealth and power, potentially resulting in increased inequality and instability."
- Potential exploitation by totalitarian regimes
- Corporate power may impeded effective governance
"The increasing power and influence of large corporations may make effective governance difficult. There exists a power asymmetry between corporate entities profiting from LLMs and other social groups (e.g. civil society...
- Global AI Divide
"General- purpose AI research and development is currently concentrated in a few Western countries and China. This ‘AI Divide’ is multicausal, but in part related to limited access to computing power in low- income count...
- Market concentration risks and single points of failure
"Market power is concentrated among a few companies that are the only ones able to build the leading general- purpose AI models. Widespread adoption of a few general- purpose AI models and systems by critical sectors inc...
- Market concentration and single points of failure
"Market shares for general- purpose AI tend to be highly concentrated among a few players, which can create vulnerability to systemic failures. The high degree of market concentration can invest a small number of large t...
- Global AI R&D divide
"Large companies in countries with strong digital infrastructure lead in general- purpose AI R&D, which could lead to an increase in global inequality and dependencies. For example, in 2023, the majority of notable gener...
Incidents in the same risk subdomain
- Coupang Allegedly Tweaked Search Algorithms to Boost Own Products
- Amazon Allegedly Tweaked Search Algorithm to Boost Its Own Products
- Gmail’s Inbox Sorting System Reportedly Reduced Visibility of Political Emails and Campaign Calls-to-Action
- Google Fined for Changing Shopping Algorithms in EU to Favor Own Service
- Amazon India Allegedly Rigged Search Results to Promote Own Products