AI incident #1001 ·

LLM Scrapers Allegedly Target Multiple Open Source Projects Disrupting the FOSS Ecosystem

What happened

In mid-March 2025, KDE's GitLab infrastructure was reportedly disrupted by aggressive AI web scrapers originating from Alibaba IP ranges. These bots allegedly ignored robots.txt and spoofed browser headers, which in turn purportedly overwhelmed the site and caused outages for developers. Similar incidents reportedly affected other FOSS projects like GNOME, SourceHut, and Fedora. The scraping is allegedly tied to large language model training, and reportedly imposes real costs and delays.

Only the incident metadata is stored here. The underlying news reports are on the AI Incident Database (CC BY-SA 4.0); use the links above to read them.

News reports (2)

Coverage catalogued by the AI Incident Database. Titles link to the original publisher; the text is not reproduced here.

Who was involved

Alleged harmed party
Sysadmins, Sourcehut, Read The Docs, Linux Weekly News, Kde, Inkscape, Gnome, Foss Projects And Communities, Fedora, Diaspora, Curl

Classification (MIT AI Risk Repository taxonomy)

Causal entity
AI
Intent
Intentional
Timing
Post-deployment
Harm level
Sectors
Countries

Risk entries describing this failure mode

Entries from the MIT AI Risk Repository coded to subdomain 6.1.

  • Monopolisation

    "Monopolisation - Abuse of market power through the control of prices, thereby limiting competition and creating unfair barriers to entry."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Power concentration

    "Power concentration - Amplification of concentration of economic and/or political wealth and power, potentially resulting in increased inequality and instability."

    A Collaborative, Human-Centred Taxonomy of AI, Algorithmic, and Automation Harms (Abercrombie2024)

  • Potential exploitation by totalitarian regimes

    The Rise of Artificial Intelligence - Future Outlooks and Emerging Risks (Allianz2018)

  • Corporate power may impeded effective governance

    "The increasing power and influence of large corporations may make effective governance difficult. There exists a power asymmetry between corporate entities profiting from LLMs and other social groups (e.g. civil society...

    Foundational Challenges in Assuring Alignment and Safety of Large Language Models (Anwar2024)

  • Global AI Divide

    "General- purpose AI research and development is currently concentrated in a few Western countries and China. This ‘AI Divide’ is multicausal, but in part related to limited access to computing power in low- income count...

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Market concentration risks and single points of failure

    "Market power is concentrated among a few companies that are the only ones able to build the leading general- purpose AI models. Widespread adoption of a few general- purpose AI models and systems by critical sectors inc...

    International Scientific Report on the Safety of Advanced AI (Bengio2024)

  • Market concentration and single points of failure

    "Market shares for general- purpose AI tend to be highly concentrated among a few players, which can create vulnerability to systemic failures. The high degree of market concentration can invest a small number of large t...

    International AI Safety Report 2025 (Bengio2025)

  • Global AI R&D divide

    "Large companies in countries with strong digital infrastructure lead in general- purpose AI R&D, which could lead to an increase in global inequality and dependencies. For example, in 2023, the majority of notable gener...

    International AI Safety Report 2025 (Bengio2025)

Incidents in the same risk subdomain

All incidents in this subdomain