| |
Creepy Crawlies
AI crawlers are consuming massive resources on git.kernel.org by inefficiently scraping HTML-rendered commits instead of simply cloning repositories, with 14 CPU cores dedicated solely to serving scrapers at any given time. The Linux kernel's publicly available git history is highly valuable training data for language models because it's guaranteed to be free of AI-generated content, making it a prime target for LLM training. The scrapers' approach generates billions of duplicate requests across the site's 922 kernel forks, creating significant "background radiation" of system load that exceeds all legitimate access combined.
Read Full Article →
← More Tech news