| |
More than 340 local news outlets, including sites owned by major newspaper chains like McClatchy, Advance Local, and Tribune Publishing, are now blocking the Internet Archive's web crawlers to prevent AI companies from scraping their content for training data. The blocking has escalated since January 2026, when major publishers first restricted access over AI concerns, though no publisher has confirmed actual scraping has occurred. Researchers, historians, and working journalists argue that restricting the archive threatens the long-term preservation of news content and undermines access to primary source materials.
Read Full Article →
← More Tech news