ThinkPatternGet the app
Story
TECHNOLOGY · APR 13, 2026

News Publishers Block Internet Archive to Stop AI Scraping

Major media organizations and Reddit are blocking the Wayback Machine's web crawlers to prevent AI companies from scraping content in violation of copyright law.

A growing number of major news organizations and digital platforms are blocking the Internet Archive and its Wayback Machine web crawlers. Analysis by Originality AI identified 241 news sites across nine countries restricting these bots, with USA Today Co. owning 87% of those sites. The New York Times has implemented hard blocking, while Reddit barred the crawler in August 2025 to protect content it is now licensing.

Publishers argue that AI companies use the nonprofit repository as a backdoor to scrape data for large language models, bypassing direct site blocks to compete with original news sources. USA Today Co. stated its restrictions target general scraping bots rather than the archive specifically. In response, a coalition including the Electronic Frontier Foundation and Fight for the Future gathered signatures from over 100 journalists, including Rachel Maddow, to urge the preservation of digital news records.

Mark Graham, director of the Wayback Machine, described the archive as collateral damage in the copyright war between publishers and AI firms. He criticized the irony of outlets using the archive for their own investigative research while simultaneously blocking it. Graham asserts that the archive has controls to limit AI abuse and warns that locking down the public web hinders society's ability to understand global events. He remains in negotiations with publishers to restore access.


Reported across 6 outlets
Actors
USA TODAY Co.The New York TimesRedditMark Graham

Keep reading in the app

The full story and every source, free in the app.

Download on the App StoreComing soonGoogle Play