A growing number of news websites are instructing Internet Archive crawlers not to capture their pages. Nieman Journalism Lab’s May 20, 2026 analysis found 382 sites with at least one such robots.txt directive, up from 241 in its January analysis. But the headline needs qualification: 342 of the 382 were local news sites, 93% were in the United States, and a robots.txt instruction shows a requested crawler policy—not that every bot was technically prevented from accessing a page.
What the latest count actually shows
Nieman Lab analyzed journalist Ben Welsh’s database of 1,167 news-site robots.txt files for its January review, then checked additional files for its May update. The updated sample covered news sites in ten countries. It identified 382 websites disallowing at least one of seven Internet Archive-associated crawler names; 141 sites were added to the January count of 241.
The distribution matters. Of the 382 sites, 342 were local news outlets, although many belonged to large chains. That makes “major publishers” an incomplete description of the trend. The sample is not a census of every news organization, and its figures describe the sites and method Nieman Lab examined rather than the entire news industry.
Why a robots.txt rule is not proof of a successful block
Robots.txt is a public set of instructions asking automated crawlers what they may access. Nieman Lab counted sites that disallowed at least one name associated with Internet Archive crawling; it did not establish that every directive stopped a crawler in practice.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
There is also an attribution complication. Wayback Machine founder Mark Graham told Nieman Lab that Wayback does not use the names “ia_archiver,” “ia_archiverbot” or “ia_archiver-web.archive.org.” The analysis nevertheless included “ia_archiver-web.archive.org” because publishers were blocking it under the assumption that the Internet Archive used it. A listed disallow rule therefore should not automatically be described as a verified Wayback block.
Why publishers say they are restricting the crawlers
Publishers gave related but not identical explanations. The recurring concern is that artificial-intelligence companies could obtain archived journalism for model training without permission or payment. Other stated goals include protecting the commercial value of reporting, preserving licensing leverage and requiring AI products to identify and link to the publisher that produced the information.
Protection from broader third-party use
Advance Local spokesperson Christine deWit described the policy this way: “This is part of a broader effort to protect the value of our published work from unfair third‑party use. This decision is not specific to the Wayback Machine.” That explanation treats the crawler rule as part of a wider rights and commercial strategy, not as a judgment about one archive alone.
Rank #2
Attribution and links back to original reporting
The Baltimore Banner said it was particularly concerned about whether AI products would direct users to the original reporting. Its chief technology officer and AI strategist, Biswajit Ganguly, told Nieman Lab: “The threat is definitely not the Internet Archive,” distinguishing the nonprofit archive from the AI services publishers fear may reuse journalism without adequate attribution.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →What has not been demonstrated
Nieman Lab reported that, as of May 20, 2026, no news publisher had confirmed to it that an AI company had already scraped its content from the Wayback Machine. AI reuse is therefore a stated risk motivating these policies, not a confirmed event in the cases reported.
Why archived news matters
Online journalism is unusually easy to alter or remove. A later version may have a different headline, wording, correction note, data table or publication date; a site migration can break old links; and a publication can shut down entirely. Archived pages give readers and researchers a way to examine what was publicly available at an earlier time.
Journalists use local-news archives when reporting follow-up stories and checking the history of public claims. Nieman Lab described articles lost during site migrations and the disappearance of a defunct publication’s archive as examples of why independent preservation matters.
Edward McCain, a journalism librarian at the University of Missouri, told Nieman Lab: “Blocking the Internet Archive’s web crawlers threatens one of the most effective ways that we capture and store news content for the long term,”
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchInternet Archive Europe argues that blocking reduces access to a public historical record. That is the Archive’s advocacy position, not an independent measurement of the effect. In its June 9, 2026 commentary, Internet Archive Europe said the Wayback Machine held more than one trillion archived web pages and preserved permanent citations for nearly five million news articles referenced on Wikipedia. It also said more than 250 journalists had signed an open letter about the issue.
Can readers still find an old news article?
Sometimes, but no single route is complete. Work through the options in this order:
Rank #4
- Check the publisher’s own archive. Search by the article’s headline, author, date or topic. This may provide the current page or a publisher-preserved earlier version, but availability and retention are controlled by that organization.
- Search the Wayback Machine. Enter the article’s page address and inspect available capture dates. A missing capture can mean the page was never archived, the crawl was disallowed, the page required access the crawler could not obtain, or the capture is otherwise unavailable.
- Use a library or institutional database. ProQuest and LexisNexis are paid services that may be available through public libraries, universities or individual subscriptions. Their holdings and access terms vary.
- Ask the newsroom or a journalist. A newsroom may retain internal copies, syndicated versions or migration backups that are not publicly indexed.
- Check contemporaneous citations. Government filings, court records, academic papers and other articles may quote or describe the original reporting, although a secondary reference is not the same as a preserved page.
None of these methods is established as a complete replacement for open web archiving. A publisher-controlled archive can disappear with the publisher, while a commercial database can be inaccessible to readers without a subscription or library access.
How the main preservation choices compare
| Option | Access | Breadth and continuity | Who controls retention | Cost and technical capacity | Earlier versions and citations |
|---|---|---|---|---|---|
| Wayback Machine | Generally public when a capture exists | Large, cross-site collection, but coverage depends on crawling, permissions and available captures | Internet Archive, subject to site rules and operational decisions | Readers do not need a paid subscription; the archive operates large-scale crawling and storage infrastructure | Can expose earlier page states and stable archive references; availability is not guaranteed |
| Publisher’s own archive | Usually public for pages the publisher chooses to retain; access policy varies | Potentially strong for that publisher, but continuity depends on migrations, budgets and whether the organization remains active | The publisher | Requires newsroom systems, maintenance and preservation planning; reader cost not stated | Can preserve authoritative versions, but earlier revisions may not be exposed |
| ProQuest or LexisNexis | Subscription or institutional access | Broad licensed collections, with coverage varying by title and date | The database provider and its licensing agreements | Paid service; access may come through a library, university or individual subscription | Useful for retrievable published records; version history and permanent public links are not established for every item |
| Newsroom archiving strategy | Public access depends on what the newsroom publishes | Can target the outlet’s own work and records, but scope and continuity vary | The newsroom or its chosen preservation partners | Requires staff time, policy, storage and technical capacity; no standard cost is established | Can support reliable internal records and citations if designed to retain versions, but practices differ |
What the restrictions mean for accountability
Fewer independent checks on changing stories
When an independent capture is unavailable, readers may have difficulty determining how a story changed, whether a correction was added, or what evidence was presented at a particular moment. That affects accountability for journalism and for the public officials, companies and institutions covered by it.
Recommended Free Tools
More fragile links in public records
Researchers, courts, educators and journalists often cite web pages that were accessible when they wrote. If the publisher later removes or restructures the page, the citation may no longer resolve. Open archives help preserve a reference outside the publisher’s own systems; blocking them shifts more responsibility to institutions that may have fewer resources or shorter retention policies.
Best Value
A trade-off between rights protection and public memory
Publishers are responding to a real policy dispute over unauthorized AI use, compensation and attribution. Yet a rule aimed at limiting third-party reuse can also reduce the independent preservation of journalism. The policy choice is not simply “archive versus no archive”: it determines who can access historical reporting, under what terms, and for how long.
What newsrooms and researchers can do
Newsrooms
- Maintain a first-party archive that records publication dates, corrections and major revisions.
- Document migration plans so old URLs, metadata and attachments remain findable.
- Set crawler policies deliberately and distinguish AI-service controls from preservation arrangements.
- Participate in professional training and shared preservation programs where resources allow.
Nieman Lab reported a December partnership among the Internet Archive, the Poynter Institute and Investigative Reporters and Editors. The initial cohort included 33 local and national outlets, with an aim to train 300 newsrooms by the end of 2027.
Researchers and readers
- Save the article’s headline, author, date and page address when you rely on it.
- Record the date you accessed a page and preserve any correction or revision notice.
- Use more than one preservation route for consequential work, because each option has different gaps.
- Label an archived copy as an earlier version rather than assuming it is the publisher’s current text.
The bottom line on “major publishers” blocking Wayback
The documented increase is substantial within Nieman Lab’s sample: 382 news websites had robots.txt rules disallowing at least one Internet Archive-associated crawler in May 2026, compared with 241 in January. Most were local outlets, and the rules are evidence of requested restrictions—not proof that every Wayback capture was technically blocked. Publishers cite AI scraping fears, commercial protection and attribution, but no publisher had confirmed AI scraping of its Wayback captures to Nieman Lab by May 20, 2026. The immediate consequence is a more fragile historical record of journalism, with readers increasingly dependent on publisher archives, paid databases and newsroom preservation practices that do not fully replace an open, independent archive.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




