Fall ResetAmazon USFall reset deals: check better picks before checkoutAmazon US: today's deals, useful picks and quick comparisons.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanFall ResetAmazon USWork and home upgrades are worth comparing todayAmazon US: today's deals, useful picks and quick comparisons.See Picks×
Skip to content
Blog

Meta May Have Torrented More Than 80 Terabytes of Data From Pirated-Book Repositories for AI Training

By TheFinanceBase Team6 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Short answer: Authors suing Meta say court evidence shows the company used BitTorrent to obtain more than 80 terabytes of data from repositories associated with unauthorized books and other publications while developing its Llama AI models. Those figures describe alleged data volume—not a verified count of unique books—and they do not establish that every file was used for training or that a court found Meta liable for piracy.

The lawsuit concerns several legally distinct acts: downloading copies, using material to train a model, and possibly uploading file pieces to other torrent users. Meta won summary judgment on the named authors’ training-copying claim in June 2025, but the court did not declare all AI training on copyrighted works lawful. A separate dispute over alleged uploading and distribution remained active in the court’s March 25, 2026 order.

What the reported 80-terabyte figures mean

The headline figures come from claims and evidence presented in Kadrey v. Meta, a copyright case brought by authors including Richard Kadrey, Sarah Silverman and Christopher Golden. In February 2025, plaintiffs described internal records showing at least 81.7 terabytes torrented through Anna’s Archive, including at least 35.7 terabytes associated with Z-Library and LibGen. They also described an earlier episode involving about 80.6 terabytes from LibGen. Ars Technica’s account of the figures summarizes the plaintiffs’ claims; these are not a final judicial measurement or finding of liability.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not add 81.7 TB and 80.6 TB together as if they were necessarily separate, non-overlapping collections. They refer to different torrenting episodes or datasets, and the available figures do not establish how much overlap there may be. Later litigation materials also referred to 134.6 TB of aggregate download activity through July 2024, but that is a different total and should not be read as 134.6 TB of unique books used in training.

A terabyte measures data, not titles. The transfers could include duplicates, different editions or formats, metadata, compressed archives, scientific papers, images, incomplete files or material that was never added to a training set. The figures alone do not tell us how many files were copyrighted, how many belonged to the plaintiffs, or what portion—if any—went into a particular Llama training run.

How torrenting fits into the dispute

BitTorrent is a peer-to-peer file-transfer protocol. Instead of downloading a whole file from one central server, a client can obtain pieces from multiple users. Torrent clients can also upload pieces to other peers while downloading, depending on their configuration and use. That two-way traffic is why the lawsuit separates acquiring files from potentially sharing them.

The authors allege that Meta’s torrenting could have caused its systems to upload pieces of copyrighted works to other users. Meta has argued that it took precautions intended to prevent “seeding,” or uploading downloaded material. Whether any protected pieces were actually uploaded, what those pieces were, and what legal consequences follow are separate questions. A torrent transfer does not by itself prove that Meta intentionally published complete books or that another user obtained a complete copy from Meta.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The court’s June 25, 2025 order explains the distinction between downloading and uploading and treats them as legally distinct acts. In broad terms, the dispute can be separated this way:

Conduct at issue Question it raises
Copying works into a dataset Was the copying infringement, or did it qualify as fair use on the evidence and claims before the court?
Using copies in model development Which material was actually used, and what evidence bears on market harm or other fair-use factors?
Model outputs Do outputs reproduce protected expression or substitute for the authors’ works?
Uploading torrent pieces Did Meta distribute protected material, and what evidence supports that claim?
Facilitating further sharing Could Meta be liable for contributing to other users’ infringement?

What the lawsuit says about Meta’s data choices

The plaintiffs’ account is that Meta’s AI organization considered using LibGen, a large shadow-library repository associated with unauthorized copies of books and other publications. Reporting on internal communications says employees raised concerns about the material’s provenance and discussed whether it could be licensed or otherwise used. The plaintiffs also say communications show that Mark Zuckerberg approved proceeding with LibGen-derived material despite internal objections. The Guardian and TechCrunch reported on those allegations and communications.

That account should not be stretched into a finding that Zuckerberg personally ordered illegal downloads or knew the details of the torrent configuration. Approval of a data-use decision, knowledge about a repository’s provenance, and knowledge of particular downloading or uploading mechanics are not automatically the same thing. The court filings and reporting describe contested evidence, not a final finding against an individual.

Meta’s position has included that the plaintiffs’ training-copying claim failed on the evidence presented, that the authors had not shown relevant market harm, and that Meta took precautions against seeding. Those are litigation arguments and positions, not proof that every disputed act was lawful. TorrentFreak reported Meta’s contention that it made sure not to seed the books in a February 2025 report.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the June 2025 ruling decided—and what it did not

On June 25, 2025, Judge Vince Chhabria granted Meta summary judgment on the named authors’ claim that copying their books for LLM training infringed copyright, based on the record in that case. The court emphasized the plaintiffs’ lack of sufficient evidence that Meta’s copying caused or would cause the relevant market dilution. The ruling therefore turned substantially on the evidence the plaintiffs had developed about market harm.

It was not a blanket ruling that AI training on copyrighted works is always fair use. It did not decide every author’s potential claim, establish that all material acquired by Meta was lawfully obtained, or resolve the separate allegation that Meta’s torrenting systems uploaded copyrighted pieces. The order itself distinguishes the training-copying issue from the distribution question. Read the court’s order.

Where the case stood in March 2026

In an order dated March 25, 2026, the court allowed the authors to file an amended complaint based on newly produced evidence concerning Meta’s uploading and distribution activity while torrenting. The order said the distribution claim had not been resolved on summary judgment and identified separate distribution and contributory-infringement theories still at issue. It also denied class discovery for the time being. The March 25 order shows why Meta’s win on the named plaintiffs’ training claim did not end the whole dispute.

As of that order, the court had not entered a final ruling resolving those separate distribution and contributory-infringement theories. The procedural snapshot here is dated to March 25, 2026; it should not be mistaken for a claim that the case could not have changed afterward.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What remains unestablished

  • The number of unique titles represented by the reported data volumes.
  • What fraction of the acquired data consisted of copyrighted works, as opposed to other material or duplicates.
  • Which specific files were selected for, and actually consumed by, each training run.
  • Whether Meta’s systems uploaded copyrighted pieces, and whether any other user completed a download from those pieces.
  • Whether the alleged conduct caused measurable market harm and what the final outcome of the remaining claims will be.

The distinction between acquisition and training matters: obtaining a dataset does not by itself prove that every file entered a model’s training corpus. Likewise, training-related copying, possible reproduction in outputs, and torrent-based distribution are different factual and legal questions.

What readers can accurately say

The court record supports saying that plaintiffs presented evidence of very large-scale torrenting from repositories associated with unauthorized books and publications while Meta was developing AI models. It does not support saying that a court found Meta “guilty of piracy,” that every downloaded file was a book, or that all of the data was used to train Llama. Meta prevailed on the named authors’ training-copying claim at summary judgment, while the separate distribution allegations remained unresolved in the court’s March 2026 order.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Written by TheFinanceBase Team

The Team behind TheFinanceBase.

Add your note

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.