Free tools Windows power users keep installed
One-click scans. No signup required.
Yes. Microsoft was sued on June 24, 2025, in the U.S. District Court for the Southern District of New York. In Bird et al. v. Microsoft Corp., No. 1:25-cv-05282, a group of named authors alleges that Microsoft used approximately 200,000 books from the Books3 collection to train its Megatron-related language models. The authors characterize those copies as pirated. The allegations have not been adjudicated, and the case remained unresolved as of August 18, 2026.
The lawsuit is a standalone action focused on Microsoft’s own Megatron work—not the separate author litigation involving OpenAI in which Microsoft was later added.
What the authors allege
According to reporting on the complaint, the plaintiffs say Microsoft copied copyrighted books without permission, obtained them through a repository they describe as pirated, and used the text in training the Megatron-Turing Natural Language Generation model family. They also allege that the resulting system could generate text reflecting expressive features such as syntax, voice, style, themes, or passages from the books.
Those are allegations in a complaint, not findings that Microsoft infringed. The legal theory can concern several different acts: downloading or storing a book, preparing a dataset, making training copies, adjusting model parameters, and generating an output that resembles source text. A plaintiff does not necessarily have to show a verbatim generated book to argue that unauthorized copies were made during dataset preparation or training.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Books3 and The Pile are not the same thing
The complaint reportedly identifies Books3, a collection of roughly 200,000 books, as the relevant source. Books3 was associated with The Pile, a much broader dataset assembled by EleutherAI. Calling Books3 allegedly pirated does not establish that every item in The Pile was pirated.
Dataset membership also is not conclusive proof that Microsoft obtained a particular file, included it in a final training run, or caused the model to memorize it. The authors would still need to connect their individual works to the alleged copying and to Microsoft’s model development.
Reuters described the allegation and the model at its report on the filing. Additional coverage appears in Bloomberg Law and Sherwood News.
Who sued Microsoft?
Reported plaintiffs include authors from both fiction and nonfiction, among them:
- Kai Bird
- Jonathan Alter
- Mary Bly
- Eugene Linden
- Daniel Okrent
- Hampton Sides
- Jia Tolentino
- Victor LaValle
- Rachel Vail
- Simon Winchester
Coverage describes the action as seeking to represent a broader group of affected copyright owners. The complete plaintiff list and each author’s ownership, registration, and inclusion allegations would be determined from the filed pleadings and subsequent court filings.
Rank #2
What Megatron is—and what this case is not
Megatron-Turing Natural Language Generation was developed through Microsoft and NVIDIA research. It is a large language model designed to generate responses to prompts. The lawsuit concerns Microsoft’s alleged use of books in that Megatron-related training pipeline; it is not a claim that Megatron is the same product as ChatGPT or Microsoft Copilot.
Microsoft’s partnership and investment relationship with OpenAI can make the company’s name appear in other AI copyright cases. That relationship does not automatically make Microsoft responsible for OpenAI’s alleged conduct.
What legal issues will matter
Reproduction during training
The central copyright question is whether making and using copies of protected books to prepare and train an AI model infringes the authors’ exclusive reproduction rights. The dispute is about the underlying copies as well as any later output.
Fair use
Microsoft could argue that training is a transformative, technologically intermediate use and therefore fair use. The authors can argue that the copying was unauthorized, consumed expressive works, and may enable competing text generation. Fair use is a fact-specific defense; neither “AI training is illegal” nor “AI training is always fair use” is settled law.
Source provenance
The alleged acquisition method is especially important. Copying a lawfully obtained book for a potentially transformative purpose presents a different factual question from copying a book obtained from a repository the plaintiffs say was pirated. The authors may contend that unlawful acquisition weighs against Microsoft’s fair-use defense and supports remedies.
Rank #3
Proof concerning each work
The plaintiffs would generally need to establish valid rights in their works, unauthorized copying attributable to Microsoft, and viable claims under applicable registration, timing, standing, and causation rules. If they pursue class treatment, they would also have to satisfy the separate requirements for class certification.
Possible defenses—unless confirmed in Microsoft’s actual filings—include that the plaintiffs cannot show their specific books were used in the relevant training run, that Microsoft did not create or knowingly use every alleged source file, that the model does not retain protected expression in the claimed manner, and that individualized issues defeat a proposed class.
Why the 2025 Anthropic ruling is relevant but not controlling
Reuters reported that a June 2025 Anthropic decision treated training on lawfully acquired books differently from the alleged use of pirated copies. That distinction gives the Microsoft plaintiffs a potentially important argument about provenance, but the decision did not decide Microsoft’s liability.
The defendants, datasets, acquisition facts, claims, and procedural postures differ. The Anthropic ruling was not blanket approval for AI training and does not establish that Books3 use by Microsoft was infringing.
What the authors are seeking
- Statutory damages: Reuters reported that the complaint seeks up to $150,000 per infringed work where the statutory requirements are met.
- An injunction: The plaintiffs seek an order stopping continued infringement or further use of the allegedly infringing material.
- Other relief: They seek additional remedies the court considers appropriate.
$150,000 is a potential statutory maximum per work in appropriate circumstances, not an automatic payment to each author. The available amount can depend on issues such as registration, willfulness, the infringement proven, and the court’s damages findings.
Rank #4
How this differs from other AI copyright cases
| Case or litigation | Main target | Primary issue | Relationship to Bird v. Microsoft |
|---|---|---|---|
| Bird v. Microsoft | Microsoft | Megatron-related models; Books3 and The Pile allegations | Standalone author action focused on Microsoft |
| Authors Guild/OpenAI litigation | OpenAI, with Microsoft later added in separate actions | Alleged use of fiction and nonfiction books to train OpenAI systems | Separate litigation, consolidated for pretrial purposes according to the Authors Guild |
| Anthropic author litigation | Anthropic | Claude training and acquisition of books | Provides fair-use and provenance context, but is not controlling here |
| NVIDIA author litigation | NVIDIA | NeMo Megatron tools and alleged dataset use or facilitation | Related technology and dataset questions involving a different defendant |
The Authors Guild’s overview of the separate proceedings is available at authorsguild.org/advocacy/artificial-intelligence/.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Microsoft’s reported response
Reuters reported that Microsoft did not immediately respond to a request for comment when the lawsuit was first reported. That is not an admission of liability. No later position—such as a verified denial, dispute over Books3, or assertion that the authors’ books were not in the relevant training set—should be attributed to Microsoft without a filing or statement supporting it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Current status
As of August 18, 2026, no final judgment resolving the authors’ claims against Microsoft had been identified in the available case reporting. A litigation tracker reports that proceedings were stayed on September 9, 2025 and describes the matter as remaining at the pleading stage as of May 2026. A stay pauses proceedings; it is not a dismissal or a ruling that Microsoft won.
For updates, the Mishcon de Reya tracker and Manuscript Report tracker provide reported status information, which should be checked against the latest Southern District of New York docket before relying on it.
What the case could mean for authors and AI developers
Training-data licensing
A ruling could influence whether companies license books directly, document permissions, or accept greater litigation risk when assembling large corpora.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Best Value
- Anti -Artificial Intelligence. Human intelligence as a collegiate design inspired by the original intelligence. Vintage varsity styling highlights creativity, curiosity, critical thinking, and authentic human ideas in the age of artificial intelligence.
- Clean university inspired graphic for programmers, engineers, educators, students, artists, writers, designers, creators, and anyone who believes human intelligence remains timeless, original, and worth celebrating.
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
Provenance and recordkeeping
Companies may face stronger pressure to track where each dataset component came from, whether files were lawfully obtained, and which versions entered a training run.
Separate analysis of training and outputs
Courts may treat the creation of training copies, model behavior, and allegedly substitutive or memorized outputs as related but distinct questions. Evidence about outputs can help show memorization or market effects without being the only evidence of copying.
Authors’ bargaining position
The dispute highlights practical concerns about consent, compensation, and whether creators can obtain meaningful relief when their works are used at scale without a license.
Bottom line
Microsoft was sued, but the lawsuit does not prove that Microsoft used pirated books or infringed the plaintiffs’ copyrights. Its significance lies in the focused question of whether alleged use of books from a pirated source changes the copyright and fair-use analysis for Megatron-related AI training. That question remains unresolved in this case.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




