Fall ResetAmazon USFall reset deals: check better picks before checkoutAmazon US: today's deals, useful picks and quick comparisons.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanFall ResetAmazon USWork and home upgrades are worth comparing todayAmazon US: today's deals, useful picks and quick comparisons.See Picks×
Skip to content
Blog

OpenAI Told Parliament That Modern AI Could Not Be Trained Without Copyrighted Material

By TheFinanceBase Team8 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The headline is a provocative paraphrase, not a literal admission that OpenAI has a legal right to copy anything for free. In a 2024 submission to a House of Lords committee, OpenAI argued that it would be impossible to train today’s leading general-purpose AI models using only public-domain material. It said modern users’ needs require access to contemporary copyrighted expression and argued that copyright law does not categorically prohibit training.

That is a commercial and legal position—not a court ruling. As of August 16, 2026, the UK still had no definitive ruling on whether training a generative AI model on copyrighted works without a licence infringes copyright.

What OpenAI actually argued

The underlying controversy began with a written submission reported in January 2024. OpenAI said that “it would be impossible to train today’s leading AI models without using copyrighted materials.” It also argued that books and other works in the public domain alone would not produce systems capable of meeting contemporary users’ expectations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpenAI’s position had three connected parts:

  • Much of the modern internet and other contemporary human expression is protected by copyright.
  • Training a model is a legally distinct use from distributing a verbatim copy of a book, article, photograph, song or software file.
  • Copyright law, in OpenAI’s view, does not categorically forbid training on copyrighted works.

The important qualification is what OpenAI did not say. It did not establish that every copyrighted work may be copied without permission, that commercial success creates a copyright exception, or that a model’s outputs may freely reproduce protected expression. The original reporting is available from Futurism.

#1 Best Overall

“Copyrighted” does not automatically mean “infringing”

A dataset can contain many kinds of material:

  • Public-domain works: Material no longer protected by copyright or never protected in the relevant jurisdiction.
  • Licensed works: Content used under negotiated terms, although a licence to access a website is not necessarily a licence to train an AI model.
  • User-provided material: Content supplied under terms that may permit particular forms of processing.
  • Facts and ideas: These generally receive less or no copyright protection, while the expressive way an author presents them may be protected.
  • Government and factual material: Treatment varies by country and by the specific work.
  • Web content with mixed status: A single page may combine uncopyrightable facts, protected writing, photographs, code and third-party material.

Public accessibility also does not eliminate copyright. A page being online, or not being blocked by a technical measure, does not by itself answer whether copying and training are lawful.

Nor is public-domain-only training technically impossible. OpenAI’s narrower claim was that such a restriction would not produce leading, broadly capable models at contemporary quality levels. That is an argument about capability, data abundance and commercial competitiveness—not proof of legal necessity.

The copyright dispute is really several disputes

“Was AI training legal?” is too broad a question. Courts may need to examine separate stages of the process:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Acquisition: Was the material obtained lawfully, or through piracy, unauthorized access or a breach of applicable terms?
  2. Intermediate copying: Did downloading, storing, preprocessing or tokenizing a work create copies that engage copyright?
  3. Training: Is the use sufficiently transformative or covered by a statutory exception?
  4. Model weights: Do the resulting weights themselves contain or reproduce protected expression?
  5. Outputs: Does the system generate new material, memorize passages, imitate protected characters, or substitute for the market for the original work?

The House of Lords committee’s 2026 report said that large-scale copying and processing during training may engage the reproduction right, even where model weights do not contain human-readable books or articles. It also noted that rightsholders often cannot determine whether their works were included because developers do not provide enough training-data transparency. See the committee’s discussion of training and copyright.

Those issues can produce different outcomes. A court might find that a particular acquisition method was unlawful while reaching a different conclusion about training. It might permit some training uses but impose liability for memorized passages or particular outputs. A ruling about model weights would not necessarily decide whether the training data was lawfully copied.

What creators and publishers object to

Authors, publishers, journalists, photographers, musicians and other creators argue that large-scale uncompensated copying can transfer value from the people who produced the work to the companies building commercial models.

Their objections include:

  • AI-generated material may compete with markets that financed the original works.
  • Outputs can substitute for licensed articles, images, music, code or other content.
  • Creators usually cannot verify whether their work was ingested.
  • Opt-out systems place the administrative burden on individual rightsholders.
  • Opaque datasets make attribution, compensation and enforcement difficult.
  • A model can memorize and reproduce protected expression even if most responses are novel.

Litigation involving The New York Times, the Authors Guild and named authors illustrates these allegations, but allegations in a lawsuit are not proof that every training act by OpenAI infringed copyright. The relevant legal questions remain dependent on the dataset, conduct, jurisdiction and evidence.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What “fair use” does—and does not—settle

In the United States, fair use is a fact-specific doctrine rather than a general exemption for AI training. Courts consider:

  • the purpose and character of the use, including commerciality and transformation;
  • the nature of the copyrighted work;
  • the amount and substantiality of what was used; and
  • the effect on actual or potential markets.

A commercial use does not automatically fail fair use, and the fact that a model does not contain a readable copy of a book does not automatically end the analysis. Conversely, calling a process “transformative” does not guarantee that it is lawful.

In later written evidence to the Lords committee, OpenAI referred to two US federal opinions that it characterized as finding AI training to be fair use. That was OpenAI’s description of selected decisions, not a universal rule covering every model, dataset or method. The evidence is available as a parliamentary PDF.

Why OpenAI says licensing-only systems could hurt competition

OpenAI’s strongest non-legal argument is economic. Licensing millions or billions of individual works may be expensive and difficult where ownership is fragmented. A mandatory licensing system, it argues, could favor established companies with large budgets, content libraries or distribution platforms while raising barriers for startups.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

OpenAI repeated that argument in its later Lords evidence, saying licensing mandates could entrench incumbent platforms and make market entry harder.

The counterargument is equally important: an expensive or difficult business model does not create a legal entitlement to free inputs. Copyright owners argue that compulsory unpaid access could undermine the licensing markets and creative incentives that produced the material in the first place.

OpenAI’s position Rightsholders’ position What is established
Broad access to contemporary data is needed for leading models. That access involves copying and monetizes creative work without permission. The legal result depends on facts and jurisdiction.
Licensing could entrench large incumbents. Free access could destroy licensing markets. Both are policy concerns, not automatic legal conclusions.
Training is different from distributing a copy. Training still requires copying and may cause market substitution. Courts must examine the full pipeline.

Where the UK stood by August 2026

The UK debate moved away from simply assuming that a broad commercial text-and-data-mining exception with an opt-out would be adopted.

The House of Lords committee said the UK had still not issued a ruling on the specific question of whether training a generative AI model on copyrighted works without a licence infringes the reproduction right. Its recommendations favored protecting incentives to license, improving transparency and strengthening enforcement rather than weakening copyright protections. Read the committee’s recommendations and conclusions.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In a May 15, 2026 announcement, the committee said the government no longer had a preference for a broad copyright exception based on an opt-out mechanism and urged mandatory transparency requirements for large AI developers. That development should not be confused with a final judicial answer: parliamentary recommendations and government policy are not the same as a court ruling.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Why the Getty–Stability AI case did not settle training legality

The Getty–Stability AI proceedings show why legal headlines need careful reading.

Getty abandoned its primary UK copyright claim after accepting that there was no evidence Stability AI’s model had been trained or developed in the UK. The court separately rejected a claim concerning whether model weights made available in the UK were themselves infringing copies, and permission to appeal was later granted on that secondary issue.

That outcome did not declare all AI training lawful. A case may end because of territoriality, evidence, pleading or the particular legal theory before a court reaches the central question. The Lords committee discussed this distinction in its training and copyright section.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What a workable compromise could include

The debate does not have to be limited to “free scraping” versus “no AI.” Possible approaches include:

  • negotiated licences with publishers, stock libraries, music companies and rights organizations;
  • collective or extended collective licensing;
  • opt-in marketplaces for training data;
  • mandatory summaries, registers or audits of training sources;
  • machine-readable rights reservations;
  • compensation funds or usage-based royalties;
  • public-domain and openly licensed datasets;
  • smaller domain-specific models trained on narrower licensed collections;
  • retrieval systems that access licensed databases at query time rather than absorbing an entire corpus; and
  • creator-controlled APIs with clearly defined permissions.

Each option has trade-offs. A crawler-blocking service may help control access but does not compensate creators for past use. A certification program may signal sourcing practices but is not a binding licence. A collective licensing body may be useful for common categories of content but unsuitable for bespoke datasets. Synthetic data can reduce reliance on human works, but quality limitations and recursive training problems remain possible.

For businesses assessing an AI provider, the useful questions are practical: Does the service grant permission or merely block access? Does it cover training, retrieval and fine-tuning? Can usage be audited? Are compensation terms enforceable? Which jurisdiction governs the arrangement? The Copyright Licensing Agency, Cloudflare and Fairly Trained illustrate different categories of licensing, access-control and compliance-related services, but none by itself resolves the global legal dispute.

What readers should take away

OpenAI’s statement was partly technical, partly economic and partly legal. It argued that public-domain-only training would not support leading modern general-purpose models, that training is not automatically equivalent to publishing a copy, and that licensing mandates could make the market less competitive.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Creators and publishers challenge each part of that position. They argue that training requires copying, that commercial models can compete with the works that supplied their training signal, and that opaque datasets prevent meaningful control or compensation.

The accurate conclusion is narrower than the headline: OpenAI argued that access to copyrighted material is commercially important for building leading AI systems, but it did not prove that all such copying is lawful, free or exempt from licensing obligations. As of August 16, 2026, the core UK question remained unresolved.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Written by TheFinanceBase Team

The Team behind TheFinanceBase.

Add your note

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.