October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Former OpenAI Researcher Said He Was Disgusted by Copyrighted-Data Use in AI Training

Suchir Balaji’s 2024 essay argued that copyrighted-data use in commercial AI training raises fair-use concerns. It also cautioned that no broad answer applies to every work or model.
From TheFinanceBase Team4 min to read

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Suchir Balaji, a former OpenAI researcher, argued that using copyrighted works to train commercial generative AI could harm creators and the markets for their work. His October 2024 essay did not claim that all AI training is unlawful: it said fair use must be assessed case by case. A Futurism report, citing a New York Times interview, described his objections and reported OpenAI’s opposing position. Neither account establishes a court’s answer to whether a particular use is fair use.

What Balaji objected to

In a report published October 24, 2024, Futurism described Balaji as a former OpenAI researcher who had worked at the company for four years. The report, attributing its account to a New York Times interview, said his work included collecting and organizing web-gathered data used to train large language models.

According to that account, Balaji’s view changed as ChatGPT became a commercial product after its November 2022 release. He was concerned that AI products could produce material reflecting or mimicking copyrighted source works, and that the resulting business model could damage the broader internet ecosystem. Futurism quoted him saying, “If you believe what I believe, you have to just leave the company,” and describing the model as “not a sustainable model” for “the internet ecosystem as a whole.”

Those quotations convey Balaji’s position; they are not findings that OpenAI infringed copyright. Nor does the report establish that he personally gathered every kind of material used in training or that all of the data at issue was copyrighted.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What Balaji argued about fair use

In his October 23, 2024 essay, “When does generative AI qualify for fair use?”, Balaji began from the premise that training generative models involves copying copyrighted data. He applied the four statutory fair-use factors to that activity and argued that commercial use and possible substitution for original works raise concerns. His analysis is an argument about how the factors should apply, not a ruling by a court.

Balaji expressly limited his conclusion: “Because fair use is determined on a case-by-case basis, no broad statement can be made about when generative AI qualifies for fair use.” That qualification matters. Whether a particular use is fair cannot be resolved merely by labeling it “AI training,” “commercial,” or “transformative”; the relevant facts and work may differ.

How the four fair-use factors frame the dispute

Balaji’s essay considers the four factors used in a fair-use analysis. The table summarizes the competing questions his argument raises; it does not predict how a court would weigh them.

Factor Question raised by the dispute Balaji’s argument, as presented in his essay
Purpose and character of the use Does the purpose and character of copying support fair use, and how does commercial use matter? He emphasizes that a commercial system may produce outputs that compete with or substitute for copyrighted works.
Nature of the copyrighted work What kind of work was used, and how does that nature bear on the analysis? His essay includes this factor in the framework, but the available account does not establish the character of each work in a particular training set.
Amount and substantiality used How much of a work was copied, and how important was the portion used? Training involves copying in his account, but the essay says the training data is not publicly known; the amount used cannot be assessed here for each work.
Effect on the potential market or value Does the use affect the market for the original or its value, including a potential licensing market? He argues that outputs can substitute for originals and that a market for data licenses bears on potential market harm. He also notes that effects vary by source and cannot be answered directly for every work from the information available.

The factors are considered together in a case-specific analysis. The concerns in Balaji’s essay—authorization, whether training is sufficiently distinct or transformative, how much is copied or reflected, and possible substitution or licensing-market effects—are questions for evaluating evidence, not conclusions that the use is unlawful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

OpenAI’s reported response

Futurism reported that OpenAI told The New York Times it builds its “AI models using publicly available data, in a manner protected by fair use and related principles” and described that approach as critical for “US competitiveness.” This is OpenAI’s reported position, not a judicial determination. The two accounts present competing arguments and do not resolve which position is legally correct.

What the report does—and does not—establish

  • It reports Balaji’s objections to copyright practices in AI training and describes his prior work at OpenAI, attributing those details to the New York Times interview.
  • His essay argues that training involves copying copyrighted data and discusses all four fair-use factors, while expressly rejecting a categorical answer.
  • The essay says the training data is not publicly known and that market effects vary by source. That limits what can be concluded about any individual work from his discussion.
  • The cited accounts do not establish that a court found OpenAI’s training unlawful or categorically protected by fair use. Futurism’s 2024 report refers to copyright lawsuits but does not establish their current status.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase09 OCT 267 minMortgage Escrow FAQs: Taxes, Insurance, Shortages, and Refunds
  2. The Money DeskBlogTheFinanceBase09 OCT 265 minHow Mortgage Escrow Accounts Work and What Homeowners Pay For
  3. The Money DeskBlogTheFinanceBase09 OCT 265 minHow to Read a Stock Chart, Volume and Market-Cap Data
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.