On January 29, 2025, OpenAI accused Chinese AI company DeepSeek of possibly using outputs from OpenAI models to help build a competing system. OpenAI said it had seen “some evidence” of model distillation and that using its outputs to develop a rival violated its terms of service. The public, however, had an immediate response: OpenAI itself argues that training AI on huge amounts of other people’s copyrighted material can be fair use.
That made the accusation look hypocritical. The irony is real, but the two situations are not automatically the same legally. OpenAI alleged output-based training, not wholesale theft of source code or model weights, and the public record did not include enough technical evidence to prove what DeepSeek did.
What OpenAI actually accused DeepSeek of doing
Contemporaneous reports said OpenAI had detected activity consistent with model distillation by groups based in China and believed DeepSeek may have used OpenAI model outputs to improve its own systems. The allegation surfaced shortly after DeepSeek released its R1 reasoning model on January 20, 2025, a release that drew global attention for its performance and efficiency claims.
OpenAI did not publish the account list, prompt corpus, dataset, forensic report or reproducible analysis behind the claim. Reporting described “some evidence,” but did not establish that DeepSeek itself authorized the activity, that the activity materially improved R1, or that DeepSeek copied OpenAI’s weights or source code.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
See the contemporaneous accounts from Semafor, Axios and Forbes.
What model distillation means
Distillation is a standard machine-learning method, not a synonym for illegal copying. A larger “teacher” model generates answers, explanations, classifications or other outputs. A smaller “student” model is then trained or fine-tuned to reproduce useful capabilities from those examples. The approach can transfer behavior at lower cost and is used for legitimate compression, specialization and research.
The legal question depends on how the outputs were obtained and used. Relevant possibilities include a contract breach, unauthorized API access, trade-secret misuse, copyright infringement or no actionable violation at all. Similar answers alone cannot prove distillation: models trained on overlapping public knowledge can converge on similar wording, and a large API query volume can also reflect benchmarking, red-teaming or ordinary product development.
Technical discussion of distillation and R1 appears in this paper. Research on memorization and copyrighted text in language models is available at this study.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Why DeepSeek’s R1 made the allegation consequential
R1 was presented as a strong reasoning competitor and performed competitively on some evaluations. Its release challenged the assumption that frontier-level progress necessarily required the largest budgets and the most computing resources. That triggered intense market and policy discussion about training costs, hardware access, efficiency and the possibility that synthetic data or teacher-model outputs had contributed to its capabilities.
Those questions do not prove that R1 was trained on OpenAI outputs. They explain why any evidence of systematic distillation would matter: it could affect how developers price APIs, protect model access and compete on capability.
Rank #3
Why OpenAI could object under its terms of service
OpenAI’s reported terms prohibited using output from its services to develop models that compete with OpenAI. If a user accepted those terms and then used OpenAI’s API or another service for that purpose, OpenAI could potentially pursue a contractual claim even if the output itself was not protected by copyright.
A contract theory would still require proof of the relevant account, user or organization, the conduct, assent to the terms and applicable law. It would not automatically establish copyright infringement, trade-secret misappropriation or ownership of DeepSeek’s model. Terms can also differ among consumer accounts, API customers, enterprise customers and downstream recipients.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Why the internet mocked OpenAI
The reaction was driven by a perceived double standard. OpenAI has argued in its litigation with The New York Times that training AI systems on copyrighted material can be protected by fair-use principles. Critics therefore condensed the dispute into a simple contrast: OpenAI says it may learn from the internet, but objects when a competitor allegedly learns from OpenAI’s outputs.
Rank #4
Futurism collected reactions from commentators including Ed Zitron, Jason Koebler and Gary Marcus, along with the broader online meme cycle: “the company that learned from everyone else is angry that someone learned from it.” The mockery reflected public-relations frustration, not a survey showing that every observer reached the same conclusion. OpenAI’s “open” branding also drew criticism when compared with DeepSeek’s more open release of some model materials, although “open source” is imprecise unless the weights, code, data and license are specified.
OpenAI’s own explanation of its position in the New York Times dispute is at OpenAI’s litigation page. Court-related background is available through the published case materials.
Why the hypocrisy criticism is persuasive—and incomplete
The strongest criticism is about consistency. OpenAI presents large-scale learning from copyrighted works as potentially transformative while insisting that a rival’s use of OpenAI-generated material violates a bright boundary. That contrast makes the complaint look selective, particularly when the alleged conduct would help a lower-cost competitor challenge OpenAI’s business.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
The legal comparison, however, involves different inputs, access methods and proof requirements:
| Issue | OpenAI training disputes | Alleged DeepSeek distillation |
|---|---|---|
| Material at issue | Human-created books, articles, websites, code, images or other works alleged by rights holders | Outputs generated by OpenAI systems |
| Primary theory | Copyright infringement and fair use | Contract breach, unauthorized use and potentially other claims |
| Access | Often alleged to involve public or scraped material | Allegedly obtained through authenticated accounts or API access |
| Protectability | Original human works generally receive copyright protection | AI outputs may receive limited or no copyright protection absent sufficient human authorship |
| Evidence needed | Copies made, memorization, market effects, licenses and fair-use factors | Account ownership, prompts, output patterns, training records and terms acceptance |
| Core uncertainty | Whether particular training uses are fair | Whether the alleged distillation occurred and what legal claim follows |
Fair use is a fact-specific litigation position, not a universal permission slip. It does not resolve questions about verbatim memorization, market substitution, licensing markets or contractual promises. Conversely, a contract restriction can be broader than copyright law. A court finding that some OpenAI training is fair use would not excuse a competitor from breaching an OpenAI agreement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What remains unproven
- Which accounts, users or organizations allegedly generated the relevant outputs.
- How many queries were made and over what period.
- What outputs were collected and whether they entered a training dataset.
- Whether multiple teacher models were used, making attribution difficult.
- Whether DeepSeek itself, a contractor, an affiliate or an unrelated third party controlled the activity.
- Whether any alleged data materially affected R1’s performance.
- Whether OpenAI could prove a contract, trade-secret or copyright claim in a particular jurisdiction.
Publicly accessible model outputs can complicate trade-secret theories, while cross-border entities, users and servers can complicate jurisdiction. Copyright in an AI-generated answer may also be limited where human authorship is insufficient.
The clearest way to understand the controversy
Several claims can be true at once:
- OpenAI made a credible-sounding allegation based on activity it said was suspicious.
- Distillation is a legitimate technique in many settings and is not inherently unlawful.
- OpenAI’s terms could prohibit using its outputs to build a competing model.
- OpenAI’s public reporting did not disclose enough evidence to independently verify the allegation.
- OpenAI’s fair-use arguments made the complaint look morally inconsistent to many observers.
- That inconsistency does not prove that OpenAI and DeepSeek engaged in legally identical conduct.
The episode was therefore both a genuine dispute about model access and an unusually effective public-relations joke. OpenAI may have had a legitimate contractual or competitive concern, but its aggressive defense of learning from other people’s work made it easy for critics to ask why its own boundary should be treated differently.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




