Reddit CEO Steve Huffman said Microsoft and other AI-search companies should negotiate with Reddit before using its content commercially. The July 2024 dispute was about more than a disclosed AI-training run: Reddit objected to commercial crawling and uses across search and AI products. Huffman did not announce a price for Microsoft, and his demand was not a court ruling that Microsoft had infringed copyright.
What did Reddit’s CEO want Microsoft to pay for?
In an interview published July 31, 2024, Reddit CEO Steve Huffman criticized Microsoft and other AI-search companies for using Reddit content without what Reddit considered acceptable commercial terms. He said Reddit had difficulty blocking Microsoft’s crawlers and indicated that companies using Reddit content commercially would need to negotiate with Reddit or risk losing access. Engadget’s account of the interview does not identify a specific invoice, price, or universal royalty plan for Microsoft.
Huffman’s argument was that Reddit’s conversations have commercial value and that public visibility should not automatically make large-scale commercial use free. Reddit wanted control over access and terms; the disagreement was a business and access-policy confrontation, not a final legal finding that Microsoft had stolen content.
Why Microsoft was singled out
Reddit said Microsoft was crawling or accessing its content and objected to how it was being used in search and AI-related services. The concern extended across search infrastructure and AI-search experiences, including retrieval and summaries, rather than being limited to one identified model-training dataset. Reddit also said it struggled to block unwanted crawlers without disrupting ordinary visitors, search visibility, or approved partners.
#1 Best Overall
The timing drew attention because Microsoft AI chief Mustafa Suleyman had recently described publicly available web content as effectively available for AI training under a “social contract” of the open web. That remark is useful context for the debate, but it should not be mistaken for a formal Microsoft legal position that all public material is free of copyright or contractual limits.
Search, retrieval, and training are different uses
“Using Reddit data for AI” can describe several distinct activities. A crawler collects pages; a search index stores and ranks them; a search product may show links or excerpts; an AI system may retrieve material when answering a query; and a model developer may use posts for training, fine-tuning, or evaluation. A company might also redistribute Reddit-derived information to customers or other services.
The July headline’s focus on “AI training” can therefore obscure the immediate commercial dispute: Reddit objected to uses across search and AI infrastructure, not just a publicly documented training run. Whether a particular use requires a license depends on its facts, applicable terms, and law.
Rank #2
- Ideal for Gifting
- Ideal for a bookworm
- Compact for travelling
What Reddit’s policies say about commercial and research access
Reddit’s Public Content Policy announcement of May 9, 2024 drew a distinction between preserving access for researchers and responsible noncommercial users and opposing unauthorized commercial bulk collection. It also identified restricted uses involving user identification, ad targeting, background checks, facial recognition, government surveillance, and nonpublic material.
Recommended Free Tools
On June 25, 2024, Reddit said it would update robots.txt and continue rate-limiting or blocking unknown bots and crawlers, while preserving certain good-faith noncommercial access, including for researchers and the Internet Archive. Its Developer Terms prohibit using Reddit services or data—including through APIs, indexing, caching, or crawling—to train large language, AI, or other algorithmic models without Reddit’s permission.
These policies make the commercial distinction clearer, but a policy statement does not by itself settle every legal question about a company’s past conduct. Reddit’s access rules, a user’s copyright, and the lawfulness of a particular copying or search use are related but separate matters.
Rank #3
What robots.txt can—and cannot—do
Robots.txt is a technical instruction for web crawlers, not a universal copyright statute or court order. Ignoring it does not automatically prove infringement. It can still matter operationally, contractually, or as evidence depending on the facts. Reddit also uses rate limits and blocking, but distinguishing ordinary browsers, search engines, approved partners, researchers, and unwanted bots is technically difficult.
Who owns a Reddit post?
Reddit does not simply own every post. Under the user agreement dated August 16, 2024, users retain whatever ownership rights they have in their content while granting Reddit a broad license to use, distribute, syndicate, and otherwise exploit it, including for AI and machine-learning training.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →That license is between users and Reddit; it does not automatically give every outside company unrestricted rights. A post may include another person’s photograph, article, or other copyrighted work, or contain personal information. Reddit’s platform and database interests, its contractual rights, individual users’ copyrights, and third parties’ rights in reposted material are not interchangeable. Nor does the legal status of obtaining data automatically determine the status of training on it or redistributing it.
Rank #4
How Reddit was building a data-licensing business
Reddit’s 2024 SEC filings described data licensing as an emerging revenue stream. The company discussed API access to real-time and historical discussion, search and analysis, machine-learning and AI training licenses, and bespoke arrangements with large commercial partners. The filings also warned that the business was new and uncertain, acknowledged that some companies might use publicly accessible Reddit data without a license, and disclosed an FTC inquiry into Reddit’s licensing or sharing of user-generated content for AI training. The inquiry was not an adverse final ruling. Reddit’s filing
In a separate IPO filing, Reddit reported $203 million in aggregate contract value for data-licensing agreements entered into in January 2024, with terms ranging from two to three years. That figure is not Microsoft revenue and does not disclose Microsoft’s terms. The filing
Reddit announced a partnership with OpenAI on May 16, 2024, involving access to and display of Reddit content through OpenAI products. Reuters also reported a Reddit–Google AI-content licensing deal in February 2024. These agreements illustrate the negotiated-access model; they do not establish that every company received the same terms or that Microsoft had an equivalent agreement. Reddit said API access remained free for noncommercial use under its published conditions. Reddit’s OpenAI announcement; Reuters’ Google deal report
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBest Value
- It can be a gift option
- Comes with secure packaging
- Helpful in various ways
What each side has at stake
Reddit’s case for licensing
- Its large corpus of human conversation can improve commercial search and AI products.
- Uncontrolled crawling may consume infrastructure resources and weaken Reddit’s ability to negotiate terms.
- Agreements can set expectations for privacy, deletion, attribution, access, and permitted uses.
- If some companies pay while competitors scrape freely, paying partners may be disadvantaged.
Why a company might resist a license demand
- Public web pages have long been crawled and indexed for search.
- Whether AI training on copyrighted material is fair use in the United States may depend on the facts.
- A search service may link to or retrieve material at query time rather than permanently train on every page.
- A platform’s demand for payment does not itself establish that a license is legally required for every use.
These are general arguments in the dispute, not a detailed legal defense attributed here to Microsoft.
What remains unresolved legally
The dispute did not settle whether training on copyrighted text is fair use in the United States, whether copying during scraping is legally significant, or whether a site’s terms bind an automated actor that never agreed to them. It also did not determine whether robots.txt creates enforceable obligations in a particular case, whether Reddit has rights to license every item users post, or whether AI summaries substitute for or transform the material they use.
Other open questions include whether model outputs infringe posts, how privacy and deletion requests can be honored after training, and whether commercial access to public data should be addressed through copyright, contract, privacy or competition law, or new legislation. Later complaints by Reddit alleging unauthorized commercial use are allegations in litigation, not established findings. Reddit’s complaint
Why the dispute matters beyond Microsoft
Reddit was trying to establish a market in which large AI and search companies negotiate for valuable data while researchers and some noncommercial users retain limited access. That approach may give a platform more control over privacy safeguards and technical access, but it can also concentrate access and revenue among a few large partners. Restricting crawlers may reduce unwanted collection while also affecting discovery through search and referral traffic.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For Reddit users, licensing announcements describe agreements between Reddit and companies, not direct per-post royalties. Reddit’s policies identify protections and limits, but they do not establish that users can opt out of every licensing use or that deleting a post automatically removes information already incorporated into a model. Reddit’s Public Content Policy says nonpublic material such as private messages and mod mail will not be licensed or distributed.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




