Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Has AI Run Out of Training Data? What Elon Musk’s Claim Really Means

Musk’s “all human knowledge” claim overstates the evidence. Public human-written text may become a constraint, but AI still has other data sources and ways to improve.
From TheFinanceBase Team7 min to read

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Not literally. Elon Musk said on January 8, 2025, that AI had exhausted “basically the cumulative sum of human knowledge” for training, but the evidence supports a narrower concern: high-quality, publicly available human-written text may become harder to obtain at the scale frontier AI models need. That is a potential constraint on one route to improvement—not proof that AI has run out of data or that progress has stopped.

What Musk said—and what he meant

During a livestreamed conversation with Stagwell chairman Mark Penn on X, Musk said AI had exhausted the cumulative sum of human knowledge available for training, adding that this had happened “basically last year”—in 2024. He argued that AI systems would increasingly need to generate synthetic data for themselves, while acknowledging a difficulty: generated material can contain hallucinations, so it is hard to know whether it is reliable training input. TechCrunch’s account of the conversation and The Guardian’s report document the claim.

“All human knowledge” is not a measurable data category. The technically defensible version is that developers may be approaching the effective supply of high-quality, non-duplicative, publicly available human text that is usable for large-scale pretraining. That is much narrower than saying every source of AI training data has been consumed.

Which part of AI training could face a shortage?

AI development uses several kinds of training and information, and a shortage in one does not mean a shortage in all the others.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
  • Pretraining exposes a model to large datasets so it learns broad patterns in language, code and other data. The data-supply concern is mainly about this stage, especially text.
  • Fine-tuning and instruction tuning use more focused examples to shape a pretrained model for particular tasks, formats or behaviors.
  • Reinforcement learning improves behavior using human feedback, AI feedback, rules or outcomes from an environment.
  • Inference-time compute gives a model more computation after it receives a prompt—for example, time to reason, search or check intermediate steps. It changes how a model uses its capabilities, not what was in its original pretraining corpus.
  • Retrieval-augmented generation supplies relevant external information when answering, rather than requiring every fact to be stored in the model’s weights.

A pretraining bottleneck can therefore coexist with continued gains from better post-training, retrieval, tools, architectures and reasoning methods.

Why public text could become a constraint

Language-model performance has historically improved as developers increase model size, training data and computation. Research on scaling laws describes empirical relationships among those factors; the Kaplan et al. study examined these relationships, while the Chinchilla paper found that compute-efficient training generally calls for scaling model size and training data together.

If available compute grows faster than the supply of useful text, developers cannot indefinitely maintain the same balance by adding more tokens. The problem is not a shortage of raw bytes on the internet. It is a shortage of material that is sufficiently useful, diverse, reliable, non-duplicative and available for a given company to use.

What the data-supply estimate actually says

Epoch AI estimated the effective stock of publicly available human-generated text at about 300 trillion tokens, with a broad 90% confidence interval of 100 trillion to 1 quadrillion tokens. Its model suggested that, under the compute and scaling assumptions it examined, text data could become a significant constraint around 2028. Epoch AI’s analysis is an estimate of an effective stock, adjusted for quality and repetition—not a count of every word online and not evidence that the supply was already exhausted in 2024.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The dates describe different things: Musk’s statement pointed to 2024 as the year the data had supposedly been exhausted; Epoch AI projected a possible future constraint around 2028 under modeled assumptions. A forecast is not an observed failure point, and the estimate depends on how much data is usable and how quickly training demand grows.

Why “the internet is exhausted” is misleading

Several distinct thresholds get collapsed into the word “exhausted.” A company may have downloaded or indexed material without being able to use it legally, affordably or effectively. Nor does a dataset become fully learned simply because it has been collected: better filtering or training methods may extract more value from it, though repeated exposure eventually has diminishing returns.

  • Publicly accessible does not mean permission to use it is settled for every company or purpose.
  • Machine-readable does not mean clean, relevant or free of duplication and spam.
  • Large in volume does not mean rich in new, high-quality training signal.
  • Published online does not include private records, unpublished expertise or skills embodied in people and organizations.

Data outside public text includes books and journalism that may require licenses; customer or enterprise records subject to privacy and contractual limits; expert demonstrations; and material in languages or formats underrepresented in common web collections. AI systems can also learn from images, video, audio, speech, scientific and biological data, sensor readings, robotics demonstrations and simulated environments. Epoch AI has discussed multimodal data as one possible way scaling could continue, while noting that its supply and transferability are less certain than for web text. Its analysis of scaling through 2030 considers those possibilities.

Can synthetic data fill the gap?

Synthetic data is produced or transformed by an AI system, simulator or program rather than collected directly from people or the physical world. It can mean generated question-and-answer pairs, code and tests, simulated robot trajectories, self-play records or artificial images. Its usefulness depends heavily on whether its quality can be checked.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where it has a stronger case

Synthetic examples are most useful when there is an independent way to evaluate them. A compiler can run generated code; a theorem prover can check a proof; a game engine can score a move; and a mathematical solver can verify an answer. In these cases, a system can generate examples, attempt solutions and learn from outcomes without treating every generated explanation as true.

That is different from training indiscriminately on AI-written prose about open-ended facts, history or social judgments. Without a trusted reference or reviewer, a model’s confident error can become another model’s training example.

What model collapse means

Research has warned that repeatedly training models on generated material can cause “model collapse”: rare patterns may disappear, and outputs may become less diverse or less representative of the original data distribution. The risk is not automatic. It depends on how examples are generated, filtered, mixed with human data and validated. Synthetic data grounded in external verification differs from an unfiltered stream of AI-written text, but it can still inherit the generator’s biases and errors. See the research on model collapse.

Practical safeguards include deterministic validators, trusted ground-truth sources, human review, adversarial testing, provenance tracking and keeping synthetic data labeled separately. These controls reduce risk; they do not make every generated example reliable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What developers can use instead of more public text

A data constraint would change the economics and mix of AI development rather than force a single replacement for web text.

  • Improve data efficiency: better filtering, deduplication, curriculum design and training methods can make existing corpora more useful.
  • License or collect proprietary material: books, journalism, specialist databases, customer data with permission and expert-created examples can offer differentiated signal, but rights, privacy, cost and processing matter.
  • Use human feedback and demonstrations: people can provide task-specific examples and preferences, though this work can be expensive, slow and subjective.
  • Expand reinforcement learning: models can learn from verifiable rewards, tools and task outcomes, not only from absorbing more text.
  • Use simulation and self-play: environments such as games can generate extensive experience with measurable outcomes, but simulated results may not transfer cleanly to the real world.
  • Train on more modalities: video, audio, sensors, scientific data and robotics can support capabilities that text alone cannot, but collecting and processing them creates its own challenges.
  • Spend more computation at answer time: longer reasoning, tool use and checking can improve some tasks without requiring a proportionate increase in pretraining text.

Each option has a different bottleneck. Licensed material may cost more but offer clearer provenance; proprietary data may differentiate a company but be difficult to collect; synthetic data is scalable but needs verification; simulation can generate experience but may omit real-world details. None demonstrates that a particular AI model has run out of data.

What the claim could mean for businesses and the public

If generic public text becomes less useful at the margin, the value of verified, specialized and legally usable datasets may rise. That could increase demand for licensing, expert annotation and data-provenance systems, while putting more attention on consent, privacy, copyright disputes and the terms under which publishers and creators’ work is used. The economics could favor organizations with differentiated datasets, but acquiring data is not enough: it must be relevant, usable and governed appropriately.

For AI users, a pretraining constraint would not by itself predict a sudden halt in products or capabilities. It could make brute-force scaling less effective or more costly, while progress shifts toward better data curation, post-training, retrieval, reasoning and multimodal systems. The effect on model quality, release schedules and costs depends on which capability is being developed and which resource—data, compute, chips, energy or expertise—is limiting it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to tell whether a model is actually data-constrained

A claim that a particular model has hit a data wall needs more than a general forecast. Useful questions include:

  • Is the claimed shortage about raw data or high-quality, legally usable material?
  • Does it concern text pretraining, or a different modality or post-training stage?
  • Is the model undertrained relative to its size and compute?
  • Can the proposed synthetic examples be independently checked?
  • Does additional data improve the target capability, rather than only reduce training loss?
  • Could the bottleneck instead be compute, chips, energy, deployment costs or the availability of expert feedback?

Without answers to those questions, “AI has exhausted its data” is a slogan, not a diagnosis.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.