The foundation models that power today's AI revolution were built on a simple premise: scrape the internet, feed the text into neural networks, watch intelligence emerge. OpenAI, Anthropic, Google, and Meta collectively ingested trillions of words—Reddit threads, Wikipedia articles, digitized books, news archives, forum posts—to train systems that now write code, draft essays, and pass professional exams. But that well is running dry. Researchers estimate the stock of high-quality, human-written text suitable for training will be substantially depleted within the next few years, and no amount of compute can manufacture what does not exist.

The industry has known this was coming. Early language models trained on relatively small corpora—millions of documents—and saw returns diminish quickly. Scaling laws discovered around 2020 showed that bigger models trained on more data produced better results, which triggered a land grab for text. Companies signed deals with publishers, licensed academic databases, controversially scraped social media platforms, and pushed legal boundaries around fair use. The low-hanging fruit—public domain literature, open-access journals, Creative Commons blogs—was picked clean years ago. What remains is either legally contested, paywalled, or lower quality.

The synthetic data trap

The obvious solution is to generate training data synthetically: have AI write text, then train the next generation of AI on that output. Some labs are already doing this, using models to produce question-answer pairs, rewrite existing content in new styles, or simulate conversations. The problem is model collapse. When AI trains predominantly on AI-generated text, errors compound, diversity shrinks, and performance degrades across generations. It is the machine learning equivalent of inbreeding. Early research on this phenomenon showed that after several cycles of synthetic retraining, models lose the ability to represent rare concepts, hallucinate more frequently, and produce increasingly homogeneous outputs. The tail of the distribution—unusual ideas, niche knowledge, creative phrasing—vanishes first.

Some researchers argue that carefully curated synthetic data, mixed with remaining human text and filtered aggressively, can extend the runway. Others believe the solution lies in multimodal training—teaching models on video, audio, and sensor data where human-generated content is still abundant. But video is harder to learn from than text, and the semantic richness of a single well-written essay may exceed hours of raw footage. The most optimistic camp believes that reinforcement learning from human feedback, where models improve through interaction rather than passive reading, can break the data ceiling entirely. That approach has produced some of the most capable systems to date, but it is expensive, slow, and requires an army of human labelers.

What comes after the text runs out

If the data wall is real and no workaround emerges, the implications are significant. The pace of capability improvement could slow sharply, turning the exponential curve of the past few years into something closer to linear. Companies that stockpiled proprietary data—legal databases, medical records, internal corporate communications—would gain a durable advantage. The open-source AI movement, which relies on publicly available datasets, would struggle to keep up. And the legal fights over training data, already intense, would become existential. Publishers and platforms that once tolerated scraping might demand compensation or cut off access entirely, knowing they control a scarce resource.

There is also a darker possibility: that the industry charges ahead with synthetic data anyway, accepting model collapse as the price of continued growth. If the degradation is gradual and unevenly distributed—if models stay good at common tasks but lose competence in specialized domains—the decline might not be obvious until it is too late to reverse. We could end up with a generation of AI systems that are fluent, confident, and subtly less capable than their predecessors, trained on an increasingly incestuous corpus of machine-generated text that drifts further from human thought with each iteration.

Our take

The data exhaustion problem is a reminder that AI progress, for all its appearance of inevitability, rests on finite resources. The internet was a one-time gift—a multi-decade accumulation of human writing, argument, and creativity, digitized and made searchable just in time to train the first language models. That gift is nearly spent. What comes next will require either a fundamental change in how models learn, a legal and economic framework that unlocks new data sources, or an acceptance that the current paradigm has natural limits. The industry has grown used to scaling its way out of problems. This one may not yield to brute force.