A former OpenAI researcher is betting that training data will become one of the next major constraints in AI development, The Decoder reports. The argument is that simply scaling models and compute may not be enough to keep improving performance.
That view reflects a growing concern across the industry. Large models have already consumed vast amounts of public text, code and media. If the easiest data sources are exhausted or legally restricted, labs may need to spend far more on curated, licensed or newly generated data.
The Decoder reports a prediction that $100 billion could flow into training data. Whether that number proves accurate, the direction is plausible: data quality, provenance and task-specific coverage are becoming strategic assets.
The practical implication is that AI competition may shift from who has the biggest cluster to who can produce the best learning material for the next generation of systems. That could benefit data providers, domain experts and companies with proprietary workflows.