AI Training Data Is Running Out in China

AI training data missing from the stripped shelves of a vast Chinese state archive hall

Chinese accounts for just 5.2% of the Common Crawl corpus while English takes 43.2%. Chinese specialists now point to AI training data, not silicon, as the ceiling their industry is about to hit. Beijing answered with a national plan whose deadline lands in 2028.

Key Takeaways

  • Chinese makes up 5.2% of Common Crawl data against 43.2% for English.
  • The global stock of public human-written AI training data could be fully consumed within six years.
  • The National Data Administration is targeting a validated national dataset ecosystem by 2028, from manufacturing to low-altitude aviation.

Have an AI Sum Up This Article

ChatGPT

Five Percent of the World’s Crawl

The imbalance is blunt and it fits in two numbers. Inside Common Crawl, the open archive that underpins a large share of today’s language models, Chinese holds 5.2% of the volume and English holds 43.2%.

That ratio stops being trivia the moment you train a general-purpose model. A corpus eight times smaller in the target language forces you either to recycle the same text more aggressively or to lean on translation, with the loss of naturalness that comes with it.

Thin access to quality data has already started to strain the development of Chinese-language models. This has moved out of the research paper and into the engineering backlog.

The dominant story for the past two years pointed somewhere else entirely. US export controls on advanced chips absorbed the attention, and the Chinese answer was built on that same ground, all the way to in-house silicon projects such as the chip DeepSeek is building to cut its reliance on Nvidia.

Data offers no such exit. You can design a replacement accelerator in a few years. You cannot retroactively manufacture twenty years of high-quality Chinese web text that nobody ever wrote.

That is what makes this bottleneck different in kind. The constraint moves from a field where engineering eventually finds an answer into one where no obvious answer exists. AI training data does not come out of a fab.


AI training data

Beijing Plans Its Corpora Like It Plans Factories

The response came from the top. In early June, the National Data Administration unveiled a national plan to expand the supply, circulation and commercialisation of high-quality training data.

The target year is 2028. By then, China wants a broad ecosystem of validated datasets spanning scientific research, manufacturing, agriculture, energy, transport, finance, healthcare, education and e-commerce.

The list also reaches into emerging fields. Embodied AI, autonomous driving, low-altitude aviation and biomanufacturing appear explicitly, which gives away the intent: manufacture signal in places the open web never produced any.

The plan insists on multimodal coverage, from text and code to images, audio and video. It follows the same logic as work already under way, including the Chinese model Fysiverse, which builds real physics directly into its training instead of inferring it from written descriptions.

All of it sits inside the AI Plus strategy, which aims to push artificial intelligence through the whole economy. Data gets treated as public infrastructure, on the same footing as a power grid or a port.

For teams building on Chinese models, that 2028 deadline is worth writing down. Sector datasets will land in waves, and whichever verticals get served first will decide which professional use cases take off ahead of the rest.


More articles on Horizon


The Wall American Labs Also See Coming

The squeeze does not stop at China’s border. Research institute Epoch AI estimates that the stock of publicly available human-written text will be fully consumed by models within six years, across every player at once.

Andrej Karpathy, an OpenAI co-founder, flags the same wall arriving before the decade closes. He draws the direct consequence from it: a capability plateau, for lack of fresh raw material to ingest.

That shared diagnosis explains how aggressively American labs are buying data right now. Licensing deals for archives, catalogues and proprietary feeds keep stacking up because the open web no longer feeds the next generation on its own.

So the gap between the two camps sits in method rather than in diagnosis. One side buys corpora on a private market, the other organises their production through sector planning.

Each approach carries a distinct failure mode. A licensing strategy scales only as far as the counterparties are willing to sell, and prices climb every time a rival signs first. A planning strategy delivers volume on schedule but inherits whatever quality the mandated pipeline produces, which is rarely the same thing as usable training signal.

The release cadence out of China makes the trade-off urgent. When one vendor ships successively larger generations, as Alibaba did with Qwen 3.8, its biggest model to date, appetite for data grows faster than the corpus behind it.

Our read is that the next competitive shift will not show up in a benchmark table. It will show up in each camp’s ability to lock down corpora the other side cannot reach, whether that comes from a signed contract or a five-year plan.

Follow the story on Horizon.

Comments

No comments yet. Why don’t you start the discussion?

    Leave a Reply

    Your email address will not be published. Required fields are marked *