What happened
At the 2026 China International Big Data Industry Expo, Liu Liehong, director of the National Data Administration, outlined “three new” developments: data empowerment of AI has entered a new stage, AI has become a fresh tool for unlocking the potential of data resources, and tokens are opening a new route for releasing the value of data elements. He described a cycle in which scenarios pull in data, data drives models, models empower applications, and applications create value.
The expo also offered evidence that data supply is increasingly oriented around specific tasks. By August 2026, more than 126,000 high-quality datasets had been built, with total volume exceeding 1,815 PB, up more than 89% compared with March. In 2025, data used for AI training and inference reached 199.48 exabytes, and inference data, at 101.34 exabytes, exceeded training data for the first time.
Industry executives and experts told China Business Journal that companies are shifting toward whole-chain coordination. Embodied intelligence is creating new data requirements, trusted data spaces are being positioned as infrastructure for safe data circulation, and token metering, pricing, and settlement are helping turn model calls into tradable services.
Why it matters
The focus on tokens gives AI model usage a measurable unit of account, which could make data, computing power, and model services easier to trade and settle. If token-based measurement becomes established, a more standardized market for AI services may emerge.
The remarks also suggest that data value creation is becoming a system-level challenge. Trusted data spaces and high-quality datasets are treated as necessary links between raw data and model development, indicating that data governance and commercialization are advancing together.
Key facts
Liu Liehong proposed three “new” directions at the 2026 expo, positioning tokens as a new path for data element value release.
In 2025, AI training and inference data reached 199.48 exabytes; inference data overtook training data for the first time.
Embodied intelligence data production grew 477.78% year on year in 2025.
By August 2026, China had built over 126,000 high-quality datasets with a total volume of more than 1,815 PB.
China's annual token call volume in 2025 was about 21,100 trillion, with daily calls rising from more than 1 trillion to 100 trillion.
Tokens were described as a key yardstick for metering, pricing, trading, and settling AI services.
What to watch next
Token competition is expected to move from generic tokens toward scenario tokens, which depend on access to compliant, high-value data for secondary model development and solving real-world problems.
Demand for unified procurement, call tracking, and centralized settlement may grow, especially as scattered individual token purchases create new governance and reimbursement burdens.
Trusted data spaces may become a core part of the data-to-model pipeline, as security requirements expand from simply preventing leaks to enabling data to be used without being seen or taken.
