ai china

China’s public-data market now has a permitted-return formula

7 sources 5 primary sources August 3, 2026

Text
Visitors walk across a blue carpet beneath large bilingual signs at the entrance to the 2025 China International Big Data Industry Expo.

Visitors enter the professional exhibition at the 2025 China International Big Data Industry Expo in Guiyang on August 27, 2025. The venue makes suppliers and products visible; it does not show that a dataset is licensed, interoperable, or repeatedly used. Photo: People’s Daily Online/Tu Min.[7]

At the entrance to Guiyang’s 2025 big-data expo, people walk between signs large enough to turn an intangible industry into a physical destination. Inside, 375 companies displayed data infrastructure, large AI models, and laboratory systems.[7] It is a persuasive picture of supply. It is also a useful warning: an exhibition can count providers and products, but it cannot show whether a buyer has permission to use a dataset, whether the fields arrive in a workable schema, or whether the same feed will still be maintained next year.

China’s latest national data census has produced an even larger supply picture. Its table counts 114,400 “high-quality datasets” at the end of 2025, up 61.13% in a year, while their reported volume reached 908.30 petabytes, up 142.58%. The amount of data used for AI training and inference reached 199.48 exabytes.[1] Those figures make the next bottleneck easier to see. China no longer needs only more data inventory; it needs a reliable route from institutional custody to repeated, lawful use.

As of 2026-08-03T12:36:53Z UTC, that route is taking the shape of regulated infrastructure. Public data used for governance or public-interest purposes is free. Public-data products used for industrial or sector development may charge a service fee, but under a government-guided ceiling. The operator’s permitted profit is tied to cost and a benchmark government-bond yield, not simply to what the most eager AI company will bid.[3]

This is a consequential market design. It can restrain monopoly rents and make a public asset broadly useful. It can also create a slow, locally administered access layer in which the decisive commercial advantage belongs to whoever can clear authorization, package a product, meter calls, document provenance, and keep the feed current. China’s AI data market will therefore be measured less by the next hundred thousand dataset labels than by renewal, reuse, and the breadth of its buyer base.

The inventory is an estimate, not a catalog

The 2025 National Data Resources Survey is unusually useful because it publishes both its headline totals and its method. The appendix lists 48,148 effective observations across 14 types of respondent, including government departments, research bodies, central enterprises, platform companies, data exchanges, local data groups, trusted data-space operators, other companies, and industry associations. Enterprise totals were estimated with stratified sampling; extreme values were trimmed, and growth rates used a fixed group observed in both periods.[1]

That makes the report more substantial than a collection of launch announcements. It does not make every aggregate an audited registry. The 114,400 figure is a national estimate under a policy category, not a list of 114,400 artifacts that a developer can inspect and download.

The official construction guide defines a high-quality dataset as processed data that can be used directly to develop or train AI and can improve model performance. It asks for static qualities such as accuracy, completeness, consistency, timeliness, diversity, authenticity, and compliance, plus dynamic testing of whether the data improves a representative model.[4] That is a demanding definition. The census does not publish artifact-level provenance, evaluation results, license terms, or user counts behind its aggregate.

China’s own guide also names the problem plainly. Published in August 2025, it identified weak cross-region sharing, incomplete authorization-platform coverage, unclear operating boundaries, high annotation and governance costs, long value-conversion cycles, and data exchanges that had not yet formed a market at scale.[4] The end-2025 inventory growth is therefore evidence of mobilization, not proof that those frictions disappeared.

Public data reaches the market through an operator

Public data is only one part of the 908.30-petabyte high-quality total, and the categories cannot be divided into each other. The survey separately reports 37.84 petabytes of open public data and 7.59 petabytes under authorized operation, the latter up 53.96% year over year. It also counts 10,200 authorized public-data products and services across industry, education, science, and health.[1] Those numbers describe a delivery channel, not a subset whose utilization can be inferred by comparing it mechanically with all high-quality datasets.

The distinction matters. Open data can be released for broad access. Authorized operation covers public resources held by local governments or national sector authorities that are assigned to a qualified legal entity for governance, development, and fair provision of products or technical services to the market.[2]

The operator is not supposed to dump raw records into commerce. The national trial rules require a resource directory, update frequency, quality information, security controls, product and service lists, cost and revenue accounting, disclosure, and an exit mechanism. Operators are chosen through competitive procedures; agreements generally run no longer than five years. Unpublished raw public data is to be strictly controlled from entering the market directly, while products and services must be registered and disclosed.[2]

In practical terms, the market unit is less likely to be “a government database for sale” than a bounded query, verification result, scored record, API call, or derived service. That protects sensitive source material and can make permissions auditable. It also means the AI user depends on an intermediary for latency, schema stability, coverage, correction, and continuity. The quality of that intermediary becomes part of the quality of the data.

The price is administered by design

China’s January 2025 pricing notice separates public purpose from commercial purpose. Products used for public governance and public-interest undertakings are free. Products used for industrial or sector development may carry a public-data operating-service fee under government-guided pricing.[3]

For paid products, regulators calculate a maximum permitted revenue based on reasonable operating costs, taxes, and permitted profit. Covered costs can include platform construction and operation, transmission, aggregation, storage, governance, labor, and resource acquisition, net of government subsidies. Permitted profit equals operating cost multiplied by a rate set at the prior year’s average yield on ten-year government bonds plus no more than six percentage points. The authorization body then establishes ceiling charges, while the operator sets actual prices at or below them.[3]

The mechanism is closer to a regulated service than a scarcity auction. Prices may be expressed per product, call, period, or amount of data. Reviews occur at least every three years, annual revenue is monitored, and a deviation of more than 10% from permitted revenue can trigger a change to the ceiling.[3]

This arrangement has a clear strength: a local operator cannot turn exclusive proximity to public records into an unconstrained toll. It has an equally clear tradeoff. Cost-based pricing rewards documented expenditure more directly than downstream usefulness. A carefully maintained, narrow dataset that saves an AI user substantial work may be inexpensive to operate; a sprawling platform may have a large cost base before it has many active customers. Regulators will need to distinguish necessary governance from overhead and useful demand from catalog volume.

Guangzhou provides an early field test. Its August 2025 trial used a fixed annual basic technical-service fee plus usage charges that decline by tier as calls increase. The city exposed more than 5,000 catalog entries and let registered data merchants inspect AI-generated simulated data before applying. At the time of the official account, more than 50 development applications had been approved and 12 public-data products had been listed after compliance review.[5] Those are concrete steps from directory to transaction, but they are still pipeline measures. They do not reveal renewals, call frequency, customer concentration, or whether a product improved an AI system.

Demand is growing—and concentrating

The national survey offers the sharpest counterweight to the supply celebration. Only 11.65% of sampled enterprises had purchased data in 2025, although their spending grew 22.36%. Among leading platform companies, by contrast, 88.89% were buyers. External data made up 74.36% of the data used in their AI development and training, and their average data-purchase spending was 60 times that of other companies; average purchased volume was 115 times greater.[1]

That is a market signal, not a verdict that small firms ought to buy more. Many companies generate adequate internal data or do not have an AI use case that justifies an external feed. It does show that demand, integration capacity, and purchasing power are concentrated in the same organizations. Beijing and Shanghai each had data-buying rates above 30%, as did finance and software and information-technology services.[1]

An earlier study of all 31 provincial-level regions found a three-tier geography in China’s data-factor market and highlighted regional imbalance, incomplete public-data institutions, and insufficient on-exchange transaction scale.[6] Its evidence predates the 2025 national pricing and authorization rules, so it should be read as a baseline rather than a judgment on the new system. The new rules address several of those institutional gaps. They have not yet demonstrated that a product authorized in one city can be discovered, contracted, evaluated, and reused by an AI developer in another without rebuilding the transaction.

That is the real macro spread: China has national-scale supply targets and increasingly specific local operating machinery, while effective demand remains clustered around large platforms, central enterprises, finance, software, Beijing, and Shanghai. If the middle layer works, it can lower the fixed cost of lawful access for everyone else. If it fragments, each locality may produce a full catalog, a designated operator, and a tariff sheet without producing a national market.

The next receipts are usage receipts

Four disclosures would show that inventory is becoming infrastructure.

First, operators should report the conversion funnel: catalog entries, qualified applications, approved products, paying users, active calls, renewals, and retired feeds. A dataset that is listed once but never used is supply policy, not market depth.

Second, products need service evidence—update punctuality, error and correction rates, schema changes, authorization time, availability, and provenance incidents. Those measures reveal whether a model builder can depend on a public-data product after the pilot.

Third, buyer concentration should fall for the right reason. Growth in small and midsize users, cross-region buyers, and repeated sector use would show that the operator layer is reducing fixed costs. A rising transaction total driven by a few platform companies would instead confirm that the market is scaling around incumbents.

Fourth, AI outcomes need a visible boundary. A vendor claiming improvement from an authorized dataset should state the task, dataset version, model, evaluation set, baseline, and deployment conditions. More petabytes do not establish better performance, and a derived verification API may create more value than a much larger raw corpus.

The thesis has a straightforward falsifier. If active customers broaden, products renew, cross-region reuse becomes routine, service quality stays measurable, and model outcomes improve under disclosed tests, then the permitted-return system will have turned public data into accessible infrastructure. If inventories and product counts keep rising while buyers remain concentrated and usage stays undisclosed, the system will have priced the route without proving that traffic exists.

Guiyang’s expo entrance makes the data economy look like a place one can simply walk into. The harder work starts beyond the sign: permission, packaging, price, maintenance, and reuse. China has chosen to govern those steps like a public-service market. Its next important number is not how much data exists, but how often a qualified user comes back.

Sources

  1. National Data Administration, National Data Resources Survey Report (2025) (published April 2026; national inventory, circulation, AI-use, purchasing, public-data, enterprise-value, and survey-method figures).
  2. National Development and Reform Commission and National Data Administration, Implementation Specification for Authorized Operation of Public Data Resources (Trial) (effective March 1, 2025; operator selection, agreement, disclosure, registration, security, and raw-data boundaries).
  3. National Development and Reform Commission and National Data Administration, “Notice on Establishing a Price-Formation Mechanism for Authorized Operation of Public Data Resources” (January 2025; free public-interest use, government-guided industrial fees, permitted revenue, profit rate, ceilings, and reviews).
  4. National Data Administration and partner institutions, Guidelines for Building High-Quality Datasets (August 2025; definitions, static and dynamic quality, inventory baseline, construction method, and acknowledged market and governance bottlenecks).
  5. National Data Administration, “Guangzhou advances implementation of a public-data authorized-operation price mechanism” (September 4, 2025; city-level fee design, catalog, application, and product-listing figures).
  6. Ren Ming et al., “Promoting Data Element Marketization: Progress, Problems and Solutions,” Documentation, Information & Knowledge 41, no. 5 (2024; 31-region study of market development, regional imbalance, institutional gaps, and exchange-trading scale).
  7. People’s Daily Online, “Professional exhibition of 2025 China International Big Data Industry Expo opens in Guiyang” (August 28, 2025; event context and Tu Min documentary photograph used as the article image).
Previous MindSpeed-LLM now speaks FSDP2; its support table still says ‘Test’

Recommended In ai china

Matched by subject and format