ai china

China’s AI copyright defense now starts in the training log

6 sources 2 primary sources September 8, 2026

Text
Five Supreme People’s Court officials seated at a press conference beneath a blue screen naming the court’s opinion on AI-related disputes.

Supreme People’s Court officials present the Opinion on Adjudicating AI-Related Dispute Cases on September 7, 2026. Photo by Wang Shengxiang, via the Supreme People’s Court.[1]

A screenshot of an AI output can show resemblance. It cannot, by itself, show how that resemblance arose. Was the claimant’s work in a training corpus? Did a user steer the model toward it? Was the result repeatable, or was one unusually close image selected from hundreds? Which checkpoint, filter, adapter, and service route produced it?

On September 7, 2026, China’s Supreme People’s Court put those questions into a national judicial-policy document. Its 24-article Opinion on Adjudicating AI-Related Dispute Cases covers personality rights, personal information, consumer protection, automated vehicles, intellectual property, evidence, and AI-assisted court filings. For model builders, its sharpest operational change sits in Article 12: when a developer raises a non-infringement defense in an AI copyright dispute, the court should require support concerning the sources of training data, records of the training process, the model’s operating mode, and, where necessary, relevant scientific theory.[1][2]

That is not a demand that every model publish its weights or corpus. It is a case-triggered evidentiary route. The claimant still has work to do, and the court still decides what evidence matters. But a developer can no longer assume that a polished model card and a present-day demo will reconstruct a disputed system after the fact. The defensible object is becoming a chain of records that begins before training and ends at the particular output.

The delta is courtroom provenance

The Opinion is guidance for China’s People’s Courts, not a statute, a formal judicial interpretation, or a dedicated national AI law. The Court itself says China has not yet enacted a specific AI law and describes the document as rules and guidance drawn from existing civil, copyright, personal-information, competition, consumer, and procedural law. It is also careful about different levels of control: general and specialized models, open and closed systems, and developers, providers, and users do not automatically carry the same responsibility.[1]

Training-data governance was already part of China’s regulatory baseline. The 2023 Interim Measures for Generative AI Services require public-facing providers within their scope to use lawfully sourced data and foundation models, respect intellectual property and personal-information rights, and explain training-data sources, scale, types, labeling rules, and algorithm mechanisms when regulators lawfully inspect them.[4] In April 2026, the Supreme People’s Court’s five-year intellectual-property plan said courts would prudently adjudicate disputes over large-model training corpora and AI-generated content.[3]

The September Opinion makes a specific litigation use explicit at national level. Provenance had already surfaced in individual courts; now the guidance says that, in a civil case, it may be required to support a developer’s own non-infringement position. That standardization addresses a practical problem Chinese courts had already identified: in September 2025, the Beijing Internet Court said AI disputes increasingly required examination of generation processes, algorithms, models, and data sources across a chain of trainers, developers, providers, and users, making technical fact-finding unusually difficult.[6]

One claim now travels in two directions

Consider the use case the Opinion implicitly sketches. A rightsholder alleges that an image generated by a particular service is substantially similar to a protected work and sues the developer or provider. The Court’s official Q&A says the claimant should first establish the preliminary facts: that the challenged content came from that AI and that it is substantially similar to the earlier work. Opacity does not erase that opening burden.[2]

The evidentiary direction can then reverse. If the developer says the output did not infringe, Article 12 directs the court to require supporting material on training-data sources, training-process records, model operation, and, where necessary, scientific basis. The Court explains this allocation through information asymmetry: those facts are generally inaccessible to the rightsholder but controlled by the developer.[1][2]

This is not automatic liability for similarity. Article 12 tells courts to consider the service and industry, who supplied relevant training data, each party’s participation, safeguards, and profit. It separately addresses the user who knew or should have known of an earlier work and generated something substantially similar without a reasonable defense. Responsibility is meant to follow control, knowledge, participation, and precautions—not the convenient fiction that “the AI did it.”[1]

The sequence matters. A claimant does not win merely by naming a famous model. A developer does not win merely by pointing to the user’s prompt. And a user does not become invisible because the service supplied the model. The output is the start of the inquiry; the production chain decides where it goes.

Four records sit behind one defense

Article 12 names four categories but does not prescribe a database schema or retention period. The following operational design is therefore an inference from the Opinion, not a checklist contained in it.

Training-data sources need receipts, not labels. “Public web,” “licensed data,” or “synthetic data” is too coarse to explain a challenged work. A useful provenance record would tie a versioned corpus slice to its acquisition route, source snapshot, governing terms, rights or permission basis, collection date, exclusions, and later removals. If a vendor or upstream foundation model supplied part of the data path, the record needs to preserve that dependency rather than flatten it into “third party.” The goal is not to prove that a corpus was generally careful; it is to locate the relevant path for a specific claim.[1]

A training-process record needs lineage. A source manifest says what entered consideration, not what reached a run. The defensible chain would connect corpus versions, filtering and deduplication rules, opt-out or removal actions, training and fine-tuning jobs, checkpoint lineage, adapters, and release candidates. This distinction is especially important when a work was removed after one checkpoint but before another. The model named in a complaint is a time-bounded artifact, not a product family in the abstract.

The operating mode needs the serving stack. The same foundation model can behave differently behind retrieval, a system prompt, a fine-tune, a safety filter, a reranker, or a regional API route. A record that stops at the checkpoint cannot explain which components participated in the output. “Model operating mode” should push teams to preserve the deployed combination: model and adapter identifiers, routing logic, enabled tools or retrieval, generation settings, and the content filters active at the relevant moment. Those fields are a practical reading of the Court’s category; the Opinion does not enumerate them.[1]

Scientific basis, where necessary, needs an explanation that can survive challenge. A benchmark slide is not a causal account of why a work could or could not influence an output. Depending on the dispute, a court may need expert evidence about memorization, similarity, data filtering, model architecture, or the effect of a prompt. Article 17 allows courts to use assessors, appraisers, expert assistants, and technical investigation officers for specialized questions. A useful technical explanation must remain linked to the actual system and records in the case, not a generic paper about how diffusion or language models usually work.[1][2]

These four layers reward mundane engineering discipline. A reproducible corpus build, a signed training manifest, an immutable deployment identifier, and a legible change log may matter more in court than a broad declaration that the company respects copyright.

Prompt replay tests the output, not just the corpus

The Opinion does not reduce an output dispute to training data. Article 18 tells courts evaluating AI-generated infringement evidence to consider the prompt’s design and influence, the similarity between the generated and protected works, consistency across repeated tests, and the effects of training, algorithm design, and content filters. These are factors, not a fixed test or scoring formula.[1]

That makes the generation session a second evidence surface. A litigation-grade replay bundle would ideally preserve the service and model version, exact prompt sequence, negative prompts, seed where exposed, generation settings, filter decisions, timestamps, output identifiers, and unselected comparison runs. Again, this is an operational inference. The Opinion explicitly names prompts, repeat testing, training, algorithms, and filters, but it does not require every field in that bundle.[1]

Repeatability is especially revealing. A single close output may show that a service can produce the contested expression under one sequence of instructions. It does not show how readily or consistently the system does so. Conversely, a failed replay on today’s model does not disprove an output generated by last month’s checkpoint. Preserving the deployed system’s identity is what makes “try it again” an experiment rather than theater.

China’s generated-content labeling rules already create one adjacent output-provenance layer. Since September 1, 2025, the rules have required explicit labels in specified cases and metadata labels carrying information such as provider and content identifiers; certain delivery paths without an explicit label also require recipient logs to be kept for at least six months. Those labels can help connect content to a provider; they do not establish what was in training or why a result resembles a particular work.[5] The courtroom chain therefore has two distinct halves: identify the output, then explain its production.

Preservation cannot mean “log everything”

Article 17 gives missing records consequences, but only conditionally. When a party controls documents or electronic data, refuses to submit them without a proper reason, and the opposing party says the contents would be unfavorable to the controller, a court may accept that contention. This is neither an automatic sanction nor an instruction to retain every byte forever. It means an unjustified refusal to produce controlled evidence, coupled with the opponent’s adverse-content assertion, may not be treated as neutral.[1]

The obvious engineering response—collect every prompt and keep it indefinitely—would create a different compliance problem. The 2023 Interim Measures say providers must protect user inputs and usage records, must not collect unnecessary personal information, and must not unlawfully retain input data that can identify users. The same measures require providers to preserve records connected to discovered unlawful use and regulatory reporting.[4] Evidence readiness therefore has to be scoped: retain what a lawful purpose requires, separate content from identity where possible, apply access controls, document deletion, and preserve a targeted legal hold when a dispute becomes foreseeable.

The Opinion does not announce public disclosure of a full corpus, source code, weights, or every user prompt. It gives courts tools to preserve and collect key technical evidence and to use specialists when necessary.[1] For an operator, the design problem is controlled producibility: can the company locate the relevant, lawfully retained record, authenticate it, explain its lineage, and provide it through the process the court orders?

The central training question remains open

The most important boundary appears in the official Q&A. The Supreme People’s Court says the Opinion does not decide two contested questions: whether AI-generated content is copyrightable as a general matter, and how to characterize the use of another person’s work without permission in model training. The Court says views remain divided and more practical experience is needed.[2]

That restraint prevents a common misreading. Article 12 does not declare all unlicensed training infringing, nor does it create a blanket training exception. It specifies responsibility factors and evidence for cases brought under existing law while leaving the substantive training rule unsettled. Records do not answer the legal question by themselves. They let a court find the facts to which an eventual or case-specific rule can be applied.

The April 2026 five-year plan had already asked courts to examine human prompting and modification when assessing the legal attributes of generated content and to explore ownership rules.[3] The September Opinion narrows the immediate advance: before a national answer on copyright status or training use exists, the judiciary is standardizing what it needs to see.

The first applied cases will set the real specification

Four details now deserve close attention. First is granularity: whether courts accept source categories and aggregate audits, or require work-level lineage. Second is replay: how judges control model version, prompt variation, sampling, and selection when testing consistency. Third is role allocation: how obligations travel across a foundation-model developer, a fine-tuner, an API provider, and a product that adds retrieval or filters. Fourth is confidentiality: how courts obtain enough technical evidence to test a defense while protecting trade secrets, personal information, and other lawful interests.

None can be settled by a compliance slogan. They will be defined through evidence orders, expert work, reasoned judgments, and eventually guiding cases. The companies best prepared for that process will not necessarily be those with the longest policy page. They will be those that can move from a challenged output backward—through prompt, route, deployment, checkpoint, training run, and source—without breaking the chain.

China’s new Opinion turns that chain into part of the product. The training log is no longer just for debugging the model. It may be how a developer explains the model when resemblance becomes a lawsuit.

Sources

  1. Supreme People’s Court of the People’s Republic of China, “Opinion on Adjudicating AI-Related Dispute Cases” (Fa Fa [2026] No. 10, released September 7, 2026; full 24-article text and official press-conference photograph; in Chinese).
  2. Supreme People’s Court News Bureau, “Officials answer questions on the Opinion on Adjudicating AI-Related Dispute Cases” (September 7, 2026; burden sequence and expressly unresolved copyright questions; in Chinese).
  3. Supreme People’s Court, “Implementation Plan for Judicial Protection of Intellectual Property Rights by People’s Courts (2026–2030)” (April 20, 2026; prior policy baseline for training-corpus and generated-content cases; in Chinese).
  4. Cyberspace Administration of China and six other authorities, “Interim Measures for the Management of Generative Artificial Intelligence Services” (published July 13, 2023; training-source, inspection, input-record, and provider duties; in Chinese).
  5. Cyberspace Administration of China and three other authorities, “Measures for Labeling Artificial Intelligence-Generated and Synthetic Content” (published March 14, 2025, effective September 1, 2025; explicit labels, metadata identifiers, and conditional recipient-log duties; in Chinese).
  6. Supreme People’s Court / People’s Court Daily, “Beijing Internet Court reports AI cases spreading into traditional industries” (September 16, 2025; prior court account of AI fact-finding and multi-party responsibility; in Chinese).
Previous ChemELLM 3.0 Pro can plan a chemical task. The plant still gets the final word Next Momenta's R7 reached a production Cadillac. The company's 114 nominations face the licensing test

Recommended In ai china

Matched by subject and format