ai china

Xunzi’s smallest editorial task: add punctuation, keep the text

7 sources 5 primary sources September 15, 2026

Loading reads and saves…
Text
Participants at the Beijing launch of the Xunzi classical Chinese language model.

Participants at the December 2, 2023 Xunzi launch and ancient-text AI workshop in Beijing. Photograph supplied by the research project team and published by Chinese Social Sciences Net.[7]

Ask a language model to punctuate an old text and the assignment sounds almost modest: supply the pauses, identify the sentences, make the passage easier to read. But there is another obligation beneath those instructions. Every original character must survive.

Xunzi, a family of Chinese language models developed for classical texts, offers a particularly revealing example. Its promise includes helping readers enter difficult books; its punctuation experiments show how easily that help can become an unrequested edit. The useful advance is an assistant whose contribution can be separated from the document it received.

A scholarly tool with an editorial job

Wang Dongbo’s team at Nanjing Agricultural University launched Xunzi in Beijing on December 2, 2023, working with Zhonghua Book Company’s Gulian subsidiary. The university described a corpus exceeding two billion Chinese characters, including material from the Siku Quanshu, the great imperial collection of classical writings. Automatic punctuation appeared alongside translation, indexing and information extraction in the announced capabilities.[1][7]

The project released both a base model, XunziALLM, and a conversational model, XunziChat. That distinction gives researchers two different starting points: adapt a model to a specialized task, or ask an already conversational system to perform it. The repository documents variants built on Qwen, Baichuan2 and ChatGLM3, among other releases.[2]

Consider the editorial task as a sequence. A transcription gives the computer characters. Punctuation proposes how those characters belong together. Translation then proposes a reading in another linguistic form. Combining these operations in one fluent answer would make it harder to locate an error. A reader needs to know whether the machine misread the document, divided it badly or misunderstood the resulting sentence.

Five readings, then a choice

In their May 2024 EvaHan system paper, Shitu Huo and Wenhui Chen explored a two-round prompting method using Xunzi-Qwen-7B. They first generated five punctuated versions, then asked the model to integrate a preferred result. The first round used a temperature of 0.95 to encourage variation; the second used 0.1. Their reported environment was an RTX 4090 with 24 GB of memory and CUDA 12.1.[3]

On Test A, their punctuation F1 rose from the supplied Xunzi chat baseline’s 61.06 to 64.17. F1 balances how many proposed marks are correct with how many required marks the system finds; it is not the percentage of passages ready for publication. The result belongs to that model, test set and prompting procedure. The paper does not establish a corresponding saving in editors’ working time.[3]

The method suggests a useful role for variation: alternative readings can become material for comparison. Yet five outputs from one model do not amount to five independent scholarly judgments. Their agreement cannot establish which reading a difficult passage should receive. For an editor, the alternatives would be more useful if the disputed positions remained visible after the model selected its favorite.

The characters that went missing

The organizers’ overview makes the danger concrete. EvaHan’s Test A covered four genres of previously unpublished texts and evaluated eleven punctuation types. Across submitted results, Chinese-character discrepancies were generally around 1–2%, with a maximum of 8%. Models added or omitted characters while attempting to annotate them. These figures describe the evaluation’s submissions; they are not an error rate established for Xunzi alone.[4]

An invented punctuation mark and an invented character are different editorial events. A comma offers a division within the supplied wording. An added character changes that wording. If a searchable edition quietly accepts both, a later reader may have no way to tell which material came from the source and which came from the model.

The overview also reports inconsistent pairing of quotation and book-title marks.[4] Preserving the text therefore does not settle the punctuation problem. A system can retain every character while still attributing speech incorrectly or drawing the wrong boundary around a title. Mechanical fidelity is the first condition for review, not the conclusion of interpretation.

Give the model less freedom at the right moment

A separate EvaHan entry, SPEADO, shows how this requirement can shape the software. Researchers from Renmin University of China and Shanghai Midu Technology adapted Xunzi-Qwen-7B with task-specific training, retrieved reference examples and output controls.[5]

Their key restriction operated during generation: the next predicted character had to come from the permitted original input or be a punctuation mark. This is constrained decoding—limiting the choices available as an answer is produced. It addresses preservation inside the generation process instead of relying solely on an instruction to be faithful.[5]

The complete SPEADO system reported 75.24 punctuation F1 on Test A, against the same 61.06 chat baseline. Its setup used A100 GPUs. That improvement belongs to the combined method, which also included fine-tuning, example retrieval and voting; the paper does not isolate the entire gain as the effect of the output restriction.[5]

My practical reading is that an editorial interface should keep the transcription fixed and store punctuation as a separate annotation. Any proposed character correction should become its own explicit suggestion. Reviewers could then accept a pause without accidentally accepting a rewritten name. This is a design implication of the studies, rather than a claim about a deployed Xunzi editing product.

What progress would look like now

As checked on September 15, 2026, Xunzi’s official site lists newer Qwen3-based models and describes a further generation under development. It also acknowledges weaknesses in training-data quality, generalization and systematic historical knowledge.[6] The 2024 scores should therefore remain attached to their older experimental systems; they cannot grade the newer releases.

For the punctuation use case, a convincing next demonstration would follow editors through complete passages. It would report review time, remaining punctuation errors and whether the original character sequence survived, with results separated by genre. A polished sample alone cannot answer those questions.

Xunzi’s appeal is easy to understand: more people could work with texts that currently demand substantial specialist preparation. The durable contribution would be a reading aid that leaves a clear record of its decisions. Add the marks, preserve the wording, and let the reader see where interpretation begins.

Sources

  1. Nanjing Agricultural University, report on the Xunzi launch, December 7, 2023; Chinese first-hand account of the December 2 event and project scope.

  2. Xunzi development team, XunziALLM repository; base and conversational models and their model-family relationships, accessed September 15, 2026.

  3. Shitu Huo and Wenhui Chen, “Ancient Chinese Sentence Segmentation and Punctuation on Xunzi LLM,” LT4HALA 2024, pp. 242–245; prompt procedure, hardware and Table 3 Test A results.

  4. Bin Li et al., “Overview of EvaHan2024: The First International Evaluation on Ancient Chinese Sentence Segmentation and Punctuation,” LT4HALA 2024, pp. 229–236; evaluation design, paired marks and character discrepancies.

  5. Xia Tian, Yu Kai, Yu Qianrong and Peng Xinran, “SPEADO: Segmentation and Punctuation for Ancient Chinese Texts via Example Augmentation and Decoding Optimization,” LT4HALA 2024, pp. 256–260; Section 3.3 and Table 3.

  6. Xunzi project, official Chinese-language site; model catalog, acknowledged limitations and development plans, accessed September 15, 2026.

  7. Zhixiao Zhao, “Xunzi classical-text language model launch held in Beijing,” Chinese Social Sciences Net, December 18, 2023; first-hand event report and photograph supplied by the research project team.

Previous RoboTwin makes the tabletop part of the robotics test

Recommended In ai china

Matched by subject and format