The Great Daxinzhuang Pottery Puzzle Challenge began with a seductively large number: nearly 20,000 Shang-dynasty sherds, photographed on both sides and opened to a global field of competitors. It ended with a smaller and more useful one. The detailed results notice records 78 physically confirmed pairs for the best-performing team; a second team had two. No first or second prize was awarded.[1]
That outcome is not a footnote to the AI experiment. It is the experiment. Ninety-four submissions entered the competition, but only two of the 17 team result sets considered by the organizers contained explicit, traceable one-to-one proposals that advanced to physical verification. The winning tier was therefore decided not by an online similarity score, but in the Daxinzhuang Site Museum's storage room, where archaeologists put the original pieces together and inspected the joins.[1]
The lesson is sharper than “humans remain important.” In this use case, the physical object is part of the inference system. A photograph can make a candidate searchable; it cannot turn that candidate into an archaeological fact.
A storage pit is not a jigsaw box
Daxinzhuang, in present-day Jinan, was an important eastern center of the Shang world. The H690 feature at the heart of the challenge began as a storage pit and later became a refuse deposit with 14 layers. Its fill held pottery alongside animal bone and fragments of gold foil. For the competition, the organizers released front-and-back high-resolution images, unique identifiers, excavation-layer information, and archaeological notes for close to 20,000 pottery fragments.[1][2]
Those facts make the collection valuable, but they also explain its difficulty. A manufactured jigsaw starts with one complete picture, one cutting event, and every piece in the box. H690 accumulated through use, breakage, discard, burial, and recovery. The organizers warned that not every corresponding fragment could necessarily be located, and that surviving pieces may be weathered, abraded, discolored, or visually plain.[2]
At this scale, brute-force pairing is punishing. A collection near 20,000 items creates on the order of 200 million possible unordered pairs before stratigraphy, vessel type, clay body, or other evidence narrows the field. Almost all of those pairs are wrong, while the number of true joins is not known in advance. A system can therefore look accurate by rejecting nearly everything and still fail at the archaeologist's actual need: placing a manageable number of real joins near the top of the queue.
This is why “pottery recognition” is not one task. Classifying a sherd by fabric, period, or vessel family can reduce the search space. Retrieving plausible neighbors can prioritize inspection. Aligning fracture edges can estimate a pose. Reconstructing a vessel requires a consistent group rather than one attractive pair. Each stage needs a different target and a different error budget.
The leaderboard could not supply the truth
The original evaluation plan gave 70% of the score to matching accuracy, 15% to vessel completeness, and 15% to solution quality and innovation. Yet the competition's own FAQ acknowledged the central benchmark problem: the organizers did not possess complete matching data for all fragments, so the live leaderboard's displayed scores were not reliable. Final judgment would have to occur after submissions closed.[1]
That boundary matters. In a standard vision benchmark, hidden labels already exist and the platform compares predictions with them. At Daxinzhuang, many of the labels had to be produced by handling the artifacts. The competition was simultaneously testing algorithms and discovering its own ground truth.
Other archaeological vision projects show how easily a respectable score can answer the wrong question. A 2026 study of 1,864 Hemudu pottery photographs found that a ResNet-50 could classify binarized sherd silhouettes at roughly 74% accuracy even though fragment shape was not supposed to define the target fabric class. The full-image model reached about 95%, and the authors treated the shape correlation as a shortcut to suppress rather than evidence to trust.[4] For fracture matching, however, edge shape may be directly relevant. The same pixel can be signal for one task and leakage for another.
A People's Daily account of a separate oracle-bone project reported more than 90% accuracy on its evaluated matching setup and 37 newly joined pairs, using edge alignment, assembly, and texture continuity.[3] That is a credible proof that learned retrieval can help. It does not imply that the same percentage will survive an open-world collection containing millions of plausible-looking negatives, missing mates, and plain surfaces. Accuracy has meaning only inside the candidate pool and labeling protocol that produced it.
Clay gets a veto
On December 27–28, 2025, the Daxinzhuang organizers moved the qualifying proposals from screen to storage room. They first checked the documentation: catalog numbers, quantities, and the claimed relationships had to remain traceable to the submitted records. Then they placed the actual sherds together under stable light.[1]
The inspection was deliberately material. Archaeologists looked for continuous fracture topography and a natural fit. They compared clay color, texture, and temper; decoration and surface treatment; wall thickness and curvature. A proposal could be confirmed, provisionally supported, or rejected as a non-continuous break or a mismatch in fabric and form.[1]
None of those checks is ceremonial. Two photographs flatten thickness, fracture depth, and the feel of a contact surface. Lighting can move color. Scale and camera angle can distort curvature. A thin deposit can hide a continuous line. Even a very persuasive edge alignment may join two pieces that were made from different clay bodies.
Earlier 3D reassembly research makes the missing information concrete. One published method used sherd-thickness profiles because thickness is stored inside the ceramic body and may be less damaged by burial than surface color or decoration.[5] Daxinzhuang's two-face image release was a pragmatic way to make a huge collection accessible, but it could not expose every measurement needed for final confirmation. The museum table was not an old-fashioned fallback after the model finished. It was the final sensor.
A better second round would reward the funnel
The official result notice says another round is being planned.[1] Its most useful benchmark would not ask one score to summarize the entire reconstruction problem. It would measure a funnel.
First comes candidate retrieval: for each sherd with a known mate in a hidden, physically verified subset, does the system place that mate within the top 10, 50, or 100 suggestions? Recall at a fixed review budget tells an archaeologist how much searching the model actually removes.
Second comes compatibility precision: when a system asserts a high-confidence join, what share survives blind physical inspection? The evaluation set should include hard negatives from the same layer, fabric, vessel class, and color family, because easy negatives reward superficial sorting rather than fine discrimination.
Third comes group consistency. Two good pairwise scores do not guarantee that three or four fragments form one coherent vessel. Curvature, wall thickness, orientation, and archaeological context must remain consistent across the group. Systems should also be allowed to abstain. Forcing every sherd into a cluster converts honest uncertainty into false reconstruction.
Finally, the benchmark should record expert minutes per confirmed join. A retrieval model that turns 200 million theoretical comparisons into 200 strong candidates may be transformative even if it never “rebuilds a pot” autonomously. Conversely, a visually impressive assembly that takes longer to audit than manual work has not yet improved the field process.
These measures would preserve the competition's most valuable design decision: every digital claim remains linked to a cataloged object and can return to physical verification. They would also make progress legible from round to round. Better retrieval, better abstention, and less expert time are all meaningful advances even before a full vessel appears.
The AI-China signal is the verification chain
This is a different kind of AI-China story from a frontier-model launch. Its critical infrastructure is a named excavation, a museum storeroom, durable catalog numbers, standardized photography, archaeologists willing to adjudicate uncertain matches, and an open competition that published an inconvenient result.[1][2]
The absence of first and second prizes is therefore a strength. It prevents the promotional claim from outrunning the artifacts. The challenge established that a large 2D corpus can attract algorithmic work and surface real joins. It also established where the present evidence stops: most of the team result sets reviewed for on-site verification did not express explicit one-to-one pair claims, and even the best output required direct material judgment.
Daxinzhuang's productive division of labor is now visible. Let the model search at a scale no person can. Let it rank, document, and abstain. Then let clay continuity, fabric, thickness, curvature, provenance, and expert handling decide what becomes part of the archaeological record. The goal is not to remove the museum table. It is to make every hour at that table yield more history.
Sources
- BNBU Institute for Advanced Study and Shandong University School of Archaeology, “Great Daxinzhuang Pottery Puzzle Challenge” — official competition description, rules, participation totals, result notice, and December 2025 physical-verification protocol.
- Yu Qinlin, “用AI拼合陶器碎片,头奖1万美元!这项考古大赛在珠启动,” Guanhai Rongmei, July 26, 2025 — H690 context, dataset description, archaeological constraints, and the Zhu Wen photograph used above.
- Yang Qingyue, “当考古遇到人工智能,” People's Daily, January 6, 2025 — overview of archaeological AI and the reported oracle-bone fragment-matching workflow.
- Xin Yu et al., “Mitigating spurious features by contrastive learning in pottery sherd recognition,” npj Heritage Science 14, 135, March 4, 2026 — Hemudu dataset design, acquisition controls, shortcut features, and evaluation boundaries.
- Michail I. Stamatopoulos and Christos-Nikolaos Anagnostopoulos, “A totally new digital 3D approach for reassembling fractured archaeological potteries using thickness measurements,” Acta IMEKO 6(3), 2017 — physical-profile evidence for sherd reassembly.