ai china

OlympiadBench asks which exam sits behind the score

5 sources 5 primary sources October 8, 2026

Loading reads and saves…
Text
Tsinghua University main building beyond a broad lawn, with trees framing its pale stone facade.

Tsinghua University's Main Building. Campus context for one of the institutions behind OlympiadBench. Photograph from Tsinghua University's official gallery; resized.[5]

Imagine two exam papers bearing the same title. One asks for an English geometry answer from a competition; the other presents a Chinese physics problem with an accompanying figure. A higher mark on one paper tells you something about performance. It does not, by itself, identify whether language, subject, difficulty, or visual interpretation made the difference.

That distinction sits inside OlympiadBench, developed by researchers at Tsinghua University, Beihang University, and Wisdom Way AI Lab and published at ACL in August 2024.[1] Its contribution to China's AI research is an unusually inspectable exam: the questions, solutions, categories, and evaluation machinery are available for scrutiny.[2]

This is a reading of that evaluation design, checked on October 8, 2026. The model percentages discussed below belong to the original research, not a ranking of today's systems.

Bilingual does not mean the same paper twice

The released collection contains 8,476 problems, with records separating the question, worked solution, final answer, subject, language, and other annotations. Its public dataset viewer lets a reader inspect individual examples rather than accept the headline count on trust.[3]

The paper's inventory contains 2,125 English and 6,351 Chinese problems. It also separates 6,728 open-ended questions from 1,748 theorem-proving problems. These are counts for the collection, not a promise that every item participates in every reported score.[1]

The subset names expose another distinction. OE_TO_maths_en_COMP identifies open-ended, text-only English competition mathematics. OE_MM_physics_zh_CEE identifies open-ended Chinese physics with images from the college-entrance-exam category. The repository supplies the naming key and explains that its material draws on international and Chinese competitions as well as Gaokao questions.[2]

Reading those labels changes the claim one can make. A Chinese-versus-English comparison also needs an account of which examinations and subjects each pool contains. Calling both pools “bilingual” does not turn them into matched translations. My inference is that an aggregate language gap cannot isolate the cost of switching languages.

A controlled follow-up would keep the problem fixed: preserve the figures and quantities, commission a checked translation, and compare responses under the same prompt and generation budget. That would answer a narrower question. The existing collection answers a broader one about performance across different kinds of demanding school science.

The denominator moves before the model does

The original GPT-4V row makes the danger easy to see. The repository reports 17.97% in its broader benchmark table and 29.07% in its text-only table.[2] The apparent improvement comes from changing the evaluated question pool; it is not evidence that an upgraded model acquired another eleven percentage points of skill.

The paper's main experiments cover automatically scorable open-ended questions. Proofs receive separate manual sampling. Models are prompted zero-shot, with answer-format instructions; the reported overall metric is micro-average accuracy, so individual questions contribute to the total rather than every named subset receiving equal weight.[1]

Consider a hypothetical evaluation containing many questions of a type a model handles well and a smaller group it routinely misses. Its overall score can look reassuring while the smaller group remains weak. That is a property of the average, even when every answer has been graded correctly.

For someone assessing an assistant for Chinese physics exercises, the relevant evidence would therefore include that subject-and-language subset, its image requirements, and its item count. A broad average remains useful for comparison when the test stays fixed. It becomes harder to interpret when two reports quietly select different papers from the same examination cupboard.

The grading rules travel with the questions

There is a second handoff between the research release and a runnable evaluation. ModelScope's Chinese-language EvalScope documentation identifies its dataset split as train, specifies mathematical answer judging, requests a final answer inside \boxed{}, and warns that theorem-proving subsets cannot currently be evaluated automatically. It also supports numerical tolerances for approximate answers.[4]

The word train here is a storage label for the split the evaluator reads. It is not evidence that a particular tested model trained on those questions. Equally, a boxed answer is a formatting convention that helps an evaluator locate a result; the box does not certify a proof.

The dataset itself keeps worked solutions and final answers in separate fields.[3] That separation is valuable: it leaves material available for a human audit even when a scoring program needs a compact target. In the public viewer, a geometry item asking for the area of a quadrilateral includes both a numerical target and a derivation using the two constituent triangles.[3] A reviewer can inspect how the result was obtained.

I would treat reproducibility as a small bundle of evidence: the exact question IDs and dataset revision, delivered text and images, model version, prompt, generation limits, and grading implementation. Recording these together would let a later reader distinguish a better solver from a changed exam or a changed marker. This is a proposed reporting practice, not an additional OlympiadBench result.

The strongest contribution here is the ability to ask precise follow-up questions. Which Chinese physics problems improved? Did the same images reach both models? Were proofs assessed separately? OlympiadBench gives those questions concrete objects to attach to. A credible score begins with knowing which paper was on the desk.

Sources

  1. Chaoqun He et al., “OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems,” ACL, August 2024 — affiliations, Table 2, and experimental scope and metric in Section 4.
  2. OpenBMB, official OlympiadBench repository — released data, subset naming, and historical full-pool and text-only result tables; accessed October 8, 2026.
  3. OlympiadBench authors, Hothan/OlympiadBench dataset — public records, annotations, and worked examples, including geometry item 2417; accessed October 8, 2026.
  4. ModelScope, EvalScope's Chinese-language OlympiadBench documentation — split, prompts, subset definitions, and automatic-grading limitations; accessed October 8, 2026.
  5. Tsinghua University, “The Main Building,” official campus photo gallery — photograph and location provenance; accessed October 8, 2026.
Previous China’s lunar crater census keeps the uncertainty in view

Recommended In ai china

Matched by subject and format