Imagine an assistant reading a photographed invoice. It returns the correct total, and a reviewer asks it to highlight the line that supplied the number. The highlight lands on the delivery address. The answer may still be useful, but the reviewer now has to search the page to verify it. A seemingly small request—show me where—has exposed a second job hiding inside “read this document.”
That separation gives OCRBench v2 its sharpest question. Developed by researchers at Huazhong University of Science and Technology, South China University of Technology, ByteDance, and the University of Adelaide, it examines what happens when visual reading must also support localization, structure, and reasoning.[1]
This assessment was checked on October 10, 2026. The diagnostic result below comes from the paper's June 5, 2025 revision; the live leaderboard is a separate, later record.
A correct answer can leave the reviewer searching
The original OCRBench assembled 1,000 checked question–answer pairs across five components, including text recognition, document questions, information extraction, and handwritten mathematics. The project's repository describes v2 as a larger bilingual collection of 10,000 human-verified pairs, reaching across 31 scenarios. It records the v2 release on December 31, 2024.[2] The change gives evaluators more ways to ask whether successful reading survives a different demand on the same visual medium.
One historical result makes that demand tangible. In the June 2025 paper, InternVL3-14B scored 78.3% for answering VQA-with-position questions, while its answer-region intersection-over-union score was 12.9%. Intersection over union measures the overlap between the predicted and reference regions relative to their combined area. It is a geometric score, so the second figure does not mean that exactly 12.9% of answers were correct.[1]
The authors report using official pretrained weights, inference code, and default parameters for open models. Their localization prompts normalize coordinates to a 0–1000 scale. These details bound the result: it tests that model's responses under that evaluation setup, including its ability to express a location in the requested format.[1]
For the imagined invoice reviewer, the practical implication is straightforward. Returning a number and returning a usable evidence region deserve separate checks. A box around the wrong line can slow review even when the number is right. Conversely, a well-placed box still leaves the reviewer to establish whether the selected line answers the question. Each output earns a different kind of trust.
The score changes meaning with the task
The public evaluation code makes these distinctions concrete. Text grounding calls an overlap calculation. Table parsing uses TEDS, a measure of similarity between table structures. Full-page OCR combines BLEU, METEOR, F-measure, and a term based on edit distance.[3] A score can therefore reward several kinds of resemblance: where text sits, how a table is organized, or how closely a transcription follows the reference. The VQA-with-position scorer itself averages answer quality and region overlap equally; the paper’s separate diagnostic figures expose what that combined score can hide.[3]
Consider a hypothetical table whose words are all legible but whose cells have been assigned to the wrong rows. A transcript could preserve much of the vocabulary while damaging the relationship between a product and its price. That is why the scoring function belongs in the explanation of a result. The evaluator needs to reward the property the application actually depends on.
There is also a language boundary. ModelScope's Chinese-language EvalScope documentation lists 300 English VQA-with-position items, alongside Chinese tasks such as handwritten-answer extraction and full-page OCR. It lists no Chinese counterpart for VQA with position.[4] The historical localization result should therefore not be presented as a measured Chinese-language localization score simply because the overall benchmark is bilingual.
The same documentation specifies zero-shot evaluation on the test split and exposes named subsets. Its short metric label is accuracy.[4] Read alongside the authors' task-specific scoring code, that label is a reason to inspect the actual calculation before interpreting a percentage.[3] A useful report would name the language, task, scoring implementation, and input settings together. Those details tell a reader which kind of reading improved.
A leaderboard period is not a shared run date
The official private-data leaderboard adds a further distinction. Its 2026.09 description says it carries forward June results and incorporates five newly evaluated models from an October 4, 2026 batch.[5] The displayed period therefore groups measurements made at different times. It should not be read as evidence that every listed system was freshly rerun together in September.
That provenance does not by itself invalidate the comparisons. It does change what a reader should request before attributing a movement to model progress: the tested model version, whether its row was carried over, and whether the evaluation settings remained comparable. Likewise, the paper's older weakness is a diagnostic example; it cannot establish the localization ability of every newer model on the board.
For builders, I would start with a small set of their own documents and retain two outputs from every run: the answer and the evidence region. Review them separately, then examine whether a person can use them together. Add structure checks where tables matter. This would turn OCRBench v2's most useful distinction into an application test without assuming that its aggregate score predicts a particular office's workload.
The invoice reviewer needs a short path back to the page. OCRBench v2 helps make that path something an evaluator can ask for, inspect, and score.
Sources
- Ling Fu et al., “OCRBench v2,” arXiv v2, June 5, 2025 — Sections 4–5 and Appendix A.6.
- Yuliang Liu and collaborators, official MultimodalOCR repository — original benchmark components, v2 collection, and dated release history; accessed October 10, 2026.
- OCRBench v2 authors, official evaluation scripts — eval.py task dispatch and IoUscore_metric.py combined answer-and-location scoring; accessed October 10, 2026.
- Alibaba ModelScope, EvalScope's Chinese-language OCRBench-v2 documentation — subset inventory, evaluation split, zero-shot setting, and metric label; accessed October 10, 2026.
- OCRBench v2 authors, official project leaderboard — private-data evaluation and the provenance note for the 2026.09 board; accessed October 10, 2026.
- China University of Geosciences (Wuhan), School of Computer Science, report on its anniversary forum, December 23, 2025 — Yuliang Liu’s talk and event photograph.