A video assistant can guess why a character leaves a room while getting the preceding action wrong. If an evaluation counts each answer separately, that successful guess still earns a point. Video-MME-v2 makes the relationship between answers part of the test: can the system recognize the evidence, follow the sequence, and sustain a conclusion across related questions?[3]
The project belongs to an important part of China’s AI research ecosystem: building the instruments used to judge models. Nanjing University’s Chinese-language faculty profile lists the Video-MME series among Chaoyou Fu’s representative work.[1] The original benchmark, launched in June 2024, offered 900 videos and 2,700 questions, with results separated by video duration and subtitle setting.[2] The second version, launched on April 7, 2026, changes the unit of assessment to groups of four questions.[3]
This article examines the April 2026 paper and the public evaluation materials checked on September 26, 2026. Its subject is how the test works, rather than a claim about today’s leading model.
A correct ending needs support
The bilingual project page provides a revealing geometry example. Two questions establish separate component values. A third combines them to calculate an area; a fourth derives the final result. The declared dependency structure is [[1,2],3,4]: the opening steps are parallel, and the later ones depend on them.[6]
That arrangement makes a familiar classroom problem visible in model evaluation. A student might reach the expected answer through an invalid intermediate calculation. Awarding the final point records success, but leaves unanswered whether the method would survive a small change in the problem. Related questions create opportunities to expose that fragility.
Video-MME-v2 distinguishes two group types. Consistency groups test a capability across related questions. Coherence groups test answers linked by a reasoning structure.[4] The distinction matters because counting correct responses and checking dependencies are different operations.
The released scoring code makes this concrete. A consistency group with three correct answers receives 56.25 out of 100, calculated as (3/4)² × 100; ordinary accuracy would be 75%. For a simple four-step chain, only the correct prefix before the first error counts. A sequence of correct, wrong, correct, correct therefore earns 6.25, while correct, correct, correct, wrong earns 56.25. Both have the same ordinary accuracy. These are worked examples from the evaluator’s rules, not measured model results.[5]
The code also gives branching structures their own treatment. It can credit an independent opening branch even when its sibling fails. Applying the simple-chain rule indiscriminately would misdescribe the benchmark.[5]
There is a boundary to the claim. These scores evaluate the pattern of submitted answers under the annotators’ dependency structure. They do not reveal the model’s private reasoning process. My reading is that this is a stronger behavioral test of whether answers support one another, while leaving room to question whether each annotated dependency is well chosen.
Subtitles change what the model receives
The repository offers materially different inputs: a fixed 64-frame sample or one frame per second; visual frames alone or frames with subtitles; a direct-answer prompt or a reasoning prompt. Subtitle text can arrive as one concatenated block or be interleaved with frames using timestamps.[3]
Consider a hypothetical cooking clip. The camera shows a covered pan while a voice announces that the heat has been turned off. A transcript supplies information the sampled picture may not contain. In another clip, the cook says one thing while visibly doing another. There, synchronized evidence matters because the task is to recognize a mismatch. Neither example is a benchmark item; they illustrate why an input setting changes the question an evaluation can answer.
The April paper reports that thinking modes helped with textual cues but could reduce performance when only visual frames were supplied. Its “with subtitles” condition also includes raw audio for models capable of accepting it.[4] That column is therefore broader than a uniform transcript added to every system.
An improvement under that condition can reflect better use of speech, text, images, or their combination. To attribute it specifically to visual reasoning requires a controlled comparison. The practical implication is to keep the modality setting attached to every quoted score.
The headline number has a denominator
The paper’s April snapshot reports Gemini-3-Pro at 66.1% average accuracy and 49.4 on the grouped scale in the condition with audio/subtitles. The human baseline is 94.9% and 90.7, respectively. Gemini’s listed sampling rate is one frame per second; other systems use different frame budgets according to their input constraints.[4]
Those figures establish a gap on this curated assessment. They do not mean the model understands 49.4% of arbitrary videos. Nor does the decline from accuracy to grouped score represent additional questions answered incorrectly: the same predictions are being aggregated differently.
A lower grouped score is partly the intended consequence of the metric. The useful question is whether it distinguishes failure patterns that matter. An assistant assembling an incident report, for example, must connect the observed action to the correct sequence before attributing a cause. For that imagined task, isolated correct answers would offer incomplete reassurance.
A convenient runner can report a different measure
There is one more distinction at the tooling layer. EvalScope’s Chinese-language integration guide describes a default zero-shot evaluation with accuracy as its primary metric. It uses public video URLs by default, allows the official archived MP4 files instead, and makes subtitle inclusion configurable.[7]
That is a useful way to run the dataset, but its documented primary result is not automatically the grouped score used in the benchmark paper. The benchmark’s own standalone evaluator explicitly calculates grouped ratings.[5] A report should identify which measure was produced before comparing it with a published table.
My conclusion is that Video-MME-v2’s most useful contribution is an inspectable test of connected answers. To evaluate a proposed improvement, preserve the model version, frame selection, audio or subtitle treatment, prompt, and scoring implementation. Then examine which groups improved. A system that repairs the missing first step has achieved something different from one that becomes better at guessing the ending.
Sources
- Nanjing University, Chaoyou Fu’s faculty profile (Chinese); institutional affiliation and the Video-MME research series, checked September 26, 2026.
- Video-MME team, original benchmark repository; June 2024 launch, dataset size, duration categories, and subtitle evaluation settings.
- Video-MME-v2 team, official repository; April 2026 launch, grouped evaluation, frame configurations, subtitle alignment, and prompting, checked September 26, 2026.
- Chaoyou Fu et al., “Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding,” arXiv:2604.05015v1, April 6, 2026; group definitions, Table 1, input conditions, and thinking-mode analysis.
- Video-MME-v2 team, standalone evaluation code, revision 28fc3bc; consistency scores and chain/branch handling in the scoring functions.
- Video-MME-v2 team, bilingual project page; geometry example and its declared question dependencies, checked September 26, 2026.
- EvalScope, Video-MME-v2 integration guide (Chinese); zero-shot defaults, primary accuracy metric, media source selection, and subtitle option, checked September 26, 2026.
- ScareCriterion12, photograph of Nanjing University’s Gulou campus, March 12, 2018; Wikimedia Commons, CC BY-SA 4.0.