At the end of a 65-second CGTN report from May 2019, a sentence appears over a colon graphic and an endoscopic screen labeled in Chinese “non-adenomatous polyp”: “The diagnosis accuracy of the technology can reach 96.93 percent.” The visuals point to the colorectal module, and CGTN’s written companion says Miying helps a doctor locate a polyp and quantify its risk. But neither the closing caption nor the narration says whether 96.93 measures localization, classification, a frame, a lesion, an examination, or a patient. No test population accompanies it. The companion adds a company-attributed comparison with physicians but supplies neither the physician cohort nor the evaluation behind that comparison.[1][2]
Tencent’s own earlier announcement offers a revealing match, although neither publication explicitly links the two figures. In July 2018, the company reported the same 96.93% for real-time colorectal-polyp localization on what it described as a large-sample, multi-source, multicenter test. That release separately gave 97.20% for distinguishing colorectal adenocarcinoma and said the system processed ten images per second.[3] On the CGTN screen, the identical decimal sits under a broader label: Tencent’s release says localization accuracy; the video caption says diagnosis accuracy.
That change in wording is the reason to watch this short artifact closely. The video does preserve something important: an early view of a Chinese medical-AI product being designed as a visual aid inside clinical work, with doctors still making the judgment. But it also shows how an interface, several cancer-screening applications, and a percentage whose precision exceeds its disclosed protocol can create a much broader impression of capability. The annotations below treat the recording as evidence of product presentation—not as a clinical validation study.
0:00–0:15 — one brand contains several tasks
The report opens with a clinical scene, then moves to a futuristic blue interface. Outlined organs and medical images make “Tencent Miying” look like a single machine for screening high-risk cancers. Yet the clip and companion article move among digestive-tract imaging, cervical screening, diagnosis, prediction, and lesion location.[1][2] Those verbs are not interchangeable. Neither are the images, labels, reference standards, or consequences of an error in each application.
The distinction matters before any metric appears. A system can mark a region in an endoscopy frame, classify an already selected image, estimate disease risk, or support a final diagnosis. Each is a different prediction problem. A percentage measured for one cannot migrate to another merely because both sit behind the same blue product interface.
The screen also performs rhetorical work. It gives heterogeneous modules a common visual language: image on one side, machine output on the other, colored marks and confidence values between them. That consistency may be useful to an operator. For a viewer, however, it can conceal the evaluation boundary. A unified interface does not imply a unified model, dataset, or level of evidence.
0:15–0:40 — “preliminary” is the clip’s most important word
Product manager Liu Juncan describes a computer analyzing medical images to help doctors diagnose and predict disease. In the digestive-tract example, the system locates a suspected lesion in real time so the doctor can reach a preliminary result. The report calls this help for a doctor, not a final decision by the software.[1]
An independent March 2019 field report from hospitals in Deqing makes that arrangement physical. Its photograph shows a Miying analysis display beside the ordinary gastroscopy monitor—the same side-by-side layout used for this article’s lead image. The report says Miying sounds a prompt when it detects a possible polyp and the doctor then decides whether to act. It also records practical limits. In the same deployment, staff reported false-positive lung-nodule prompts and said local accuracy had not yet been checked against pathology; for endoscopy, they stressed that results remained dependent on the operator’s technique.[7]
That is a more useful model of the product than “AI diagnoses cancer.” It is a second visual channel inside an existing procedure. Its value depends not only on whether a marked box corresponds to a lesion, but also on whether the prompt arrives in time, whether the doctor understands it, whether distracting marks are tolerable, and whether the combined workflow finds clinically relevant disease more reliably than the unaided workflow.
The report then says pilot hospitals send doctors’ feedback to Tencent so the algorithm can be optimized.[1] This is the company’s description of iterative development, not independent evidence that the process occurred exactly as summarized. Nor does it establish that every revised version inherits an earlier score. Once feedback changes the model, training data, threshold, or interface, the version attached to a result becomes part of the claim.
0:40–1:05 — the number changes verbs
The final sequence moves from cervical-screening research and hospital pilots back to a colon graphic and endoscopic image for the 96.93% statement.[1] CGTN’s companion article makes the colorectal context explicit: it says Miying can help locate a polyp and quantify its risk, then repeats 96.93% as a level the technology company says it can reach.[2] The task family is therefore recoverable even though the caption leaves it implicit.
The unresolved issue is the metric. Tencent’s 2018 announcement assigns the identical figure to real-time colorectal-polyp localization and gives adenocarcinoma discrimination a different 97.20% result.[3] The video instead calls 96.93% “diagnosis accuracy.” CGTN does not say its number came from the earlier release, so the match is evidence of changed public framing, not proof of a provenance chain.
Even the longer Tencent announcement does not provide enough information to reconstruct the result. It calls the test large-sample, multi-source, and multicenter without defining “multi-source”; the press release does not identify the denominator, define what counted as a correct localization, report uncertainty, describe patient-level separation between training and test data, or give a blinded external comparison.[3] The precision of “96.93” therefore exceeds the precision of the public evidence envelope around it.
This is not a claim that the figure was false. It is a claim about what the figure can support. Localization accuracy might be calculated per still image, per video frame, per lesion, or per examination. A result can change substantially when near-duplicate frames from one procedure are separated differently, when negative cases are added, or when the test moves from curated images to uninterrupted clinical video. Without the unit and protocol, 96.93% cannot be translated into the chance that a patient receives a correct diagnosis.
A later study turns one percentage into several evaluations
A 2021 peer-reviewed paper offers a much fuller evaluation of a related colon-polyp localization system. It does not establish that its model was the identical checkpoint shown in 2019, so its figures should not be treated as a retrospective validation of the video. The team trained and tested the system using more than 71,000 images from 20 centers, then separately evaluated it on 47 complete unaltered colonoscopy videos and in a prospective single-center study at Shanghai Changhai Hospital. Four authors listed Tencent Healthcare affiliations, and the paper reports Tencent funding, prototype-software, computer, and GPU support; this was a disclosed Tencent-supported study, not an independent replication.[4]
Performance changed with the evaluation layer. On the still-image test set, the paper reported 95.0% sensitivity and 99.1% specificity. On full videos, frame-level sensitivity was 92.2% and specificity 93.6%. In prospective use, CADe sensitivity was 98.4%—185 of 188 polyps detected—with an average of 2.2 false-positive prompts per withdrawal.[4]
The prospective comparison was a single-arm, self-controlled observational design, not a randomized assisted-versus-unassisted trial. The AI monitor sat behind the colonoscopist, while an observer recorded whether the clinician or CADe detected each polyp first. In that constructed comparison, the combined colonoscopist-plus-CADe tally was 0.90 polyps per colonoscopy versus 0.82 for colonoscopists’ detections alone; adenomas were 0.32 versus 0.30. The corresponding changes in the share of procedures with at least one polyp or adenoma did not reach statistical significance: 45.9% to 48.3% for polyps and 22.0% to 23.9% for adenomas.[4]
Those numbers answer different questions. The image test asks whether selected frames are recognized. The full-video test introduces motion, fluid, folds, glare, and long stretches without a lesion. The prospective study asks what happens when model and clinician observations are combined during a procedure. Folds, light reflections, fecal fluid or bubbles, and normal anatomical structures were documented sources of false positives. In subgroup analyses, the increase in polyp or adenoma counts was not statistically significant when bowel preparation was inadequate or withdrawal lasted under six minutes; that association does not by itself establish causation. The authors called for multicenter randomized trials.[4]
This layered reporting is far less memorable than one accuracy number, but it is more informative. It identifies failure modes, connects model behavior to workflow, and distinguishes technical detection from an examination-level outcome. It also shows why the CGTN clip’s underspecified label matters: “96.93% correct under what unit and protocol?” is not pedantry. It determines whether the number describes a frame, a marked region, a lesion, an examination, or a patient.
Regulation restores the verb: mark, then assess
In June 2023, China’s National Medical Products Administration announced approval of Tencent Healthcare’s colon-polyp electronic-endoscope detection software. The regulator’s English notice describes a bounded function: the software displays and marks suspected polyp regions in colonoscopy images, and doctors make the assessment using those images and the patient’s condition.[5]
Four years after the video, that language restores the human and the verb. The product marks; the doctor assesses. Regulatory clearance does not prove that every use improves patient outcomes, and the 2023 software should not be assumed to be unchanged from the 2019 demo. But the approved indication supplies what the closing caption lacked: an intended user, an input, an action, and a decision boundary.
The contrast is instructive. “Diagnosis accuracy” lets one number stand for an entire care process. “Marks suspected regions for a doctor” describes a tool. The second formulation is less cinematic and much easier to evaluate.
How to watch an accuracy claim
The 2025 STARD-AI reporting guideline was written for diagnostic-accuracy studies involving AI. It asks authors to make the intended use, input data, participant and dataset selection, reference standard, analysis, workflow integration, intended end user and required expertise, and limitations visible.[6] Applied to a short product video, its logic becomes a six-question pause button:
- What exact task does the model perform?
- Is the unit an image, frame, lesion, examination, or patient?
- Who or what supplies the reference answer?
- Was the test separated by patient, hospital, device, time, and software version?
- Is the comparison model-only, doctor-only, or doctor with assistance?
- Does the endpoint measure a box on a screen or an improvement in care?
The clip and companion article answer only fragments of that list. They show the intended interaction, identify the colorectal application, and name hospital pilots; they do not disclose the evaluation protocol behind the final figure.[1][2] Tencent’s 2018 release supplies the localization label but still leaves important design details unreported.[3] The 2021 paper supplies the most detailed peer-reviewed evidence here, while also exposing the design caveats and failure modes that a one-minute demonstration naturally edits out.[4]
What the recording preserves
CGTN’s report remains valuable because it captures Tencent Miying during a pilot and pre-market period, four years before the 2023 authorization described here. The preliminary result and doctor-feedback loop in CGTN, together with the side-by-side display and real-time prompt documented in Deqing, show a company trying to fit computer vision into an existing procedure. These are product-design facts that a benchmark table would miss.[1][2][5][7]
The video package also preserves a subtler act of evidence compression. The colorectal task remains visible in the imagery and companion text, but the caption’s generic “diagnosis accuracy” blurs localization and diagnosis. The identical decimal appears in Tencent’s earlier localization release, yet neither CGTN source says that test produced the caption, and none of the three provides the protocol needed to interpret it. The precision survives; the evaluation boundary does not.[1][2][3]
Restoring that boundary changes the story. Miying was not a machine that had become 96.93% correct at medicine. It was a family of imaging tools whose public materials showed a colon-polyp aid and an unusually precise claim without a reproducible unit or protocol. Later research on a related system separated image, video, and clinical-workflow evaluations instead of asking one percentage to represent them all.[4] The durable lesson is not to distrust every demo. It is to keep the endpoint, unit, and version attached to the number.
Sources
- CGTN, “Future of health care: Screening high-risk cancers with AI,” official video, May 23, 2019.
- Pan Zhaoyi, CGTN, “How has cutting-edge tech changed China’s healthcare sector?”, published May 22 and updated May 23, 2019.
- Guo Yuhui, Tencent Internet+, “Tencent Miying releases China’s first real-time AI colorectal-tumor screening system” (original Chinese title: “腾讯觅影发布国内首个结直肠肿瘤实时筛查AI系统”), hosted by Tencent Healthcare, July 8, 2018.
- Sheng-Bing Zhao et al., “Establishment and validation of a computer-assisted colonic polyp localization system based on deep learning,” World Journal of Gastroenterology 27, no. 31, 2021.
- CCFDIE, “Computer-aided Detection Software for Electronic Endoscope for Colon Polyps Approved for Marketing,” National Medical Products Administration, June 1, 2023.
- Viknesh Sounderajah et al., “The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence,” Nature Medicine 31, 2025.
- Fu Mengwen, TMTPost via Phoenix Technology, “探访德清县医共体:医疗AI如何下基层?” (“Inside Deqing’s county medical community: How does medical AI reach primary care?”), March 21, 2019; source of the lead photograph.