The most revealing person in this roughly 80-second video is not synthetic. He is Qiu Hao, the Xinhua presenter whose appearance and voice became source material for one of the news agency's first “AI anchors.” By putting the human model and his screen likeness in the same short report, the clip makes an important dependency visible: the product begins with a person before it can imitate a presenter.[1][3]
That observation prevents a larger category error. A synthetic anchor can turn an approved text into a repeatable audiovisual performance. It does not thereby find the story, interview a witness, authenticate an image, reconcile conflicting accounts, decide that a claim is ready to publish, or accept responsibility for a correction. Those are reporting and editorial functions. The anchor is a delivery layer.
Watch the clip twice. On the first pass, compare the human and synthetic performances: face, voice, gaze, pace and the authority implied by the familiar studio form. On the second, ignore the novelty and ask what must have happened off-screen before each rendered sentence could appear. The video is short; the production boundary it exposes is not.[1]
First pass — find the human inside the product
The clip's framing supplies a better starting point than the phrase “AI anchor.” It identifies Qiu and explains that Xinhua's on-screen presenters were modelled on flesh-and-blood anchors. This is not a disembodied system inventing a broadcast identity. A real worker supplies the vocal and visual reference; a newsroom supplies the institutional setting; software supplies the reproducible performance.[1]
At about 0:18, the edit places Qiu in the foreground with his synthetic counterpart visible on a monitor behind him. Around 0:38, the interview framing returns attention to the human presenter. Near 0:57, the avatar fills the frame beside a 365 × 24 graphic. In less than a minute, the sequence moves from source person to copy to the promise of continuous availability. It visualizes the product's two inputs—borrowed identity and supplied text—before making the scale claim.[1]
The distinction matters because likeness carries accumulated cues of credibility. A suit, desk, steady gaze and measured delivery are not evidence that the entity on screen investigated the words it is reading. They are part of a visual grammar audiences already associate with news. The synthetic version can reproduce that grammar at scale while remaining causally distant from the reporting behind the script.
Xinhua's own archival photograph makes the deployment context concrete. It shows visitors beside a large display running the anchor at the fifth World Internet Conference in November 2018, where Xinhua and Sogou presented the system. Xinhua described the anchor as a digital double with the reporting capability of a human presenter, but the image itself documents a public product launch, not an autonomous reporting process.[2] Read the institutional claim and the photographic evidence separately.
Now replay the human-versus-avatar comparison with five questions in mind:
- Whose face and voice are being reproduced?
- Who selected and verified the facts in the script?
- Who approved the final wording and pronunciation?
- Is the synthetic nature of the presenter obvious throughout the clip?
- Who corrects the record if the delivered statement is wrong?
Only the first question can be answered from appearance. The other four require provenance outside the frame.
Second pass — follow the text, not the face
A Chinese technical account published days after the launch interviewed Sogou managers about the production chain. Their description begins with entered news text. The system predicts features such as prosody and emphasis, synthesizes speech and lip movement, generates matching facial expressions, and coordinates the streams on a timeline. The account also says the 2018 system was strongest in relatively serious presentation settings and had more difficulty with highly expressive emotion.[3]
Every named step is a transformation of presentation. Text becomes timing; timing guides voice; voice, lips and expression become a synchronized video. The pipeline can make delivery faster and more consistent, but no step in that description establishes whether the input sentence is true. A flawlessly synchronized falsehood remains false. A visible lip mismatch may reveal a rendering defect; a plausible delivery can conceal an upstream editorial defect.
This is the central annotation for the clip: realism is an output property, not a reporting method. The better the face and voice become, the easier it is to mistake confidence of delivery for confidence in the underlying evidence. A newsroom therefore needs two separate quality systems. One tests facts, sourcing, context and editorial judgment. The other tests pronunciation, timing, visual artifacts, labelling and playback. Passing the second cannot substitute for passing the first.
The unit being scaled is a render
The commercial logic becomes clearer when the anchor is treated as a rendering service. A human presenter normally has to enter a studio, read the script, record another take when necessary and repeat the process for each language or format. Once a synthetic identity has been built, approved text can generate another video without repeating all of that performance labor. The initial Sogou account emphasized rapid video generation and the possibility of continuous operation given sufficient server capacity.[3]
Xinhua's own follow-up provides a first-party measure of that throughput. In February 2019, three months after launch, the agency said its synthetic anchors had produced more than 3,400 reports totalling over 10,000 minutes. At the same event, Xinhua and Sogou introduced a standing version with gestures and a female anchor, and described changes to waveform-based audio, expression synthesis, lip motion and matching gestures to meaning.[4]
Those numbers show that the format moved beyond a single demonstration. They do not measure factual accuracy, audience trust, editorial independence, labor saved, or performance against a human control group. Nor are they an independent audit; they are Xinhua's report about its own system. The defensible conclusion is narrower: the organization demonstrated repeatable production and distribution at substantial volume.
Scale also moves the operational bottleneck upstream. If rendering becomes cheap, a newsroom can publish more presentational variants than editors could once record. That increases the need for version control over scripts, names and numbers; an approval trail tied to each output; persistent disclosure that the presenter is synthetic; and a correction mechanism capable of finding every derivative. Faster delivery raises the value of editorial controls because a mistake can be reproduced just as efficiently as a correct bulletin.
A face needs a provenance record
The clip is especially useful because it shows Qiu as the origin of the avatar's identity. That relationship should not disappear once viewers meet only the synthetic version. A robust provenance record would identify the human model, record the scope and duration of consent, state which organization controls the likeness, name the system and version used for each render, preserve the approved script, and log the responsible editor.
Disclosure should also survive reuse. A label placed only in an opening frame can vanish when a clip is excerpted, cropped or reposted. A more resilient design combines an on-screen identifier with metadata and a linked record. The objective is not merely to help a careful viewer “spot the AI.” It is to make origin and accountability retrievable even when perceptual detection becomes unreliable.
This governance layer is not an argument against automation. It is what allows automation to remain a tool of a newsroom instead of becoming an alibi for the newsroom. A 2025 declaration adopted by press and media councils in South-East Europe and Türkiye, reported by UNESCO, puts the general principle clearly: AI should support journalism rather than replace human judgment or editorial responsibility, and AI-generated content should be transparent and labelled.[6] The declaration is regional guidance, not Chinese law, but the separation of tool capability from editorial accountability travels well.
Later evidence tests communication, not just resemblance
The 2018 clip invites a visual test: how close does the synthetic presenter look and sound to Qiu? Later research suggests that resemblance is only part of audience experience. A 2026 peer-reviewed qualitative study asked 11 Chinese news consumers to watch hard-news segments delivered by Xinhua's virtual anchor Xin Xiaomeng, drawn from broadcasts between January and June 2025. Participants reported problems involving stress, intonation and rhythm; nine of the 11 expressed concerns about human connection, aesthetic quality or the social role of news presentation.[5]
That study needs a firm boundary around it. Eleven purposively recruited consumers in one language and state-media setting cannot represent all Chinese viewers, much less global audiences. The design used subtitles, and the authors say it could not isolate their compensating effect on comprehension. Its interviews identify perceived problems and plausible mechanisms; they do not prove that a particular vocal feature caused trust to fall.[5]
Within those limits, the finding sharpens the first viewing. A presenter does more than pronounce each word. Stress can identify what matters in a sentence; rhythm can separate clauses; tone can help an audience interpret significance. Synthetic delivery may be technically continuous while communication remains effortful. Eliminating the need for another studio take is not the same objective as preserving the meaning a skilled presenter adds.
The useful boundary
The Xinhua-Sogou launch mattered because it productized a visible part of news production early: the conversion of text into a recognizable presenter-led video. The correct lesson is neither that the avatar was “just a gimmick” nor that it had become a reporter. It was a real production system aimed at a real cost and distribution problem, wrapped in occupational language broader than the mechanism described by its makers.[2][3][4]
On a final replay, replace the question “Can this AI replace the anchor?” with three narrower ones. Which task is being automated? Which evidence demonstrates that task? Which human or institution remains answerable for the result? For this system, the evidenced task is synthetic delivery of supplied text. The reporting chain remains upstream, and accountability remains with the publisher.
That is why the title's distinction is more than semantics. The synthetic anchor can deliver the bulletin. The newsroom still has to report it.
Sources
- South China Morning Post, “Meet the real-life news presenter behind the world's first AI anchor,” video, November 29, 2018.
- Xinhua, “Xinhua domestic photographs of the week: ‘AI synthetic anchor,’” November 9, 2018.
- Zhidx via Phoenix Technology, “With Xinhua releasing an AI virtual anchor, Sogou wants to use this technology to ‘clone’ humans,” November 12, 2018.
- Xinhua, “Media integration advances as Xinhua's AI synthetic anchors receive a major upgrade,” February 19, 2019.
- Jiaxian Li and Xing Pang, “The anomaly of Chinese AI news anchors: a study of speech irregularities and their impact on news communication effectiveness,” Frontiers in Computer Science, May 29, 2026.
- UNESCO, “AI and Media Ethics: Press councils from South-East Europe and Türkiye adopt landmark declaration,” June 11, 2025.