ai china

Galbot put the shelf before the humanoid

9 sources 5 primary sources July 25, 2026

Text
A Galbot G1 robot uses a gripper to pick up a grilled sausage inside a FamilyMart convenience store in Beijing.

Galbot G1 works at a FamilyMart in Beijing's Zhongguancun district on June 10, 2026. The documentary photograph shows the company’s real operating thesis: a wheeled, dual-arm robot performing a bounded retail handoff rather than imitating a person on an open street.[7][9]

On June 10, 2026, a customer inside a FamilyMart in Beijing’s Zhongguancun district asked a robot for a grilled sausage. The machine turned toward the warmer, closed a gripper around the food, placed it in a tray, and directed the customer to the register. Xinhua’s report is a small piece of field evidence, not proof of general-purpose intelligence. It is useful precisely because the task is so ordinary.[7][9]

The robot was Galbot G1, built by Beijing-based Galbot. Its public story can look like a familiar humanoid pitch: a young, heavily funded company, an imposing machine, a stack of vision-language-action models, and expansion claims across retail, manufacturing, and healthcare. A closer reading reveals a more disciplined strategy. Galbot has organized its body, research program, and early deployments around one repeatable loop: travel through a controlled indoor space, locate an object among many packages, grasp it, and complete a handoff.

As of July 25, 2026, that loop is the strongest reason to take the company seriously. It is also the right boundary for judging it. Galbot has assembled more public evidence for a robot that can work a shelf than for a humanoid that can work anywhere.

The body is already a job description

G1 is humanoid above the waist and deliberately unlike a person below it. Galbot’s developer manual describes a four-wheel omnidirectional base, two seven-degree-of-freedom arms, a lifting torso, cameras, force sensors, lidar, and ultrasonic sensors. The machine is 1.73 meters tall at maximum extension, weighs 92.5 kilograms, carries 5 kilograms at each wrist, reaches vertically from floor level to 2.1 meters, and travels at up to 1.5 meters per second.[2]

Those specifications read less like an attempt to reproduce the human body than a plan for reaching most of a commercial shelf. Wheels surrender stairs, curbs, and rough ground in exchange for a broad stable base, predictable motion, and less energy spent balancing. The lifting torso substitutes a vertical rail for knees. Dual arms preserve the useful human part of the arrangement: one arm can stabilize, hold a tray, or clear space while the other manipulates an item.

The operating limits make the trade explicit. Galbot’s manual specifies indoor use, warns against uneven or wet floors and unstable Wi-Fi, asks people to remain 1.5 meters away during operation, and lists an eight-hour working duration measured under laboratory conditions at 60% average speed. Wired charging takes 3.5 hours.[2] An “eight-hour robot” headline would therefore overstate the evidence. What the manual actually establishes is a machine designed for engineered interiors, with safety space and network conditions that an operator must maintain.

That is not a weakness hidden in the fine print. It is the commercial premise. A convenience store, front-end warehouse, or pharmacy can be measured and rearranged. Galbot is choosing to narrow the environment before trying to widen the robot.

The research stack begins with the hand

The same narrowing appears in Galbot’s most legible research artifact. GraspVLA, published by researchers from Galbot, Peking University, the University of Hong Kong, and the Beijing Academy of Artificial Intelligence in May 2025, is not a universal robot policy. It is a foundation model for grasping.[3]

Its data strategy is the important part. The team generated SynGrasp-1B, a billion-frame synthetic action dataset with photorealistic rendering and extensive variation, then coupled that action training with Internet image-and-language data. The model joins a vision-language system to an action expert and uses progressive action generation to produce motion. In the authors’ real-world test sets, GraspVLA achieved roughly 90% success across the reported categories; the paper also shows adaptation to selected new tasks with much less annotation than full action demonstrations required.[3]

Those results have a firm evaluation boundary. The experiments test grasping setups chosen by the authors, not continuous store operation, complete customer orders, or unattended shifts. The paper itself notes that generalization to web-derived categories improves more slowly as synthetic training data scales.[3] The careful conclusion is that Galbot has a plausible way to reduce the cost of teaching a robot about long-tail objects. It has not shown that synthetic pretraining removes the need for field data, exception handling, or human recovery.

Galbot’s commercial layer turns that research direction toward crowded shelves. The World Robot Conference’s official exhibitor profile describes GroceryVLA as a retail-specific end-to-end model built on the company’s grasping work, alongside TrackVLA for navigation. It says the stack is already used in commercial and industrial deployments.[5] This is useful first-party product evidence, but it does not provide the experimental detail of the GraspVLA paper. GroceryVLA should be treated as a deployed system claim whose reliability still needs independent measurement.

Navigation widens the laboratory, not yet the product proof

Galbot’s navigation research is broader than its retail footprint. NavFoM, developed with Peking University, BAAI, the University of Science and Technology of China, Zhejiang University, and other collaborators, was published in September 2025 and later as an ICLR 2026 paper. The model was trained on eight million navigation samples spanning wheeled robots, quadrupeds, drones, and vehicles, with tasks including language-directed navigation, object search, target tracking, and autonomous driving.[4]

That cross-embodiment design matters strategically. A research program that can reuse representations across bodies may be more valuable than one model tied to one chassis. It suggests Galbot wants navigation to become shared intellectual infrastructure while grasping and post-training become specific to a job.

But NavFoM should not be used to smuggle a general-autonomy claim into the dossier. Its public evaluations are research evaluations. G1’s own manual lists a sensor-rich product system, including lidar, multiple cameras, inertial sensors, and ultrasonic sensors.[2] The available sources do not show exactly how NavFoM, TrackVLA, conventional planning, maps, and safety controls divide responsibility in every deployed store. The research direction is credible; the production architecture remains only partly visible.

The store closes the loop

Retail does more for Galbot than supply a photogenic demonstration. It creates repeated, structured interaction with a difficult object distribution: reflective bottles, deformable bags, boxes packed tightly together, transparent containers, hot-food equipment, trays, and customers who interrupt the scene. The floor plan can stay bounded while the object catalog keeps changing. That is a useful compromise between laboratory order and open-world disorder.

The deployment record is becoming substantial enough to examine. A Beijing municipal portal reported in September 2025 that Galbot robots had worked for more than 150 days in Haidian shops and completed tens of thousands of delivery orders, with the system expanding to more than ten locations in the city. The report describes pharmacy work that includes inventory, replenishment, picking, delivery, and packing.[6] In June 2026, Xinhua reported a different scale: more than 150 “Galaxy Space Capsule” retail stores across over 20 cities, alongside the G1 working inside the Zhongguancun FamilyMart.[7][9]

These figures should not be blended into one audited fleet metric. The municipal account and Xinhua report describe different formats and dates; both rely at least partly on company or official statements. Neither discloses intervention rate, successful orders per operating hour, mean time between failures, site retention, or labor required behind the scenes. A store count proves distribution. It does not by itself prove autonomous uptime.

Still, there is a meaningful sequence in the public record. Galbot was established in May 2023 and says it now operates research centers in Beijing, Shenzhen, Suzhou, and Hong Kong, with joint laboratories or research centers involving Peking University and medical partners.[1] By June 2025, Beijing’s science and technology commission reported that the company had raised a new RMB 1.1 billion round, had deployed its retail solution in nearly ten Beijing stores, and had formed a joint venture with Bosch-linked Boyuan Capital to explore industrial commercialization.[8] By mid-2026, state news photographs showed the product inside a global convenience-store chain, performing a customer-facing task.[7][9]

The strongest interpretation is not that Galbot has solved humanoid robotics. It is that the company has connected four things unusually early: a hardware body optimized for indoor reach, a synthetic-data grasping program, a broader navigation research lane, and a retail format capable of generating operational repetitions. Each layer gives the next one something it needs.

What the dossier still cannot answer

The most important missing number is the human-intervention rate. A robot can complete thousands of orders while still requiring frequent resets, remote assistance, shelf preparation, or technician visits. None of the public sources used here provides a common denominator that separates autonomous task completion from assisted recovery.

The second missing comparison is economic. Galbot has not publicly shown, in a form available in these sources, whether a G1 store beats a vending system, automated storage-and-retrieval cabinet, fixed arm, or human-plus-software workflow on total cost per completed order. A mobile dual-arm robot is more flexible than those alternatives. Flexibility only becomes commercial value if it exceeds the maintenance, safety-space, charging, and supervision costs of a more constrained machine.

The third gap is transfer. GraspVLA and NavFoM make a case for reusable learned components, while Galbot’s company materials name industrial and healthcare ambitions.[1][3][4] The decisive evidence would be the same base policies moving between retailers, pharmacies, warehouses, and factories with limited site-specific retraining—and with reliability disclosed under comparable conditions. Without that evidence, “general-purpose” remains a direction rather than a measured property.

These gaps do not erase the retail record. They define the next proof. Galbot’s thesis will strengthen if deployments publish autonomous completion rates, recovery categories, maintenance hours, and repeat-site economics. It will weaken if apparent generality resolves into heavy store-specific engineering or invisible human operation.

For now, the company’s most interesting decision is anatomical and organizational at once. Galbot did not wait for a robot that could walk anywhere. It put wheels beneath two arms, trained the hands on synthetic abundance, and went looking for shelves where repetition could become data. The humble handoff at a convenience-store counter is not the end of the humanoid story. It is the company’s chosen place to begin proving one.

Sources

  1. Galbot, “About Us”—official company history, research-center footprint, joint research relationships, and stated application areas.
  2. Galbot Developer Platform, “GALBOT G1”—official product manual covering dimensions, payload, sensors, omnidirectional base, battery test conditions, charging, and operating limits.
  3. Shengliang Deng et al., “GraspVLA: a Grasping Foundation Model Pre-trained on Billion-scale Synthetic Action Data,” arXiv:2505.03233 (May 2025)—training data, architecture, evaluation setup, scaling behavior, and adaptation results.
  4. Jiazhao Zhang et al., “Embodied Navigation Foundation Model,” ICLR 2026—NavFoM’s eight-million-sample, cross-embodiment and cross-task navigation design.
  5. World Robot Conference, “Beijing Galbot Co., Ltd.”—official exhibitor profile describing G1, GraspVLA, GroceryVLA, TrackVLA, and claimed commercial deployment.
  6. Open Beijing, Haidian District, “Galbot humanoid robot has efficiently been ‘on duty’ in Haidian shops for more than 150 days” (September 2025)—reported pharmacy workflow, order volume, and Beijing deployment footprint.
  7. Xinhua, “Out of the laboratory: how robots get ‘on the job’” (June 12, 2026)—Zhongguancun convenience-store field report, company deployment count, and Guo Xing’s documentary photograph used as the article image.
  8. Beijing Municipal Science & Technology Commission and Zhongguancun Administrative Commission, “Galbot completes a new RMB 1.1 billion financing round” (June 24, 2025)—founding timeline, financing, retail deployments, and industrial joint venture.
  9. Beijing Daily, “‘Spring Festival Gala model’ Galbot robot officially starts work at this convenience store” (April 17, 2026)—identification of the Zhongguancun Dinghao site as a FamilyMart and description of its regular in-store role.
Previous Triton-distributed is turning communication into compiler territory Next China moves its AI policy to the checkout counter

Recommended In ai china

Matched by subject and format