What happened

The 2026 World Robot Conference in Beijing Yizhuang gathered more than 300 enterprises and 3,000 exhibits, but the hottest topic was not robots' physical dexterity—it was what kind of 'brain' should control them. Practitioners at the event argued over whether VLA models, world models, or some unified architecture will prove to be the ultimate form of embodied intelligence.

The debate atmosphere cooled compared with recent years: VLA branding was less visible, and 'end-to-end' is no longer treated as a universal answer. Instead, 'world model' emerged as the new buzzword, though companies defined it in varied ways—from generating future states to predicting causal chains to reasoning in latent space. Tsinghua researcher Su Hang argued that VLA and world models are not mutually exclusive; Xinghaitu's CEO said the two share the same Transformer-based core and will eventually converge. Several firms, including Qinglang Intelligent, Zhicheng AI, Wudong Power and Unitree, showcased different world-model implementations.

Other players pushed further: the Beijing Humanoid Robot Innovation Center unveiled Pelican-Unify, a unified-representation embodied world model that merges visual understanding, action control and world prediction; Zhipingfang open-sourced a brain-like model inspired by the cortex-cerebellum-spinal cord; and Shenyi Robot proposed a self-evolving 'hybrid physical expert model.' The clearest takeaway from the event, according to observers, was that no single route has won—but the direction is sharpening, with data becoming the industry's central concern.

Why it matters

The shift in conversation from robot limbs to robot brains signals that embodied intelligence is entering a software-and-data-driven phase. If hardware has been pushed to its limits, the next competitive frontier lies in how robots perceive, reason and act in the physical world—and that will be decided by algorithmic architecture plus real-world data.

The fragmentation of world-model definitions and the persistence of multiple routes mean huge uncertainty for investors and developers. Companies are betting on different technical bets: some fuse VLA with world models, some pursue unified representations, and some look to brain-derived structures. The stakes are high because the winning architecture could become the standard platform for humanoid robots in homes and factories.

Although many executives expressed grand visions—from liberating billions of people from repetitive labor to achieving a unified AI—the industry is still lacking shared definitions and evaluation benchmarks. The real competition, as several speakers stressed, is about who can efficiently collect and use data to close the loop from perception to action in authentic environments.

Key facts

At WRC 2026, more than 300 enterprises and over 3,000 exhibits were showcased in Beijing Yizhuang.

Qinglang Intelligent released KOM 3.0, described as the world's first service-industry VLA architecture integrating a latent-space world model.

Zhicheng AI is pursuing a JEPA-based physical world model for 'mental simulation' before task execution.

The Beijing Humanoid Robot Innovation Center's Pelican-Unify topped the WorldArena global evaluation list in May 2026.

Zhipingfang open-sourced NeuroVLA, a brain-like architecture modeled on the cortex, cerebellum, and spinal cord.

Industry consensus is heading toward data-centered practical challenges rather than ideological route debates.

What to watch next

Whether world-model evaluation standards will be unified—without them, comparisons between different systems remain ambiguous and slow the industry's progress.

Which companies can prove their architecture works in real homes and factories, not just shows and arenas, by running data loops smoothly in daily operation.

The potential emergence of hybrid designs: many industry insiders predict the final robot brain may combine VLA-style action generation with world-model-style prediction and correction, rather than picking one exclusive model.

Sources