Dialogue with Tencent Chief Scientist Zhang Zhengyou: AI Remains a "Brain in a Vat"; Tencent Tairos Aims to Solve Embodied AI Deployment Challenges

Wallstreetcn
2026.07.19 23:56

Embodied AI is still in its early stages

Author | Huang Yu

Over the past two years, embodied AI has gained significant momentum, with "accelerated commercialization" becoming a major hallmark of industry development. At the 2026 World Artificial Intelligence Conference (WAIC), robots showcased an impressive array of capabilities.

However, beneath this fervor, the actual capabilities of embodied AI still fall short of external expectations, and there are issues of homogenization in deployment scenarios and physical forms.

During a media exchange session at WAIC, Zhang Zhengyou, Chief Scientist at Tencent, Director of Tencent Robotics X Lab, and Director of Futian Lab, told media outlets including Wall Street News that embodied AI is still in its infancy. With few genuine real-world deployments, companies are searching for scenarios where embodied AI can truly make an impact, which often leads them to similar conclusions.

Nevertheless, Zhang believes this is not necessarily a bad thing and urges confidence in deployment. "When any industry starts, many companies participate, but survival of the fittest will eventually occur. Therefore, I do not think this will hinder the industry's development."

Regarding deployment scenarios for robots, Zhang repeatedly expressed optimism about elderly care, although he also acknowledged it as the most challenging sector.

As for hyper-realistic humanoid robots, which have recently attracted considerable attention, Zhang told Wall Street News that Tencent has not entered this specific field. In his view, emotional value does not necessarily need to be provided through hyper-realism; one must think from first principles.

Tencent Robotics X Lab was established as early as 2018, dedicated to researching and applying cutting-edge robotics technologies.

Although the current wave of embodied AI development is hot, Tencent consistently emphasizes its boundaries of participation. In early 2025, Ma Huateng, Chairman and CEO of Tencent, stated that Tencent aims to be a partner for all robot manufacturers rather than replacing them by manufacturing hardware, aligning with Tencent's overall strategic goals.

Against this backdrop, during the 2025 WAIC, Tencent officially launched Tairos, its open platform for embodied AI. It is the first domestic software platform for embodied AI to offer large models, development tools, and data services in a modular manner, opening up to the robotics industry via a plug-and-play approach.

It is reported that over the past year, Tencent Robotics X Lab has collaborated with robotics companies such as Unitree, Agibot, Dobot, Songyan Dynamics, Leju, and Huayan, as well as scenario partners including Bosch, Lens Technology, Jingdezhen, and Dunhuang, to explore the deployment of embodied AI across different robots and industrial scenarios.

During the 2026 WAIC, Tencent introduced new models and intelligent agent achievements in its embodied AI series, systematically closing the loop of "perception–body–action" for the first time.

Zhang explained that this product line follows a clear thread: first enabling intelligence to understand the physical world, then allowing it to imagine target states, and finally translating ideas into stable actions—integrating these capabilities into continuously online, reusable intelligent agents.

The embodied foundation models released by Tencent this time include Hy-Embodied-VLM-1.0, Hy-Embodied-RxBrain-1.0, and Hy-Embodied-VLA-0.5, all built upon the Hunyuan large model.

Specifically, the VLM model acts like the "right brain," responsible for understanding images, space, and scenes; the RxBrain model serves as the embodied "brain," unifying cognition, planning, and imagination of future states; and the VLA model connects the "cerebellum" to the body, converting high-level goals into continuous, correctable actions.

According to reports, the core competitiveness of Hy-Embodied-VLA-0.5 lies not in training a larger model, but in constructing a complete learning stack where "data–model–training–deployment" work synergistically, accumulating over 10,000 hours of human demonstration data through a sub-millimeter high-precision UMI acquisition system.

Zhang pointed out that true intelligence requires connecting language, vision, spatial cognition, body control, and environmental feedback, validated within the closed loop of "perception–body–action."

This essentially reflects the basic approach of Tencent Robotics X in exploring native embodied intelligence: moving from disembodied to embodied, and from fragmented modules to an integrated closed loop.

In Zhang's view, today's most sophisticated artificial intelligence is essentially still a "brain in a vat." Over the past few years, the progress of large models has been astonishing. However, once they enter the real world, problems arise: they may not truly understand object positions, spatial relationships, and affordances; nor can they easily act, receive feedback, and make timely adjustments in a constantly changing environment.

"It is like a brain kept in a vat—rich in knowledge but lacking a body and the closed loop for direct interaction with the world. We call this state 'disembodied intelligence.'"

The scarcity of embodied AI data is a key factor hindering embodied AI from becoming smarter.

Zhang pointed out that in the digital pyramid of embodied AI, the bottom layer consists of internet data, the next layer comprises first-person operational videos, the second layer includes data collected using sensor-equipped gloves, and the top layer is teleoperation data. Although teleoperation data has low collection efficiency, it is considered the most effective data for training robots.

Furthermore, Zhang stated that data required for robot training includes internet videos, first-person operational videos, data collected with sensor gloves, and teleoperation data. To make robots smarter through training, these four types of data need to be integrated.

Unlike the large language model technology stack, which has converged into a widely accepted scope, Zhang noted that there is no consensus yet on technical solutions for embodied AI models. Last year, the VLA (Vision-Language-Action) paradigm gained traction in the industry, while this year discussions focus on world models. Hy-Embodied-RxBrain-1.0 represents Tencent's exploration in world understanding models.