Dialogue with Aiou Intelligent Co-founder Ding Zhezhang: "Human Data" Becomes the Core Fuel for Embodied Pre-training

Dialogue with Aiou Intelligent Co-founder Ding Zhezhang: "Human Data" Becomes the Core Fuel for Embodied Pre-training

```

Author | Huang Yu

Although embodied intelligence has experienced rapid growth in recent years, whenever the issue of data supply for embodied intelligence is discussed, the industry is often given a cold shower.

As a data service provider for embodied intelligence, Ding Zhezhang, co-founder of i-io Intelligence, recently admitted during a media exchange with Wallstreetcn and others that when the company first entered the industry in 2023, they found that all types of embodied intelligence data were in extremely short supply. In recent years, there has been some progress in basic capabilities, but there is still a scarcity of data for training proprietary tasks.

It is precisely due to the scarcity of high-quality embodied intelligence data that humanoid robots still appear clumsy when completing some actions that seem very simple.

To solve the bottleneck of data supply for embodied intelligence, a new paradigm has emerged in the industry.

Ding Zhezhang said that in 2025, the "concentration" of real robot data would be very high. The construction of domestic data collection sites, key research papers, and the emphasis of all companies have all revolved around real robot data collection. However, since the end of 2025, another paradigm has been discovered that is easier to scale and more efficient in data collection—human-centered (human data) collection.

In this context, Ding Zhezhang observed that before 2025, the majority of data used by some companies to train large embodied models consisted mainly of real robot data, or "10-20% real robot data + 80-90% simulated synthetic data." By 2026, the concentration and acceptance of human data will be significantly improved, while the voices around synthetic simulation will diminish to some extent.

The move from simulated synthetic data to human data may seem like a structural adjustment of data sources, but behind it lies a new phase for embodied intelligence training systems, the data industry chain, and even commercialization paths.

A competition around "training fuel" has quietly begun.

Human Data Steps into the Spotlight

If embodied intelligence models are likened to "robot students" in the process of maturing, data is their core textbook for understanding the real world, and the complete training process consists of three core steps: pre-training, post-training fine-tuning, and operational deployment.

Unlike large language models, which mainly rely on internet text data, embodied intelligence needs to learn the "logic of action" in the real world. This type of data source is extremely scarce, and the greatest bottleneck in the industry now is how to collect high-quality training data for embodied intelligence at scale.

Ding Zhezhang explained the progress of data acquisition in recent years with a simple example: If you want to train a simple "help me get a cup" task, data is easy to find and training is fast; but if you want it to fetch a pair of shoes from the bottom shelf, the data is very scarce, and training is difficult.

Meanwhile, Ding Zhezhang revealed that among the data requests his company receives, there are two most sought-after categories. One is data on dexterous hand's long-range, fine operations. Previously, dexterous hands were not mature or easy to use, so it was harder to collect high-quality task data with them. But with recent hardware iterations, the demand for this type of data has increased further.

The second is data on whole-body movement of robots. Ding Zhezhang points out that starting more than half a year ago, some basic models for whole-body remote operation and whole-body control emerged in the industry, enabling robots to initially take in whole-body input and follow accordingly. To take the next step, more data on robot whole-body control is needed.

Currently, the data used to train embodied intelligence models in the industry mainly falls into three categories: Robot Data, Human Data, and Synthetic Data.

Among them, robot data comes from the operation of robots in real environments. It can fully record the robot's joint movements, sensor feedback, and interaction processes with the physical world. It is considered the most valuable and closest data to real deployment scenarios.

Human data is collected through cameras, smart glasses, and other devices from a human-first perspective to record human operation processes, helping robots learn how humans complete tasks. Synthetic data, meanwhile, is rapidly generated in large quantities through simulated environments or generative models, aimed at resolving the difficulty and insufficiency of robot data collection.

Ding Zhezhang pointed out that the reason the concentration of real robot data is high in 2025 is because, last year, everyone believed that robot data was relatively easy to scale—robot manufacturers could output robots in batches, and the data could be iteratively generated in batches.

However, robot data is expensive and inefficient to acquire, and is limited by the number of robots. For foundational models in embodied intelligence that require hundreds of thousands or even millions of hours of training data, relying solely on robots for collection is almost impossible to meet the demand.

Ding Zhezhang observed that over the past six months, the data demands of embodied intelligence model manufacturers have shifted more toward human data, using it as the core "fuel" for pre-training. This method makes data collection easier, scales larger, and naturally fits the needs of high computing power, cloud storage, and large pre-trained models.

At the same time, Ding Zhezhang pointed out that the influence of synthetic data is also diminishing. Compared to synthetic data, human data is richer in scenarios, physical attributes, and generalizable tasks. The biggest challenge for synthetic data is reconstructing physical laws and broadening task domains in simulated environments. Now, with device-based, non-robotic human-centric collection, simply equipping people with a device and placing them in any scenario can collect scene data.

"So, some of the synthetic data previously used just for sheer volume has been replaced with non-robotic, human data."

Of course, every embodied intelligence company has its own strategy for choosing data types.

Ding Zhezhang pointed out that there are still some companies mainly using synthetic data. Some use a hybrid approach: hybrid of "synthetic + human data" for pre-training, then real robot data plus some synthetic edge/corner cases for post-training.

Ding Zhezhang said the biggest advantage of synthetic data is that training, iteration, and testing in simulation does not damage the robot itself, so for situations involving cerebellar control or dangerous environments, simulation is still used to accelerate training.

"But overall, the trend is: One, human-centric data's concentration and acceptance are significantly rising; Two, real robot data is forever indispensable—whatever the path, real robot data is needed to ensure effect in physical-world deployment," said Ding Zhezhang.

However, Ding Zhezhang also emphasized that although human data is easier and faster to collect, the process of "transformation to the robot" on the backend is more crucial. The overall cost from collection to training is actually about the same. So, if you only look at collection volume, human data is much more accessible, but taking into account post-processing and transformation, the cost catches up.

The Flywheel Hasn't Started but Industrial Opportunities Are Emerging

Although the data acquisition paradigm for embodied intelligence keeps evolving, the entire industry remains at a very early stage—the data flywheel has yet to start, and the scarcity of high-quality data will persist for a long time.

Against this backdrop, although the robot market has heated up over the past year, Ding Zhezhang believes there is still a gap between expectations for what embodied robots can do and their actual capabilities.

"This year, many more people approached us to talk about robot deployment, but have there really been a lot more large-scale deployments? I have my doubts."

In his view, the biggest bottleneck for improving current robot capability is the lack of continuous data iteration in real scenarios.

Therefore, i-io Intelligence is more optimistic about a gradual commercialization path: robots enter the field with partial autonomy, humans complete the rest of the tasks by remote operation, and as robots keep accumulating real-world operational data, their autonomous capabilities are gradually improved.

This approach is quite similar to the development path for autonomous driving.

True rapid iteration in autonomous driving was not achieved by having a lot of data all at once, but rather by mass-produced vehicles continuously transmitting real road data, forming a "collection—training—deployment—recollection" data flywheel.

Embodied intelligence, for now, is still before the data flywheel phase.

Ding Zhezhang expects that in the next two to three years, as more and more robots enter real-world scenarios and accomplish tasks via "partial autonomy + partial manual teleoperation," real-world data will continue to flow back, and the industry will gradually build its own data flywheel.

Although data scarcity means embodied intelligence is still far from widespread adoption, a flood of opportunities is emerging in the industry value chain.

According to Wallstreetcn, as demand for human data collection continues and grows, many companies that previously worked in traditional or large model data services are suddenly pushing hard into embodied intelligence data services. This is because the threshold for human data collection is much lower than for real robot data.

Huang Yang, Deputy Director of Heterogeneous Computing at Tencent Cloud, also revealed that by 2026, Tencent Cloud's embodied intelligence-related computing power consumption will increase by 4 to 5 times compared with 2025, most of it due to data cleaning. "The computational needs for model training itself haven't changed much; the recent spike has all happened at the data collection stage."

Huang Yang said the demand for data processing in embodied intelligence is huge. When choosing a cloud partner, customers will focus on the capability and cost-effectiveness of sustainable future supply. Tencent Cloud can provide a variety of chip specifications and levels, and also offer software to help customers improve computational efficiency. Overall, the plan for i-io Intelligence can lower costs by more than 30%.

Additionally, data service providers such as i-io Intelligence are becoming more closely integrated with cloud vendors.

It is reported that i-io Intelligence's collaboration with Tencent Cloud has already reached the foundational level of the data platform—Tencent Cloud provides fundamental storage and computing power, while i-io Intelligence manages the frontend entry and management. Together, they support the full chain from data collection, cleaning, annotation, to training scheduling. The i-io Intelligence data platform tool has also been integrated into the Tencent Cloud application portal.

In scenarios where i-io Intelligence conducts cross-border remote teleoperation, Tencent Cloud uses its TRRO solution to provide stable and secure transmission of control and image streams, offering foundational support even under weak network conditions, making it possible to have "a human in Shenzhen and a robot in the US" for international teleoperation.

A cloud vendor insider pointed out that the more the industry turns to "human-centered" data, the more dependent it will be on underlying infrastructure.

In other words, when the industry uses human data as the fuel for pre-training, cloud storage, computing, and transmission stacks become the fundamental conditions enabling this data to truly "run."

The embodied intelligence race is no longer just for end-product robot manufacturers. Data service providers, cloud platforms, and computing power vendors have all joined the battle. This new round of competition around data will also become a key variable determining the speed of commercialization and industry landscape of embodied intelligence.

 

Risk Warning and DisclaimerThe market involves risks, and investment requires caution. This article does not provide individual investment advice, nor does it take into account any particular user’s specific investment objectives, financial situation or needs. Users should consider whether any opinions, viewpoints or conclusions herein fit their particular circumstances. Investment based on this article is at your own risk. ```