Mifeng Yao Maoqing: Embodied data will inevitably converge towards leading companies.
On August 31, at the MEgo data acquisition factory on Dieqiao Road in Pudong, Shanghai, the 20,000th MEgo ontology-free data acquisition device rolled off the production line and was delivered to JD.com. At the same event, MEgo announced that it had accumulated over 1 million hours of ontology-free data production and signed a business cooperation agreement with Tencent Robotics X Lab for ontology-free data.
Founded in February this year, MEgo entered the mass production stage in June and generated millions of hours of data in the following three months.
In his speech, Yao Maoqing said that the first pole of the physical AI data grid has been erected today. He then spent nearly an hour answering a series of real-world questions facing the industry: how to view the closure of physical machine data collection centers, whether objectless systems will replace physical machines, and what the final competitive landscape of this industry will be.
01 The absence of a physical body does not replace the real device.
The first question during the exchange was about the closure of the humanoid robot data collection center in Shijingshan District, Beijing.
Yao Maoqing said this was an isolated case. His reasoning was that well-managed data collection plants were still operating at full capacity, and Mifeng's own real-device data collection plant was producing data in two shifts. The recently released Foundation Models for Antminer, Xiaomi, and Zhiyuan all used hundreds of thousands of hours of real-device data, which is exactly what all the data collection plants in China are currently producing, and many of them don't even meet the standards.
The relationship between the two types of data is hierarchical, not a substitution.
Ontology-free data is suitable for learning general representations during the pre-training stage, while real-device data is suitable for post-training. However, for specific tasks, real-device data with the corresponding ontology is indispensable for practical application.
He placed the other two types of data sources further down the list. Internet video is currently a crutch; it's the only option for acquiring knowledge and representations when first-person perspective data is insufficient. Once ontology-free data reaches tens of millions of hours, it can be gradually discarded. Simulation data accounts for a small proportion of the training phase, being used more for evaluation due to cost. Simulation requires purchasing digital assets, modeling, and GPU rendering, while with the large-scale operation of ontology-free devices, the cost of a single hour of real data is already lower than that of simulation.
To date, Mifeng has accumulated hundreds of thousands of hours of real device data.
02. As demand converges, the industry will move towards oligopoly.
Since the beginning of the year, what customers want has changed.
Yao Maoqing said that at the beginning of the year, the situation was quite diverse, with everyone iterating, trying, and converging. Even using mobile phones and action cameras to collect data could lead to sales. By the second half of the year, metrics such as 60FPS, 1080P, dual-lens cameras, and depth perception gradually became the consensus, and demand was mainly defined by leading clients. There are already single clients ordering 10,000 sets of equipment right away, and there aren't many companies in China capable of producing that many units.
His answer to the problem of inconsistent data standards was that there was no solution.
When a model company uses first-person perspective data from several suppliers simultaneously, the effectiveness of mixed training is difficult to guarantee. Yao Maoqing gave the example of autonomous driving: slightly modifying the ISP algorithm—the step of converting CMOS raw data to RGB images—is imperceptible to the human eye, but the trained model may look completely different. Neural networks remain sensitive to data signals.
The only solution he offered was centralized purchasing. This was also the situation he observed: as several companies began to increase their volume, clients abandoned the practice of buying a little from each company to test the waters, and instead focused on exclusively supplying one or two companies.
When pressed on whether the industry would become an oligopoly, his answer was definitely yes.
The reason lies in the investment structure. Tens of thousands of sets of equipment represent hundreds of millions of yuan in investment, requiring the storage of hundreds to thousands of petabytes of data. Building a data center also involves an investment of hundreds of millions to billions of yuan. This business demands engineering capabilities, operational expertise, and financial strength simultaneously.
The acceptance criteria were also raised during the discussion. Yao Maoqing said that one million hours is the effective data that can be put on the shelf after deduplication. The operational information in the clips must be effective, and there cannot be any human slacking off; the average frame rate of the nominal 60 frames must be at least 59.98; the spatially extracted handpose, poorly done ones are a few centimeters, while customers require 7 millimeters for the head, and some even require 5 millimeters, and it must be verifiable by the true value read by the encoder of the robotic arm and dexterous hand.
Scene distribution is another hurdle. He said that many companies have data that is all about home and folding clothes, with each task only lasting one or two hours, which cannot meet the needs of top users.
To judge the quality of a data company, he said: blindly collecting 1 million hours of data is not difficult; you can just repeat the same tasks, but it's almost worthless. You need to look at the diversity of the scenarios covered, whether there are detailed quality verification reports, and finally, how much data has been signed off on and revenue confirmed.
03 From millions to hundreds of millions of hours, the path is crowdsourcing
Reaching AGI requires hundreds of millions of hours of data, a scale Yao Maoqing repeatedly mentioned that day. Going from 1 million to 100 million is a 100-fold increase; he said linear replication is not a viable path, and a way to achieve a "dimensionality reduction attack" is needed.
The answer is crowdsourcing.
The path is from centralized to semi-centralized, and finally to a fully inclusive model. The benchmarks are delivery riders and ride-hailing drivers; anyone can participate, turning their skills, time, and experience into quantifiable metrics for payment. Mifeng is already testing related products.
Of the current 20,000 sets of equipment, some are used by Mifeng for crowdsourced data collection, while others are sold to users both domestically and internationally. The equipment is no longer just in the hands of professional data collection plants and BPO companies; retirees and stay-at-home mothers in the community are also contributing data production capacity.
Operating a device-less data acquisition facility is more difficult than operating a physical device data acquisition facility. Physical device data acquisition is something that can be completed in a controlled manner with a site, some scenes, and a group of people. What a device-less facility needs to do is to efficiently transfer equipment assets between different scenes, industries, and cities. The number of personnel is much larger and requires training, and they also need to be able to actually enter those real-world scenes.
This is also where the difference lies compared to the previous generation of data service companies. Yao Maoqing said that the annotation work of the Scale AI generation could be completed by sitting in front of a computer. Someone with a decent computer could be trained for half a day to a day and then start working. The chain of physical AI is much longer. The scene provider needs to provide the scene, enter it, and then collect the data. After collecting the data, there is still storage, slicing, annotation, and handpose extraction. Customers have to find suppliers one by one upstream, which is a very unfriendly process. Mifeng provides a one-stop turnkey solution. Customers provide the distribution of their needs, and Mifeng delivers data that can be directly used to train models.
Figure previously launched its Index platform and announced a $1 billion investment. Yao Maoqing couldn't say how many hours the other party actually collected data; based on public information, the scenarios seem to be primarily domestic services, while Mifeng's strategy covers a wider range of scenarios in China.
There is currently no standard price for data pricing; it depends on the scarcity of the scenario, the difficulty of access and acquisition, and delivery time requirements. Distribution rules are also still developing, and there are two models in the industry: one is selling usage rights while ownership remains with the data provider; the other is exclusive purchase, where MiBee needs to remove the data from its own servers.
These are all still undecided. Yao Maoqing's assessment of the industry's pace is that, from a commercial perspective, data will be the sector that achieves scale and industrialization before models themselves in the next few years. Models are the beginning, but data will define the final outcome.
This article is from the WeChat official account "Hard AI" . For more cutting-edge AI information, please go here.

Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.