Ant Lingbo CEO: Robots Don't Need to Resemble Humans to Be Valuable

Deep News
Sep 15

"Robots don't need to replace people." This was the central message from Zhu Xing, CEO of Ant Lingbo Technology, during his appearance at the INCLUSION Conference on Bund in mid-September. Lingbo, the embodied intelligence company under Ant Group that is currently advancing independent fundraising, is charting a course that prioritizes practical value over mimicking human form.

Zhu Xing addressed external speculation about the company's financial needs, clarifying that the decision to spin off from Ant Group was driven by two key factors: talent incentives and corporate governance. He emphasized that "this industry requires long-term commitment and will experience many cycles," reflecting a patient, strategic approach to the sector's inevitable ups and downs.

Over the past year, Lingbo's LingBot series of models has made a significant name for itself in the industry. The company has released a suite of innovations including the first embodied-native world action model, LingBot-VA 2.0, and the first video generation foundation model designed for embodied intelligence, LingBot-Video. Additionally, they've developed LingBot-VLA 2.0, which is compatible with robots from over a dozen mainstream brands, and recently open-sourced the compact LingBot-World 2.0 model. According to Zhu Xing, while overseas peers may not recognize the name "Ant Lingbo," they certainly know LingBot. He noted that when international competitors release new achievements, they often include Lingbo's models as a baseline for their evaluations.

Rejecting the Humanoid Obsession

The most crowded direction in embodied intelligence is humanoid robotics, but Zhu Xing firmly believes that "robots don't need to replace humans or look like humans." Lingbo's focus is instead on the tangible commercial and social value that embodied intelligence can create as a new technology. This philosophy shapes their technical roadmap, with Zhu Xing insisting that physical intelligence must be trained from scratch—from 0 to 1—based on the actual demands of robots completing tasks in the physical world. This is precisely why Lingbo remains committed to an embodied-native approach and undertakes large-scale pretraining from the ground up.

During their 1.0 phase, the team experimented with fine-tuning general-purpose video generation models built for the digital world. The results were disappointing. "After all that fine-tuning, we ended up destroying the original model's prior knowledge and capabilities," Zhu Xing explained. The fundamental problem is that models from the digital realm prioritize aesthetics and spectacular special effects, whereas robots require adherence to physical laws and real-time responsiveness. He illustrated the point with a vivid analogy: "A bottle of mineral water dancing in the air and landing gracefully in my hand might look impressive, but it violates physical laws and holds no meaning for a robot trying to complete a task."

In terms of form factor, Lingbo equally rejects the humanoid obsession, instead championing cross-morphology generalization. Heavy-duty operations and lightweight sorting don't require the same hardware; and when robots eventually enter homes, different forms will be needed for elderly care versus childcare. This refusal to fixate on humanoid shapes also stems from a clear-eyed assessment of the industry's maturity. Embodied intelligence is still in its infancy. "Everyone needs to be patient. The good news is that we're starting to see limited, real-world deployments," Zhu Xing said.

Prioritizing Economic Viability in Deployment

At the conference, Lingbo showcased the application scenarios currently being rolled out: pharmacies, logistics sorting, and industrial loading/unloading operations. The pharmacy use case is particularly pragmatic. Large drugstores face intense pressure during daytime meal-delivery peaks when pharmacists are swamped, and the situation is even more grueling at night, especially for stores with high delivery volumes. Robots can offer much-needed assistance in these demanding conditions.

For Lingbo, pharmacies serve as an ideal "outpost" for entering more complex scenarios. A typical pharmacy contains thousands of SKUs, testing a machine's ability to grasp and generalize; additionally, close physical interaction between machines and humans demands exceptional safety protocols. Zhu Xing argues that while it's possible to create dazzling robot demos, spectacle doesn't equal real value. What he truly prioritizes is authentic deployment that creates a virtuous cycle of data generation and model iteration.

He outlines two clear conditions for successful deployment: capability is the necessary prerequisite, initially favoring relatively safe, closed environments with short task chains; cost is the sufficient condition, and the economics must make sense. His projection is that from the second half of this year through next year and into 2026, robot adoption in productivity scenarios will experience rapid growth.

Lingbo places an enormous emphasis on data, having built its own data infrastructure tailored to model training needs. While the industry loves to boast about accumulating hundreds of thousands of hours of data, Lingbo prioritizes data quality and distribution, and began custom data collection early on. Their approach is methodical: the model team first defines what data and distributions are needed, conducts pilot collection rounds, develops a standard operating procedure, and then scales up with suppliers via streaming delivery rather than batch processing. Through this rigorous data funnel—from raw information to training-ready datasets—they've boosted their usable data rate from around 15% to over 90%.

Regarding commercialization, Lingbo remains deliberately restrained. Zhu Xing asserted, "We won't over-prioritize commercialization just to chase a better revenue number—that would be penny-wise and pound-foolish." However, he acknowledges that only when customers genuinely pay for their solutions will the company face the true test of its offering. The "war of a hundred models" in embodied intelligence is just beginning, and Zhu Xing predicts the field will consolidate to just a few players, much like the large language model market. Whether Lingbo earns a seat at the big table will depend on how deep they can dig their data moat. More critically, it hinges on solidifying pretraining geared toward the physical world. The truly difficult challenges today have shifted from solving a single problem to solving many: once a scenario works, the question becomes how to replicate success across more scenarios at lower cost. Lingbo's explorations in pretraining and embodied-native technology are ultimately aimed at answering this fundamental question.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10