Drooid Logo
Back to story perspectives

Full Breakdown

Ant Group's Robbyant Launches Open-Source AI Models for Robotics

1/30/2026, 11:21:29 PM

Advancements in Embodied Intelligence

Ant Group, a prominent Chinese fintech company, has made significant strides in the field of robotics by open-sourcing its artificial intelligence models through its subsidiary, Robbyant. This initiative aims to enhance machine intelligence capable of performing complex tasks in real-world environments. The latest offerings include the LingBot-VLA, a vision-language-action model designed to serve as a "universal brain" for robots, facilitating scalable deployment across various industries. Robbyant's CEO, Zhu Xing, emphasized the necessity for "highly capable and cost-effective foundation models" to enable widespread adoption of embodied intelligence.

Key Features of LingBot-VLA

The LingBot-VLA model has been developed to support robots in understanding and executing tasks with minimal programming requirements post-deployment. It has been trained on over 20,000 hours of real-world data and has demonstrated superior performance in benchmark tests, particularly on the Galaxea GM-100 benchmark, which evaluates capabilities against 100 common real-world jobs. This model's ability to transfer knowledge across different robotic configurations minimizes the need for hardware-specific training, streamlining the deployment process for developers.

Introduction of LingBot-Depth

In addition to LingBot-VLA, Robbyant has also introduced LingBot-Depth, a high-precision spatial perception model aimed at improving robots' depth sensing and 3D environmental understanding. This model has shown significant advancements in overcoming challenges associated with depth data acquisition, particularly in complex optical environments. Collaborating with Orbbec, a leader in robotics and AI vision, Robbyant has integrated LingBot-Depth into next-generation depth cameras, enhancing the operational capabilities of robots in various settings.

Technical Innovations and Collaborations

LingBot-Depth utilizes advanced techniques such as Masked Depth Modeling (MDM) to reconstruct missing depth information, thus improving the reliability of robots in environments with challenging optical conditions. The model was trained using approximately 10 million raw samples, ensuring robust performance across diverse scenarios. Zhu Xing highlighted the importance of reliable 3D vision for the advancement of embodied AI, stating that the collaboration with Orbbec aims to lower barriers to advanced spatial perception.

Broader Implications and Future Directions

Robbyant's efforts are not limited to technical advancements; the company envisions creating intelligent robotic companions that can assist in everyday tasks, including elderly care and medical assistance. By open-sourcing its models and collaborating with hardware partners, Robbyant aims to foster innovation in the robotics industry, ultimately enhancing the integration of intelligent systems into daily life.

Official Statements & Responses

Zhu Xing remarked, “Reliable 3D vision is critical to the advancement of embodied AI. By open-sourcing LingBot-Depth and collaborating with hardware pioneers like Orbbec, we aim to lower the barrier to advanced spatial perception.” Len Zhong from Orbbec noted, “Robbyant's work in spatial intelligence models complements Orbbec's expertise in 3D vision chips and robotic vision systems.”

Criticism & Opposition

While the advancements in AI and robotics are notable, some experts have raised concerns about the limitations of current humanoid robots, which often rely on preprogrammed routines. Critics argue that without overcoming these limitations, the potential for economic productivity in robotics remains constrained.

Conflicting Reports & Gaps

There are no significant conflicting reports regarding the capabilities of the LingBot-VLA and LingBot-Depth models; however, the broader implications of their deployment in various industries remain to be fully explored. Further information on the practical applications and user experiences of these models is anticipated as they are integrated into real-world scenarios.