Xiaomi Launches Robot Foundation Model Trained on 100K Hours of Data
Original: Xiaomi-Robotics-1
Why This Matters
Xiaomi's large-scale pre-training approach signals a potential inflection point for robotics foundation model scaling.
Xiaomi released Xiaomi-Robotics-1 on July 16, 2026, a robot foundation model pre-trained on over 100,000 hours of real-world manipulation trajectories across 1,700+ scenarios, combined with 7,200+ hours of in-house real-robot data collected in real homes.
Xiaomi unveiled Xiaomi-Robotics-1, a ready-to-use robot foundation model designed to overcome the data scarcity bottleneck that has historically limited robotics policy scaling. The model follows a two-stage training paradigm modeled after large language models. In the pre-training stage, the model ingests over 100,000 hours of embodiment-free UMI trajectories spanning household, commercial, industrial, and outdoor environments across more than 1,700 scenarios. An automated annotation pipeline powered by a vision-language model (VLM) segments long video clips and labels each segment with language descriptions of gripper and object state transitions—making large-scale labeling feasible. Xiaomi reports that pre-training exhibits clean scaling behavior: validation action error decreases consistently as data volume and model size increase. The post-training stage then aligns the model with real robot embodiments and natural-language instruction following. Xiaomi collected over 7,200 hours of in-house real-robot data in actual homes, covering tasks such as tidying sofas, sorting shoe cabinets, and storing kitchenware. Additional data comes from filtered open-source robot datasets and manually annotated high-quality UMI data. After post-training, the model can be deployed out-of-the-box for a wide range of mobile manipulation tasks and was evaluated in unseen environments with unseen object instances.