機器人與具身 AI
HiFi-UMI Trains Robots Directly on Millimeter-Accurate Wearable Data, Eliminating the Need for Real-Robot Demonstrations in Post-Training
HiFi-UMI combines offline stereo-inertial SLAM, microsecond-level synchronization, and six-view bimanual capture in a portable data system, enabling human-only manipulation data to be used directly for robot post-training. Real-world success rates across three VLA/world-action models approached an in-situ teleoperation baseline, though the results remain limited to controlled tasks evaluated by a single team.

Robot policies typically cannot complete their final training using human egocentric video alone. Wearable Universal Manipulation Interfaces (UMIs) are easy to scale, but suffer from trajectory drift, inaccurate relative poses between the two hands, camera desynchronization, and visual occlusion. The industry therefore tends to use them for pre-training, followed by calibration with a small amount of real-robot teleoperation data. HiFi-UMI attempts to remove this “real-robot anchor” by raising data quality to a level suitable for direct deployment.
The system reconstructs trajectories using head-mounted offline stereo-inertial SLAM, directly measures the relative pose between the two grippers instead of estimating it after the fact, and uses a shared GPIO trigger to keep sensor synchronization error below 40 microseconds. Each hand is equipped with two wide-angle cameras, providing six views in total and an approximately 200-degree field of view per hand. The paper reports an end-effector localization error of about 3 millimeters within the workspace. The data also undergoes automated trajectory reconstruction, simulation replay, and quality validation, reducing the likelihood that flawed demonstrations enter the training set.
Using only HiFi-UMI demonstrations for post-training, the team trained StarVLA-QwenPI, OpenPI-π0.5, and LingBot-VA. Their real-world success rates differed from the in-situ teleoperation baseline by −2.5, +3.1, and −0.6 percentage points, respectively. The best policy achieved an 85% success rate on a precision insertion task. Pre-training on 4,000 hours of data collected with the same system reduced action error by 41% across ten unseen tasks and further increased StarVLA-QwenPI’s real-world success rate by 18.1 percentage points.
The publicly released HiFi-UMI-2K dataset contains more than 2,000 hours of data, 482,000 replayable demonstrations, and over 110 scenes, and is licensed under CC BY 4.0. From an engineering perspective, the key question is whether third parties can reproduce these results using different robot arms, control frequencies, and camera calibrations—and whether long-term SLAM drift, contact forces, and end-effector differences will still require a small amount of real-robot calibration. The current figures do not directly demonstrate that costly robot-collected data can be replaced wholesale by wearable data collection.