具身 AI 與機器人
FolDeX Evaluates Long-Horizon Garment Folding With More Than 2,000 Hours of Real-World Robot Data
FolDeX turns garment folding into a real-world benchmark spanning more than 20 tasks and over 10 robot embodiments, specifically testing whether heterogeneous data can be reused across tasks, environments, and hardware. The team also offers remote evaluation on physical robots, but the data license, full download process, and first set of comparable baselines remain to be confirmed.

Embodied AI is often evaluated in simulators or on short-horizon rigid-object manipulation tasks. But fabric continuously deforms and causes occlusions, while every bimanual action changes the states that remain reachable afterward, making the sim-to-real gap particularly pronounced. FolDeX therefore uses garment folding as its primary task and says it is built entirely from physical-robot data, comprising more than 2,000 hours of manipulation experience, over 20 tasks, and more than 10 robot embodiments to evaluate long-horizon deformable-object manipulation.
Rather than comparing only the success rates of individual policies, the benchmark is designed around four data-reuse scenarios: leveraging human interventions and failure-recovery trajectories collected during deployment; transferring across garment categories and between rigid- and deformable-object tasks; reusing data after changes to lighting, backgrounds, and tabletop configurations; and transferring across different robot embodiments. These dimensions directly address the costliest question facing robotics teams: whether existing real-world experience can be carried into new environments, instead of collecting demonstrations from scratch whenever the garment, camera, or robot arm changes.
The team has also proposed a remote real-robot evaluation platform where external researchers can submit policies for testing with held-out objects, controlled initial states, and standardized execution procedures. Compared with having each group purchase its own hardware and report results independently, this could reduce bias caused by hardware differences and cherry-picked successful trials. However, the platform’s scheduling capacity, retry rules for failed runs, and safety constraints could also affect reproducibility.
Current information comes primarily from the preprint and the official challenge page. A complete dataset inventory, licensing terms, sensor-synchronization format, and independent validation of multiple public baselines are still lacking. More than 2,000 hours describes the dataset’s scale; it does not mean the data covers the long-tail variation found in real homes. Key issues to watch include whether the data becomes practically downloadable, whether the evaluation server remains continuously available, and how large the real performance gaps are for VLA or world-action models in cross-embodiment transfer and failure-recovery tracks.