2026/09/04
Debates over technical routes are still ongoing in the embodied‑intelligence sector, yet the whole industry has reached an unprecedented consensus on the data shortage crisis. One fundamental challenge lies in the unstable foundation of large‑language‑model architectures within the physical world. The physical environment is non‑discrete and filled with intricate forces and interactive relationships, data of which remains severely scarce today.
Videos or simulation‑generated data alone are far from sufficient to enable robots to comprehend and replicate human motions. High‑quality, multi‑modal motion data consistent with real‑world physical laws is urgently required. This forms the background and core value of MotionDecode, a high‑quality multi‑modal motion‑data system for embodied‑intelligence training jointly launched by CHINGMU and EnDecodeX.

Scarce real‑world data and inherent defects of simulation data: Globally, publicly‑available real motion datasets for embodied‑intelligence research total less than tens of thousands of hours, an order‑of‑magnitude gap away from the 100‑million‑hour target for model training. Meanwhile, simulation data suffers from the Sim‑to‑Real Gap due to deviations from real‑world physical rules.
Adaptation difficulties caused by human‑robot motion heterogeneity: Human movements rely on sophisticated musculoskeletal coordination. Robots of diverse configurations with varying degrees‑of‑freedom and joint types require loss‑less mapping of human motion trajectories onto mechanical structures. Discrepancies brought by motion heterogeneity directly impair imitation‑learning efficiency.
Easy multi‑modal data capture, hard multi‑modal alignment: Multi‑source sensors for vision, tactile feedback and proprioception operate at different sampling frequencies and time references. Collected data suffers from timestamp drift and spatial‑coordinate misalignment. Misaligned data introduces noise and negatively interferes with model training.
Insufficient customized data‑production capacity: Custom dataset generation, from scenario setup and sensor selection to quality inspection, lacks standardized toolchains and mature workflows, resulting in long cycles, high costs and unstable data quality, which slows algorithm iteration in vertical application scenarios.
At present, MotionDecode holds a data reserve of 3,000 hours, with an expected annual output of up to 500,000 hours. To accelerate industrial progress, CHINGMU has officially rolled out the MotionDecode Data Open Initiative. The first‑phase release of 1,000‑hour high‑quality human‑motion datasets is open for free global access via application upon scanning the QR code. Datasets will be continuously updated, and additional sample data will be gradually published on GitHub and Hugging Face. Please follow CHINGMU’s official website for further announcements.

🌐 MotionDecode Homepage: https://chingmudata.github.io/MotionDecode/
🧠 MotionDecode Datasets: https://huggingface.co/datasets/CMRobot/MotionDecode
Expanded multi‑modal coverage: Synchronous capture of body motion, high‑resolution hand data (including force feedback), ego‑centric first‑person view, exo‑centric third‑person view, 6‑DoF object poses and environmental scene information.
Diverse application scenarios: Industrial manufacturing, household services, supermarket retail, medical rehabilitation, logistics & warehousing, ball‑game interaction, digital entertainment and other typical use‑cases, including over 500 human‑robot tasks and more than 200 tools and objects, providing abundant training materials for vertical‑scenario algorithms.
Developer‑friendly compatibility: Supports more than 10 common data formats such as BVH, FBX and CSV, and adapts to mainstream robot‑simulation frameworks including MuJoCo, Isaac and URDF, lowering integration barriers and allowing algorithm teams to focus on model development.

On the application side, motion‑retargeting algorithms resolve the motion‑heterogeneity challenge. Captured motion data can be adapted to mainstream simulation environments (MuJoCo, Isaac, URDF) and robot platforms of different mechanical configurations. One set of high‑quality motion data can be learned and reused by multiple robot models, greatly improving the value of motion‑data assets.



