Home>Document Center>常见问题
Category Directory
常见问题
IROS 2025 Zhejiang University: Research on Dextrous Manipulation of Dual Robots

# Research on the Skilled Operations of Dual Robots 

The dexterous operation of dual robots, due to its high degree of freedom coordination, multimodal perception fusion, and complex skill combination, has become an important research direction in the field of robotics. Although the traditional single-robot operation technology can utilize human demonstration data to guide reinforcement learning methods, it often has difficulty generalizing when it comes to coordinated tasks involving multiple sub-skills. Especially in dynamic scenarios, when the state of the object being operated changes, the operating hand needs to respond according to the real-time three-dimensional posture changes. However, existing methods lack explicit perception of the object's posture and size, resulting in uncoordinated hand movements in dynamic states. Moreover, successful operation tasks require time-aligned combinations of sub-skills, such as the operating hand must establish a stable grip before the active hand starts to rotate. However, existing methods lack mechanisms to coordinate these interdependent sub-processes. 

To address these issues, a research team led by Sun Zhengnan from the Industrial Control Institute of the School of Control Science and Engineering at Zhejiang University, including Ye Qi, has conducted the study titled "VTAO-BiManip: Dual Robot Dextrous Manipulation Based on Masked Visual-Tactile-Action Pre-training and Object Understanding". The aim is to achieve human-like dual robot manipulation through multimodal pre-training and curriculum reinforcement learning. This research outcome has been published in the renowned robotics conference IROS.

Figure 1: The left image presents an example of the bottle cap rotation dataset; the middle image shows the simulation results of the bottle cap task using Shadow Hand; the right image demonstrates the real-world deployment using Leap Hand.


# 1. Research Plan:

The research team designed and constructed a VTAO (Visual-Tactile-Action-Object) data acquisition system to capture the demonstration data of human hand operations. This system consists of: 1) Hololens2 for visual data collection; 2) Gloves for tactile and movement collection of both hands; 3) Qingtong visual motion capture system for precisely capturing hand movements and the 6-degree-of-freedom posture of objects; 4) Personal computer for data collection and alignment. Through this system, 216 human hand operation trajectories were collected, involving 26 different bottles, each trajectory lasting 6-17 seconds, totaling 61,684 frames of time-aligned multimodal data.

Figure 2: VTAO data acquisition system for human hand operation


Based on the MAE (Masked Autoencoder) architecture, the research team proposed the VTAO-BiManip pre-training framework. This framework adopts an encoder-decoder structure: the encoder processes visual, tactile and current action inputs, and projects different modal information into latent tokens through an attention mechanism; the decoder reconstructs the original perception modalities, predicts subsequent actions, and estimates the object state. This joint reconstruction mechanism achieves cross-modal feature fusion during the pre-training process, enabling the model to recover the complete environmental interaction from partial perceptual observations.


Figure 3: Overview of the VTAO-BiManip pre-training framework and its downstream hand operation and sub-skill collaboration curriculum reinforcement learning



To address the challenges of multi-skills learning, the research team adopted the PPO reinforcement learning strategy and introduced a two-stage course reinforcement learning framework: In the first stage, the bottle was fixed on the table, encouraging the left hand to learn to grasp the bottle and the right hand to learn to unscrew the cap; in the second stage, the bottle was released, forcing both hands to collaborate in the learning of operational skills.

Scheme Innovation and Advantages:

1. The introduction of action prediction and object understanding as supplementary modalities has enabled cross-modal feature fusion. 

2. A VTAO data acquisition system was designed, capturing multimodal demonstration data of human hand operations. 

3. A two-stage curriculum reinforcement learning framework was proposed, effectively addressing the challenges in the learning of dual sub-skills. 

4. The effectiveness of the method was verified in both simulation and real environments, with a success rate exceeding that of existing methods by more than 20%.



# 2. Experimental Verification:

This experiment deployed the task of rotating bottle caps with both hands in the Isaac Gym simulation environment. 15 different bottles from ShapeNet were selected, of which 10 were used for training and 5 for testing. The experiment utilized an Intel Xeon Gold 6326 and NVIDIA 3090 system. During the pre-training phase, the AdamW optimizer was used with a learning rate of 2e-5, and the mask ratios were set at 0.75 for vision, 0.5 for touch, and 0.5 for action. In the reinforcement learning stage, a two-stage curriculum learning approach was employed. The first stage involved 1000 iterations of training, and the second stage involved 3500 iterations, totaling approximately 62 hours.

Figure 4: Qualitative results of different pre-training methods in the simulation. Left figure: Training process; Right figure: Evaluation results


By comparing the baseline methods of different modal configurations, the experimental results show: 1) The multimodal joint pre-training method VTA significantly outperforms the single-modal pre-training methods V, T, A, as well as the multimodal method VTA-Scr without pre-training; 2) The performance improvements from VT to VTA (adding action modal prediction) and VTO (adding object understanding) indicate that both action prediction and object understanding can significantly enhance visual-tactile fusion; 3) VTAO achieves the highest success rate by combining the two modalities, indicating that the tasks of action prediction and object understanding have complementary benefits.


Figure 5: Training process of the ablation study. Left image: Verification of manual operation; Middle image: Ablation of action prediction range; Right image: Ablation of action token quantity


The ablation study further validated the key design choices: 1) Using both hands' data compared to using only the right hand's data, the success rate can be significantly higher regardless of which modality is removed; 2) The action prediction mechanism significantly improves the performance of downstream tasks; 3) The object understanding module is crucial for achieving dynamic operation strategy adjustments based on the state of the target object.


# 3. Experimental Results:

This study demonstrates that the VTAO-BiManip framework successfully achieves dual-handed dexterous manipulation through the combination of multimodal pre-training and curriculum reinforcement learning. This method achieved a success rate of 63% (for objects already present) and 51% (for objects not present) in the simulation environment, outperforming existing methods by more than 20% in terms of performance. The research team also verified the effectiveness of the method in a real environment, as shown in Figure 6.


Figure 6: Our Operating Platform


# IV. References:

1 Liu Y, et al. M2VTP: Masked Multi-modal Visual-Tactile Pre-training for Robotic Manipulation. ICRA 2024.

2 Chen Y, et al. Towards human-level bimanual dexterous manipulation with reinforcement learning. NeurIPS 2022.

3 Qin Y, et al. DexMV: Imitation learning for dexterous manipulation from human videos. ECCV 2022.

4 Rajeswaran A, et al. Learning complex dexterous manipulation with deep reinforcement learning and demonstrations. RSS 2018.

5 He K, et al. Masked autoencoders are scalable vision learners. CVPR 2022.

6 Chen Y, et al. Visuo-tactile transformers for manipulation. CoRL 2022.

7 Li H, et al. See, Hear, and Feel: Smart Sensory Fusion for Robotic Manipulation. CoRL 2023.

8 Zhao T Z, et al. Learning fine-grained bimanual manipulation with low-cost hardware. arXiv 2023.

9 Kataoka S, et al. Bi-manual manipulation and attachment via sim-to-real reinforcement learning. arXiv 2022.

10 Chen Y, et al. Sequential dexterity: Chaining dexterous policies for long-horizon manipulation. arXiv 2023.

11 Nair S, et al. R3M: A universal visual representation for robot manipulation. arXiv 2022.

12 Brohan A, et al. RT-1: Robotics transformer for real-world control at scale. arXiv 2022.

13 Xiao T, et al. Masked visual pre-training for motor control. arXiv 2022.

14 Yuan W, et al. GelSight: High-resolution robot tactile sensors for estimating geometry and force. Sensors 2017.

15 Makoviychuk V, et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv 2021.


# 5. Original Link:

https://arxiv.org/abs/2501.03606

https://suzend.github.io/VTAO-Bimanip/

CHINGMU × Xiying Group: A New Chapter in Digitalization of Film and Television

In the field of film art, when the movements of characters are smooth and fluid, the expressions of digital characters are vivid and expressive, and the virtual shooting seamlessly integrates with the real scenes, the success of such works is inseparable from the strong support of motion capture technology.


Currently, the Chinese film and television industry is accelerating its transformation towards digitalization and intelligence. Technological innovation has become the core driving force for the industry's breakthrough and upgrading. Against this backdrop, on October 22nd, CHINGMU Qingtong Vision and Xiying Group, which carries the profound cultural heritage of western China's films, officially reached a strategic partnership. This strong alliance between technology and film and television, with high-precision optical motion capture technology as the link, not only injects core power into the digital upgrade of Xiying Group, but also marks a crucial step for the Chinese film and television industry in its digital transformation process, where high-tech is implemented in the industry and empowers the upgrading of creative production.


The signing ceremony site

As the first national high-tech enterprise in China to achieve full-stack self-research and development of optical motion capture technology, Qingtong Vision has always focused on deepening the research and development of motion capture technology and innovative services in the digital entertainment field. Relying on its self-developed motion capture system and real-time rendering engine, Qingtong Vision can accurately capture dynamic data of limbs, expressions, hands, as well as complex rigid and flexible bodies, helping the fields of film, animation, and stage achieve efficient production and immersive experiences of "what you see is what you get". 


In this collaboration, Qing Tong Vision will provide comprehensive technical support to the Digital Motion Capture Studio of Xiying Group, covering virtual production, CG animation, virtual rehearsals and other film creation scenarios. It will build a full-chain technical support from action data collection to digital asset generation. Coinciding with the industrial opportunity of the 2025 Xi'an International Virtual Reality Film Festival, both parties will jointly explore the content creation and ecosystem construction of virtual reality films, promote technological upgrades in aspects such as dynamic shaping of virtual characters and immersive interaction, and provide core support for creating a "technology - content - industry" closed loop for the film festival.


The blue iris motion capture technology made an impressive debut and received widespread praise. 

This collaboration between Qingtong Vision and Xiying Group holds significant milestones for the digital development of China's film and television industry. On one hand, it breaks the long-standing industry situation where high-end motion capture technology relies heavily on imports. By adopting a model where domestic self-developed technologies serve China's film and television industry, it fully demonstrates the maturity of domestic motion capture enterprises in core technology research and the ability to adapt to the industry, providing the industry with a safe and controllable domestic technology option. On the other hand, this collaboration has established a collaborative innovation model between technology enterprises and film and television groups. It upgrades motion capture technology from a single tool to a full-process creative empowerment system, promoting the transformation of film and television production from traditional experience-driven to modern data-driven, and setting an innovative benchmark for the deep integration of technology and industry in the industry.


Mr. Zhang Haiwei, CEO of Qingtong Vision, Accepted an Interview and Talked About Industry Technologies

As a leading global service provider in optical motion capture, Qingtong Vision has long established a full-process digital solution capability covering solution design, digital asset production, character creation, scene construction, motion capture studio rental, performance capture assistance, and data restoration & processing.

At present, its ecological cooperation landscape covers top domestic companies and renowned platforms including Baidu, Tencent, Huawei, Douyin, Mengniu, XPeng Motors, and CCTV. It provides optical motion capture technology support and customized solutions in fields such as virtual production, CG animation, digital humans, and stage art.

In addition, Qingtong Vision has continuously deepened industry-university cooperation, establishing partnerships with more than 300 domestic universities and colleges, including Tsinghua University, Peking University, Communication University of China, and Shanghai Art & Design Academy. It provides cutting-edge technical support for film and television creation research and intelligent teaching, offering sustained impetus for discipline development and talent cultivation in educational institutions.


The signing with Xingying Group marks another significant practice of CHINGMU Qingting Visual Technology in empowering the digitalization of the film and television industry. In the future, the two parties will carry out deeper cooperation in areas such as film content creation and technological research and development, jointly exploring new paths for the digitalization of China's film and television industry. And Qingting Visual Technology will continue to promote the application innovation of optical motion capture technology in the entertainment field, helping Chinese film and television works shine with even more brilliant light and shadow charm on the global stage.


CHINGMU × UITM: Overseas Empowerment of Sports Science

Under the global trend of upgrading global sports research towards more refined and quantitative analysis, universities and research institutions around the world are accelerating their exploration of the integration path between technology and academic research. On September 18th, CHINGMU Qing Tong Vision and the Sports and Leisure Science College of Universiti Teknologi MARA (UiTM) in the Shahabang campus of UiTM officially signed a cooperation memorandum. This collaboration not only marks an important milestone for Qing Tong Vision in the field of overseas sports science, but also provides crucial technical support through customized optical motion capture solutions. It injects new technological impetus into UiTM's sports analysis laboratory, helping it enhance efficiency and effectiveness in sports research and practical applications.


The dean of the Sports and Leisure Science College of Malaysia's Malaya Institute of Craft Technology, Associate Prof. Dr Raja Mohamed Firhad (left) 

Signed the cooperation memorandum with Zhang Haiwei (right), the founder and CEO of Qingtong Vision 

The signing ceremony was presided over by Associate Prof. Dr Raja Mohamed Firhad, the dean. It is worth noting that this cooperation agreement was signed at a time when the NAPESS International Conference (National Association of Physical Education and Sports Science) first left India and landed in Malaysia. At the event, over 100 experts and representatives from the fields of education and research from China, Malaysia, India, Thailand, Japan, Ukraine, etc. gathered to exchange views on cutting-edge topics such as sports science, training load monitoring, movement performance optimization, physical education, and youth sports talent selection. As a technical partner of UITM, Qing Tong Vision made a simultaneous appearance at the NAPESS conference, highlighting the technical advantages and application scenarios of its core products such as optical motion capture software and motion capture cameras. The system can integrate force platforms and electromyography devices, relying on real-time high-precision motion capture capabilities, multi-modal motion data analysis, and archive management, to help researchers conduct technical movement review and scientific training in a precise and data-driven manner.


Ding Guanhong, the product manager of Qingtong Vision, gave an introduction to the optical motion capture system at the NAPESS international conference. 

On the other hand, in response to the demand for the convenience of motion capture, Qingtong Vision can also provide a markerless motion capture solution. By combining the AI reference camera R3 with the motion capture software, it can achieve high-precision motion capture data collection in a smooth and convenient manner, building a "technology + research" collaborative bridge for the development of sports science. Zhang Haiwei, the founder and CEO of Qingtong Vision, said: "This is an important step in Qingtong Vision's global development. We look forward to injecting more technological power into sports science and sports health through this cooperation." 

After the meeting, the Qing Tong Visual Technology team conducted a customized motion capture training session for the core team members of the UITM Motion Analysis Laboratory. Through a combination of theory and practice, they provided detailed explanations on the principles and characteristics of optical motion capture equipment, the hardware setup process, as well as the parameter adjustment techniques for motion science scenarios, the linkage application of motion capture data and scientific research analysis software, etc. This helped the team quickly grasp the key technical points and laid a solid foundation for the subsequent application of motion capture technology in daily scientific research.


Ding Guanhong, the product manager of Qing Tong Vision, conducted motion capture training for the UITM Motion Analysis Laboratory team.

A group photo of Qing Tong Vision with the team members of the UITM Motion Analysis Laboratory at the University of Technology Malaysia (UITM) 

Starting from the open and cross-border hosting approach of the first NAPESS international conference, to the new chapter of technological collaboration initiated by CHINGMU, the Eye Vision and the Malaysian University of Technology and Arts (UITM), this collaboration not only injected new vitality into the academic exchanges of sports research, but also built a new bridge for cooperation between China and Malaysia in the fields of technology and sports education. 

In the future, as the Qing Tong visual motion capture technology gradually becomes operational in the UITM Sports Analysis Laboratory, both parties will further deepen their cooperation in areas such as course teaching, research projects, and technology transfer, jointly exploring more application scenarios of smart sports, and making motion capture technology a booster for global sports health and sports research and innovation development.


A virtual reference trajectory scheme for tracking control of wheeled mobile robots with slip disturbances
Motion capture equipment procurement: Key points guide


       # Motion Capture Equipment Buying Guide In an era of rapid development in virtual reality, film & animation, digital twin and other industries, Motion Capture (Mocap) has become a key driver of digital transformation. Whether building metaverse scenarios, conducting scientific research experiments, or producing film visual effects, high-precision motion capture equipment serves as an indispensable "digital bridge". However, with a wide variety of mocap devices available on the market, how to make a scientific procurement decision? What core issues should be noted? This article will answer them one by one. ## 1. Clarify Your Needs First: Scenarios Determine Equipment Selection Before purchasing motion capture equipment, you must first clarify your application scenarios and technical requirements, as performance and functional demands vary across different use cases. - **Film & Animation**: Requires high precision and low latency to capture facial expressions and finger details. - **Industrial Simulation**: Requires high robustness to adapt to complex environments such as strong light and vibration. - **Scientific Research**: Requires data visualization capabilities to support kinematic and dynamic analysis. - **Virtual Human / Metaverse**: Requires multi-target synchronous capture and supports large-space multi-person interaction. Take CHINGMU’s motion capture system as an example. Its self-developed infrared optical technology supports the capture of human bodies, rigid bodies and soft bodies. It has also solved industry challenges including precise hand motion capture, quadruped animal capture and outdoor motion capture. In terms of capture range, it supports spaces from dozens of square meters to thousands of square meters, and is applicable for both underwater and above-water, indoor and outdoor environments, meeting customized demands across various industries. ## 2. Focus on Core Performance Parameters The performance of motion capture equipment directly affects data quality and user experience, so close attention must be paid to these key indicators. - **Precision**: Mainstream devices generally achieve sub-millimeter precision. CHINGMU’s mocap system features a 3D positioning accuracy of ±0.02mm, enabling stable tracking of high-speed moving objects. - **Latency**: Low latency is essential for real-time interaction scenarios, ensuring authentic and complete data with seamless synchronization, delivering reliable results for scientific research, film & animation, medical rehabilitation and other fields. - **Anti-interference Capability**: The system must maintain stable data output in complex environments with electromagnetic interference and occlusion, requiring strong robustness. - **Scalability**: Support for multi-device collaboration, compatibility and scalability greatly improve efficiency and convenience. ## 3. Choose the Right Supplier: Technical Strength Is Key Motion capture equipment involves high technical barriers, making the supplier’s R&D capability and industry experience particularly critical. Prioritize the following aspects: - **Proprietary Intellectual Property Rights**: Avoid OEM products to ensure long-term technical iteration support. Self-developed brands can also provide secondary development and customized services. - **Proven Cases**: Service records for top-tier clients and well-known projects are important benchmarks for industry-proven reliability. CHINGMU has provided solutions for thousands of renowned enterprises and institutions worldwide, covering film & television, industry, medical treatment, scientific research automation, motion analysis, virtual simulation and other fields with rich experience and advanced technologies. - **Full-process Services**: Including deployment, commissioning, training and after-sales support, reflecting a complete service system to protect your interests. ## 4. Built for Precision, Reliable Service, Your First Choice CHINGMU is a leading high-tech enterprise with computer vision and artificial intelligence at its core. It specializes in the R&D, production, sales and service of 3D intelligent perception and human-computer interaction systems, and provides customized full-process solutions for various industries. Its products and technologies have served thousands of well-known enterprises and institutions worldwide in the metaverse, digital entertainment, virtual reality, scientific research, industrial automation, bioengineering and other fields, winning wide customer trust and a strong brand reputation. With an international vision, CHINGMU’s products are exported to Asia, Europe, the Americas and beyond. The company is committed to achieving world-class technology and service quality, providing customers with full-stack digital solutions to capture every wonderful frame. ## Final Notes When purchasing motion capture equipment, consider technology, service and long-term value comprehensively. CHINGMU has always adhered to the philosophy of "Customer First, Product Core, Service Priority, Win-Win Cooperation". With internationally leading motion capture technologies, it builds a complete platform for creativity and research.新路径。

Sales Manager
Algorithm Engineer
Overseas Sales Manager
IROS 2025:Blue Eye Motion Capture Black Technology


The annual flagship event in the robotics community, IROS 2025, is fast approaching. As a standout star in motion capture technology, CHINGMU has prepared a full lineup of exciting surprises: multiple intensive technical workshops, in-depth industrial cluster visits, an immersive motion capture showcase, and blockbuster new product launches at Booth T03. Miss it, and you’ll have to wait another year! Here’s an exclusive sneak peek at the must-see highlights.

3 High-Profile Workshops: Industry Leaders Unveil Cutting‑Edge Innovations in Robotics

No more endless scrolling through papers for valuable insights! CHINGMU is sponsoring 3 workshops featuring top experts across three key fields: Marine Robotics, Active Perception Technology, and Medical Robotics. From motion capture for underwater robots and the secrets of perceptual interaction in embodied intelligence, to precision applications in medical scenarios — every session is packed with pure, high-value technical knowledge.
⏰ Important Notice: Detailed schedules and content for the workshops are listed below. Don’t miss your exclusive knowledge hub!


CHINGMU Technical Tour @ IROS 2025   Unlock the Code of Robotics Innovation Theories alone aren’t enough? Consider it done! CHINGMU has partnered with IROS 2025 to launch an exclusive **Technical Tour**. From October 21 to 23, a dedicated shuttle bus will depart on schedule every day, taking you deep into the core of Hangzhou’s robotics industry cluster. You will visit the Zhejiang Institute of Quality Science, CHINGMU’s motion capture demonstration site, BrainCo, the Humanoid Robotics Industry Innovation Center, and more to witness real-world applications and achievements. Covering everything from R&D to industrial deployment, this tour offers a one-stop overview of the **full intelligent evolution chain of robotics**. --- Live Demonstrations: Cutting-Edge MoCap Technologies on Display During the exhibition, CHINGMU will present a spectacular "MoCap+" showcase at Booth T03. Witness motion capture integrated with robotics and motion control, aerial swarms, motion replication, special-scene tracking, and more — with a variety of scenarios staged one after another, revealing the boundless potential of technological convergence. And there’s more: a **blockbuster new product launch**! Dual upgrades in precision and efficiency — the first chance to experience it is reserved exclusively for on-site visitors. --- 📍 Find CHINGMU Here! ▶ Booth No.: T03 From brainstorming workshops and on-site industry tours to live demonstrations and surprising new product launches, CHINGMU is ready to explore the infinite possibilities of motion capture technology with you every day at IROS 2025. Join us at Booth T03 from October 19 to 25. Let’s dive deep into robotics innovation, explore, and enjoy the future together!

CHINGMU Mocap System Empowers Unitree Robotics at CMG World Robot Competition
CHINGMU Enables Cross-dimensional Comedy with a Robot
UITM | CHINGMU Mocap Technology Empowers Research and Teaching Innovation
GALBOT | CHINGMU Enables Humanoid Robot High-Dynamic Tennis Rally Breakthrough
CHINGMU and AMD rewrite the world record: 100-person simultaneous real-time motion capture

Shanghai, May 31 — CHINGMU hosted its “100 People Enter the Digital World Simultaneously” real-time motion capture challenge & digital motion showcase at the CHINGMU MCP Boundless Studio in Shanghai. AMD served as the lead computing partner. The event featured a live performance by 100 motion-capture actors, demonstrating the synergy of large-scale optical motion capture, AI computing, real-time digital character animation, and embodied AI.

 

1



The entire challenge was witnessed and officially documented on-site by the Shanghai New Hongqiao Notary Public Office, which issued a notarial certificate attesting to the authenticity of all generated technical data.


The First Publicly Verified 100-Person Real-Time Motion Capture System

 

Inside a 1000㎡ capture space, 100 actors performed group movements. The system used 76 CHINGMU Kunpeng (K) Series K26 optical mocap cameras, together with a single high-performance workstation equipped with a 64-core AMD Ryzen Threadripper PRO 9985WX processor and an AMD Radeon RX 9070 XT graphics card (16GB VRAM). Operating at industrial-grade 120 fps, it captured ~5,300 markers across 100 actors, performing real-time 3D reconstruction, marker identification, and skeletal solving with an end-to-end latency under 12 ms. On-site large screens simultaneously showed live performance, real-time skeleton data, motion trajectories, and system status — visualizing how human motion enters the digital world in milliseconds.


2

3

100-person motion capture

According to publicly available sources, the previous internationally certified record for real-time multi-person motion capture stood at 19 persons, while the highest publicly demonstrated scale in China was approximately 41 persons. This challenge pushed the scale to 100 persons for the first time—a 2.3-fold increase in participant count and an exponential leap in data volume—making it the largest publicly verified real-time multi-person motion capture achievement on record, and marking a new era of industrial-scale 100-person applications for global motion capture technology.


From Dozens to 100: The System-Level Challenges

 

Scaling real-time mocap from dozens to 100 people is not linear superposition, but a comprehensive test of multi-target recognition, occlusion handling, real-time solving, system sync, and long-term stable operation.

 

Exponential Data Explosion: The system processed roughly 60,000 2D image points from 76 cameras per frame, equating to 7.2 million 2D image points per second—all within an 8 ms single-frame window, while maintaining an end-to-end latency below 12 ms.

 

4


Occlusion & Identity Recognition: With 100 performers moving simultaneously, marker loss/re-acquisition, mis-assignment, and skeletal drift grow exponentially. Error correction must compete for compute resources within the real-time window.


5

 

Stable End-to-End Operation: All 100 motion data streams must stay strictly synced. The system ran under heavy load for over one hour without crashes, frame drops, or memory leaks.


Coupled with high organizational costs and limited trial-and-error room for a 100-person team, these interconnected challenges make 100-person real-time mocap a recognized technical "no man's land."


System Synergy: Deep Integration of CHINGMU and AMD

 

The success came from collaborative operation of optical mocap systems, real-time algorithms, on-site engineering, and high-performance computing. CHINGMU handled system design and core algorithm R&D; AMD provided compute power for high-concurrency data processing and real-time visualization.

 

CHINGMU redesigned core algorithms for 100-person high-concurrency scenarios, improving key algorithmic efficiency by 300%+, and introduced a differentiated marker placement scheme to optimize large-scale marker recognition and identity stability.


7


AMD's CPU (64-core Threadripper PRO 9985WX) handles high-concurrency thread scheduling, motion solving, data reconstruction, and multi-actor sync. The GPU (Radeon RX 9070 XT) handles real-time rendering of the mocap software CMAvatar, isolating compute from rendering.


8


Through joint optimization of BIOS config, thread scheduling, and data paths, both teams achieved a 20% overall system performance boost and significantly reduced latency — enabling stable millisecond-level operation at 100-person scale.


Where Tech Meets Art: A Watchable Digital Motion Showcase


The event also featured a digital motion showcase. Professional mocap actress Xixiyu SAKANA opened with an interactive "Light Up the Digital World" performance, leading into the 100-person group capture.


9

10


One hundred dancers from Shanghai Film Art Academy, wearing mocap suits, performed the original dance "Light of Shanghai" as the core challenge — delivering smooth, stable, and visually striking motion data, transforming raw tech data into vivid artistic expression. On-site screens displayed live dance, real-time skeleton data, and system status, clearly showing how real movement becomes computable, reusable, and drivable digital assets.


11

12


Embodied AI Demo: One mocap actor used CHINGMU's system to teleoperate 6 Unitree G1 robots in a synchronized dance, together with Westlake Robot's teleop platform. This demonstrated the "human motion into robot body" pathway, highlighting how high-quality motion data powers robot control, embodied AI training, and human-robot collaboration.


13


From Technology Breakthrough to Industry Enablement

 

The core value of this challenge goes beyond breaking the world record. It fully verifies the engineering capability of motion capture systems for large-space, high-concurrency, and complex interaction scenarios, opening new application boundaries for the digital economy and embodied intelligence industries. This capability can deeply empower digital content production fields such as group animation, virtual production and digital performance, while enabling applications including robot teleoperation, humanoid robot motion learning, and embodied intelligence dataset construction, providing key data support for intelligent systems to understand real-world movements.

 

This is not only a technological milestone but also a validation of CHINGMU's ability to build large-scale real-time motion data infrastructure for the entire industry. The 100-person scale is by no means the end. As the capture space, camera arrays, and computing platforms continue to evolve, real-time multi-person motion capture has vast potential to scale even further. Moving forward, CHINGMU will continue to advance high-precision motion capture and multi-modal data acquisition, building a high-precision motion data infrastructure that connects the physical world, the digital world, and the robotics industry, and driving the development of the global digital intelligence industry.


微信图片_20260602161732_246_12


As a landmark tech event of the Information Consumption Festival 2026, this showcase provides a highly shareable visualization model for Shanghai's industrial integration in ultra-HD video, AI computing, digital content, and embodied AI — helping build an innovation hub where digital technology meets real economy. More importantly, this challenge, built on fully self-developed core technology, marks a critical leap for China's mocap industry from "domestic substitution" to "global leadership," proving that China's self-developed optical mocap systems have reached world-class status and are reshaping the global mocap landscape.

 

In the era of embodied intelligence, human motion data is becoming the core production asset driving intelligent system evolution. Moving forward, CHINGMU will continue to advance high-precision mocap and multimodal data acquisition, building a high-precision motion data infrastructure that connects the physical world, the digital world, and the robotics industry — and driving the global digital intelligence industry forward.



A Record of a Mortal’s Journey to Immortality《凡人修仙传》 | "Multi-Camera Motion Capture + Real-Time Rendering" — CHINGMU Empowers Freedom in Action Cinematography
Born to Fly | CHINGMU Virtual Production Technology Transforms Film Production
Shanghai University of Sport | Full Racket Motion Analysis for Scientific Table Tennis Training
Shanghai University of Engineering Science | Motion Capture Enables Sports Bra Shock Reduction Optimization
Huazhong University of Science and Technology | Precision Testing of Ships Under Complex Sea Conditions
Unitree Robotics | CHINGMU Empowers Efficient Robot R&D with Intelligent Motion Capture Algorithms
Copyright © 2025 SHANGHAI CHINGMU VISION TECHNOLOGY CO., LTD
沪公网安备31010602005676号 沪ICP备16005979号-1 Site Map