2022/08/02
“Public awareness of motion‑capture technology remains relatively limited,” remarked Zhang Haiwei, Founder of CHINGMU Vision.
“While motion‑capture has been widely adopted in animations, video games and VTuber productions, most ordinary audiences have little idea about what performers wear behind‑the‑scenes or how their real‑world movements are converted into on‑screen footage. This knowledge gap has given rise to numerous public misunderstandings surrounding motion‑capture technology.”

Readers who follow the virtual‑creator community are likely familiar with CHINGMU Vision. The tech‑driven enterprise devotes itself to R&D of motion‑capture and computer‑vision technology. It has carried out collaborative projects with major companies including Huawei, NetEase, Tencent and Original Force Animation, and delivered deep‑rooted technical support for well‑known virtual idols, video games and animated IPs. Its motion‑capture solutions are also deployed across industrial, sports‑science and medical sectors. As a front‑runner within China’s motion‑capture industry, CHINGMU Vision possesses rich‑hand‑on experience in both technology research and commercial implementation.
In this exclusive interview, Intelligence‑Chick sat down with Zhang Haiwei to discuss the present‑day development and real‑world applications of motion‑capture technology.
Through insights shared by this industry insider, this article aims to help readers gain an in‑depth understanding of everything that lies behind the term “motion capture”.

“I first stepped into this field because I studied computer vision during my postgraduate years. I graduated back in 2008, when barely any enterprises opened dedicated computer‑vision job vacancies,” Zhang Haiwei recalled.
Computer‑vision technology enables computers to perceive and interpret the physical world via camera‑based visual input, which forms the fundamental technical foundation for motion capture.
Few people paid attention to computer vision, let alone motion capture, at that time. Just like many new graduates, he faced difficulties finding roles matching his academic background. Even so, he firmly believed human‑computer interaction would become a major future trend. After exploring multiple paths, he decided to focus his career on motion‑capture technology.
“Back then, the motion‑capture sector remained a niche market with high technical barriers and limited competition. Meanwhile, I remained confident about the long‑term prospects of human‑computer interaction.”
Time has proven his judgment correct.
Nowadays, motion‑capture technology has permeated daily life. Beyond film, animation and gaming, it delivers irreplaceable value for high‑end medical research. It supports intraoperative navigation for sophisticated surgical procedures, optimizes painful rehabilitation treatments, generates customized recovery training plans for patients, and aids in designing well‑fitted artificial knee joints. Sports‑equipment development and athlete‑movement analysis also rely heavily on motion‑capture solutions.
For instance, CHINGMU Vision once participated in R&D for women’s sports‑bras. Motion‑capture equipment was used to measure chest vibration during exercise, evaluate shock‑absorption performance and test product adaptability across different body shapes.

(Sports‑bra)
“Motion capture first gained widespread public attention during the VR boom, and then experienced its second wave of popularity with the rise of the metaverse. Before VR became mainstream, few ordinary people had even heard of motion capture. Today its applications appear everywhere, and public recognition will only keep growing,” Zhang Haiwei stated.
Virtual‑idol culture, 3D content creation and human‑computer interaction represent major long‑term trends. Though they may not dominate the whole market, industry demand for motion‑capture services is expected to expand ten‑fold to one‑hundred‑fold. Once metaverse‑related virtual platforms are fully established, motion‑capture user‑bases could scale from tens‑of‑thousands to tens‑of‑millions or even hundreds‑of‑millions of users.
The motion‑capture industry now faces unprecedented transformation opportunities.

Motion‑capture technology has become ubiquitous across online entertainment, among which optical motion capture is the variant most‑frequently encountered by audiences, especially for 3D VTubers.
Optical motion‑capture falls into two primary categories:
Marker‑based optical capture: Performers wear dedicated suits fitted with reflective markers, the same workflow used for Hollywood movie characters and most 3D virtual avatars.
Marker‑free (AI‑powered) motion capture: No special suit or markers are required; tracking is completed purely via cameras, at the cost of slightly reduced precision, widely adopted by many Live2D VTubers.

Some creators also deploy simplified tracking hardware such as HTC Trackers for live‑streaming productions. In most industry‑related conversations, “optical motion capture” generally refers to marker‑based solutions.
Marker‑based optical‑capture cameras track individual reflective spheres. Standard motion‑capture suits are manufactured from OK‑fabric, a material commonly used for knee‑supports and abdominal binders. Reflective markers attach to the fuzzy suit‑surface via hook‑and‑loop fasteners, allowing flexible repositioning across most body regions (breathable non‑hook‑and‑loop fabric is used for armpits and crotch areas). Computers calculate marker coordinates in 3D‑space and retarget captured performer‑movement data onto virtual‑character skeletons.
While suit‑mounted markers are adequate for virtual‑idol productions, sports‑science and medical research require markers to be affixed directly onto human skin. Relative slippage between clothing and skin would otherwise introduce tracking errors and distort movement‑analysis results.

During the interview, Intelligence‑Chick raised a sensitive question: Can motion‑capture suits scratch performers’ skin?
Zhang Haiwei responded with a helpless smile.
“Concerns like this mostly stem from insufficient public understanding. Motion‑capture suits are essentially performance‑grade everyday clothing. Though their fabric surface carries a slight textured grain and is not as silky as silk, severe skin‑lacerations requiring suit‑removal for wound‑treatment are physically impossible.”
If skin injuries do occasionally occur, he explained, they are far more likely to result from accidental collisions during energetic choreography, or sharp‑edged foreign‑objects trapped inside the suit that abrade skin during movement. To date, CHINGMU Vision has not recorded such incidents among its performers.
“We once planned to shoot an educational test video, attempting to figure‑out under‑what‑circumstances a motion‑capture suit could cause severe skin‑scrapes. We even consulted dermatology experts to research how certain skin‑types might be vulnerable to friction‑related injury. After thorough preparation, we ultimately decided not to release the video, hoping related controversies would fade away.”
When asked whether this viewpoint could be published publicly, he sighed and gave his consent.

Within the ACGN market, motion‑capture is widely deployed for animation, video‑game and virtual‑idol productions.
The virtual‑idol sector has experienced explosive growth. Back in 2021, the VTuber market recorded a 350 % year‑on‑year increase. Operating costs for virtual‑idol projects have risen accordingly. The virtual idol Xingtong has long served as an industry benchmark for high‑end motion‑capture resources.
“High‑grade optical motion‑capture remains costly, due to expensive hardware and substantial R&D investment. Optical‑tracking counts as high‑precision, high‑investment technology across every one of its application fields,” Zhang Haiwei explained.
He cited Xingtong’s behind‑the‑scenes production report, estimating that its dedicated motion‑capture studio alone cost over 20 million yuan. Other mainstream virtual‑idol groups regularly invest hundreds‑of‑thousands or even millions‑of‑yuan into motion‑capture workflows.
The thriving domestic ACGN economy has opened‑up new commercial opportunities for the motion‑capture industry. Many young tech‑companies now have staff immersed in sub‑culture communities. As celebrity‑endorsement costs rise, brand‑marketers are increasingly turning toward virtual idols as a fresh, cost‑effective, fast‑track commercialization channel, and many technology‑service providers are competing to seize this market.

Despite this boom, technical barriers still stand in the way of independent VTubers. Launching a virtual‑avatar channel requires not only motion‑capture equipment, but also Live2D illustration or high‑quality 3D character‑modelling services. The time‑consuming technical‑adaptation work between illustrators, modellers and motion‑capture teams represents one of the biggest bottlenecks. At present, few unified platforms connect independent creators, 3D‑modellers and motion‑capture service‑providers. Most new VTubers source custom‑built models independently before seeking motion‑capture‑adaptation services.
CHINGMU Vision has experienced these integration‑related pain‑points firsthand. During preparations for its July‑1 new‑product launch event, roughly 80 % of its workload was spent adapting third‑party character‑models for motion‑capture workflows.

“Many modellers lack practical knowledge of motion‑capture pipelines, and additional tuning is needed to match virtual‑avatars to real‑life performer‑movement. This adaptation‑phase takes up huge amounts of manpower and time. Many new motion‑capture customers run into the problem of owning capture hardware but lacking compatible character‑models. To solve this issue, we are building an integrated marketplace platform, where users can purchase equipment, rent capture‑studio space, and modellers can publish standardized‑format avatar‑assets. This unified‑standard ecosystem will drastically cut‑down model‑adaptation cycles and lower barriers for aspiring VTubers. Our long‑term goal is to drive down overall industry costs and popularize virtual‑broadcast workflows.”
CHINGMU Vision’s new‑product launch highlighted innovative applications of optical motion‑capture for metaverse and animation industries.
A long‑standing debate within the animation community argues that heavy reliance on motion‑capture workflows stifles animators’ creative imagination, even though motion‑capture greatly accelerates the industrialization of China’s domestic animation sector. Zhang Haiwei acknowledged that his team has encountered workflow‑related conflicts between motion‑capture technicians and hand‑key animators.

(A Record of a Mortal’s Journey to Immortality adopted motion‑capture technology supported by CHINGMU Vision.)
“Motion capture delivers clear cost‑savings and efficiency‑gains, yet we have repeatedly seen situations where company‑leadership wishes to adopt motion‑capture pipelines, while animators resist the change. Some animators worry motion‑capture technology will replace their jobs. Others are reluctant to spend extra‑time learning entirely‑new workflows that differ drastically from traditional key‑frame animation. These frictions have slowed industry‑wide adoption.”
“Such resistance mostly comes from misunderstandings. Motion‑capture is just a production tool. Raw performer‑movement data is never applied one‑to‑one onto virtual characters. Animators still retain full creative control to refine, exaggerate or modify captured‑motion data in post‑production. Many animators already physically act‑out character‑movements before animating, then reference footage of their performance. Motion‑capture simply helps them transfer those self‑recorded movements into digital environments faster, leaving more time for artistic refinement.”
“In well‑established western animation‑programs, animators are formally taught performance skills, because excellent animators must understand human‑body movement and emotional‑expression mechanics.”

“Technology inspires art rather than replacing it. Motion‑capture‑workflows boost animation‑studio productivity and profitability, bringing tangible economic benefits to China’s hard‑pressed animation‑industry. Artistic vision always remains the decisive factor in finished animated works; motion‑capture is purely an auxiliary production‑tool.”
CHINGMU Vision has identified four major bottlenecks restricting mainstream adoption of traditional optical‑motion‑capture hardware:
High cost: Professional optical‑capture systems usually cost hundreds‑of‑thousands or millions‑of‑yuan.
Large‑space requirements: Legacy optical‑tracking systems frequently demand capture‑zones of 50 m² or above, raising rental‑cost barriers for small teams.
Poor portability: Conventional cameras are permanently wall‑mounted, requiring complete disassembly, transportation and recalibration for off‑site shoots.
Steep‑learning curve: Operation traditionally demands dedicated, highly‑trained motion‑capture technicians.
To break through these limitations, CHINGMU Vision migrated its camera hardware from FPGA to ARM‑based architecture, becoming the first motion‑capture vendor worldwide to deploy ARM‑powered professional tracking‑cameras. This hardware‑upgrade drastically cut equipment‑costs down to the ten‑thousand‑yuan tier. New‑optimized algorithms enable high‑precision tracking within compact 3 m × 3 m ~ 5 m × 5 m indoor‑spaces. Remote real‑time data‑transmission supports cross‑location collaborative performances between geographically‑separated teams.
Drawing‑on years‑of‑experience operating public VR‑attractions, CHINGMU Vision prioritized simplified, beginner‑friendly software operation.
“Some professional clients once commented our software interface felt ‘too simple, almost non‑professional‑looking’. But that easy‑to‑use, one‑click workflow is exactly our competitive advantage.”
The newly‑launched Prometheus system cuts enterprise‑grade virtual‑idol motion‑capture operating‑costs by roughly 90 %, delivering capture‑quality matching mainstream commercial virtual‑idol productions within small‑space environments.

(CHINGMU Vision’s ten‑thousand‑yuan‑level optical‑motion‑capture camera X1)
Nevertheless, marker‑free AI motion‑capture remains a more budget‑friendly starting‑point for independent VTubers, since even affordable optical‑tracking systems carry a five‑figure price‑tag. CHINGMU Vision’s Prometheus product‑line is primarily targeted toward enterprise‑scale virtual‑idol teams, given the relatively‑low success‑rate of independent VTuber projects in China.
“Competition within the virtual‑creator industry continues to intensify. Creators aiming for long‑term commercial success will eventually need comprehensive, high‑quality technical support.”
Looking toward metaverse‑development, Zhang Haiwei offered a vivid metaphor: “Motion‑capture is the key that unlocks the metaverse. It transfers our real‑world movements, facial‑expressions and emotions into our digital‑avatars. Someday, if motion‑capture evolves into full‑blown emotion‑capture technology, users will experience far‑fewer limitations inside virtual worlds, and science‑fiction‑style metaverse‑scenarios similar to Sword Art Online or Ready Player One may finally become reality.”

When asked to share entertaining behind‑the‑scenes stories from virtual‑idol collaborative projects, Zhang Haiwei paused for a moment.
“Honestly, no fun anecdotes immediately spring to‑mind. Every performer works incredibly hard, spending long hours on repeated rehearsals and live‑shows. Everyone in this industry remains down‑to‑earth and fully committed to their craft.”
Beneath the glamorous digital‑avatars admired by millions, virtual‑idol performers lead lives much like everyone‑else.
As the metaverse wave sweeps across industries, many stakeholders have grasped the key to virtual worlds —‑ yet countless doors still remain locked.
CHINGMU Joins OpenLET Community to Co‑Build an Open Ecosystem for Human Motion Data
On August 27 CHINGMU Shanghai Vision Technolog...

