Home>Support&Download>Company Updates
Resources and Downloads
Company Updates

Talent‑behind‑the‑avatar and the virtual model cannot share the same screen, yet real‑human and virtual‑human co‑streaming is totally achievable! / Company Updates

2022/09/23



A so‑called “dox‑reveal” video posted by Bilibili creator SwP_Forever finally brought him his first video with over 10 000 views. Revealing the real‑life talent behind a virtual avatar is generally regarded as a major taboo among VTubers. SwP_Forever has also stressed that he will avoid blurring the boundary between his real‑world identity and his virtual persona, since maintaining an independent, pure‑image avatar is what wins VTubers affection from audiences.




Follow: Bilibili ID‑‑SwP_Forever


Though talents seldom appear on‑camera during standard VTuber streams, live broadcasts featuring real‑human and virtual‑avatar on‑stage interaction are growing increasingly popular. This streaming format helps virtual idols and VTubers break through their original community boundaries and reach wider audiences.

Professional real‑time mixed‑reality live streaming relies on complete workflows and high‑end equipment. The manpower, hardware and financial investment for one single high‑grade live event can easily reach nearly one million yuan, not including avatar‑model production fees.


CHINGMU Vision Virtual‑Streaming Motion‑Capture Studio


Avatar‑image design comes first. Two mainstream avatar styles dominate today’s market: anime‑style avatars and hyper‑realistic high‑fidelity digital humans, whose costs range widely from several hundred to hundreds of thousands of yuan. The core factor behind this price gap lies in realism requirements. Hyper‑detailed avatars with skin pores, fine vellus hair and rich textures on costumes and headdresses inevitably come with a high price tag.


Anime‑Style Virtual Girl Group: SiXi Maruko




6380584450459595273775936

Hyper‑realistic Digital‑Human Created for Durex & Volvo Valentine’s Day Campaign

Careful viewers may notice that ultra‑high‑precision digital‑human avatars often show reduced realism during live‑streaming compared with pre‑recorded videos and posters. This limitation stems from current rendering hardware and software capabilities. Rendering a million‑polygon avatar together with elaborate virtual environments places huge demands on server performance. Most live‑streaming‑ready digital‑human models today therefore contain between 300 000 and 500 000 polygons.

Once 3D modelling is finished, the avatar requires rigging, physics‑simulation solving and motion‑capture data input before it can move naturally. Skeleton rigging largely determines whether the avatar will suffer from mesh clipping during movement. Physics‑solvers control the natural swing of hair, clothing and fabrics. Motion capture, meanwhile, acts as the soul that brings digital‑human characters to life.


Maya Skeleton‑Rigging Work


A digital human is far more than a virtual shell. The behind‑the‑scenes talent endows it with voice, tone, movements and personality traits. Motion‑capture hardware serves as the essential medium connecting real‑world performers and their virtual avatars.

Optical motion‑capture cameras are permanently installed inside dedicated streaming studios. CHINGMU Vision’s professional virtual‑broadcast studio is equipped with 30 MC1300 optical motion‑capture cameras, supporting multi‑performer live‑stream sessions. Motion‑capture data is processed by a dedicated capture PC; an Intel i7 processor is generally sufficient for this workload.


CHINGMU Vision Virtual‑Streaming Studio —‑ MC1300 Optical Motion‑Capture Cameras


Apart from the motion‑capture workstation, a separate rendering computer is required. Facial‑capture, hand‑tracking and body‑motion data are transmitted via a network switch to the rendering engine for character and scene visualization. Popular software includes Unreal Engine and Unity. It is recommended that rendering workstations be fitted with an Intel i9 or higher‑grade CPU and an RTX 3080 Ti graphics card.

To composite real‑world talent footage into a virtual environment for real‑human‑virtual‑avatar co‑broadcasts, real‑time chroma‑keying technology is required. Hardware requirements include a green‑screen backdrop, professional studio lighting and a physical camera for filming the live‑action performer. Studio lighting not only illuminates the talent, but also ensures even brightness across the green‑screen background to deliver cleaner keying results. Chroma‑key compositing can be completed within Unreal Engine. Operators then match focal length and camera angles between virtual and physical cameras to align depth and spatial positioning, creating the illusion that the real‑life performer and digital avatar occupy the same virtual space.

At this stage, the foundational scene environment for real‑human‑virtual‑avatar streaming is ready. Adding an audio‑capture sound‑card completes the setup, and the stream can go live. Professional multi‑shot broadcasts benefit greatly from a video switcher, which connects the rendering PC, audio‑card, live‑action camera and streaming‑server. Chroma‑key processing can be performed directly on‑board the switcher, and broadcast signals can be split to push live feeds simultaneously to multiple streaming platforms.


Virtual‑Live Streaming Workflow Diagram


Real‑human‑virtual‑avatar live‑stream production covers numerous technical links and requires diverse hardware. Only a small number of service providers are capable of integrating the full industrial workflow. As an industry pioneer in virtual‑broadcast solutions, CHINGMU Vision delivers full‑stack virtual‑live‑streaming services. Join CHINGMU’s official community for further information on virtual‑streaming projects.


Copyright © 2025 SHANGHAI CHINGMU VISION TECHNOLOGY CO., LTD
沪公网安备31010602005676号 沪ICP备16005979号-1 Site Map