VRA-L: A real-time system for generating expressive avatar animations and object detection based on RGB cameras and monocular videos
Zhihao Lai, Qingbei Zhao · 2024
To achieve low-cost, markerless motion capture, the system uses MediaPipe for pose estimation and Kalidokit for real-time 3D character animation. It employs a key-value pair mapping technique for versatile skeleton control and camera simulation for user interaction. The system integrates Graph Convolutional Networks (GCN) for feature learning and Long Short-Term Memory (LSTM) networks for processing temporal sequences. By extracting key skeletal points from hands, face, and pose, and combining them with GCN and LSTM, it accurately captures human poses. On an RTX 4070 device, it showed a 30% reduction in acceleration error and a 65% reduction in network parameters and computational load, achieving 55-60 frames per second for 3-Dimension pose estimation. Spherical linear interpolation is applied to enhance motion smoothness, providing a good user experience with real-time action recognition and display.