Precise finger articulation in sign language video generation
Po-Hsuan Hsiao, Jia-xuan YU, Kuan-Lun Chen, Po-Hsuan Fan Chiang, Jen‐Tzung Chien, Hsin‐Te Wu, Shang-Tse Chen, Jyh‐Shing Roger Jang · Enterprise Information Systems · 2026
Sign language relies on coordinated facial expressions and hand gestures, posing challenges for generative video systems that typically emphasize facial or body realism alone. This paper presents a virtual sign language video generation framework based on Stable Video Diffusion (SVD) that produces full-length signing animations from a single portrait. Temporal consistency is improved via frame refinement, while SignPose and PoseNet enhance motion accuracy and detail. A hand-region enhancement strategy further improves finger reconstruction. Experimental results demonstrate clearer gestures, more natural facial expressions, and realistic, comprehensible virtual sign language videos.