Long Term Arm and Hand Tracking for Continuous Sign Language TV Broadcasts
Patrick Buehler, Mark Everingham, D.P. Huttenlocher, A. Zisserman · 2008
The goal of this work is to detect hand and arm positions over continuous sign language video sequences of more than one hour in length. We cast the problem as inference in a generative model of the image. Un-der this model, limb detection is expensive due to the very large number of possible configurations each part can assume. We make the following con-tributions to reduce this cost: (i) using efficient sampling from a pictorial structure proposal distribution to obtain reasonable configurations; (ii) iden-tifying a large set of frames where correct configurations can be inferred, and using temporal tracking elsewhere. Results are reported for signing footage with changing background, chal-lenging image conditions, and different signers; and we show that the method is able to identify the true arm and hand locations. The results exceed the state-of-the-art for the length and stability of continuous limb tracking. 1