Vision and learning for intelligent human-computer interaction
Thomas S. Huang, Ying Nian Wu · 2001
It was a dream to make computers see. The research in computer vision provides promising technologies to capture, analyze, transmit, retrieve and interpret visual information. However, due to the richness and large variations in the visual inputs, the practice of many statistical learning techniques for visual motion capturing and recognition are confronted by some similar problems, such that making intelligent and visually capable machines is still a challenging task. This dissertation concentrates on two important problems: capturing and recognizing human motion in video sequences, which are crucial for the research and applications of intelligent human computer interaction, multimedia communication, and smart environments. This dissertation presents three effective techniques for visual motion analysis tasks: non-stationary color model adaptation for efficient localization, multiple visual cues integration for robust tracking, and learning motion models for capturing articulated hand motion. Besides, this dissertation describes a novel statistical learning method, the Discriminant-EM (D-EM) algorithm, in the framework of self-supervised learning paradigm. D-EM employs both labeled and unlabeled training data and converges supervised and unsupervised learning. Many topics in the dissertation is unified by the four problems of self-supervised learning, i.e., transduction, co-transduction, model transduction and co-inferencing. Extensive experiments and two prototype systems have validated the proposed approaches in the domain of vision-based human computer interaction.