Automatic isolation of region of interest for multi-pose Audio Visual Speech Recognition
Amarsinh Bhimrao Varpe, Prashant Borde, Ramesh R. Manza, Pravin Laxmikant Yannawar · 2014
Automatic Speech Recognition (ASR) by machine has attracted many researchers to design robust recognition system. Latest advances in speech recognition by incorporation of visual cues for enhancement of speech gives rise to Audio Visual Speech Recognition system. Multi-pose AVSR is gaining attention due to its robust feature extraction capability. This process is contributing in design of robust AVSR. This paper presents effect of lip-color based automatic isolation of region-of-interest over `Viola-Jones' algorithm for multi-pose audio visual speech recognition system. The `Viola-Jones' algorithm was widely used for detection of face components (eyes, nose and mouth) and offers accurate face detection for full frontal visual stream whereas lip-color based isolation scheme was applicable for all visual streams. The efficiency of method in detection of ROI for multi-pose AVSR is measured as 95.8% for full-frontal streams, 91.8% for 45° streams and 63.2% for side face streams.