Semantic segmentation of videophone image sequences
Peter J. L. van Beek, Marcel J. T. Reinders, Bülent Sankur, Jan C. A. van der Lubbe · Proceedings of SPIE, the International Society for Optical Engineering/Proceedings of SPIE · 1992
A system for segmentation of head-and-shoulder scenes into semantic regions, to be applied in a model-based coding scheme on video telephony, is described. The system is conceptually divided into three levels of processing and uses successive semantic regions of interest to locate the speaker, the face and the eyes automatically. Once candidate regions have been obtained by the low level segmentation modules, higher level modules perform measurements on these regions and compare these with expected values to extract the specific region searched for. Fuzzy membership functions are used to allow deviations from the expected values. The system is able to locate satisfactorily the facial region and the eye regions.