Detecting Scenes in Lifelog Videos based on Probabilistic Models of Audio data
Kiichiro Yamano, Katunobu Itou · The Journal of the Acoustical Society of America · 2008
Lifelog videos are recorded every activity in everyday lives. To utilize them efficiently, it is required to be indexed automatically. To index continuous shots of the lifelog, significant scenes are detected automatically. For detection, many researches employ image features such as color and edge, however, the accuracy is insufficient. In this study, we propose probabilistic models for scene detection from lifelog video. In this method, mel-frequency filter bank output of audio tracks of the lifelog videos is modeled statistically with hand-labeled training data. We tested the proposed method to use train station scene. We collected 11 hour sound data for such scenes. To analyze them, we defined seven categories, such as stopping trains, passing trains, starting trains, waiting, and so on. Our method achieved to 100% for waiting scene and 18.4% in average.