Violent Scenes Detection with Large, Brute-forced Acoustic and Visual Feature Sets

Florian Eyben, Felix Johannes Weninger, Nicolas H. Lehment, Gerhard Rigoll, Björn Wolfgang Schuller · mediaTUM – the media and publications repository of the Technical University Munich (Technical University Munich) · 2012

This paper describes the TUM approaches for violent scenes detection in movies, submitted for the MediaEval 2012 Affect Challenge.Score fusion is used to fuse Support-Vector Machine (SVM) confidence scores assigned to short fixed length windows within each movie shot.SVM predictors for acoustic and visual channels are trained.For the acoustic channel, a large set of acoustic features based on the set from the INTERSPEECH 2012 Speaker Trait Challenge is employed.A comprehensive set of common video low-level descriptors such as optical flow, gradients, and hue and saturation histograms is used for the visual channel.

Read the paper · More papers on PaperTik