Visual comparison of statistical feature aggregation methods for video-based similarity applications
Adolfo Almeida, Pieter de Villiers, Allan De Freitas, Mergandran Velayudan · 2020
Video data is increasingly being provided by multiple sources. These sources are used to obtain reliable feature information within the context of training data shortages. As such it becomes crucial to interpret and refine such information. Recent research on video content analysis often use deep-learning features due to their outstanding performance in different domains. These features are aggregated over time to create a video-level descriptor. In this research, we explore the potential of statistical feature aggregation methods in combining deep-learning features that represent the visual content of videos. In particular, the contributions of this paper are two-fold: Firstly, we compared statistical feature aggregation methods by selecting movie sequels, calculating their in-sequel-mean-distance and out-of-sequel-mean-distance as well as the Bhattacharyya distance between them and using Principal Component Analysis (PCA) and T-distributed Stochastic Neighbour Embedding (t-SNE) to visualise their distribution in the feature space. This is important to easily comprehend the video content without the need to watch them and understand how well the statistical feature aggregation methods represent the videos. Secondly, the performance of these aggregation methods is explored in the context of a content-based video retrieval task which is evaluated in terms of relevance. The observed results show that the statistical feature aggregation method based on variance outperforms the methods based on maximum, mean, median, median absolute deviation and interquartile range. This outcome is supported when interpreting the t-SNE visualisation method as well as Bhattacharyya distance calculations. Overall statistical feature aggregation methods which measure the spread of a distribution achieve better performance compared to methods that are a measure of location.