Self-Supervised Multi-Task Learning Using StridedConv3D-VideoSwin for Video Anomaly Detection
Rahman Indra Kesuma, Bambang Riyanto Trilaksono, Masayu Leylia Khodra · 2024
Video Anomaly Detection (VAD) is the problem of identifying anomalous events in the temporal space from the given video data. The studies in VAD domain are still challenging because the definitions of anomalies are quite varied, the data limitations to describe anomalous events, and the anomalous data tends to be relatively close to normal data. The limitations of labeled data have made several studies focus learning only on dominant data, normal events, which is often known as the One-Class Classification (OCC). Self-Supervised Multi-Task Learning (SSMTL), one of the architectures in OCC, still uses a pure CNN-based encoder, which does not maximize the global information usage in feature extraction from each data input. Meanwhile, in the OCC development, many studies have utilized transformers which have better effectiveness, especially in adversarial attacks. On the other side, CNN has advantages regarding considering inductive bias in learning local spatial context. Therefore, this study aims to use Video Shift-Window Transformer (VideoSwin) architecture for video data feature extraction by implementing 3D Convolutional Neural Net (Conv3D) in video patch embedding, StridedConv3D-VideoSwin. The application of StridedConv3d-VideoSwin in SSMTL provides performance improvement for predicting anomaly scores on the pedestrian activity domain. These results consider the detection with Macro and Micro AUC metrics on the CUHK Avenue, UCSD Ped2, and ShanghaiTech datasets. Furthermore, it provides less memory usage even though the floating point operations increases by 74.42%.