Towards Good Practice for Action Recognition with Spatiotemporal 3D Convolutions

Kensho Hara, Hirokatsu Kataoka, Yutaka Satoh · 2018

The purpose of this study is to explore good practice for training convolutional neural networks (CNNs) with spatiotemporal three-dimensional (3D) kernels. Recently, 3D CNNs in the field of action recognition are rapidly developed, and the performance levels of them have improved significantly. However, to date, conventional research has mainly focused on their architecture, and has not sufficiently explored their training configurations. We conduct various experiments with different training configurations on Kinetics, UCF-101, and HMDB-51 datasets to share the knowledge of 3D CNNs for the research community. According to the results of those experiments, the following conclusions could be obtained. (i) Data augmentation by spatiotemporal random cropping improved the performance levels. (ii) Data augmentation by multi-scale spatial cropping increased the accuracies in most cases whereas multi-scale temporal cropping decreased them. (iii) A corner cropping strategy, which is previously shown as a good method for two-stream 2D CNNs, resulted lower accuracies for 3D CNNs compared with simple random cropping. (iv) Freezing early layers of 3D CNNs improved the performance levels when fine-tuning 3D CNNs on a relatively small dataset.

Read the paper · More papers on PaperTik