Effectiveness of Automatic Caption Generation Method for Video in Japanese

Masafumi Matsuhara, Jin Tsushima · 2023

In recent years, the surge in the Internet and social networking services (SNS) has boosted video and media use, expanding opportunities for information communication through video. Deep learning has enhanced video caption technology, proposed for describing human actions. Using natural language, machines can interpret data understandably, aiding information sharing between systems and users. Studies using surveillance cameras have been proposed to detect anomalies in daily activity videos, with a focus on anomaly detection rather than leveraging the findings. In this paper, to capture and utilize human actions, we propose a method for automatic video caption generation using the Encoder-Decoder model and Attention in Japanese. We conducted experiments in which we trained videos of everyday actions and generated captions using the trained model in our proposed method. Experimental results show that the proposed model learns correctly and captures the sequential features of the videos. In addition, BERTScore effectiveness in quantitatively evaluating Japanese captions was tested. The effectiveness of BERTScore in evaluating Japanese sentences is proven in our proposed method.

Read the paper · More papers on PaperTik