Sequence-Aware Learnable Sparse Mask for Frame-Selectable End-to-End Dense Video Captioning for IoT Smart Cameras
Syu-Huei Huang, Ching-Hu Lu · IEEE Internet of Things Journal · 2023
In recent years, Artificial Intelligence of Things (AIoT) has been widely adopted across various smart systems, accelerating the development of edge computing. Nevertheless, existing research on end-to-end dense video captioning falls short in one of two areas: 1) either it does not prioritize global information or 2) it tends to focus on irrelevant details. Our study proposes an end-to-end dense video captioning model with sequence-aware learnable sparse mask. This model results in improved focus on essential information in a video while ignoring irrelevant details, thus enhancing the quality of caption generation. In addition, existing video captioning research which uses all input video frames are frequently hampered by redundancy and thus generate incorrect captions. To overcome this issue, we propose a lightweight frame selection model that primarily utilizes our proposed lightweight attention-enhancement residual gated network to achieve the desired accuracy with a smaller computational cost. The effectiveness of our proposed approaches was tested and compared to existing models. Our model achieved higher accuracy compared to previous studies, and the lightweight frame selection network resulted in higher efficiency while generating more accurate captions after frame selection.