Simple2In1: A Simple Method for Fusing Two Sequences from Different Captioning Systems into One Sequence for a Small-scale Thai Dataset
Wuttinan Longjaroen, Thodsaporn Chay-intr, Kotaro Funakoshi, Ananlada Chotimongkol, Sasiporn Usanavasin · 2023
The increasing number of Deaf and Hard of Hearing (DHH) individuals has amplified the need for quality captions, especially in real-time. Previous studies have successfully explored various methods to enhance caption quality, including alignment-based and neural network approaches. While alignment-based methods represent a conventional approach, neural network models are inherently more complex yet offer superior performance. However, these neural models demand large datasets, posing challenges for languages such as Thai with limited datasets available. In this paper, we propose a simple but effective method to improve the quality of Thai captions on a small-scale Thai dataset. Our method utilizes a pre-trained mT5 model to generate a single sequence from two sequences from two different captioning systems. Despite a small dataset and limited input sequences, our method shows potential in improving six evaluation metrics, surpassing all baseline models.