Personalised video summarisation using video-text multi-modal fusion
Rakhi Akhare, Subhash K. Shinde · International Journal of Computational Vision and Robotics · 2025
Video summarisation techniques have evolved in recent years, mostly focusing on visual material and ignoring user preferences. In this work, the topic of query-focused video summarisation is addressed. Long videos are given as input, and the goal is to produce a query-focused video summary using the user's sentences rather than keywords. The two parts of the proposed personalised video summarisation (PVS) system are the query-relevance computation module and the feature encoding network. In order to provide a customised video summary, the suggested end-to-end approach combines encoded visual and textual information and assigns a query relevance score. The suggested PVS model is tested using the fast-text and Resnet embeddings on the video-query dataset. In comparison to various combinations of language and vision models, the suggested PVS model performs better and achieves an accuracy of 0.53%. This study assists the research community to work in the field of multimodal video summarisation.