SMCA: Movie Trailer Audience-Suitability Rating Classification using Staged Multimodal Cross-Attention
Kézia Victória Galdino de Castro, Carl Mitzchel Padua, Edjin Jerney H. Payumo, Nathaniel David P. Samonte, Jessie James P. Suarez · 2025
Movie trailers play a pivotal role in capturing audience attention and shaping movie expectations. Given their potential to expose young viewers to inappropriate content, the Motion Picture Association’s Classification and Rating Administration (CARA) assigns independent ratings to trailers, separate from the movies they promote. This study proposes a Staged Multimodal Cross-Attention (SMCA) framework to enhance the classification of movie trailers into green-band (suitable for general audiences) or red-band (restricted for younger audiences). The approach addresses the limitations of prior models, such as the Gated Multimodal Unit (GMU), by improving cross-modal alignment between text, audio, and visual features.The proposed system extracts features using pre-trained models: BERT for text, Vision Transformer (ViT) for video frames, and log-mel spectrograms processed with ViT for audio. These features are combined through a two-stage cross-attention mechanism, facilitating targeted interactions between modality pairs before integrating them into a unified representation. A multilayer perceptron (MLP) classifier predicts the final trailer rating.Experiments on the Multimodal Movie Trailer Dataset (MMTrailer) demonstrate the SMCA’s superior performance. The A-V-T configuration (audio-visual-text fusion) achieved an average accuracy of 83.02%, F1-score of 65.80%, and precision of 68.06%, surpassing GMU and other baseline models. The T-VA configuration (text-visual-audio) yielded the highest accuracy (84.68%) and precision (77.18%). Notably, the T-V-A configuration exhibited the lowest standard deviation for precision, indicating stable performance. Error analysis revealed that misclassifications often stemmed from nuanced tone shifts or misalignment of thematic content. Overall, the SMCA framework offers an effective and robust solution for the nuanced task of movie trailer classification, demonstrating improvements in accuracy, precision, and cross-modal interaction stability.