Harris Hawks-Optimized Audio, Structural, and Video Feature Fusion for Multimodal Video Classification Using LSTM-RF Framework
International journal of intelligent engineering and systems · 2025
The exponential growth of multimedia data and the ease of data-sharing have made indexing and searching relevant content in large databases a persistent challenge.In this study, we propose a novel content-based YouTube data classification system designed to address the complexities and diversity of web video content.Given the substantial number of categories and their frequent overlaps (e.g., a home video featuring a sports event), classifying such multimedia data remains a significant research challenge.Our system employs a hybrid feature set combining structural, keypoint-based, video, and audio features.Structural features include Average Shot Length (AvgShotLength) and Shot Temporal Activity, while Extreme Features (X-Feat) are extracted for keypoint-based representation.For audio, this paper proposes a novel Harris Hawks Optimization (HHO) to refine Mel Frequency Cepstral Coefficients (MFCC) and delta coefficients.Video features are enriched with Histogram of Oriented Gradients (HOG), Color Correlogram.Classification is performed using a two-classifier system comprising Long Short-Term Memory (LSTM) networks for temporal data and Random Forests for feature-based decisions, with final predictions determined via majority voting.The proposed system was evaluated on the YouTube dataset, achieving a peak accuracy of 98.53%, demonstrating its effectiveness in managing the diversity and overlap inherent in web video classification.These results indicate that a specific classification framework using specific feature sets, or even specific features, will be more effective for its target genre.