Sentiment analysis using Kernel Variance Projection and LR–BiLGMP–Skd deep learning model

Alaa Abdullah Al-Saadi, Chee‐Onn Chow, Wei Ru Wong, Anis Salwa Mohd Khairuddin · Egyptian Informatics Journal · 2026

Dimensionality reduction (DR) techniques are critical for preprocessing large-scale datasets in tasks such as sentiment analysis (SA). While Kernel Principal Component Analysis (KPCA) captures latent nonlinear structures, it heavily relies on the kernel function and often struggles with large datasets and noise, which may degrade classification performance. This limitation arises because the eigenvectors are derived from the kernel function and are constrained within a restricted solution space. To address these limitations, we propose a novel Kernel Variance Projection (KVP) algorithm that preserves local data structure while extracting informative representations from high-dimensional data through a two-stage process. First, a PCA-based kernel function maps the data into a linear space, enabling incomplete Cholesky decomposition for efficient scalability on large datasets. Second, the resulting eigenvalues and eigenvectors are used to construct an orthogonal projection matrix via the Moore–Penrose pseudoinverse, reducing Word2Vec embeddings from 300 to 10 dimensions and significantly lowering computational complexity. Building on this optimized feature representation, we develop the Learning Rate–BiLGMP–Scheduler (LR–BiLGMP–Skd) model, which enhances regularization in the embedding layer and dynamically adjusts the learning process. By integrating BiLSTM and Global Max Pooling layers, the model complements the proposed KVP algorithm, improving accuracy, F1 score, consistency, regularization, and computational efficiency. Comparative experiments against PCA, KPCA, and UMAP using ConvBiLSTM and LR–BiLGMP–Skd models on Tweets, Sentiment140, and IMDB datasets demonstrate that KVP improves model consistency and regularization while achieving superior classification performance. Across datasets, KVP achieves strong results in accuracy (96% on Tweets, 78.84% on Sentiment140, and 88.39% on IMDB) and F1 score (95.1% on Tweets, 78.32% on Sentiment140, and 88.43% on IMDB) when combined with the LR–BiLGMP–Skd model. Compared with baseline approaches, KVP consistently improves accuracy by approximately 1%–2% and F1 scores by approximately 1%–3%, while maintaining lower loss values (0.15–0.46) and reducing training epochs by approximately 5–7, highlighting its effectiveness in improving both regularization and scalability. The source code for the proposed KVP method is available at: https://github.com/Tigerlion-7/kvp-dr-sentiment-analysis.git .

Read the paper · More papers on PaperTik