A Hybrid Shifted Inverted Euclidean-Linear Kernel with Stacked Classifiers for Tamil Slang Classification

Ramkumar. R, Sureshkumar Nagarajan · 2024

This study introduces a novel SVM-based method for classifying Tamil slang using a Hybrid Shifted Inverted Euclidean-Linear Kernel. The kernel integrates Eigen value shifting and clipping to ensure stability and positive definiteness that combines distance-based similarity with regularization techniques. Applied to voice samples from four Tamil dialects such as Chennai, Kovai, Madurai and Nellai, features like Mel Frequency Cepstral Coefficients, Pitch, and Energy were extracted and fused. The hybrid approach, enhanced with a stacked classifier system, outperforms traditional kernel functions (Linear, Radial Basis Function, polynomial, sigmoid) and machine learning algorithms (K Nearest Neighbor, Logistic Regression, Naïve Bayes, Decision Tree, Random Forest). Results show that the Combined Kernel achieves the highest accuracy, balanced precision, and recall, offering a scalable method for dialect classification with applications in speech recognition and linguistic analysis.

Read the paper · More papers on PaperTik