A Word2Vec-Based Machine Learning Framework for Binary Sentiment Classification
Mohit Kharbanda, Arshi Husain, Virendra P. Vishwakarma · Zenodo (CERN European Organization for Nuclear Research) · 2026
Automatically judging whether a piece of text expresses a positive or negative opinion– sentiment analysis– remains a core problem in natural language processing, and movie reviews are a widely used testbed for it. This paper builds a sentiment classification pipeline around Word2Vec embeddings and evaluates it on a corpus of 40,000 labelled movie reviews spanning both sentiment classes. Reviews are first cleaned, tokenized, filtered for stop-words, and lemmatized; the resulting tokens are then mapped into 100-dimensional Word2Vec vectors that serve as numerical input to the classifiers. Using a stratified 80:20 split (32,000 training instances and 8,000 held-out test instances), four classical supervised learners– Support Vector Machine (SVM), Random Forest (RF), K-Nearest Neighbors (KNN), and XGBoost– are trained and compared on accuracy, precision, recall, F1-score, and ROC-AUC. SVM comes out ahead of the other three, reaching 85.51% accuracy, 85.38% precision, 85.66% recall, an F1-score of 85.52%, and a ROC-AUC of 93.04%. XGBoost is the runner-up at 82.99% accuracy, with Random Forest close behind at 82.50% and KNN trailing at 79.66%. These results indicate that a Word2Vec-plus-SVM pipeline is a strong, computationally light baseline for binary movie-review sentiment classification and a reasonable starting point for future work on richer text representations.