Comparative Study of Static and Contextual Text Vectorization for Sentiment Analysis

A. Bhargavi · International Journal for Research in Applied Science and Engineering Technology · 2025

Sentiment analysis, a core task in Natural Language Processing (NLP), relies heavily on effective text representation techniques to capture semantic and syntactic nuances. This study presents a comparative analysis of widely-used vectorization methods—Bag of Words (BoW), Term Frequency–Inverse Document Frequency (TF-IDF), Word2Vec, GloVe, BERT, and RoBERTa—in the context of sentiment classification. Using the IMDb movie reviews dataset, each method is evaluated based on classification performance, using accuracy and F1-score as primary metrics. Results demonstrate that while deep contextual embeddings such as BERT and RoBERTa achieve the highest accuracy—RoBERTa in particular offering enhanced contextual sensitivity—simpler representations like TF-IDF provide competitive results with significantly lower computational overhead. The findings highlight the trade-offs between accuracy and efficiency, offering practical guidance for embedding selection in sentiment analysis applications.

Read the paper · More papers on PaperTik