Extracting Gender Textual Nuances Using Text Similarity for Gender Classification Improvement

Haneen Tamim Abd Ali, Dhamyaa A. Nasrawi · 2024

Businesses can target various customer classifications with customized marketing efforts and product offers for better customer satisfaction and participation by using gender classification in the text. This study’s goal is to use text to classify humans into male and female groups based on gender. This study employs frequency and similarity methods to evaluate the extraction of gender-specific features from English texts on using one-topic TripAdvisor dataset and multitopic Twitter dataset. The features have been organized into an array for implementation by using machine learning classifiers like Support Vector Machine, Random Forest, and Logistic Regression The investigation looks to developments in language use that are unique to gender. Comparing the results of the study to those of previous research, the Random Forest classifier yielded the highest accuracy of 87.7% on multitopic Twitter dataset, whereas logistic regression yielded the best accuracy of 74.5% on one-topic TripAdvisor dataset. The study’s findings indicate that the highest level of accuracy was achieved with a multitopic dataset containing a diverse range of words, as compared to a one-topic dataset, which yielded fewer results.

Read the paper · More papers on PaperTik