Gender Classification on Twitter Based on Feeds and User Descriptions Using Xlnet-Fasttext
Vito Rozaan Alandeta, Derwin Suhartono · Informatica · 2024
Gender falsification in social media content is an increasingly troubling challenge, with users often choosing to hide their true gender identity or pretend to be members of a different gender. This can lead to negative consequences, including the spread of disinformation, discrimination and online security risks. To overcome this problem, this research proposes a text classification-based solution to identify gender fakes in social media texts. This method involves extracting linguistic features from texts, such as word usage, sentence structure, and language patterns that can provide clues to the author's gender. Therefore, this research aims to introduce a new transformers-based approach that uses XLNet and is also modified with additional Fasttext embedding. Modifications were made to the embedding section which can increase XLNet's understanding of text context in carrying out text classification. The results of this research are that baseline XLNet gets a fairly good performance score in gender classification based on Twitter feeds, namely with accuracy, precision, recall and f1-score of 0.704, 0.770, 0.598, 0.674 respectively, while XLNet-FastText gets the respective scores. -respectively 0.714, 0.770, 0.609, 0.680. And for gender classification based on user account descriptions, baseline XLNet gets scores of accuracy, precision, recall, f1-score of 0.705, 0.771, 0.598, 0.674 respectively while XLNet-FastText gets scores of 0.724, 0.751, 0.6324, 0.686 respectively.