Authorship Attribution of Small Messages Through Language Models

Antônio Theóphilo, Anderson De Rezende Rocha · 2022

Social media platforms brought numerous benefits to our society, but several wrongdoings such as racism, misogyny, hate speech, anti-science statements, and large-scale misinformation came alongside. Authorship attribution is a forensics tool that can help fight against these misconducts when applied to the small texts posted on these platforms. In this work, we exploit the recent developments in language models to tackle the problem of authorship attribution of small messages. Training one model per suspect, we devise a generative approach and compare it against a state-of-the-art discriminative method. Our results show that generative and discriminative features are complementary and can be leveraged to improve the results of current methods. Finally, we propose a strategy to use discriminative and generative models jointly and draw future research paths.

Read the paper · More papers on PaperTik