Writing Style Author Embedding Evaluation

Enzo Terreau, Antoine Gourru, Julien Velcin · 2021

Learning authors representations from their textual productions is now widely used to solve multiple downstream tasks, such as classification, link prediction or user recommendation.Author embedding methods are often built on top of either Doc2Vec (Le and Mikolov, 2014) or the Transformer architecture (Devlin et al., 2019).Evaluating the quality of these embeddings and what they capture is a difficult task.Most articles use either classification accuracy or authorship attribution, which does not clearly measure the quality of the representation space, if it really captures what it has been built for.In this paper, we propose a novel evaluation framework of author embedding methods based on the writing style.It allows to quantify if the embedding space effectively captures a set of stylistic features, chosen to be the best proxy of an author writing style.This approach gives less importance to the topics conveyed by the documents.It turns out that recent models are mostly driven by the inner semantic of authors' production.They are outperformed by simple baselines, based on state-of-the-art pretrained sentence embedding models, on several linguistic axes.These baselines can grasp complex linguistic phenomena and writing style more efficiently, paving the way for designing new style-driven author embedding models.(Goodread corpus).A simple classification model (e.g., SVM, MLP) is trained to predict the class for each author through its embedding.The accuracy score then allows to compare each method.Maharjan et al. (2019) use authorship attribution to evaluate the quality of their model.Authorship attribution consists in predicting the author of a given document.It requires that document and author representations lie in the same space.It is performed either by clustering or simply by computing the cosine similarity between a document embedding and each author's embeddings to get either an accuracy score or a coverage error.This task could be a reference to evaluate author embedding.Being able to perfectly associate an author with its production ensures that the method efficiently captures each author's writing habits and characteristics.However, one of the biggest issues of authorship attribution is the lack of interpretability.It fails to reveal if a given author embedding method is more based on the content/topics or on the author writing style.To fully understand the distinction between writing style and content we can mention the book Exercises in Style of French author Raymond Queneau, who wrote the same story 99 times, but in 99 different ways (see Table 1).Although the story which is told remains the same, the choice of words and complexity of each sentence strongly differs.A way to get around this issue is to evaluate one's method on at least two datasets with various profiles.Sari et al. (2018) show that the decisive features to discriminate the author of a document can either be topic based or style based, depending on the dataset under study.Using at least two datasets when evaluating author embedding methods is a good step to better understand the model capacities.

Read the paper · More papers on PaperTik