Content-centric age and gender profiling Notebook for PAN at CLEF 2013
Wee-Yong Lim, Jonathan Wee Pin Goh, Vrizlynn L. L. Thing · 2013
Abstract Author profiling can be considered a form of text analysis of which the objective is to ascertain characteristics of the author behind a text sample. This paper describe the design and implementation of an approach for determining the age group (10s, 20s, or 30s) and gender (male/female) of text samples for the author profiling task in PAN 2013. Evaluation is then based on the compounded accuracy in determining the correct age group and gender of authors of samples in a test corpus. The training corpus provided for this task contains English and Spanish text samples from online contents (e.g. blogs, chats) of authors. Content in each sample are split into one or more “conversations”, of which are all wholly attributed to a specific author. To the best of our knowledge, interweaving re-sponses of other person(s)(if any) are filtered, focussing the scope of the analysis to the writing style and content present in individual author’s sample. The under-lying research in this work is, thus, the empirical investigation of features that can be extracted from the text samples, that are helpful in identifying the gender and age group of an author based purely on characteristics present within his/her text samples. Main contribution in this work is a concise content-based feature based on sim-ilarity scores between given text samples and corpora of the different classes. This feature is compared and used with some common style-based, vocabulary and idiosyncrasies features. Results from experiments on a balanced subset of the PAN 2013 authorship profiling training corpus paint a clear contrast between the content-based feature and the other features, favouring the former for both the English and Spanish samples. Ultimately, 24 five-fold cross validation tests were ran on the different feature sets on the balanced corpus, with the best accuracies for simultaneous gender and age group classification at 48.52 % and 61.23 % for the English and Spanish samples respectively, in contrast to a baseline of 16.67%. 1