Language variety identification in Spanish tweets
Wolfgang Maier, Carlos Gómez‐Rodríguez · 2014
We study the problem of language variant identification, approximated by the problem of labeling tweets from Spanish speaking countries by the country from which they were posted.While this task is closely related to "pure" language identification, it comes with additional complications.We build a balanced collection of tweets and apply techniques from language modeling.A simplified version of the task is also solved by human test subjects, who are outperformed by the automatic classification.Our best automatic system achieves an overall F-score of 67.7% on 5-class classification.