Classifying English Documents by National Dialect
Marco Lui, Paul F. Cook · 2013
We investigate national dialect identifica-tion, the task of classifying English doc-uments according to their country of ori-gin. We use corpora of known national origin as a proxy for national dialect. In order to identify general (as opposed to corpus-specific) characteristics of national dialects of English, we make use of a va-riety of corpora of different sources, with inter-corpus variation in length, topic and register. The central intuition is that fea-tures that are predictive of national ori-gin across different data sources are fea-tures that characterize a national dialect. We examine a number of classification ap-proaches motivated by different areas of research, and evaluate the performance of each method across 3 national dialects: Australian, British, and Canadian English. Our results demonstrate that there are lex-ical and syntactic characteristics of each national dialect that are consistent across data sources. 1