Statistical Analysis of Varieties of English
Christopher F. H. Nam, Sach Mukherjee, Marco Schilk, Joybrato Mukherjee · Journal of the Royal Statistical Society Series A (Statistics in Society) · 2012
Summary Linguistic corpora are databases of text which are linguistically marked up or otherwise structured and designed to be representative of a specific language. The growing availability of such corpora has brought with it opportunities for statistical analysis. The paper develops and uses statistical approaches to address questions pertaining to an important linguistic phenomenon: the use of different syntactic alternatives. We present a model-selection-based approach for determining possible driving attributes affecting verb complementation for written sentence constructions using the verb ‘give’ in three varieties of English. We are interested in explaining the choice of alternatives in terms of a variety of sentence level linguistic features such as the meaning of the verb, in addition to the country of origin.