Statistical Analysis of Varieties of English

Christopher F. H. Nam, Sach Mukherjee, Marco Schilk, Joybrato Mukherjee · Journal of the Royal Statistical Society Series A (Statistics in Society) · 2012

Summary Linguistic corpora are databases of text which are linguistically marked up or otherwise structured and designed to be representative of a specific language. The growing availability of such corpora has brought with it opportunities for statistical analysis. The paper develops and uses statistical approaches to address questions pertaining to an important linguistic phenomenon: the use of different syntactic alternatives. We present a model-selection-based approach for determining possible driving attributes affecting verb complementation for written sentence constructions using the verb ‘give’ in three varieties of English. We are interested in explaining the choice of alternatives in terms of a variety of sentence level linguistic features such as the meaning of the verb, in addition to the country of origin.

Read the paper · More papers on PaperTik