Good reasons for noting bad grammar : empirical investigations into the parsing of ungrammatical written English
Jennifer Foster · Trinity's Access to Research Output (TARA) (Trinity College Dublin) · 2005
This thesis is concerned with the parsing of ungrammatical written English sentences. Over a period of eighteen months, a 20,000 word corpus was developed which consists of ungrammatical sentences which were noticed while reading a variety of English texts. Each sentence in this corpus was corrected, producing a second corpus of grammatical sentences. This thesis argues that the compilation of such a corpus is a useful computational linguistic resource, describes the methodological decisions which were made in compiling the corpus, presents and discusses the results of a small questionnaire study which was used to investigate the reliability of the corpus data, analyses the differences between the ungrammatical and grammatical corpus, and makes use of the corpus in three separate parsing studies.