Statistical Representation of Grammaticality Judgements: the Limits of N-Gram Models

Alexander Clark, Gianluca Giorgolo, Shalom Lappin · 2013

We use a set of enriched n-gram models to trackgrammaticality judgements for different sorts ofpassive sentences in English. We construct these models by specifying scoring functions to map thelog probabilities (logprobs) of an n-gram model fora test set of sentences onto scores which dependon properties of the string related to the parametersof the model. We test our models on classificationtasks for different kinds of passive sentences.Our experiments indicate that our n-gram modelsachieve high accuracy in identifying ill-formed passivesin which ill-formedness depends on local relationswithin the n-gram frame, but they are far lesssuccessful in detecting non-local relations that produceunacceptability in other types of passive construction.We take these results to indicate some ofthe strengths and the limitations of word and lexicalclass n-gram models as candidate representations ofspeakers’ grammatical knowledge.

Read the paper · More papers on PaperTik