Matrix:a statistical method and software tool for linguistic analysis through corpus comparison

Paul Rayson · Lancaster EPrints (Lancaster University) · 2003

This thesis reports the development of a new kind of method and tool (Matrix) for advancing the statistical analysis of electronic corpora of linguistic data.First, we describe the standard corpus linguistic methodology, which is hypothesis-driven.The standard research process model is 'question -build -annotate -retrieve -interpret', in other words, identifying the research question (and the linguistic features) early in the study.In recent years corpora have been increasingly annotated with linguistic information.From our survey, we find that no tools are available which are datadriven on annotated corpora, in other words, a tool which assists in finding candidate research questions.However, Matrix is such a tool.It allows the macroscopic analysis (the study of the characteristics of whole texts or varieties of language) to inform the microscopic level (focussing on the use of a particular linguistic feature) as to which linguistic features should be investigated further.By integrating part-of-speech tagging and lexical semantic tagging in a profiling tool, the Matrix technique extends the keywords procedure to produce key grammatical categories and key concepts.It has been shown to be applicable in the comparison of UK 2001 general election manifestos of the Labour and Liberal Democratic parties, vocabulary studies in sociolinguistics, studies of language learners, information extraction and content analysis.Currently, it has been tested on restricted levels of annotation and only on English language data.This thesis describes a method of statistical profiling within corpus linguistics.It focuses on a method (called Matrix, together with a piece of software implementing this method 2 ) that was developed by the author and has already been used in both academic and commercial contexts.The Matrix software is principally a tool for use with annotated corpora and has the novel ability to perform statistical comparisons of corpora at multiple levels of annotation, including the lexical level.The frequency profile is the first port of call when investigating corpora and leads on to other research activities such as concordancing, and collocation analysis.These tasks can be applied to aid investigation and understanding of bodies of text in areas such as language teaching, linguistic research, content analysis, software engineering, machine translation and lexicography.The next three sections of this first chapter introduce the notion of statistical profiling and annotation within the field of corpus linguistics.We then outline the objectives of the study.The final section describes the structure and content of the remaining chapters of this thesis. Corpus linguisticsA corpus is defined in the Concise Oxford English Dictionary as a 'body, collection of writings'.Aston and Burnard (1998: 4) note that the second edition of the Oxford English Dictionary lists five distinct senses for the word.Only two of these particularly refer to language.However, preliminary standards guidelines have distinguished between the terms corpus and collection or archive, of which only corpus is related to some linguistic purpose (Sinclair, 1996).The most commonly agreed upon plural of corpus is corpora.3 There is no accepted minimum or maximum size for a corpus, or specification of what it should contain.A corpus could contain 2 The choice of the name Matrix relates to the appearance of the output of the tool: a matrix in mathematics is a rectangular array of elements set out in rows and columns.No link is intended to the film starring Keanu Reeves.3 The frequency and acceptability of other plural forms of the word corpus (e.g.corpuses) have been much debated on the CORPORA electronic mailing list.Aston and Burnard (1998: 63-73) devote ten pages to the question.Frequency-sorted word lists have long been part of the standard methodology for exploiting corpora.Sinclair (1991: 30) writes, "Anyone studying a text is likely to need to know how often each different word form occurs in it".Tribble and Jones (1997: 36) outline a pedagogical methodology for using texts in the language classroom, proposing that the most effective starting point for understanding a text is a frequency-sorted word list.The frequency list records the number of times each word occurs in the text; it can provide interesting information about the words that appear (or do not appear) in a text.The list can be arranged in order of first occurrence, alphabetically or in frequency order.First-occurrence order serves as a

Read the paper · More papers on PaperTik