A Shallow Syntactic Analyser to Extract Word Associations from Corpora
Roberto Basili · Literary and Linguistic Computing · 1992
Many recent studies on lexical acquisition are based on the extraction of word associations from corpora. Often associations are extracted by the cooperative effort of a syntactic and a extracted by the cooperative effort of a syntatic and a statistical processor, to reduce the number of ‘accidental’ associations. Due to computational complexity requirements that impose severe constraints on the completeness of the adopted grammar, the performances of the syntactic analysers adopted in these studies are usually quite poor. In this paper we describe a ‘shallow’ syntactic analyser that exhibits high linguistic performances with respect to the aforementioned objectives, at a reasonable computational cost. The analysis is performed in three steps: First, a (general purpose) morphologic analyser tags words by the appropriate part of speech. Secondly, a segmentation algorithm isolates text units (e.g. phrases). The segmentation algorithm is tuned to the specific sublanguage under examination. Finally, a syntactic analyser extracts binary and ternary relations between words (e.g. subject-verb, noun-preposition-noun, etc. ), called ‘elementary syntactic links’ (esl). The grammar is a discontinuous grammar (DG) where skip rules are used to detect non-adjacent attachments. The analyser has been experimented on two corpora, an economic enterprise database and a legal domain, which exhibit very different linguistic styles. Several performance measures have been applied to the analyser, such as complexity, processing time, accuracy and precision. In evaluating the accuracy and precision, the reference performer is the full set of esl that would be generated by a complete grammar of the target sublanguage, rather than the opinion of a linguist, which is usually taken as a reference in other similar studies.