Automatic Extraction of the Phraseology of a Legal Subdomain

Fabienne Fritzinger, Ulrich Heid, Nadine Siegmund · 2010

Many domains of expertise are rather fragmented, and even if different subdisciplines share the core of a specialised language, their terminology and their phraseology may differ considerably. This situation can also be found in juridical language. There is a core common to most subdisciplines, but each legal domain also has its terminology and phraseology. If this is accepted for terms, the fact that the same differentiation can be found in specialised phraseology, is much less acknowledged. However, if a jurist is supposed to produce a text on a domain that falls outside his specialisation, he may still know the terminology to be used, but likely much less the specialised phraseology of the ‘foreign‘ domain. A good dictionary of juridical language should thus provide clear labels of subdomains wherever possible. In this paper, we address this issue in the context of the extraction of single and multiword terms and of specialised collocations from corpus data. We intend to answer the following methodological and technical research questions: if we use simple frequency-based extractors on juridical texts from different domains, can these provide the phraseology of a given juridical subdomain by contrasting multiwords from different domains? And can we identify phraseological sequences longer than two items by systematically analysing the context of word pairs? We also wish to contribute to a more detailed description of the use of adverbs in juridical phraseology.

Read the paper · More papers on PaperTik