Multi-word Prediction for Legal English Context: A Study of Abbreviated Codes for Legal English Text Production

Pertti Väyrynen, Kai Noponen, Tapio Seppänen · 2008

Here, we investigate using strings of expressions longer than a single orthographic word in English word prediction in the legal English domain. The goal of the kind of prediction strategy, called multi-word prediction, is to speed up performance of humans in text production by means of word prediction. Accuracy of two prediction techniques was preliminarily estimated on a simulation without using human subjects with the lexicon of 7,009 multi-word units of legal English. The results show that the average of 70% of characters can be saved for the units in the lexicon in the best-case performance. An improvement in performance actually gained with a real text mainly depends on length and token frequency of units predicted. We also show how the length of multiword units predicted appear to be related to the code lengths used in their prediction and how this finding can be utilized to practical advantage in multi-word prediction.

Read the paper · More papers on PaperTik