How Fully Does a Machine-Usable Dictionary Cover English Text?
Geoffrey R Sampson · Literary and Linguistic Computing · 1989
Computational linguistics applications involving unrestricted natural language need to exploit the resources of information about NL vocabulary contained in ordinary printed dictionaries. I examined the completeness of coverage of one standard English dictionary for which a machine-usable version is widely available by running it over a c 50,000-word cross-section of written English. The article analyses the gaps revealed some, particularly in the area of proper names, are quite predictable, but other biases in coverage, notably against technical vocabulary and against derived forms with negative meanings, were less expected. The algorithms for matching target words against dictionary entries were considerably complicated by the attempt to provide for variant usage with respect to capitalization, hyphenation, and diacritics. A machine-usable dictionary could be turned into a more convenient tool by reducing all headwords to a base orthographic form lacking these features, and instead indicating them by codes in the body of an entry