A hybrid approach to the identification and expansion of abbreviations

Janine Toole · 2000

This paper introduces a two-stage system for identifying and expanding abbreviations. It is based on a hybrid architecture where rule-based and statistical methods are combined. The first task of the system is to differentiate abbreviations from other types of unknown words such as names and misspellings. The second task of the system is to identify the intended complete word. The system is evaluated using data from the Air Safety Reporting System (ASRS) database: a domain where document retrieval is directly impacted by the large amount of unknown words, of which abbreviations are a frequent class. Introduction Natural language text is not always ideal for information retrieval (IR). Many information-rich documents contain misspellings, abbreviations, and other misleading variants of the key words that are necessary for quality information retrieval. For example, retrieval of records from the Air Safety Reporting System (ASRS) database is complicated by the fact that approximately e...

Read the paper · More papers on PaperTik