Extracting and organizing acronyms based on ranking

Weijian Ni, Yalou Huang · 2008

The paper addresses the problem of automatically extracting and organizing acronyms and expansions e.g., ‘ROM’ and ‘Read Only Memory’) in text. To deal with the problem, we propose a two-step approach based on ranking. In the first step, for each occurrence of acronyms in text, we rank the expansion candidates around the acronym and extract the top ranked ones. In the second step, expansion candidates collected in the first step are organized before presented to end users. ‘Organize’ here means grouping expansions and then ranking them according to their correctness and popularity. In this way, the numbers of expansions in the results which users need to examine will be drastically reduced. Experimental results based on real-world dataset show that our approach can always rank correct and popular expansions to the top. Experimental results also show that the trained ranking models are generic and perform well on different domains.

Read the paper · More papers on PaperTik