A New Approach for Domain New Words Detection

Liu Hua · Zhongwen xinxi xuebao · 2006

The paper puts forward a new method for domain new words detection,which directly extracts labeled by specialist in web pages,and stored them in classified wordlist according to the column of source web page.The simple approach can detects new words and clusters quickly.Using the approach,from 6 hundred million web pages covering 15 domains,we extracted 229237 words,including 175187 new words,the new words ratio is 76.42%.New words are mostly Named Entities,which have steady structure and integrated meaning,and are conducive to ambiguity and unknown words in Chinese word segmentation.They will be useful for text representation,such as text categorization and key words indexing.

Read the paper · More papers on PaperTik