Named entity recognition in titles of Chinese videos from the web

Quan Qi, Jing Dong · 2011

Video title is the essential source of information used to describe video content in text-based video recommendation system. Automatically recognizing the named entities such as organizations, brand, and people in the titles of video from the Web is a difficult task. Owing to the characteristics of title-naming and Chinese video-up-loaders' habits, titles' length is always very short, and they can hardly provide enough information for recognizing. The abbreviations can always be found in titles also makes recognizing performance not satisfactory. In this paper, we propose a method which utilizes title' search results returned by search engine as the supplementary of titles. To overcome shortcomings of returning no relevant search results, we also utilize topic model to generate relevant context of titles. In experiments, 62305 Chinese automotive video titles are crawled from web for testing our method. The result shows that named entities in the informal form could be recognized in most titles, and the performance of named entity recognition is satisfactory.

Read the paper · More papers on PaperTik