A General Framework for Building Applications with Short and Sparse Documents

Journal Ijmer, Syed Jani Basha, Sayeed Yasin · 2014

with the explosion of e-commerce and online communication and publishing, texts become available in a variety of genres like Web search snippets, forum and chat messages, blogs, book and movie summaries, product descriptions, and customer reviews. Successfully processing them, therefore, becomes increasingly important in many Web applications. However, matching, classifying, and clustering these sorts of text and Web data pose new challenges. Unlike normal documents, these text and Web segments are usually noisier, less topic-focused, and much shorter, that is, they consist of from a dozen words to a few sentences. Because of the short length, they do not provide enough word co- occurrence or shared context for a good similarity measure. Therefore, normal machine learning methods usually fail to achieve the desire accuracy due to the data sparseness. To deal with these problems, we present a general framework that can discover the semantic relatedness between Web pages and ads by analyzing implicit or hidden topics for them.

Read the paper · More papers on PaperTik