Quick-and-clean extraction of linked data entities from microblogs

Oluwaseyi Feyisetan, Elena Paslaru Bontas Simperl, Ramine Tinati, Markus Luczak–Roesch, Nigel Shadbolt · 2014

In this paper, we address the problem of finding Named Entities in very large micropost datasets. We propose methods to generate a sample of representative microposts by discovering tweets that are likely to refer to new entities. Our approach is able to significantly speed-up the semantic analysis process by discarding retweets, tweets without pre-identifiable entities, as well similar and redundant tweets, while retaining information content.

Read the paper · More papers on PaperTik