Tressel: Semantic mark-up of RSS feeds

Brian McLernon, Nicholas Kushmerick · 2006

www.smi.ucd.ie/tressel The recent explosion in the popularity of RSS and other syn-dication technologies has led to a wealth of information be-ing published from an increasingly diverse range of sources. However this popularity makes it difficult for users to find interesting documents, and this challenge is compounded by the fact that most RSS clients offer very modest personal-ization capabilities. We propose Tressel, a collaborative RSS aggregator that extracts semantically meaningful passages of text from RSS feeds. Tressel employs a semi-supervised machine learning algorithm to identify semantic information. The algorithm learns from a small amount of training data provided by the user. Tressel is collaborative in that the input of each user is used to benefit all of the users. One challenge with such an open system is the danger that users could (intentionally or accidentally) introduce noisy training data. In this paper, we describe Tressel’s archi-tecture and adaptive information extraction algorithm, and then report on experiments which demonstrate that we can reliably detect noisy training data. 1.

Read the paper · More papers on PaperTik