Collaborative Web Data Record Extraction
Gengxin Miao, Firat Kart, L.E. Moser, Peter Michael Melliar-Smith · 2009
This paper describes a Web service that automatically parses and extracts data records from Web pages containing structured data. The Web service allows multiple users to share and manage a Web data record extraction task to increase its utility. A recommendation system, based on the probabilistic latency semantic indexing algorithm, enables a user to find potentially interesting content or other users who share the same interests with the user. A distributed computing platform improves the scalability of the Web service in supporting multiple users by employing multiple server computers. A Web service interface allows users to access the Web service, and allows programmers to develop their own applications and, thus, extend the functionality of the Web service.