A Simple and Efficient Approach for Retrieving Live HTML-based Internet Information

Nader Mohamed, Jameela Al‐Jaroodi · 2007

Abstract: Dynamic HTML documents on the Internet contain useful information that can be reused by other applications. Unlike XML documents, the problem with HTML documents is that they do not have any semantics for the data in the page. Although a programmer can write a program that retrieves a peace of information from a specific HTML document available on the Internet, it will be very difficult to write several different programs to retrieve information from different dynamic HTML pages with varying formats. This paper develops a simple and generic approach to retrieve dynamic HTML Internet-based information. In this approach, several techniques that can be used to retrieve data from dynamic HTML documents are developed. These techniques were developed as a Java class which programmers can use to integrate and reuse HTML Internet Information for other real-time applications that need this information for their operations. The integrated Internet information can be weather information, stock prices, and top news. A number of experiments were conducted to measure the performance of these techniques. In addition, the paper discusses a number of applications that may benefit from this approach.

Read the paper · More papers on PaperTik