Three level method using machine learning and rule based approach for extracting web-table information
Sungwon Jung, Sung-Shin Lim, Hyuk‐Chul Kwon · 2005
Generally, Authors of HTML documents use various methods to clearly convey their intention. The table is the preeminent method among these, because the table contains meaningful data displayed in a structure with rows and columns. However, on the Internet, tables are used for the purpose of the knowledge structuring as well as design of documents. It is not easy task to distinguish those two tables because HTML does not separate presentation and structure. This makes information extracting from those tables more difficult. Therefore, in this paper, we are firstly interested in classifying tables into two types: meaningful tables and decorative tables. After that we extract information from meaningful tables.