A basic study on attribute name extraction from the web
Fumitaka Nakane, Masanori Otsubo, Yoshinori Hijikata, Shogo Nishida · Conference proceedings/Conference proceedings - IEEE International Conference on Systems, Man, and Cybernetics · 2008
A large number of semistructured documents exist on the Web. We can find pages that contain keywords by using a search engine. But when we want to obtain information about an object like a notebook computer with 1 GB memory, a method is needed that automatically extracts attribute name (in this example, ldquomemoryrdquo) and attribute value (in this example, ldquo1 GBrdquo). In the past, many researchers examined extracting attribute values corresponding to each attribute name. This paper discribes a method that extracts schemas (sets of attribute names) using bootstrapping algorithm.