Multi-media based web mining for an information resource
H S Tan, Susan Ella George · WIT transactions on information and communication technologies · 2002
In this paper we examine multi-media based web mining with view to establishing a web-based information resource that is able to retrieve documents based upon multi-media features within the document. A multi-media based ‘search engine’ potentially far exceeds what could be achieved fi-om a purely textual one, but before such a comprehensive engine could be developed, a data mining process is required to extract appropriate features from unstructured web documents. Firstly, this paper discusses the need for visualization of search results and describes some user interface work that has been conducted towards improving the linear list presentation that is commonly found for presenting search results. Secondly, the principles of the self-organising feature map (SOM) are described including the method of training before providing examples of how the SOM can be used for visualization and how it has been used in various search-related applications. The SOM provides a topological ordering of the clustering. Thirdly, we present how the SOM has been used to present the results of data organisation using the clustering ability of the SOM, operating with appropriate features extracted fi-om the web-pages. The extraction of appropriate features fi-om documents that can then be used to ‘index’ the document within the multimedia search space.