Web-knowledge-based learning for ill-posed inverse problems

V. Rao Vemuri, Na Tang · 2006

In machine learning, challenges occur frequently for real-life problems, because most of real-life problems are ill-posed. Here, ill-posed problems refer to the application domains where the given data is not high-quality enough (incomplete, insufficient or noisy) to build an accurate predictive model. From a mathematical perspective, a standard method of solving ill-posed problems is via regularization, a mathematical artifice that strives to introduce some additional information about the solution. The purpose of this dissertation is to explore and extend the concept of regularization to problems arising in machine learning, by incorporating additional domain into learning models. The additional can be specified by domain experts as well as through a discovery process from the web. The focus of this research therefore is to explore methods of mining available on the web and methods of incorporating the mined into learning models. Toward this goal, we first construct a general framework of web-based learning and demonstrate two case studies, one in data-poor domain and one in data-rich domain. We then examine the methods of discovery to obtain related knowledge through the web. Both unsupervised and semi-supervised text mining algorithms are proposed and compared to classical information retrieval methods. Finally we explore methods of incorporating prior into learning models as additional or supplemental information. A Bayesian learning algorithm is proposed to smoothly combine prior and statistical data. The experiments show that we can efficiently and effectively obtain through the web by using unsupervised or semi-supervised information retrieval methods. Furthermore, learning with the supplemental (either from domain experts or from the web) and data outperforms learning with pure data.

Read the paper · More papers on PaperTik