Learning from data with complex interactions and ambiguous labels
Soumya Shubhra Ray, Mark W. Craven, David Page · 2005
Abstract In this thesis, we develop and evaluate machine learning algorithms that can learn effectively from data with complex interactions and ambiguous labels. The need for such algorithms is motivated by such problems as protein-protein binding and drug activity prediction. In the first part of the thesis, we focus on the problem of myopia. This problem arises when greedy learning strategies are applied to learn from data with complex interactions. We present skewing, our approach to alleviating myopia. We describe theoretical results and empirical results on Boolean data that show that our approach can learn effectively from data with complex interactions. We investigate the effects of various parameter choices on our approach, and the effects of dimensionality and class-label noise. We then propose and evaluate a variant that scales better to high-dimensional data. Finally, we propose and evaluate an extension that is able to learn from non-Boolean data with similar complex interactions as in the Boolean case.