Machine learning for signal processing: geometry, kernels, and symbolic sequences
Dan Schonfeld, Liuling Gong · 2010
Machine Learning is programming computers to optimize a performance criteria using example data or past experience. There are many successful applications of machine learning in various domains, e.g., speech and handwriting recognition, customer behavior study. The kernel trick, by mapping the non-linear problem from the input space to a new space (called the feature space) through a non-linear transformation, is applicable to both regression and classification/clustering problems. In this thesis, we focus on the problems of regression and classification/clustering based on kernel methods. A unified approach of kernel density estimation, namely the Space Kernel Analysis (SKA) is proposed. The foundation of SKA is based on the definition of space kernel which characterizes a space-dependent similarity between two vectors. It is a generalization of the classical kernel which operates in the feature space on pairwise vectors. The properties of the SKA are analyzed and the relationships between the SKA and several other kernel density estimation techniques are examined. The problem of learning from high-dimensional noisy data is especially studied. A new learning approach based on projections onto multi-dimensional ellipsoids (POME) is introduced, which is applicable to unsupervised clustering, semi-supervised clustering and classification in high-dimensional noisy data. The performances of POME-based algorithms are tested and compared with other well-known algorithms in classification and clustering. While the distance between any two points in a numeric space can be directly computed by the kernel function, the metric space of the symbolic data is hard to define. The similarity metric on strings, namely the string kernel, is an essence issue in text mining and has various applications, e.g., text categorization and protein classification. We define the string kernel as the function that measures the similarity between two strings. According to whether the kernel is decomposable as the inner product of two feature vectors, the string kernel falls into two categories. Different string kernels are analyzed and the consistency properties between them are studied.