Kernel Machines for Multi-Class Classication: A Joint Kernel Approach
Antolin Thomas Janssen · 2009
Pattern analysis is about the automatic detection of patterns in data, and plays an important role in many modern artificial intelligence and computer science problems. With patterns we mean any relations, regularities or structures that are present within a source of data. By detecting significant patterns in the data, a system can make predictions about new data coming from the same source. We can say that the system has acquired generalisation power by learning a pattern in the data from that source. There are many important problems that can only be solved using this approach, ranging from bio informatics to web retrieval. In recent years, pattern analysis has become a standard method for software engineering, and is used in many commercial products. This thesis describes several algorithms and compares their performances on different datasets. This to research the possibilities that we have when we want to extract possibly interesting relationships between datapoints. The improvement of these kind of techniques can be important for future research because with increasing computational power, increasingly larger datasets can be analyzed, so that we can find relationships we never knew existed. One of these datasets is called SCOP, or the Structured Classification of Proteins, which we can analyze to better understand how our body works and how different proteins are related. To detect these relationships we use an algorithm called Regularized Least Squares. Also its performance is compared to a very similar algorithm, called SVMLight, and several other machine learning techniques, namely Naive Bayes, KStar and Random Forest. We show that the performance depends on the problem that is analyzed, where things like linear separability and the number of examples in the dataset have a notable influence. Nevertheless, the results are stable on most datasets, resulting in a prediction system that can be considered as a reliable advise-tool for making decisions about that dataset. A rapidly developing field within this research is multiclass classification, which extends the binary classification into a method that can handle problems with multiple classes. The multiclass extension for the algorithm that we consider is a joint kernel approach which also provides stable results and could be optimized to improve performance even more.