Probabilistic models in machine learning

Hisashi Kobayashi, Brian L. Mark, William Turin · Cambridge University Press eBooks · 2011

Introduction Machine learning refers to the design of computer algorithms for gaining new knowledge, improving existing knowledge, and making predictions or decisions based on empirical data. Applications of machine learning include speech recognition [164, 275], image recognition [60, 110], medical diagnosis [309], language understanding [50], biological sequence analysis [85], and many other fields. The most important requirement for an algorithm in machine learning is its ability to make accurate predictions or correct decisions when presented with instances or data not seen before. Classification of data is a common task in machine learning. It consists of finding a function z = G( y ) that assigns to each data sample y its class label z . If the range of the function is discrete, it is called a classifier , otherwise it is called a regression function. For each class label z , we can define the acceptance region A z such that y ∈ A z if and only if z = G( y ) . An error occurs if the classifier assigns a wrong class to y . The probability of classification error is called the generalization error in machine learning, where Z denotes the actual class to which the observation variable Y belongs. The classifier that minimizes the generalization error is called the Bayes classifier and the minimized ε( G ) is called the Bayes error . In practical applications, we generally do not know the probability distribution of ( Y , Z ).

Read the paper · More papers on PaperTik