On Data Classification Efficiency Based on a Trade-off Relation between Mutual Information and Error Probability

Mikhail Lange, Andrey Lange, S. V. Paramonov · 2020

We propose a data classification model which yields an average mutual information between a set of objects and a set of class-label decisions as a function of error probability. Optimization of the model consists in minimization of the average mutual information by conditional distributions for the decisions subject to a given constraint on the average error probability. It is equivalent to calculating the rate-distortion function in a scheme of coding the source class labels with a given fidelity when a set of the class labels and a set of the objects are connected by an observation channel with known class-conditional probability distributions. Given set of the objects and known observation channel, a lower bound to the rate-distortion function is calculated. This bound is independent on a decision algorithm and yields a potentially achievable error probability subject to a fixed value of the average mutual information. The obtained bound is useful for evaluating an error probability redundancy of any decision algorithm with given discriminant functions.

Read the paper · More papers on PaperTik