Sensibilité de la Valeur Espérée du Risque Empirique Induit par l’Algorithme de Gibbs dans le Problème d’Apprentissage Supervisé

Samir M. Perlaza, Iñaki Esnaola, H. Vincent Poor · HAL (Le Centre pour la Communication Scientifique Directe) · 2022

An explicit expression for the sensitivity of the expected empirical risk (EER) induced by the Gibbs algorithm (GA) is presented in the context of supervised machine learning. The sensitivity is defined as the difference between the EER induced by the GA and the EER induced by an alternative probability measure on the models. When several datasets are available, the sensitivity plays a central role to determine whether or not a lower EER might be observed by aggregating several datasets. Necessary and sufficient conditions for decreasing the EER by dataset aggregation are presented. Such conditions, which are on the GA parameters and the referencemeasures assumed for each constituent dataset, boils down to the evaluation of the sign of the sum of some relative entropy terms. From this perspective, sensitivity appears as (a) an alternative metric to evaluate the generalization capabilities of the Gibbs algorithm; and (b) a theoretical ground to study the use of several datasets describing the same phenomenon, yet subject to different data acquisition systems, i.e., datasets with different statistical properties.

Read the paper · More papers on PaperTik