An overview of inference methods in probabilistic classifier chains for multilabel classification

Deiner Mena, Elena Montañés, José Ramón Quevedo, Juan José del Coz · Wiley Interdisciplinary Reviews Data Mining and Knowledge Discovery · 2016

This study presents a review of the recent advances in performing inference in probabilistic classifier chains for multilabel classification. The interest of performing such inference arises in an attempt of improving the performance of the approach based on greedy search (the well‐knownCCmethod) and simultaneously reducing the computational cost of an exhaustive search (the well‐knownPCCmethod). UnlikePCCand asCC, inference techniques do not explore all the possible solutions, but they increase the performance ofCC, sometimes reaching the optimal solution in terms of subset 0/1 loss, asPCCdoes. Theε‐approximate algorithm, the method based on a beam search and Monte Carlo sampling are those techniques. An exhaustive set of experiments over a wide range of datasets are performed to analyze not only to which extent these techniques tend to produce optimal solutions, otherwise also to study their computational cost, both in terms of solutions explored and execution time. Onlyε‐approximate algorithm withε=.0 theoretically guarantees reaching an optimal solution in terms of subset 0/1 loss. However, the other algorithms provide solutions close to an optimal solution, despite the fact they do not guarantee to reach an optimal solution. Theε‐approximate algorithm is the most promising to balance the performance in terms of subset 0/1 loss against the number of solutions explored and execution time. The value ofεdetermines a degree to which one prefers to guarantee to reach an optimal solution at the expense of increasing the computational cost.WIREs Data Mining Knowl Discov2016, 6:215–230. doi: 10.1002/widm.1185 This article is categorized under: Technologies > Classification Technologies > Machine Learning

Read the paper · More papers on PaperTik