Multimedia Classifiers: Behind the Scenes

Manjunath Ramachandra Iyer · 2021

This tutorial provides an in-depth understanding of the art and science behind the decision-making in a multimedia classifier. A multimedia classifier typically takes image, text, waveform, ordinal number or categorical data or their combination as the input and produces a single output indicating the class of the input pattern. Such a piece of AI ML system is extensively used as a decision-making element in several autonomous systems. The yardstick used by the human experts for decision making for the same input pattern often differs from the system, still producing the same output. In some cases, the outputs differ for the same input data and throws open a question on the reliability of the model. If such models are used in critical applications, which is often the case in an autonomous system, adequate mitigations for minimizing impact of the misjudgment has to be taken. It calls for ripping open the decision making process in the black box classifiers. Unwinding the black box is the need of the hour for the regulatory bodies as well. EU region has already made it mandatory to provide the details of the decision-making mechanism if it involves some form of AI ML components. This tutorial throws light on the decision making process in a classifier that may be used for a variety of applications. More than one technique to get a glimpse of the classifier in action would be discussed. The explanation can come in the form of ta heatmap indicating the relevant features influencing the decision making process, patterns learnt by the neurons or Textual description of the attributes of the input. The bottom line is these explanations are to be consistent. The mechanism to achieve the coherent explanation would be detailed in this tutorial.

Read the paper · More papers on PaperTik