iClass - Applying Multiple Multi-Class Machine Learning Classifiers combined with Expert Knowledge to Roper Center Survey Data *

Marmar R. Moussa, Marc Maynard · LWA · 2015

As one of the largest public opinion data archives in the world, Rop- er Center (1) collects datasets of polled survey questions as they get released from numerous media outlets and organizations with varying degrees of format ambiguity. The volume of data introduces search complexities over survey questions asked since the 1930s and poses challenges when analyzing search trends. Up to this point, Roper Center question-level retrieval applications used human metadata experts to assign topics to content. This has been insuffi- cient to reach required levels of consistency in catalogued data, and provides an inadequate base for creating an advanced search experience for research clients. The objective of this work is to combine the human expert teams' knowledge of the nature of the poll questions and the concepts and topics these questions express, with the ability of multi-label classifiers to learn this knowledge and apply it to an automated, fast and accurate classification mecha- nism. This approach cuts down the question analysis and tagging time signifi- cantly as well as provides enhanced consistency and scalability for topics' de- scriptions. At the same time, creating an ensemble of machine learning classifi- ers combined with expert knowledge is expected to enhance the search experi- ence and provide much needed analytic capabilities to the survey question data- bases. In our design, we use classification from several machine learning algo- rithms like SVM and Decision Trees, combined with expert knowledge in form of handcrafted rules, data analysis and result review. We consolidate this into a 'Multipath Classifier' with a 'Confidence' point system that decides on the rel- evance of topics assigned to poll questions with nearly perfect accuracy.

Read the paper · More papers on PaperTik