Predicting Task Completion from Rich but Scarce Data.

José P. González-Brenes, Jack Mostow · 2010

We present a data-driven model for predicting task completion in Project LISTEN’s Reading Tutor, which takes turns picking stories and listens to the child read aloud [1]. However, children do not always finish stories, and we would like to understand why, or at least detect when they are about to stop. So our EDM challenge is to learn a model to predict task completion – a widely used metric of dialogue systems ’ performance. Such a model could help detect imminent disengagement in time to address it, and identify factors that influence task completion, including tutor behaviors, thereby providing useful guidance to make tutors engage students longer and more effectively. The richness of multimodal tutorial interaction over time makes the space of possible features to describe it large relative to the amount of data. When the number of features is large compared to the amount of data, classifier learners tend to overfit the data, so we need a method that learns robust models from few training examples with many features. Consider the supervised learning problem with training data S = {(x (i) , y (i))}, i = 1…n, where each data point is a p-dimensional vector x (i) , and y (i) is its label. The number of features p may exceed the number of data points (p>> n). A binary logistic regression model has the following form, where the vector θ contains the p parameters of the model: 1 p(y =1 | x;θ) =

Read the paper · More papers on PaperTik