Generative Data Imputation for Sparse Learner Performance Data Using Generative Adversarial Imputation Networks

Liang Zhang, Jionghao Lin, John Sabatini, Diego Zapata‐Rivera, Carol M. Forsyth, Yang Jiang, John Hollander, Xiangen Hu, Arthur C. Graesser · 2025

As learners engage with Intelligent Tutoring Systems (ITSs) by responding to a series of questions, their performance data, such as correct or incorrect responses, is crucial for assessing and predicting their knowledge states through analysis and modeling. However, data sparsity, often arising from skipped or incomplete responses, poses challenges for accurately assessing learning and delivering personalized instruction. To address this, we propose a generative data imputation method based on Generative Adversarial Imputation Networks (GAIN) to complete missing learning performance data. Our approach employs a three-dimensional (3D) framework structured by learners, questions, and attempts, with an adaptable design along the attempts dimension to manage varying sparsity levels. Enhanced by convolutional neural networks in the input and output layers and optimized with a least squares loss function, our GAIN-based method aligns the input and output shapes with the dimensions of question-attempt matrices across the learners' dimension. Extensive experiments on datasets from three types of ITSs, including AutoTutor Adult Reading Comprehension (ARC), ASSISTments and MATHia, demonstrate that our approach generally outperforms baseline methods, e.g., tensor factorization-based methods and other Generative Adversarial Network (GAN) variants, in imputation accuracy across different setting of maximum attempts. Bayesian Knowledge Tracing (BKT) modeling further validates the imputed data's efficacy by estimating learning parameters, including initial knowledge P (L0), learning rate P (T), guess rate P (G), and slip rate P (S). Results reveal that the imputed data not only enhances model fit but also closely aligns with the original sparse distributions by capturing underlying learning behaviors, indicating greater reliability in learner assessments. Kullback-Leibler (KL) divergence measurements of all these learning parameters confirm that the imputed data effectively preserve essential learning characteristics, maintaining low divergence

Read the paper · More papers on PaperTik