A Probabilistic Model for Knowledge Component Naming.
Cyril Goutte, Serge Léger, Guillaume Durand · NPARC · 2015
Recent years have seen significant advances in automatic identification of the Q-matrix necessary for cognitive di-agnostic assessment. As data-driven approaches are intro-duced to identify latent knowledge components (KC) based on observed student performance, it becomes crucial to de-scribe and interpret these latent KCs. We address the prob-lem of naming knowledge components using keyword auto-matically extracted from item text. Our approach identifies the most discriminative keywords based on a simple proba-bilistic model. We show this is effective on a dataset from the PSLC datashop, outperforming baselines and retrieving unknown skill labels in nearly 50 % of cases. 1. OVERVIEW The Q-matrix, introduced by Tatsuoka [9], associates test items with attributes of students that the test intends to as-sess. A number of data-driven approaches were introduced to automatically identify the Q-matrix by mapping items to latent knowledge components (KCs), based on observed stu-dent performance [1, 6], using, e.g. matrix factorization [2, 8], clustering [5] or sparse factor analysis [4]. A crucial issue with automatic methods is that latent skills may be hard to describe and interpret. Manually-designed Q-matrices may also be insufficiently described. A data-generated descrip-tion is useful in both cases. We propose to extract keywords relevant to each KC from the textual content corresponding to each item. We build a simple probabilistic model, with which we score keywords. This proves surprisingly effective on a small dataset obtained from the PSLC datashop. 2. MODEL We focus on extracting keywords from the textual content