Explainable Machine Learning for Predicting Computational Thinking Proficiency in K-12 Students Using National-Level Data

Hyuna Noh, Hansung Kim, Hyeoncheol Kim · IEEE Access · 2025

This study aimed to identify the key factors predicting computational thinking (CT) attainment among Korean elementary- and middle-school students and to provide educational implications for promoting CT education. Using national-level digital literacy assessment data collected by the Korean government (14,468 elementary- and 16,622 middle-school students), we developed machine-learning-based predictive models and applied Shapley Additive Explanations (SHAP), an explainable AI (XAI) technique, to evaluate feature importance and interaction effects. The SHAP analysis revealed that information and communications technology (ICT) literacy levels, coding skills, and perceptions of the future occupational usefulness of digital devices were common positive predictors of CT attainment for both groups. For elementary students, Internet search proficiency and grade level exerted additional positive effects, whereas the time spent on digital devices for out-of-home learning purposes, leisure-centered use, and high autonomy negatively affected CT attainment. Among middle-school students, the ability to create presentation materials emerged as a key positive factor, whereas the effects of grade level and information-search ability were relatively limited. SHAP dependence analysis demonstrated that combinations of ICT literacy levels, grades, coding skills, and perceptions of occupational usefulness consistently enhanced CT attainment, whereas the combination of high digital autonomy with occupational usefulness perceptions hindered CT attainment. These findings highlight the need to systematically strengthen ICT and coding education, establish stage-specific strategies, and provide guidance for goal-oriented and structured digital use. Ultimately, this study contributes to academic significance by applying XAI methods to national-level CT data, enabling interpretable analysis of feature importance and interaction effects, and uncovering complex relationships among variables that traditional statistical approaches could not fully reveal.

Read the paper · More papers on PaperTik