Advancing clustering methods in physics education research: A case for mixture models
Minghui Wang, Meagan Sundstrom, Karen Nylund‐Gibson, Marsha Ing · Physical Review Physics Education Research · 2025
Clustering methods are often used in physics education research (PER) to identify subgroups of individuals within a population who share similar response patterns or characteristics. Among these, k -means (or k -modes, for categorical data) is one of the most commonly used clustering methods in PER. This algorithm, however, is distance-based rather than model-based: it relies on algorithmic partitioning and assigns each individual to one subgroup through hard assignment. Researchers must also conduct analyses to relate subgroup membership to other variables. Mixture models are a model-based alternative that offers several statistical tools for choosing the optimal number of subgroups, accounts for classification errors by assigning individuals probabilities of belonging to each subgroup rather than hard assignment, and allows researchers to directly integrate subgroup membership into a broader latent variable framework. In this paper, we outline the theoretical similarities and differences between k -modes clustering and latent class analysis (one type of mixture model for categorical data). We also present parallel analyses using each method to address the same research questions in order to demonstrate these similarities and differences. We provide the data and code to replicate the worked example presented in the paper for researchers interested in using mixture models.