Unsupervised streaming feature selection based on dynamic pseudo-labeling
Hui Tang, Yunyun Zhang, Peng Zhou · The Computer Journal · 2025
Abstract Streaming feature selection (SFS) is a vital research area in machine learning and data mining that focuses on dynamically selecting optimal feature subsets from continuously arriving feature streams. In real-world applications such as competent healthcare and financial risk control, features are generated dynamically, while acquiring label is often costly or infeasible. Then, unsupervised streaming feature selection (USFS) methods were developed to address this. Recent advances aside, existing USFS approaches still struggle with adapting to evolving streaming features importance, sensitivity to data distribution changes, and decoupling from sample structure modeling. To overcome these challenges, we propose a novel online USFS algorithm based on dynamic pseudo-labels—introducing the pseudo-label mechanism to the USFS task for the first time, to our knowledge. Our approach constructs a feature evaluation framework that integrates pseudo-labels with mutual information, enabling a closed-loop iterative feature selection process and adaptive pseudo-label optimization. Specifically, the algorithm employs a two-stage clustering strategy comprising initial clustering followed by pseudo-label-driven optimization clustering, dynamically generating and updating pseudo-labels to enhance the accuracy and reliability of feature selection progressively. Experimental results demonstrate that proposed method outperforms SFS algorithms under unlabeled conditions, confirming its effectiveness and broad applicability.