Online data stream classification with incremental semi-supervised learning
H. R. Loo, Muhammad Nadzir Marsono · 2015
This paper proposes an online data stream classification that learns with limited labels using selective self-training. Data partitioning steps are proposed to improve stream mining efficiency. Simulation on Cambridge and KDD'99 datasets shows up to 99.3% average accuracy for 10% labeled data and 98.4% for 1% labeled data. Data partitioning also speeds up classification process by 80% with only 0.2% reduction in accuracy.