A Comparison Between the K-Nearest Neighbors Algorithm and Logistic Regression in the Field of Cell Type Annotation
Lezhou Wen · Theoretical and Natural Science · 2025
With the advancement of single-cell sequencing technologies, high-capacity gene expression data have made cell type annotation across diverse cell populations feasible. However, the high-dimensional and complex nature of these datasets poses challenges for algorithm selection, as traditional manual annotation methods have become inadequate. Machine learning algorithms offer a robust alternative, yet choosing the optimal algorithm remains a critical step. This study provides a detailed analysis of two classical machine learning algorithms--k-Nearest Neighbors (KNN) and Logistic Regression and compares their strengths and limitations in cell type annotation from the perspective of algorithmic principles and data characteristics, aiming to offer practical guidance for selecting machine learning approaches. KNN, a distance-based non-parametric method, excels in small-sample and nonlinear scenarios but suffers from the "curse of dimensionality" in high-dimensional spaces, requiring efficiency optimization via dimensionality reduction or locality-sensitive hashing. In contrast, LR, relying on linear assumptions, performs well with large-scale, high-dimensional data through regularization to prevent overfitting, yet its performance declines with small samples or nonlinear distributions. Each algorithm has its own benefits; the choice between algorithms should consider factors such as sample size, feature dimensionality, data quality, interpretability, and the alignment between the true data distribution and the algorithm’s inherent assumptions.