Homogeneity Pursuit in Clustered Data Analysis When Cluster Sizes Are Small

Yan Sun, Liming Tan, Wenyang Zhang, Zhenyu Zhu · Journal of Business and Economic Statistics · 2025

Clustered data analysis is an important topic in data science. A well established approach is to assume all clusters share the same unknown parameters of interest, and the difference between different clusters is formulated and accounted for by cluster effects. Whilst this approach works very well in many issues, such as exploring the global impact of an explanatory variable on the response variable, it does not provide much insight about individual attributes of each cluster. Assuming different clusters have completely different parameters would result in too many unknown parameters, which would lead to large variances of the final estimators. Following the idea of homogeneity pursuit proposed in Ke, Fan, and Wu, various modeling approaches are proposed in recent literature to group the unknown parameters and explore the individual attributes in clustered data analysis. However, most of them are either difficult to implement or require each cluster to have reasonably big cluster size. In this article, we propose a new approach, which is easy to implement and does not require any cluster to have big size, and establish its asymptotic properties without assuming the size of any cluster tends to infinity. We also conduct intensive simulation studies to show the approach works very well when sample size is finite. Finally, we apply the approach to a well-known financial dataset to show its superiority in exploring individual attributes in clustered data analysis.

Read the paper · More papers on PaperTik