Multiple Samples Clustering with Second-moment Information in Stock Clustering

Xiang Wang · 2020

The clustering algorithms that view each object data as a single sample drawn from a certain distribution have been a hot topic for decades. Many clustering algorithms, such as k-means and spectral clustering, are proposed based on the assumption that each clustering object is a vector generated by a Gaussian distribution. However, in real practice, each input object is usually a set of vectors drawn from a certain hidden distribution. Traditional clustering algorithms cannot handle such a situation. This fact calls for the multiple samples clustering algorithm. In this paper, we propose two algorithms for multiple samples clustering: Wasserstein distance based spectral clustering and Bhattacharyya distance based spectral clustering, and compare them with the traditional spectral clustering. The simulation results show that the second-moment information can greatly improve the clustering accuracy and stability. These algorithms are applied to the stock dataset to separate stocks into different groups based on their historical prices. Investors can make investment decisions based on the clustering information, to invest stocks in the same cluster and get the highest earning or to invest stocks of different clusters to avoid the risk.

Read the paper · More papers on PaperTik