Latent Cluster Discovery in a Collection of Low Count Non-Gaussian Time Series in Economics Domain
Ishani Chakraborty, Genie Rubaiyat · 2024
The main focus of this work is to discover naturally occurring clusters in behavioral time series, and then associate a numerical representation with every cluster, which could be used to further study and control cluster properties. We presented a higher order but parsimonious model similar to HMM to do predictive analysis on time series. We revised the modeling assumptions of an unpublished large language model [1] to make it more suitable to our problem, which practically transformed it into a new model, which is based on mixture transition distribution [2]. This work is the first application of this modified language model [1] in the area of finance, and in the area of customer(user) behavior study and user group discovery. The broader class of this kind of models [2], [3] has very few applications in finance at all.Our user behavior dataset, which is a collection of low count non-Gaussian time series, is collected from a loyalty program marketing company. This non-Gaussian time series data set contains bursts and flat stretches, hence it is already hard to model. But, it also poses significant modeling challenges, because it is not annotated, and is unaligned. Moreover, these time series are of short length and of unequal length. As a result, our work involved extensive pre-processing of raw data, including merging several time series using aggregate data at the data processing level to create longer time series for better estimation of parameters. We derived numerical representation of every time series from estimated parameters by learning the proposed model, and then clustered these representations. After clustering we derived a numerical representation of every cluster. We used three algorithms for clustering and three metrics to evaluate the clusters. We used cluster quality evaluation scores as measures of goodness of fit of the models that we experimented with. Although we used two other baseline models for our work, we see that the best well-separated dense latent clusters have been discovered when we used the numerical representation learned from our model. Hence, we find that our proposed model performed better than or equally well with the baseline models.