Associative and sequential classification with adaptive constrained regression methods

Wade M. Bannister · 2007

Associative classification methods use mined association rules as classifiers and have been proposed recently to provide accurate models that are more understandable to the practitioner. However, these methods have thus far been difficult to use as they continue to return large and overly complex models. This research proposes a new associative classifier using adaptive constrained regression models to select linear combinations of association rules. The norm of the coefficients is used as a constrained regression model which performs a natural variable selection by constraining the total size of the parameter estimates, forcing many parameter values to zero. To further reduce the model size, adaptive techniques are employed which evaluate the model at each step of the algorithm and halt the model building process once it has reached a critical value. In comparisons that utilize sample data, this methodology compared favorably with existing classification methods, returning substantially smaller models with only a minor performance tradeoff. Additionally, the constrained methodology is combined with the logistic penalty with similar success. Classification using time-ordered data is an area largely unexplored by researchers, but is a necessity when analyzing the volumes of transactional data found in many industries. The constrained associative classifier is extended to sequential data by exploiting the algorithmic relationship between association rule mining and sequential pattern mining. Frequent sequential patterns are mined and then the constrained adaptive methodology is applied to select patterns to be used for classifying the outcome. In tests using simulated data, the proposed methodology accurately identified embedded sequences and used them to predict a binary class variable. This sequential methodology was applied to large healthcare transaction data to identify patterns of care which impact patient outcomes. A methodology to preprocess and transform the data is developed and implemented to optimize the data structure for sequential classification. The resulting model successfully identified 29 patterns of care which could be used to identify best practices and quantify the quality of care provided by physicians.

Read the paper · More papers on PaperTik