Preprocessing and Symbolic Representation of Stock Data
Mukesh Kumar, Arvind Kalia · 2012
There has been a lot of interest in mining the time series data. Stock data mining plays an important role to visualize the behavior of financial market. In financial data mining the data is normally represented in the numeric format, however, the symbolic representation is also used to evaluate the overall impact. Time series data are difficult to manipulate, but when they are treated as symbols instead of data points, interesting patterns can be discovered and it becomes an easier task to mine them. In this paper, a symbolic representation of NSE stock data of thirteen years period i.e. from Jan. 1996 to Dec.2008 is presented. The data preprocessing is an essential part of data mining Data cleaning fills in missing values, smoothes noisy data, handles or removes outliers, resolves inconsistencies. First of all the data was normalized, Normalization was done on the dataset using min-max normalization. The data transformation steps performed include offset translation, removing of linear trend, and removing of noise using moving average smoothing method. Further a best fitting line is used to remove the linear trend from the dataset. Euclidean distance measure has been used to establish relationships among various stocks. Three symbols [up, down, neutral] have been used for symbolic representation of the data and distance is evaluated as per the matching pattern of these symbols. It has been found that symbolic representation provides an easier interpretation and helped to determine an overall pattern. Symbolic pattern is having resemblance with price change pattern in numeric representation.