A methodology and algorithm for automatically classifying text documents to strategic intents
Robin Karmakar · Chalmers Publication Library (Chalmers University of Technology) · 2011
The thesis develops and presents a strategy, methodology and tool for automatically classifying documents according to strategic intents and describes how to monitor changes in this distribution over time.The methodology uses artificial neural networks and latent semantic indexing to evaluate the similarity of documents to generate the proper allocation to a set of predefined clusters.The documents to be classified are patents and publications.The document classes are based on a company internal classification scheme, that is strategically aligned with the company's business objectives.The methodology is structured and constructed as a process considering the steps from raw data acquisition to the final output: a classification of documents and their temporal distributions.The results show that using the internally defined strategic intents supports a high-level analysis of technology areas and competitor activity within the same.The classification algorithm performed well to classify different sources of information, with best performance F -measure values in the range 0.73 -0.90 depending on the dataset used.For publications the algorithm performed well given that the number of training documents in the prelabeled set was sufficient.For patents the algorithm performs well, even if the classifier is created and trained using publications.Performance also increases with the ratio of prelabelled to unlabelled documents.The number of neurons in the hidden layer does not significantly affect classifier performance, but the number of correcting iterations does.In choosing between a high number of neurons or a high number of iterations, for a given computational effort the focus should be on increasing the number of iterations.Finally the thesis shows that company specific publishing trends can successfully be analyzed and evaluated over time using the suggested framework.