Paradigmatic and syntagmatic rule extraction for lifelong machine learning topic models
Muhammad Taimoor Khan, Shehzad Khalid · 2017
Lifelong Machine Learning (LML) based topic models are designed with an automatic learning mechanism. They are highly suitable for large-scale data with many datasets. The model processes each dataset for generating topics while it retains valuable knowledge from it as rules. The model grows in knowledge as it processes more datasets, following a continuous learning mechanism. The knowledge learned through past experience is utilized to produce better results for future datasets. Generally the learning procedure consists of an evaluation criterion that measures the confidence in a rule and a threshold that provides the par value. The existing LML based topic models learn rules with measures of co-occurrence from statistics which limits it to learning syntagmatic rules only. However, rules consisting of paradigmatic words are ignored as they do not co-occur. They are usually word synonyms and are alternately used. The proposed ParaSyn-LMLTM model learns paradigmatic rules as well while improving syntagmatic rules with techniques from linguistics. It enhances the learning capability of the model which is reflected via improved quality of topics on Chen2014 dataset.