Knowledge discovery in academic registrar data bases using source mining: Data and text

Ma. Teresa Rios, Francisco J. Cantú-Ortiz · Journal of the Association for Information Systems · 2006

In this paper we describe a knowledge-based system for extracting knowledge from academic and registrar databases using source mining, where the sources are data or text. Other sources not included in this research are image, sound or gestures. Patterns of student behaviour were obtained by examining data from student attributes such as city of birth, scholarship needs, field of knowledge and major, student gender and other attributes, for various undergraduate academic programs offered by the campus of Tecnologico de Monterrey university system across the country. These patterns proved useful in predicting student enrolment and designing advertising campaigns. We use text mining techniques to match and compare course description from universities with which we have student ex-change programs for course revalidation. Equivalence of 64 courses were obtained by using text mining techniques for matching course descriptions for universities like Michigan State, Carnegie Mellon and New Mexico State. These equivalences helped the International Programs Office in developing course revalidation. Data mining techniques employed include C4.5 decision tree learning and feed-forward neural networks as implemented in the SIPINA intelligent environment (Sipina Research). Text mining techniques utilized are based on statistical and syntactic-semantic analysis and include Clasi-Tex (Clasitex), IBM Intelligent Miner for Text (IntelligentMiner), Text Roller (TextRoller), and Free Text Technologies Master Text

Read the paper · More papers on PaperTik