TEXAS2: A System for Extracting Domain Topic Using Link Analysis and Searching for Relevant Features.
Sangwon Hwang, Yongseok Lee, Young-Kwang Nam · 2016
It is very important to understand the domain topic of software to maintain and reuse it. However, the continual development and change in its size makes it difficult to understand it. To solve this problem, researches have been recently conducted to extract the domain topic using various information search techniques such as LDA, with the researches on LDA-based techniques being especially active. However, since only unstructured information such as an identifier or note is used in most research, without including structured ones like information calling, problems in which extracted topics are different from the characteristics of the program can occur. In this paper, we propose a method to generate documents and extract topics using both structured and unstructured information. We also generate indexes based on the frequency of the identifier of the source code, and propose a system that extracts an association rule based on the simultaneous generation of the method. We as well establish a system that provides highly reliable search results to user queries by combining domain topics, indexes with scores, and the association rule information. Consequently a TEXAS2 system for this study was established and confirmed a high user satisfaction on search results to the queries in a performance test.