File clustering using naming conventions for legacy systems
Nicolas Anquetil, Timothy C. Lethbridge · 1997
Decomposing complex software systems into conceptually independent subsystems represents a significant software engineering activity that receives considerable research attention. Most of the research in this domain deals with the source code; trying to cluster together files which are conceptually related. In this paper we propose using a more informal source of information: file names. We present an experiment which shows that file naming convention is the best file clustering criteria for the software system we are studying. Based on the experiment results, we also sketch a method to build a conceptual browser on a software system. Introduction Maintaining legacy software systems is a problem which many companies face. To help software engineers in this task, researchers are trying to provide tools to help extract the design structure of the software system using whatever source of information is available. Clustering files into subsystems is considered an important part of this ac...