A Model for Selecting Relevant Topics in Documents Aimed at Compliance Processes
João Alberto da Silva Amaral, Fernando Buarque De Lima Neto · 2021
This paper proposes a semantic Natural Language Processing (NLP) approach used to assist in the automated characterization of information relevant to compliance activities. In this context, the proposed model combines two topic modeling techniques: Latent Semantic Analysis (LSA) and Latent Dirichlet Allocation (LDA), the first used to assist in the dimensionality reduction process, while the second, used to identify the number of relevant topics addressed in the processed data. Interesting results were achieved when three large databases tested the model, namely: Database of European Laws (period 1952 to 1990), Database of Audit reports issued by the State General Secretariat of Management of Pernambuco (period 2010 to 2019), and Database of Appellate Decisions issued by the Brazilian Federal Accountability Office (year of 2019). We compared the performance of three machine learning methods: K-means, LSA and LDA. In our experiments, we observed that (i) pre-processing techniques have a marked influence on the result of topic extraction; that (ii) Silhouette techniques and topic coherence produced the best value for the quantitative topics in all databases ; that LSA associated with LDA presented the best performance in all three databases, regarding the identification of relevant themes (topics) identified. Overall, the best results were obtained using the database in English (European Union Laws).