Descubrimiento de tópicos a partir de textos en español sobre enfermedades en México

Alejandro López López · 2022

In social networks there is a large amount of information that can reach to be valuable on numerous topics. For example, in the domain of diseases Many people around the world publish various information about them, including conditions, signs, symptoms, procedures, medications and treatments. This information is found in texts in a disorganized way making it difficult for readers to find valuable information and perform an analysis manual of it is a tedious, difficult and time-consuming process. For this we resort to computational systems, algorithms or methods of analysis of texts to find topics or topics of interest. That is why in this work an approach for the discovery is presented of topics from texts in Spanish on three diseases (Diabetes, Cancer and COVID-19) in Mexico, using LDA (Latent Dirichlet Allocation) algorithms widely used in the literature, and BTM (Biterm Topic Model) an alternative that groups two terms to find the topics. This approach has as hypothesis that the use of phrases over words to enter algorithms improves the coherence results of the topics. An evaluation of experimental results was carried out based on the topic coherence metric. This evaluation has shown that the use of phrases it is more effective than using single words to discover topics. In addition, they have achieved the best consistency results by disease as follows: 0.7421 with the BTM algorithm for 100 topics on COVID; 0.6755 with the BTM algorithm for 80 topics on Cancer; and 0.6357 with the BTM algorithm for 80 topics about Diabetes.

Read the paper · More papers on PaperTik