Unsupervised K-means Analysis of Tuberculosis Data in Brazil: Identifying High Prevalence States and Temporal Trends
Angelo Rossini, Domingos Alves, Victor Cassão, Newton Shydeo Brandão Miyoshi · Procedia Computer Science · 2024
This paper aims to demonstrate the findings obtained through the analysis and application of an unsupervised K-means algorithm on the SINAN database from 2001 to 2022 in Brazil, with the objective of understanding which states have the highest number of tuberculosis cases and identifying similarities among them that may contribute to a higher case rate relative to the local population. We will begin with a brief historical introduction, followed by an overview of the characteristics related to tuberculosis transmission. Subsequently, we will discuss the results obtained from the year-to-year analysis of the collected data.