Detection and Redaction of Sensitive Data in Application Logs using LLMs
Milind Dinesh, A Aswin, Jinesh M. Kannimoola · 2025
Any public organization that collects or processes client or user data is legally required to protect this information, ensuring that it is not retained, shared, or stored anywhere as plain text. Logs, which aggregate data from multiple systems and services that an organization uses, can unintentionally include sensitive information. Conventional rule-based approaches and regular expressions (regex) to extract Personally Identifiable Information from logs can sometimes be inefficient. They cannot be easily adapted to varying log formats and face high false positives. This work proposes a system using a Large Language Model fine tuned for Named Entity Recognition task to identify sensitive items in logs coupled with a redaction module that encrypts the identified entities to generate tokens that can be used to replace the sensitive data. This system can reduce false positives, be adapted to varying log formats and is more secure thus ensuring privacy without compromising practicality.