Learning to detect PII: Tabular vs. Document classification models for network traffic analysis
Rishika Kohli, Shaifu Gupta, Manoj Singh Gaur, Soma Sekhar Dhavala · Journal of Information Security and Applications · 2025
Detecting Personally Identifiable Information (PII) exfiltration from mobile network traffic is critical for preserving user privacy. Traditional approaches rely on machine learning classifiers trained on manually engineered features extracted from network packets. Deep learning offers the potential to remove the reliance on such an external feature selection process; however, its effectiveness depends significantly on how underlying packets are encoded. In this work, we investigate deep learning paradigms for PII detection with a focus on the impact of feature encoding strategies. We explore tabular modeling approaches, including both an existing architecture (FT-Transformer) and proposed modular frameworks that integrates a pretrained language model (all-MiniLM-L6-v2) for semantic feature embeddings, followed by a classifier. We also evaluate document classification modeling by applying pretrained language models such as TinyBERT directly to the raw packet content. We further demonstrate the feasibility of on-device inference by deploying trained models using ONNX and TensorFlow Lite. Finally, we recommend modeling strategies based on data size , performance , resource utilization , and generalizability , enabling model selection according to the primary requirement of the deployment scenario.