Enhanced Segment Tree Approach for Multi-Attribute Numerical Data in K-Anonymization for Data Privacy
Priyank Jain, Yashwant Aditya, Kolli Rakesh, Manasi Gyanchandani · IEEE Transactions on Privacy · 2026
This research focuses on improving the k- Anonymization for numerical attributes by optimizing the critical phases of clustering and generalization through a novel Segment Tree-Based k-Anonymization (STKA) framework. Traditional methods face significant limitations: sorted-based approaches require costly re-sorting (O(n log n)) for dynamic updates; unsorted methods rely on inefficient linear scans (O(n)) for range queries; and dynamic methods suffer from frequent recalculations causing performance bottlenecks. Beyond computational efficiency, ensuring robust data privacy while minimizing information loss remains a central challenge in anonymization. Excessive generalization often safeguards privacy at the expense of data utility, whereas insufficient anonymization risks sensitive information disclosure. The proposed STKA framework effectively balances this trade-off by achieving optimal partitioning that enhances privacy protection while preserving high data usability. The proposed STKA framework addresses the aforementioned limitations by achieving logarithmic time complexity O(log n) for both min-max range queries and updates, eliminating resorting and recalculations. Compared to traditional approaches, STKA demonstrates: (1) 67% faster build time than sortedbased methods like ISKA/IAKA when handling dynamic updates; (2) 89% improvement in query response time over unsorted methods like WF-C-means and PCAA; (3) 35% reduction in information loss compared to dynamic approaches like CASTLE and FAANST; and (4) superior scalability for both batch and streaming data with minimal computational overhead. The experimental evaluation on datasets up to 1M records shows that STKA consistently outperforms state-of-the-art methods (MRA, SKA, IAKA, ISKA, MIAE) across multiple k-values, achieving 14.8-24.1% information loss compared to 31.2-53.1% for traditional methods, thereby delivering a more privacypreserving yet information-rich anonymization framework for large-scale data environments.