Efficient Duplicate Comment Detection for Rulemaking Agencies With Unsupervised Deep Learning: A Cost-Effective and High-Accuracy Approach

Zhicheng Huang, Weixuan Dong, Jeong‐Nam Kim, James Hollenczer, Hyelim Lee, Matthew L. Jensen, Elena Bessarabova, Neil Talbert, Rui Zhu, Yifu Li · IEEE Transactions on Computational Social Systems · 2025

Government agencies tasked with soliciting comments from the American public in response to changes in rulemaking have long been interested in finding effective techniques to automatically process spam comments, including flagging and filtering duplicate comments. Duplicate submissions are problematic because they obscure genuine public input and may amplify fringe opinions nonrepresentative of the majority of the American public. Given that duplicate comments are generated from templates or predrafted texts, they can easily overwhelm the commenting process, and massive opinion spam campaigns have made manually flagging duplicate comments a costly and time-consuming task for rulemaking agencies. To help agencies in their deduplication efforts, we developed and tested a highly accurate and cost-effective technique based on unsupervised deep machine learning. Our method used the graph augmented deep learning model to embed the comment texts into vectors, and then we applied the clustering method to the embedded vectors to identify similar comments. To test the efficacy of the proposed approach, we applied our technique to the raw data from the Federal Communications Commission’s (FCC) proposed rule change on Net Neutrality in 2017. Our study results demonstrate the superiority of the proposed deep learning technique at identifying duplicate comments, relative to other state-of-the-art deduplication approaches.

Read the paper · More papers on PaperTik