Identifying Duplicate Police Reports
Alan Firmiano, Ticiana L. Coelho da Silva · 2021 20th IEEE International Conference on Machine Learning and Applications (ICMLA) · 2021
Several crimes occur every day, and the first step in investigating these crimes begins with a police report. Victims report the criminal facts, which in turn must be detailed and contain accurate information about the incident or crime (e.g., factual, accurate, clear, concise, complete, and timely). In addition, the bulletin helps to safeguard the police operation itself, showing where the series of investigative operations that police agencies have been carrying out began. In cities with high crime rates, it is unfeasible to require the police to read and analyze all reported crime narratives. However, it would be helpful if employees could identify reports with similar modus operandi. Priority legal document retrieval is an information retrieval task used to retrieve past case documents related to specific cases and guide the police on how to act. Given a police report, the main objective of this work is to determine the most similar or duplicate police report. Another method is to encode the narrative as an embedded vector. In this article, we experimented with different pre-trained representations at the sentence level. We found the one that most effectively captures the semantic attributes of police report vocabulary and recognizes repeated reports. We are also investigating whether the summarized sentences identify duplicate police reports. Finally, we compare the effectiveness of the duplicated police report with the available sentence incorporation model trained in a large corpus. Our goal is to evaluate the performance of these embedding models (the one trained with our corpus and the pre-trained) to capture duplicate narratives.