SATISFY EASY LADDERS: Accessible machine learning to facilitate analysis of larger sensitive text data samples
Nicola Fox · CrimRxiv · 2025
Large-scale data analysis has improved our understanding of crime and policing, yet computational methods largely focus on numerical data, overlooking crucial context and mechanisms. Narrative text offers depth but is often analysed manually in small samples. Machine learning (ML) techniques provide solutions, but complex algorithms demand extensive training data and expertise, restricting their use. This article explores simpler approaches with smaller training datasets, enabling broader adoption by non-specialists. To this end, this article presents a worked-through example of ML model development for classifying 193 child safeguarding practice review documents to identify cases mentioning exploitation, missing incidents, school exclusion, and special educational needs and disabilities. It details key steps, from text import, sentence splitting and data leakage avoidance to tokenisation, embedding, model evaluation, and human-in-the-loop evaluation, and presents the mnemonics ‘SATISFY’ and ‘EASY LADDERS’ for helping researchers to remember key considerations and steps. Wider application could help researchers explore mechanisms and contexts in crime and security studies at a generalisable scale.