Gaming Without Harm: AI-Driven Content Moderation to Improve Safety in Roblox

Arda C. Varol, Larissa T. Coffee · 2025

Adolescents extensively use social media and gaming platforms, which are prone to exposure to inappropriate content. Despite existing chat moderation systems, such as Roblox's hashtag-based filters, users bypass them using altered expletive variations. This project proposes an enhanced Natural Language Processing (NLP) solution to improve content moderation by detecting expletive words, and their near-identical variations. Utilizing datasets from Carnegie Mellon University and GitHub, synthetic variations of expletives are generated using the Levenshtein Edit-Distance algorithm. Then, Jaro similarity measure is employed to identify subtle linguistic modifications on expletive content. The solution's accuracy is tested against Roblox's current filtering system, evaluating improvements in detecting inappropriate content. While Roblox's chat moderation system had a success rate of 44.6% when detecting expletive content, the proposed solution (ChatGuard) achieved 82.5% accuracy. Overall, the findings result in enhancing online safety and mitigate risks to adolescent users while balancing the need for effective moderation with freedom of expression.

Read the paper · More papers on PaperTik