Trustworthiness and explainability of a watermarking and machine learning-based system for image modification detection to combat disinformation
Andrea Rosales, Agnieszka Malanowska, Tanya Koohpayeh Araghi, Minoru Kuribayashi, Marcin Kowalczyk, Daniel Blanche, Wojciech Mazurczyk, David Megías · 2024
The widespread use of digital platforms, prioritising content based on engagement metrics and rewarding content creators accordingly, has contributed to the proliferation of disinformation and its far-reaching social and political impact. In addition, digital platforms often operate as black boxes, concealing their decision-making processes from users and prioritizing investor interests over ethical and social considerations. Consequently, this has contributed to the erosion of general trust in verification systems. To mitigate this issue, our project proposes a two-stage verification system. The first stage allows media industries to watermark their image and video content. The second stage involves implementing a machine-learning-based manipulation detection system for suspicious content. We present findings from an international user experience study, where potential online news consumers verified the authenticity of images on a prototype version of our system. In this paper, we reflect on critical issues of explainability addressed by participants in our user study and how we addressed this issue in the platform’s design.