Enhancing Data Integrity: Enforcing Ownership and Authorization Protocols in Machine Learning Model Development
Pranav Tyagi, Ayush Bhatt, Ayushi Agarwal, Ketan Sharma, Rajit Nair · 2024
There is a current lack of protection for multimedia data such as images, audio, video, and text from unauthorized machine learning training along with misuse of authors rights to their work; this paper therefore proposes a novel framework with the intention of giving the users more control and authorization for their data. It uses small distortions via simple mathematical operations which are still perceivable by the human eye while adding noise which will cause the current ML and DL models to fail. For images, by using random functions, pixel colors are changed in order to maintain picture quality but shift ML attention to the introduced patterns. Environmental sounds improve the audibility of audio and the human reader does not need to guess what the author means. In terms of quality, sine, cosine, and tangent transformations can be as applied to video frames. Text is encrypted by replacing like appearing Unicode characters that changes the machine learning vectors without losing content information. The results reveal that such micro-interferences allow for addressing the problem of protecting creators from misuse and unauthorized use of material by making it impossible for ML models to learn the initial data patterns which belong to the creators.