Preserving Privacy During Reinforcement Learning With AI Feedback

David Gao, Ian J. Miller, Ali Ataeemh Allami, Dan Lin · 2024

Leveraging the scalable efficacy of reinforcement learning from AI feedback (RLAIF), large language models (LLMs) can be refined toward human intent alignment. While current paradigms employ LLMs to annotate and refine model outputs, the area of privacy preservation within RLAIF, particularly in sensitive domains like healthcare where data privacy is indispensable, remains inadequately explored. Addressing this gap, we propose PriMa, a novel data obfuscation algorithm integrated into the RLAIF pipeline, empowering organizations to harvest the benefits of RLAIF while protecting privacy concerns. Notably, this work pioneers the investigation of privacy vulnerabilities, specifically membership inference attacks, within practical RLAIF implementations. Our empirical evaluations demonstrate PriMa’s effectiveness in preventing membership inference attacks while significantly enhancing model alignment compared to the baseline RLAIF architecture.

Read the paper · More papers on PaperTik