Optimizing RLHF Reward Models with Fairness Constraints
Jingyi Zhao, Yiyong Lei, Weili Zheng, Sitong Lai, Hui Ding, Chengcheng Zhao · 2025
This paper proposes a dynamic fairness-constrained optimization method for RLHF reward models, which addresses the limitations of fixed-constraint approaches in flexibility and efficiency through an innovative constraint-weight adaptation mechanism. We design two dynamic adjustment strategies (linear and exponential), with the linear strategy ($\lambda$ranging from 0.0042 to 0.0344) coupled with early stopping demonstrating optimal performance. Experiments on a large-scale dialogue dataset (153,850 training samples) show our method reduces inter-group fairness disparity by 53.5 % (from 0.0621 to 0.0289) while cutting 65 % training time via early stopping. Ablation studies confirm the effectiveness of dynamic constraints in balancing model performance and fairness, particularly showing advantages in sensitive scenarios like medical Q&A. This work provides a crucial technical solution for developing fair and efficient RLHF systems.