Creating Collaborative Data Representations Using Matrix Manifold Optimal Computation and Automated Hyperparameter Tuning
Keiyu Nosaka, Akiko Yoshise · 2023
Balancing performance and privacy in centralized machine learning, especially in sensitive domains such as medicine and finance, remains a significant issue. Concerns about data confidentiality result in a reluctance to share confidential data even among related entities, hindering integrated analysis despite its acknowledged benefits. Recent advancements in distributed data analysis have provided a solution to the concern, particularly through the framework of federated learning. This framework enables multiple clients to train models collaboratively while preserving the confidentiality of their data. Fed-DR-Filter extends federated learning by introducing the Privacy-Preserving Data Representation (PPDR) transformation algorithm that converts private datasets into a Global Data Representation (GDR) across multiple clients [4]. The GDR retains the correlation of the original data while ensuring its confidentiality. The PPDR is obtained by adding noise to the GDR, which is then used to filter noisy labels in federated learning environments based on the preserved correlation. Another recent development in distributed data analysis is Data Collaboration Analysis (DCA) [5]. It employs Collaborative Data Representation (CDR), a data representation almost identical to GDR, to directly generate high-performance models without the need for model sharing or iterative communication among clients. The present study builds the crucial role of CDRs and GDRs in DCA and Fed-DR-Filter by proposing a novel method for CDR creation with a solid theoretical foundation. The purpose is to enhance the performance of these techniques by providing an alternative and more effective algorithm for CDR and GDR creation. The primary contributions of this study include a novel optimization for CDR creation utilizing matrix manifold optimization and automated hyperparameter tuning. An algorithm for solving this problem is proposed with its empirical evaluation of artificial and real-world datasets within the DCA setting. The result shows substantial improvement in mean recognition performance.