Privacy-preserving collaborative deep learning methods for multiinstitutional training without sharing patient data
Abstract Although deep learning models have great promise for clinical applications, there are numerous obstacles to training effective deep learning models. There is a need for collecting large quantities of diverse data, which often can only be achieved through multiinstitutional collaborations. One approach to multiinstitutional studies is to build large central repositories, but this is hindered by concerns about data sharing, including patient privacy, data deidentification, regulation, intellectual property, and data storage. These challenges have made centrally hosting data less impractical. An alternative approach is to have the data hosted locally and have the model trained in a collaborative fashion. Depending on the collaborative learning approach, model weights, model gradients, or smashed data are shared instead of raw patient data. These approaches can also reduce the communication overhead while reducing the need to share private patient data. In this chapter, we will review and compare the current techniques for distributing learning, handling data heterogeneity, and preserving patient privacy for multiinstitutional applications.
Paper
Full text
Privacy-preserving collaborative deep learning methods for multiinstitutional training without sharing patient data
Semantic Scholar · Computer Science · 2021
Abstract
Abstract Although deep learning models have great promise for clinical applications, there are numerous obstacles to training effective deep learning models. There is a need for collecting large quantities of diverse data, which often can only be achieved through multiinstitutional collaborations. One approach to multiinstitutional studies is to build large central repositories, but this is hindered by concerns about data sharing, including patient privacy, data deidentification, regulation, intellectual property, and data storage. These challenges have made centrally hosting data less impractical. An alternative approach is to have the data hosted locally and have the model trained in a collaborative fashion. Depending on the collaborative learning approach, model weights, model gradients, or smashed data are shared instead of raw patient data. These approaches can also reduce the communication overhead while reducing the need to share private patient data. In this chapter, we will review and compare the current techniques for distributing learning, handling data heterogeneity, and preserving patient privacy for multiinstitutional applications.