Privacy-preserving collaborative deep learning methods for multiinstitutional training without sharing patient data

Abstract Although deep learning models have great promise for clinical applications, there are numerous obstacles to training effective deep learning models. There is a need for collecting large quantities of diverse data, which often can only be achieved through multiinstitutional collaborations. One approach to multiinstitutional studies is to build large central repositories, but this is hindered by concerns about data sharing, including patient privacy, data deidentification, regulation, intellectual property, and data storage. These challenges have made centrally hosting data less impractical. An alternative approach is to have the data hosted locally and have the model trained in a collaborative fashion. Depending on the collaborative learning approach, model weights, model gradients, or smashed data are shared instead of raw patient data. These approaches can also reduce the communication overhead while reducing the need to share private patient data. In this chapter, we will review and compare the current techniques for distributing learning, handling data heterogeneity, and preserving patient privacy for multiinstitutional applications.

Paper

Full text

PDF

Privacy-preserving collaborative deep learning methods for multiinstitutional training without sharing patient data

Semantic Scholar · Computer Science · 2021

Abstract

Abstract Although deep learning models have great promise for clinical applications, there are numerous obstacles to training effective deep learning models. There is a need for collecting large quantities of diverse data, which often can only be achieved through multiinstitutional collaborations. One approach to multiinstitutional studies is to build large central repositories, but this is hindered by concerns about data sharing, including patient privacy, data deidentification, regulation, intellectual property, and data storage. These challenges have made centrally hosting data less impractical. An alternative approach is to have the data hosted locally and have the model trained in a collaborative fashion. Depending on the collaborative learning approach, model weights, model gradients, or smashed data are shared instead of raw patient data. These approaches can also reduce the communication overhead while reducing the need to share private patient data. In this chapter, we will review and compare the current techniques for distributing learning, handling data heterogeneity, and preserving patient privacy for multiinstitutional applications.

Similar papers

© 2026 NYSGPT2525 LLC