Unimodal Training-Multimodal Prediction: Cross-modal Federated Learning with Hierarchical Aggregation

Multimodal learning has significantly advanced the extraction of features from varied data sources, enhancing model performance. Federated learning (FL) complements this by enabling collaborative training while maintaining data privacy. The fusion of these two fields, multimodal federated learning, offers considerable promise. Yet, standard methods often incorrectly assume that each node in the FL network has a full complement of multimodal data, which is rare in real-world applications. In our study, we present a novel architecture designed to surmount these challenges, termed the Unimodal Training - Multimodal Prediction (UTMP) framework, positioned within the multimodal federated learning paradigm. Our proposed model, the HA-Fedformer, is a transformer-based model crafted to facilitate unimodal training on the client-side using exclusively unimodal datasets and to execute multimodal inference by synthesizing insights from multiple clients. Our HA-Fedformer model effectively handles non-IID data through a novel uncertainty-aware aggregation technique and layer-wise Markov Chain Monte Carlo sampling in local encoders. It also resolves misaligned language sequences via cross-modal decoder aggregation, capturing correlations between decoders trained on different modalities. Our comprehensive evaluations conducted on widely recognized sentiment analysis benchmarks demonstrate the superiority of the HA-Fedformer. The results show that our model achieves a substantial uplift in performance.

Paper

Similar papers

© 2026 NYSGPT2525 LLC