Weakly supervised segmentation with cross-modality equivariant constraints

Weakly supervised learning has emerged as an appealing alternative to\nalleviate the need for large labeled datasets in semantic segmentation. Most\ncurrent approaches exploit class activation maps (CAMs), which can be generated\nfrom image-level annotations. Nevertheless, resulting maps have been\ndemonstrated to be highly discriminant, failing to serve as optimal proxy\npixel-level labels. We present a novel learning strategy that leverages\nself-supervision in a multi-modal image scenario to significantly enhance\noriginal CAMs. In particular, the proposed method is based on two observations.\nFirst, the learning of fully-supervised segmentation networks implicitly\nimposes equivariance by means of data augmentation, whereas this implicit\nconstraint disappears on CAMs generated with image tags. And second, the\ncommonalities between image modalities can be employed as an efficient\nself-supervisory signal, correcting the inconsistency shown by CAMs obtained\nacross multiple modalities. To effectively train our model, we integrate a\nnovel loss function that includes a within-modality and a cross-modality\nequivariant term to explicitly impose these constraints during training. In\naddition, we add a KL-divergence on the class prediction distributions to\nfacilitate the information exchange between modalities, which, combined with\nthe equivariant regularizers further improves the performance of our model.\nExhaustive experiments on the popular multi-modal BRATS dataset demonstrate\nthat our approach outperforms relevant recent literature under the same\nlearning conditions.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC