Strengths
The writing is concise and the contributions are clear. The author proposes a new smoothness mechanism, generally an interesting idea, to integrate local and global information. Besides the quantitative experiments, there are some qualitative experiments such as attention maps that show the model's relative importance on different regions in a slide.
Weaknesses
There are three major concerns with this work:
1. The choice of encoder. The authors have used Resnet18 for RSNA and Camelyon16 and Resnet50 for PANDA. I could not find any justification for why different encoders have been used! Also, with the trend toward foundation models, transformer-based backbones are now of interest to the community. Therefore, for the sake of consistency, the author should use the same encoder for different datasets or report both resent50 and resnet18. And, for the sake of the generality of their work, they should add one encoder such as ViT or Swin as well (w/ ImageNet weights) to support the generality of their claims.
2. The body of research has been founded on local-to-global interactions. There are quite a few standard graph-based methodologies in the literature that the authors need to compare their work against. Currently, there are no representatives from those families of the MIL method in the benchmark. Two examples of such methods are (1) and (2).
3. There is a family of MIL methods in the literature that try to pseudo-label the instances during training, which is essentially equivalent to localization in this work. For instance, two of the most recent such methods are (3) and (4). The author should compare their method against these as they are essentially from the same family.
(1) Chen, Richard J., et al. "Whole slide images are 2d point clouds: Context-aware survival prediction using patch-based graph convolutional networks." Medical Image Computing and Computer Assisted Intervention–MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceedings, Part VIII 24. Springer International Publishing, 2021.
(2) Li, R., Yao, J., Zhu, X., Li, Y., Huang, J. (2018). Graph CNN for Survival Analysis on Whole Slide Pathological Images. In: Frangi, A., Schnabel, J., Davatzikos, C., Alberola-López, C., Fichtinger, G. (eds) Medical Image Computing and Computer Assisted Intervention – MICCAI 2018. MICCAI 2018. Lecture Notes in Computer Science(), vol 11071. Springer, Cham. https://doi.org/10.1007/978-3-030-00934-2_20
(3) Z. Shao et al., "LNPL-MIL: Learning from Noisy Pseudo Labels for Promoting Multiple Instance Learning in Whole Slide Image," 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 2023, pp. 21438-21438, doi: 10.1109/ICCV51070.2023.01965.
(4) Ren, Q. et al. (2023). IIB-MIL: Integrated Instance-Level and Bag-Level Multiple Instances Learning with Label Disambiguation for Pathological Image Analysis. In: Greenspan, H., et al. Medical Image Computing and Computer Assisted Intervention – MICCAI 2023. MICCAI 2023. Lecture Notes in Computer Science, vol 14225. Springer, Cham. https://doi.org/10.1007/978-3-031-43987-2_54