Equi-ViT: Rotational Equivariant Vision Transformer for Robust Histopathology Analysis

Vision Transformers (ViTs) have gained rapid adoption in computational pathology for their ability to model long-range dependencies through self-attention, addressing the limitations of convolutional neural networks that excel at local pattern capture but struggle with global contextual reasoning. Recent pathology-specific foundation models have further advanced performance by leveraging large-scale pretraining. However, standard ViTs remain inherently non-equivariant to transformations such as rotations and reflections, which are ubiquitous variations in histopathology imaging. To address this limitation, we propose Equi-ViT, which integrates an equivariant convolution kernel into the patch embedding stage of a ViT architecture, imparting built-in rotational equivariance to learned representations. Equi-ViT achieves superior rotation-consistent patch embeddings and stable classification performance across image orientations. Our results on a public colorectal cancer dataset demonstrate that incorporating equivariant patch embedding enhances data efficiency and robustness, achieving an accuracy (mean ± SD) of $86.8 \pm 0.59 \%$ across rotations and outperforming standard ViT performance ($83.1 \pm 6.93 \%$), suggesting that equivariant transformers could potentially serve as more generalizable backbones for the application of ViT in histopathology such as digital pathology foundation models. The code is available at https://github.com/fyc423/Equi-ViT.

Paper

References (14)

Scroll for more · 2 remaining

Similar papers

© 2026 NYSGPT2525 LLC