Dual-LVT: A Dual Attention Language-Vision Transformer for Tumor Segmentation

Accurate target volume delineation is critical for successful cancer treatment. Deep learning (DL) has driven significant advances in automatic delineation, and the emergence of large language models (LLMs) enables the use of multimodal data, such as clinical notes, to assist these algorithms. However, existing multimodal approaches often treat language as secondary to vision and fuse modalities in a coarse manner, underutilizing valuable textual information. We propose the Dual Attention Language–Vision Transformer (Dual-LVT), which incorporates language em-beddings from LLM-generated summaries of doctors' clinical notes. Dual-LVT separates modality-fused features from pure vision information and enables language features to directly generate coarse segmentation maps, guiding downstream tasks more effectively. We further introduce an Adaptive Hounsfield Unit Clipping (AdaHU) module, which dynamically adjusts HU ranges via a gating mechanism to handle scanner-induced heterogeneity. Evaluations on multiple public datasets for 2D segmentation and 3D volume delineation show that Dual-LVT achieves superior accuracy and robustness over state-of-the-art models, highlighting the potential of LLMs as primary feature extractors in automatic delineation systems. The code is available at https://github.com/Shira7z/Dual-LVT.git.

Paper

Full text

PDF

Dual-LVT: A Dual Attention Language-Vision Transformer for Tumor Segmentation

Semantic Scholar · Computer Science · 2025

Abstract

Accurate target volume delineation is critical for successful cancer treatment. Deep learning (DL) has driven significant advances in automatic delineation, and the emergence of large language models (LLMs) enables the use of multimodal data, such as clinical notes, to assist these algorithms. However, existing multimodal approaches often treat language as secondary to vision and fuse modalities in a coarse manner, underutilizing valuable textual information. We propose the Dual Attention Language–Vision Transformer (Dual-LVT), which incorporates language em-beddings from LLM-generated summaries of doctors' clinical notes. Dual-LVT separates modality-fused features from pure vision information and enables language features to directly generate coarse segmentation maps, guiding downstream tasks more effectively. We further introduce an Adaptive Hounsfield Unit Clipping (AdaHU) module, which dynamically adjusts HU ranges via a gating mechanism to handle scanner-induced heterogeneity. Evaluations on multiple public datasets for 2D segmentation and 3D volume delineation show that Dual-LVT achieves superior accuracy and robustness over state-of-the-art models, highlighting the potential of LLMs as primary feature extractors in automatic delineation systems. The code is available at https://github.com/Shira7z/Dual-LVT.git.

Similar papers

© 2026 NYSGPT2525 LLC