AttriMOT: Semantic-Aware Multimodal 3D Multi-Object Tracking with Attribute-Level Alignment

3D multi-object tracking (MOT) in complex and dynamic environments remains challenging due to the time-varying reliability of sensor modalities, severe occlusions, and the difficulty of distinguishing instances with similar appearances. Existing methods mainly rely on coarse category-level semantics or heuristic multimodal fusion strategies, which limits fine-grained instance discrimination and leads to unstable trajectory association under complex scenarios. Moreover, current 3D MOT frameworks generally lack the ability to leverage attribute-level semantic information for robust tracking and semantic-aware target retrieval. To address these limitations, we propose AttriMOT, a semantic-aware multimodal 3D MOT framework. Specifically, a category semantic anchoring and competition suppression mechanism is introduced to preserve discriminative fine-grained attribute information among visually similar instances. An attribute-level multimodal alignment module establishes structured correspondences across 3D geometry, 2D appearance, and textual semantics, enabling robust cross-modal representation learning. Furthermore, a parameter-free adaptive confidence fusion strategy dynamically balances LiDAR- and camera-derived trajectory confidence to improve tracking stability under varying environmental conditions. In addition, a semantic-aware trajectory selector is designed to support text-specified target retrieval and trajectory locking, enabling controllable semantic-guided 3D tracking. Extensive experiments on challenging 3D MOT benchmarks demonstrate that AttriMOT consistently outperforms state-of-the-art methods in tracking accuracy and robustness. In particular, AttriMOT achieves 1.33% improvement in HOTA and 0.54% improvement in MOTA compared with the best existing method, while also providing enhanced semantic controllability and text-guided tracking capability.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC