Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models

Automatic target recognition (ATR) is crucial for safety-critical tasks such as navigation and surveillance, particularly in demanding military scenarios with unfamiliar terrains, harsh environments, and unseen object categories. Conventional and open-world object detectors often fail in these contexts, lacking exposure to such novel conditions. Meanwhile, Large Vision-Language Models (LVLMs) exhibit zero-shot recognition capabilities across diverse settings yet struggle with precise localization. We address these limitations by combining the localization strength of open-world detectors with the recognition confidence of LVLMs, creating a robust pipeline for zero-shot ATR in novel domains and classes. Our study compares several LVLMs on underrepresented military vehicles, examining factors such as distance range, modality, and prompting strategies. These findings provide insights for developing more reliable ATR systems in uncharted environments.

Paper

Similar papers

© 2026 NYSGPT2525 LLC