Swarm Intelligence in Geo-Localization: A Multi-Agent Large Vision-Language Model Collaborative Framework

Visual geo-localization demands in-depth knowledge and advanced reasoning skills to associate images with precise real-world geo-graphic locations. Existing image database retrieval methods are limited by the impracticality of storing sufficient visual records of global landmarks. Recently, Large Vision-Language Models (LVLMs) have demonstrated the capability of geo-localization through Visual Question Answering (VQA), enabling a solution that does not require external geo-tagged image records. However, the performance of a single LVLM is still limited by its intrinsic knowledge and reasoning capabilities. To address these challenges, we introduce smileGeo, a novel visual geo-localization framework that leverages multiple Internet-enabled LVLM agents operating within an agent-based architecture. By facilitating inter-agent communication, smileGeo integrates the inherent knowledge of these agents with additional retrieved information, enhancing the ability to effectively localize images. Furthermore, our framework incorporates a dynamic learning strategy that optimizes agent communication, reducing redundant interactions and enhancing overall system efficiency. To validate the effectiveness of the proposed framework, we conducted experiments on three different datasets, and the results show that our approach significantly outperforms current state-of-the-art methods. The source code is available at https://github.com/Applied-Machine-Learning-Lab/smileGeo.

Paper

References (56)

Scroll for more · 38 remaining

Similar papers

Reviewer Jv7J4/10 · confidence 4/52024-07-12

Summary

Proposed a graph based learnable multi-agent framework. The framework consists of multiple stages : Forwarding: Election (K: Answer agents; R: Reviewer) -> Review -> K Discuss till a final conclusion is reached. Proposed a mechanism to learn the graph connections dynamically. The major Contributions Introduced in the paper: (A) A new swarm intelligence geo-local framework smileGeo; (B) Dynamic learning strategy; (C) A new Geo-dataset (test mainly).

Strengths

The major strengths of the proposed smileGeo frameworks are: (a) the learnable Graph based communication strategy seems works well empirically. In table 2, authors demonstrated that it helps achieve better acc, but lower average token costs. (b) The proposed method is also scalable as shown in table 3. (c) Used attention-based GNN to predict optimal connections and optimal election. Also empirically justified the effectiveness of attention based GNN. (d) Also constructed Simple rules of updating edges(connections) that works well in practice.

Weaknesses

The major weaknesses are as follows: (a) Comparisons with baselines seems unfair. (b) Missing details of the evaluation setup, metrics, etc.

Questions

Here are my questions to the authors: (a) Table 1 involves comparison between open/closed source single LVLMs with smileGeo-single. However, smileGeo appears to primarily focus on a multi-agent framework, without introducing any new single LVLM architectures. (b) The comparative results of different agent frameworks without web searching are reported in table 2. How are 'acc' and 'tks' determined for each framework? Which types of LVLMs are employed? Did they aggregate all LVLMs and calculate averages per framework, or utilize the best-performing LVLM specific to each framework? It's important not to unfairly advantage one framework over others by using superior LVLMs. (c) Question about the comparison with LLM/LVLM-based agent frameworks: For the integration frameworks you compared in Table 2, what specific LVLMs were integrated within the LLM-Blender, LLM Debate, and smileGeo frameworks? Did you use the same LVLM combinations for the different frameworks in the comparison? Different LVLM combinations may have different underlying behavior on metrics.

Rating

4

Confidence

4

Soundness

3

Presentation

3

Contribution

3

Limitations

yes, the authors adequately addressed the limitations.

Reviewer EWQq6/10 · confidence 4/52024-07-14

Summary

This works proposes a new visual geo-localization framework with multiple LVLM (Large Vision Language Model) agents. The agents communicate with each other to estimate the geo-location of the input image. A dynamic learning strategy is proposed to optimize the communication patterns among agents to improve efficiency. The method is evaluated on the proposed GeoGlobe dataset.

Strengths

+ The idea of tacking worldwide city-level geo-localization with multiple LVLM agents is very interesting. + The result is surprisingly good with zero-shot setting, which is even better than powerful close-source models. + Detailed comparison with other agent-based methods is provided. The ablation study on the number of agents is also very detailed. + The writing is easy to follow.

Weaknesses

- The authors could make the geo-localization setting more clear in the introduction, for example, the paper focuses on worldwide city-level geo-localization. There are lots of different settings for geo-localization problem and this could be confusing for some researchers. - This paper provides a comparison with three traditional geo-localization methods, i.e., NetVLAD, GeM, and CosPlace. However, these three methods are either retrieval-based landmark matching methods or fine-grained classification-based place recognition methods. It would be better to provide a direct comparison with worldwide geo-localization method on city-level setting, e.g., [A]. Although I believe LVLM-based method is better at this setting, a comparison can make it more convincing. [A] Pramanick, Shraman, et al. "Where in the world is this image? transformer-based geo-localization in the wild." European Conference on Computer Vision. Cham: Springer Nature Switzerland, 2022. - There are only two qualitative results in the appendix. Given that the accuracy is over 60%, it should be easy to find successful and failed cases to demonstrate the actual output cases of the proposed methods. It can also better illustrate how multiple agents help the geo-localization process. - There are also some existing worldwide geo-localization datasets that could be used for more comprehensive evaluation, e.g., IM2GPS3K, YFCC4K.

Questions

See the weaknesses.

Rating

6

Confidence

4

Soundness

4

Presentation

3

Contribution

4

Limitations

The authors mentioned the limitations in the checklist.

Reviewer mcak4/10 · confidence 3/52024-07-22

Summary

The paper introduces smileGeo, a novel framework for visual geo-localization, which involves identifying the geographic location of an image. The authors argue that while Large Vision-Language Models (LVLMs) show promise in this area, their individual performance is limited. SmileGeo leverages the concept of "swarm intelligence" by enabling multiple LVLMs to collaborate and refine their location predictions through a multi-stage review process. To enhance efficiency, the framework incorporates a dynamic learning strategy that optimizes the selection of LVLMs for each image. Furthermore, the paper introduces "GeoGlobe," a new dataset designed to evaluate visual geo-localization models in open-world scenarios where many images depict locations not seen during training. Experimental results demonstrate that smileGeo outperforms existing single LVLMs and image retrieval methods, highlighting the effectiveness of collaborative learning for visual geo-localization.

Strengths

* The idea of using an ensemble of networks/agents for geolocalization is interesting and novel. The authors propose a graph-based social network to enable collaboration between the agents. * The ability to search the internet and provide the agents with relevant information is interesting and improves the performance on the task of geolocalization. * The paper proposes GeoGlobe, a new dataset for benchmarking models on the task of geo-localizing landmarks. The dataset could be utilized in future for other learning based geospatial tasks.

Weaknesses

* The paper only seems to tackle the problem of geolocalizing **landmark images**. While this is a challenging problem, the current literature [1, 2, 3] has already tried to address the problem of geolocalizing arbitrary ground-level images. The latter problem requires learning sophisticated geographic and visual features. I think even searching the internet cannot effectively solve the geolocalization problem for non-landmark images. * Limited applicability: The framework is built entirely upon the capabilities of different LVLMs (e.g. GPT4, LLaVA, etc). It seems the framework cannot generalize beyond the training data used for training LLMs. * The work fails to address the practical applications and real-life use cases of the framework. Why do we require such a framework? * The limitation and failure cases are not adequately mentioned in the paper. [1] Vivanco Cepeda, Vicente, Gaurav Kumar Nayak, and Mubarak Shah. "Geoclip: Clip-inspired alignment between locations and images for effective worldwide geo-localization." Advances in Neural Information Processing Systems 36 (2023). [2] Haas, Lukas, Michal Skreta, Silas Alberti, and Chelsea Finn. "Pigeon: Predicting image geolocations." In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12893-12902. 2024. [3] Berton, Gabriele, Carlo Masone, and Barbara Caputo. "Rethinking visual geo-localization for large-scale applications." In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4878-4888. 2022.

Questions

* At present, each agent sees the same information. It might be interesting to incorporate different kinds of information that is revealed differently to the agents, such as multi-view images or panorama images. * How much compute time is used for a single inference run? * Why does the performance of some LLMs decrease with web searching?

Rating

4

Confidence

3

Soundness

4

Presentation

3

Contribution

3

Limitations

Limitations are insufficiently addressed in the paper. The future works mentioned in the conclusion are vague and fail to specify specific future directions for the work.

Reviewer mcak2024-08-11

I thank the authors for their responses. I have a few clarifying questions: [1] What is so __unique__ about the framework that it could only be employed for geolocalization? I think the proposed framework is very general and not unique to geolocalization. [2] "A single LVLM indeed faces these challenges, as many large models attempt to consume vast amounts of data for pre-training. Our motivation is to address the biases in pre-trained LVLMs by combining their strengths." How do you ensure that the LLM agents have sufficient world knowledge that is relevant for the task of geolocalization? Are there any empirical studies to prove the statement? [3] Regarding practical application: How can the framework be used in robot navigation? Robot navigation requires fast response times for decision-making at each stage. The authors have mentioned that "99% Response Time for smileGeo is less than 25 seconds." This does not make the framework scalable for real-time applications.

Authorsrebuttal2024-08-11

**Response to Q1:** Thank you for your questions. Geo-localization is a complex task that requires extensive geospatial knowledge and strong reasoning abilities. LVLMs offer a novel approach to visual geo-localization by leveraging their powerful visual question-answering (VQA) capabilities, eliminating the need for external geo-tagged image records. This motivation led us to design a discussion framework around LVLMs to fully utilize their strengths and achieve better results. Since different LVLMs have different memory and reasoning capabilities for geo-localization tasks, we designed an LLM agent selection module in the proposed framework, which can select the most suitable agents for geo-localization of the target image for discussion, thus improving the efficiency of the framework. When the selected LLM agents cannot reach a high-confidence conclusion, they can autonomously call an internet-based geographic image search tool to supplement it with additional positioning information. We also appreciate your acknowledgment of our framework’s potential as a general approach that could be extended to other fields, reinforcing that the underlying concept of the LLM-based discussion framework is versatile. In the revised version, we will mention our intention to explore its application to other areas in future work. **Response to Q2:** Thank you for your questions. This was the motivation behind our first comparative experiment, where we compared different single LVLMs in Table 1. Even without retrieval assistance, most closed-source large model agents (such as GPT-4V and Gemini-1.5-pro) and some open-source large models (like Qwen-VL) achieved higher experimental accuracy than some image retrieval-based methods (as shown in Table 3). This demonstrates that LVLMs inherently possess the ability to analyze and process geo-location data, as well as the capacity to retain geo-location knowledge. Furthermore, our framework allows LVLM agents to search the internet and obtain sufficient world knowledge directly. As mentioned in Section 4.2, "models with larger parameters, such as llava–1.6–34b, demonstrate superior reasoning abilities compared to smaller models," leading to significant improvements in accuracy and outperforming traditional retrieval-based geo-localization methods. These experiments confirm that LLM agents, particularly closed-source large models and LVLMs with larger parameters, exhibit strong memory and reasoning capabilities for geo-localization tasks, both independently and with additional geo-location information. Additionally, there are also many reports verifying that using LVLMs for geo-tagging has gained widespread acceptance; please check the following links [1-3]. [1] https://x.com/itsandrewgao/status/1785827031131001243 [2] https://lingoport.com/i18n-term/llm/ [3] https://www.assemblyai.com/blog/llm-use-cases/ **Response to Q3:** Thank you for your concerns. In this paper, we propose a framework that effectively addresses the geo-localization task, which could be a critical component of robot navigation. Robot navigation typically involves many stages, such as localization, trajectory planning, and execution of the planned route. Tasks like localization and path planning often prioritize accuracy over real-time processing, especially in city-level navigation, where incorrect trajectory planning can lead to significant resource wastage. Several studies [1][2] utilize LMM/LVLM agents for UAV dispatching, a specific aspect of robot navigation. While inference using LMM/LVLM is known to be very time-consuming, the successful application of methods in these studies indicates promising prospects for LMM/LVLM-based geo-localization. Additionally, we believe that as LLM technology advances—through methods like quantization, distillation of open-source LLM agents, and calling of more lightweight and faster closed-source LLM agents (e.g., GPT-4o-mini)—our proposed framework will soon be capable of real-time responses. We will include this prospect in the revised paper. [1] Liu, S., Zhang, H., Qi, Y., Wang, P., Zhang, Y., & Wu, Q. (2023). Aerialvln: Vision-and-language navigation for uavs. In Proceedings of the IEEE/CVF International Conference on Computer Vision (pp. 15384-15394). [2] Zhao, H., Pan, F., Ping, H., & Zhou, Y. (2023). Agent as Cerebrum, Controller as Cerebellum: Implementing an Embodied LMM-based Agent on Drones. arXiv preprint arXiv:2311.15033.

Reviewer mcak2024-08-12

I thank the authors for providing additional clarifications. However, after reading the responses, I am still unsure about the motivation for using such a general framework for the task of geolocalization. The response time of 25 seconds seems significant and the dependence of the framework on closed-sourced LLMs raises questions about whether the framework is cost-effective. Hence, I shall keep my rating unchanged.

Reviewer EWQq2024-08-12

Thanks for the rebuttal. It addresses the concerns and I will keep the rating.

Authorsrebuttal2024-08-12

As we approach the end of the author-reviewer discussion period, we respectfully wish to check in and ensure that our rebuttal has effectively addressed your concerns regarding our paper. Should you have any remaining questions or need further clarification or additional experimental results, please do not hesitate to let us know. We appreciate the thoughtful reviews and the time you’ve invested in providing us with valuable feedback to improve our work. If you believe that our responses have sufficiently addressed the issues raised, we kindly ask you to consider the possibility of raising the score.

Reviewer Jv7J2024-08-13

I appreciate the authors' response. However, I remain unconvinced by their explanation regarding the implementation of the majority rule after a certain number of discussion rounds when consensus cannot be reached. The method for determining the number of discussion rounds requires further analysis, as it is a crucial aspect of their proposed framework. This situation is likely to occur frequently in practice. Additionally, after reviewing the author's discussion with the reviewer mcak, I agree with most of the concerns raised by the reviewer mcak. I find the authors' response to be unconvincing, especially the motivation and cost issues. Therefore, I also recommend a borderline reject.

Program Chairsdecision2024-09-25

Decision

Reject

© 2026 NYSGPT2525 LLC