Smart Mobility With Agent-Based Foundation Models: Towards Interactive and Collaborative Intelligent Vehicles

This letter reports the insights gained during a Distributed/Decentralized Hybrid Workshop on Foundation/Infrastructure Intelligence (FII), where we discussed the evolving role of Foundation Models in the field of intelligent vehicles. These models, pre-trained on multimodal data, have emerged as pivotal in the landscape of intelligent vehicles by leveraging their capabilities for high-level reasoning. Ongoing research focuses on these models to further improve scene perception and decision-making, aiming to develop adaptive systems for robot navigation and autonomous driving. However, for smart mobility across the Cyber-Physical-Social space, foundation intelligence should learn human-level knowledge to perform sophisticated interactions and collaborations based on human feedback. Agent-based Foundation Models, as the new training paradigm, can generate cross-domain actions consistent with perception information, paving the way to realize interactive and collaborative agents. This letter discusses the challenges of enhancing and leveraging the scene understanding and spatial reasoning capabilities of the pre-trained foundation model for smart mobility. It also offers insights into the embodied employment of foundation and infrastructure intelligence in enhancing multimodal interactions between robots, environments, and humans.

Paper

Full text

PDF

Smart Mobility With Agent-Based Foundation Models: Towards Interactive and Collaborative Intelligent Vehicles

Semantic Scholar · Engineering · 2024

Abstract

This letter reports the insights gained during a Distributed/Decentralized Hybrid Workshop on Foundation/Infrastructure Intelligence (FII), where we discussed the evolving role of Foundation Models in the field of intelligent vehicles. These models, pre-trained on multimodal data, have emerged as pivotal in the landscape of intelligent vehicles by leveraging their capabilities for high-level reasoning. Ongoing research focuses on these models to further improve scene perception and decision-making, aiming to develop adaptive systems for robot navigation and autonomous driving. However, for smart mobility across the Cyber-Physical-Social space, foundation intelligence should learn human-level knowledge to perform sophisticated interactions and collaborations based on human feedback. Agent-based Foundation Models, as the new training paradigm, can generate cross-domain actions consistent with perception information, paving the way to realize interactive and collaborative agents. This letter discusses the challenges of enhancing and leveraging the scene understanding and spatial reasoning capabilities of the pre-trained foundation model for smart mobility. It also offers insights into the embodied employment of foundation and infrastructure intelligence in enhancing multimodal interactions between robots, environments, and humans.

Similar papers

© 2026 NYSGPT2525 LLC