Smart Mobility With Agent-Based Foundation Models: Towards Interactive and Collaborative Intelligent Vehicles
This letter reports the insights gained during a Distributed/Decentralized Hybrid Workshop on Foundation/Infrastructure Intelligence (FII), where we discussed the evolving role of Foundation Models in the field of intelligent vehicles. These models, pre-trained on multimodal data, have emerged as pivotal in the landscape of intelligent vehicles by leveraging their capabilities for high-level reasoning. Ongoing research focuses on these models to further improve scene perception and decision-making, aiming to develop adaptive systems for robot navigation and autonomous driving. However, for smart mobility across the Cyber-Physical-Social space, foundation intelligence should learn human-level knowledge to perform sophisticated interactions and collaborations based on human feedback. Agent-based Foundation Models, as the new training paradigm, can generate cross-domain actions consistent with perception information, paving the way to realize interactive and collaborative agents. This letter discusses the challenges of enhancing and leveraging the scene understanding and spatial reasoning capabilities of the pre-trained foundation model for smart mobility. It also offers insights into the embodied employment of foundation and infrastructure intelligence in enhancing multimodal interactions between robots, environments, and humans.
Paper
Full text
Smart Mobility With Agent-Based Foundation Models: Towards Interactive and Collaborative Intelligent Vehicles
Semantic Scholar · Engineering · 2024
Abstract
This letter reports the insights gained during a Distributed/Decentralized Hybrid Workshop on Foundation/Infrastructure Intelligence (FII), where we discussed the evolving role of Foundation Models in the field of intelligent vehicles. These models, pre-trained on multimodal data, have emerged as pivotal in the landscape of intelligent vehicles by leveraging their capabilities for high-level reasoning. Ongoing research focuses on these models to further improve scene perception and decision-making, aiming to develop adaptive systems for robot navigation and autonomous driving. However, for smart mobility across the Cyber-Physical-Social space, foundation intelligence should learn human-level knowledge to perform sophisticated interactions and collaborations based on human feedback. Agent-based Foundation Models, as the new training paradigm, can generate cross-domain actions consistent with perception information, paving the way to realize interactive and collaborative agents. This letter discusses the challenges of enhancing and leveraging the scene understanding and spatial reasoning capabilities of the pre-trained foundation model for smart mobility. It also offers insights into the embodied employment of foundation and infrastructure intelligence in enhancing multimodal interactions between robots, environments, and humans.