Multimodal large language models (MLLMs) have demonstrated an extraordinary capacity to bridge textual and visual inputs. Nonetheless, MLLMs still face limitations in situated physical and social interactions in sensorially rich and multimodal real-world settings, where the embodied experience of a living organism appears fundamental. We suggest that the next frontiers for MLLM development require the incorporation of both internal and external embodiment-modeling not only external interactions with the world but also internal states and drives. Here, we describe mechanisms of internal and external embodiment in humans and relate these to current advances in MLLMs in the early stages of aligning to human representations. Our dual-embodied framework proposes to model interactions between these forms of embodiment in MLLMs so as to bridge the gap between multimodal data and world experience.
Paper
References (100)
Scroll for more · 38 remaining