Toward Embodied Intelligence: An Architecture for Natural Dialogue and Action Execution in Assistive Robots
This work presents a new architecture for an assistive mobile robot designed to support the elderly and individuals with disabilities in performing daily indoor tasks. The proposed framework integrates multimodal perception, language-based reasoning, and safety-aware action planning to enable natural and effective two-way communication between humans and robots. At its core, the system utilizes large language models (LLMs) for dialogue management, contextual understanding, and reasoning over fused sensory inputs, including vision, speech, and proprioceptive data. By combining speech recognition, object detection, and local memory modules, the robot not only interprets explicit user commands but also infers implicit intentions, predicts missing information, and requests clarifications when necessary. A dedicated safety layer filters and validates action sequences before execution, ensuring reliability and user safety. The architecture further incorporates short- and long-term memory structures, enabling the robot to maintain a dialogue history and semantic knowledge of the environment. This bidirectional interaction model allows the robot to generate both natural conversational responses and executable action plans in a context-aware manner. Preliminary implementation and testing demonstrate promising performance, bridging the gap between conversational AI and embodied robotic action in real-life assistive scenarios.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex