This paper presents a novel robotic system leveraging Vision-Language-Action (VLA) models and ROS2 to autonomously address key needs in eldercare settings. The system interprets spoken commands through OpenAI's Whisper for realtime speech-to-text conversion and SpaCy for natural language understanding, detects objects via a Logitech webcam integrated with the PaliGemma vision-language model, and estimates spatial relationships using Intel's MiDaS monocular depth estimation algorithm. The robotic manipulation subsystem employs MoveIt2 for motion planning and a suction-based end effector for secure grasping, all orchestrated through a modular ROS2 architecture running on an NVIDIA Jetson AGX Orin platform. Experimental evaluations demonstrate a 90.2% task completion rate, 95.3% object detection accuracy, and 2.8-second average response latency across 150 trials with common eldercare objects. This work demonstrates the viability of edge-deployed VLA models for eldercare applications, offering a practical approach to enhancing independence through intuitive, voice-activated robotic assistance.
Paper
Full text
An Open VLA Powered Robot with ROS2 for Autonomous Elder Care Assistance
Semantic Scholar · 2025
Abstract
This paper presents a novel robotic system leveraging Vision-Language-Action (VLA) models and ROS2 to autonomously address key needs in eldercare settings. The system interprets spoken commands through OpenAI's Whisper for realtime speech-to-text conversion and SpaCy for natural language understanding, detects objects via a Logitech webcam integrated with the PaliGemma vision-language model, and estimates spatial relationships using Intel's MiDaS monocular depth estimation algorithm. The robotic manipulation subsystem employs MoveIt2 for motion planning and a suction-based end effector for secure grasping, all orchestrated through a modular ROS2 architecture running on an NVIDIA Jetson AGX Orin platform. Experimental evaluations demonstrate a 90.2% task completion rate, 95.3% object detection accuracy, and 2.8-second average response latency across 150 trials with common eldercare objects. This work demonstrates the viability of edge-deployed VLA models for eldercare applications, offering a practical approach to enhancing independence through intuitive, voice-activated robotic assistance.