An Open VLA Powered Robot with ROS2 for Autonomous Elder Care Assistance

This paper presents a novel robotic system leveraging Vision-Language-Action (VLA) models and ROS2 to autonomously address key needs in eldercare settings. The system interprets spoken commands through OpenAI's Whisper for realtime speech-to-text conversion and SpaCy for natural language understanding, detects objects via a Logitech webcam integrated with the PaliGemma vision-language model, and estimates spatial relationships using Intel's MiDaS monocular depth estimation algorithm. The robotic manipulation subsystem employs MoveIt2 for motion planning and a suction-based end effector for secure grasping, all orchestrated through a modular ROS2 architecture running on an NVIDIA Jetson AGX Orin platform. Experimental evaluations demonstrate a 90.2% task completion rate, 95.3% object detection accuracy, and 2.8-second average response latency across 150 trials with common eldercare objects. This work demonstrates the viability of edge-deployed VLA models for eldercare applications, offering a practical approach to enhancing independence through intuitive, voice-activated robotic assistance.

Paper

Full text

PDF

An Open VLA Powered Robot with ROS2 for Autonomous Elder Care Assistance

Semantic Scholar · 2025

Abstract

This paper presents a novel robotic system leveraging Vision-Language-Action (VLA) models and ROS2 to autonomously address key needs in eldercare settings. The system interprets spoken commands through OpenAI's Whisper for realtime speech-to-text conversion and SpaCy for natural language understanding, detects objects via a Logitech webcam integrated with the PaliGemma vision-language model, and estimates spatial relationships using Intel's MiDaS monocular depth estimation algorithm. The robotic manipulation subsystem employs MoveIt2 for motion planning and a suction-based end effector for secure grasping, all orchestrated through a modular ROS2 architecture running on an NVIDIA Jetson AGX Orin platform. Experimental evaluations demonstrate a 90.2% task completion rate, 95.3% object detection accuracy, and 2.8-second average response latency across 150 trials with common eldercare objects. This work demonstrates the viability of edge-deployed VLA models for eldercare applications, offering a practical approach to enhancing independence through intuitive, voice-activated robotic assistance.

Similar papers

© 2026 NYSGPT2525 LLC