Speech Recognition and LLM Performance in Elderly Care Home Conversations

Conversational robots offer promise in elderly care, but dialectal speech poses challenges for automatic speech recognition (ASR). This study evaluates a conversational robot integrating Microsoft Azure ASR and GPT-4o in real-world interactions with elderly users. Results show that ASR accuracy varied significantly (95% for standard French, 45–56% for Dutch dialects (e.g., West Flemish), often leading to transcription errors. Despite this, the LLM restored conversational coherence in 44–52% of misrecognitions, while users contributed 25–35% of repairs. Comparative ASR analysis showed Whisper’s superior dialectal robustness (28% WER) but high latency. Interaction durations ranged from 17 to 45 minutes, with participants perceiving the robot as understanding them despite ASR challenges. This study uniquely integrates ASR performance, LLM recovery, and user adaptation, highlighting the need for hybrid ASR solutions and context-aware dialogue management in elderly-care robots. Findings highlight the importance of context-aware dialogue management, hybrid ASR strategies, and user-driven conversational adaptation for effective human-robot interactions in real-world settings.

Paper

Full text

PDF

Speech Recognition and LLM Performance in Elderly Care Home Conversations

OpenAlex · AI in Service Interactions · 2025

Abstract

Conversational robots offer promise in elderly care, but dialectal speech poses challenges for automatic speech recognition (ASR). This study evaluates a conversational robot integrating Microsoft Azure ASR and GPT-4o in real-world interactions with elderly users. Results show that ASR accuracy varied significantly (95% for standard French, 45–56% for Dutch dialects (e.g., West Flemish), often leading to transcription errors. Despite this, the LLM restored conversational coherence in 44–52% of misrecognitions, while users contributed 25–35% of repairs. Comparative ASR analysis showed Whisper’s superior dialectal robustness (28% WER) but high latency. Interaction durations ranged from 17 to 45 minutes, with participants perceiving the robot as understanding them despite ASR challenges. This study uniquely integrates ASR performance, LLM recovery, and user adaptation, highlighting the need for hybrid ASR solutions and context-aware dialogue management in elderly-care robots. Findings highlight the importance of context-aware dialogue management, hybrid ASR strategies, and user-driven conversational adaptation for effective human-robot interactions in real-world settings.

Similar papers

© 2026 NYSGPT2525 LLC