Research and Implementation of a Human-Computer Interaction AI Intelligent Robot Based on Speech
With the advancement of deep learning, language model training, and edge computing, speech-based human-robot interaction (HRI) systems are expanding their applications in manufacturing and medical escort service robots. In this paper, a modular speech-driven human-robot interaction system is proposed and implemented, which integrates speech recognition (ASR), natural language understanding (NLU), dialogue management, and action execution modules. To evaluate the engineering trade-offs of different technology options, three sets of comparative trials were designed and implemented: ASR (cloud-based commercial vs. local models), NLU (traditional statistical methods vs. Transformer fine-tuning), and dialogue management (rule-driven vs. reinforcement learning). Simulation experiments were conducted in the ROS/Gazebo environment. Subjective tests were also performed, with evaluation metrics including word error rate (WER), accuracy, F1-score, response latency, task completion rate, and user satisfaction. The results show that the Transformer-based NLU is significantly better than the traditional methods in semantic parsing. ASR in the cloud has obvious advantages in recognition quality, but the local model is more suitable for real-time control scenarios in terms of delay and privacy. Dialogue management is recommended to adopt a hybrid strategy of "rule first + reinforcement learning enhancement." Finally, the problems of system engineering deployment, model compression, edge-cloud collaboration, and ethical compliance were discussed.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex