Research on Dialogue Interaction Mechanism of Multimodal Speech Agents Driven by Large Language Models
This study addresses critical limitations of traditional speech agents in cross-modal semantic understanding, context-aware management, and dynamic response generation. We propose a novel LLM-Driven Multimodal Speech Agent (LMA) dialogue interaction mechanism, featuring: (1) a hierarchical cross-modal fusion architecture with dynamic attention mechanisms; (2) a cross-modal semantic alignment method combining LLM-guided contrastive learning; (3) a context-aware dialogue state manager with memory compression and dynamic attention; (4) a reinforcement learning-based dynamic multimodal response generation strategy. Experiments are conducted on three public benchmark datasets -- Fluent Speech Commands (FSC), DailyDialog, and Multimodal-E4 -- and a real-world self-collected dataset. Results demonstrate that LMA achieves intent recognition accuracy of 92.1% (vs. GPT-4o's 91.3%), dialogue coherence of 0.91 (vs. GPT-4o's 0.87), and task completion rate of 87.3% (vs. GPT-4o's 85.6%), significantly outperforming four baseline methods and two commercial systems. The ablation study validates the effectiveness of each module.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex