The widespread adoption of voice-activated systems has revolutionized human-computer interaction, yet existing solutions predominantly support English, limiting accessibility for a linguistically diverse population like India. This research presents a novel approach to building a Multi-Lingual Wake Word Detection System with Speaker Authentication tailored to Indian languages. By integrating robust wake word recognition with speaker verification, the system ensures both accessibility and security. The architecture leverages Mel-spectrogram-based feature extraction and deep learning models such as EfficientNet and ResNet-50 with ArcFace loss to process and verify user commands. Optimized for low-resource devices, the system demonstrates high accuracy and low latency across various acoustic environments and regional dialects. This work contributes toward inclusive AI design by addressing linguistic diversity and privacy concerns in voice-enabled technologies.
Paper
Full text
Multi-Lingual Wake Word Detection System with Speaker Authentication for Indian Languages
Semantic Scholar · 2025
Abstract
The widespread adoption of voice-activated systems has revolutionized human-computer interaction, yet existing solutions predominantly support English, limiting accessibility for a linguistically diverse population like India. This research presents a novel approach to building a Multi-Lingual Wake Word Detection System with Speaker Authentication tailored to Indian languages. By integrating robust wake word recognition with speaker verification, the system ensures both accessibility and security. The architecture leverages Mel-spectrogram-based feature extraction and deep learning models such as EfficientNet and ResNet-50 with ArcFace loss to process and verify user commands. Optimized for low-resource devices, the system demonstrates high accuracy and low latency across various acoustic environments and regional dialects. This work contributes toward inclusive AI design by addressing linguistic diversity and privacy concerns in voice-enabled technologies.