Speaker identification is a critical enabler for operational intelligence and personalized customer engagement, yet it remains challenging due to background noise interference, data scarcity under limited training conditions, and the need to generalize across large and diverse speaker populations. Traditional methods relying on handcrafted acoustic features or spectrogram-based inputs often fail to capture fine temporal dynamics and generalize well to unseen conditions. To address these limitations, we propose Additive Margin with Dynamic Augmentation Network (AM-DANet), a novel one-dimensional residual network designed for robust speaker identification. The proposed framework improves data utilization through an improved preprocessing pipeline and incorporates an additive margin-based learning strategy to promote strong inter-class separability while preserving compact intra-class representations, enabling the model to learn more discriminative speaker embeddings. In addition, a dynamic batch generation scheme performs random segmentation and amplitude scaling to increase temporal and amplitude diversity, effectively mitigating overfitting in data-scarce settings. Comprehensive evaluations on the English100, TIMIT, and LibriSpeech datasets demonstrate that AM-DANet achieves state-of-the-art accuracies of 99.68%, 99.93%, and 99.14%, respectively. Further analyses, including noise robustness tests under additive Gaussian and babble noise conditions, statistical significance evaluation, computational efficiency assessment, and t-SNE visualization of the embedding space, consistently validate the model's stability, discriminative capability, and suitability for real-time deployment. These results establish AM-DANet as an efficient and reliable framework for advancing intelligent, identity-aware customer routing and engagement in real-world applications.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex