Speech-to-speech translation (S2ST) technology bridges language barriers, enabling seamless communication across cultures. This paper presents a modular system integrating Automatic Speech Recognition (ASR), Machine Translation (MT), and Text-to-Speech (TTS) for real-time translation. Leveraging Python, Google APIs, and self-supervised models, the system achieves high accuracy (ASR: 90– 95%, MT: 85–90%) and low latency (2–3 seconds). Key contributions include noise filtering, scalable architecture, and support for low-resource languages. Applications span healthcare, education, and global collaboration, emphasizing the practical significance of this innovation.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex