In this evolving, interconnected world, powerful communication over different languages is crucial for international collaborations, businesses and accessibility. Language barriers often become obstacles, hampering individual interactions as well as employee engagement in organizations. To handle this, we are presenting an AI-Powered instant voice message translator for multilingual communication, which enables users to converse in their preferred language through voice messaging, without delays. Our system is equipped with Automatic Speech Recognition (ASR) to interpret spoken words into text and then, with the help of Neural Machine Translation (NMT), the text is changed into a target language and again this translated text is converted to speech using Text-To-Speech (TTS) to produce natural speech output. Being built with a distributed and scalable architecture, this system ensures low latency processing and effective resource usage that supports real-time voice conversation. Usage of Web sockets enables two-way audio streaming, which helps in increasing responsiveness and also end-to-end encryption safeguards user's privacy. Unlike the other available translation tools, our solution allows one-to-one communication, making it best suitable for international collaborations, meetings, customer service, tourism and many more. This model achieved a Word Error Rate of 8.5%, BLEU score of 42.5 with reduced latency of less than 3 seconds and accuracy of 94%, which was much better than other available tools.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex