Replacing Human Audio with Synthetic Audio for On-device Unspoken Punctuation Prediction

We present a novel multi-modal unspoken punctuation prediction system for the\nEnglish language which combines acoustic and text features. We demonstrate for\nthe first time, that by relying exclusively on synthetic data generated using a\nprosody-aware text-to-speech system, we can outperform a model trained with\nexpensive human audio recordings on the unspoken punctuation prediction\nproblem. Our model architecture is well suited for on-device use. This is\nachieved by leveraging hash-based embeddings of automatic speech recognition\ntext output in conjunction with acoustic features as input to a quasi-recurrent\nneural network, keeping the model size small and latency low.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC