Roxi-Duplex: Low-Resource Indian-English Adaptation of a Full-Duplex Speech-to-Speech Model with Synthetic Two-Channel Training Data
Full-duplex speech-to-speech systems listen while speaking, allowing interruption, overlap, and backchannel behaviour to be represented within the generative architecture rather than added only through external turn-taking logic. Open full-duplex systems have limited support for Indian English, while suitable two-channel Indian-English dialogue data are scarce. We present Roxi-Duplex, a parameter-efficient adaptation of Moshi for Indian-English customer-support interaction. The training pipeline combines coherent scripted assistant speech generated by an Indian-English text-to-speech model with real but semantically unpaired Indian-English recordings on the user channel, then assembles both streams on a shared timeline containing gaps, overlap, and backchannels. A rank-64 LoRA adapter trained for 1,500 steps on 94 minutes of stereo audio requires 18.7 GB peak memory and approximately 17 minutes on one A100 40 GB GPU. Relative to base Moshi, adaptation increases accent-centroid similarity from 0.639 to 0.754, support-persona responses from 3% to 75%, and generations with audible speech from 18/40 to 39/40. It preserves backchannel continuation (88%) but yields more slowly to full barge-ins than a VAD-based cascade. The results show that inexpensive synthetic adaptation can transfer accent and persona while retaining duplex interaction, although semantic user–assistant alignment, human perceptual evaluation, and broader task coverage remain open limitations. This manuscript is a preprint and has not undergone formal peer review.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex