Summary
his paper identifies two significant limitations of MUSHRA evaluation for TTS systems:
1. **Reference-Match Bias**: The use of explicit reference samples can lead raters to assign lower scores to synthesized samples that deviate from the reference, even if the quality is high.
2. **Judgment Ambiguity**: MUSHRA tests lack fine-grained guidance, making it difficult for raters to provide consistent evaluations.
To address these shortcomings, the authors propose two solutions:
1. Not explicitly identifying the human reference to the rater.
2. Clearly listing 9 key aspects of TTS evaluation (such as pronunciation errors and unnatural pauses) for raters to assess, followed by calculating the final MUSHRA score using a predefined formula.
Experimental results from three TTS systems in two Indian languages (Indic and Tamil) demonstrate that these proposed solutions effectively alleviate the issues associated with MUSHRA. The results show more consistent results with one-to-one comparisons between TTS systems and reference samples assessed by CMOS, and a reduction in score vari among different raters. Additionally, the complete dataset, comprising 47,100 ratings from 471 listeners evaluating three TTS systems, is publicly available for future research in this area.
Strengths
- This paper highlights a significant issue with the foundational assumption of MUSHRA scores: it presupposes that a real reference sample should consistently receive high scores. However, this may not always be the case, as modern TTS systems can sometimes outperform human references.
- The authors conduct a thorough analysis of MUSHRA test results, providing both qualitative and quantitative evidence to demonstrate the two major shortcomings of the original MUSHRA framework.
- In addition to addressing the two primary shortcomings, the paper systematically examines various aspects of MUSHRA evaluation, including:
- A comparison of MUSHRA and CMOS results
- The sensitivity of MUSHRA scores to the number of raters and utterances
- Protocols for rejecting raters
- Different strategies for establishing anchors
- The proposed MUSHRA variants are straightforward and effectively address the identified shortcomings, enhancing the reliability of the evaluation process.
Weaknesses
- While the issues with MUSHRA are language-agnostic, this paper focuses exclusively on TTS systems for Indian languages, which are trained on relatively small datasets. The findings would be more compelling if the authors included results from widely-used datasets of high-resource languages and publicly available pre-trained English (or any other high-resource language) TTS systems with verified quality.
- The nine aspects selected for MUSHRA-DG appear to be arbitrary, as the authors do not provide justification or references for their choices, leaving unanswered why some other common aspects (such as word repetition) were not included.
- As already noted in Section 6.1 (line 410), even though the reference-matching bias issue is mitigated, discrepancies remain between the results of CMOS, MUSHRA-NMR, and MUSHRA-DG-NMR, particularly for the best-performing ST2 system in Tamil. More detailed explanations regarding these differences will be appreciated.
Questions
In line 256, the authors highlight an issue with MUSHRA tests, stating that “the individual box-plots have a high variance indicating that the same rater rates the system very differently across utterances.” However, this variance could also result from failures in the generation process, particularly since the variance in the reference (REF) is smaller than that of any TTS system. Could the authors clarify why they interpret this observation as an issue with rater consistency rather than a reflection of generation quality?