Summary
This paper outlines a roadmap for achieving superhuman speech understanding by defining five levels of speech and audio understanding tasks and establishing a benchmark based on this framework. The first level involves automatic speech recognition, a task that has been successfully completed by various multimodal large language models (LLMs). As we move up to higher levels, the tasks require more contextual information and knowledge. In addition to mapping out these levels, the paper presents preliminary benchmarks for each level, comparing existing models with human performance. This information is invaluable and includes detailed analyses highlighting how far we are from reaching the established baseline.
Strengths
- High-level (superhuman) speech understanding based on LLMs is a core topic for realizing conversational AI. However, we don't have an established roadmap and benchmark toward this goal, and this attempt is very valuable.
- The benchmark has human performance, and we can evaluate our models with its superhuman degree in each task
Weaknesses
- The paper does not address speech generation tasks (e.g., Text-to-Speech (TTS) or dialogue system responses) and lacks completeness as a roadmap or benchmark for the field.
- Certain terminology choices, such as "paralinguistics," may not align with conventional use within the speech community, leading to potential misinterpretation. For instance, "paralinguistics" typically encompasses emotional aspects of speech, as illustrated by the Computational Paralinguistics Challenge organized by ISCA Interspeech, which includes emotion recognition tasks. This divergence in terminology creates ambiguity and suggests the roadmap may lack refinement in this respect.
- Due to these terminology issues, distinctions between levels—specifically L2 and L3 tasks—are unclear. Some tasks categorized at L2, such as pitch perception, seem more complex than certain L3 tasks, such as gender recognition.
- I am skeptical of the overall categorization scheme, as some tasks listed (e.g., acoustic scene classification and L4 tasks) are not strictly "speech" tasks but are more accurately described as audio processing tasks. Additionally, L4 tasks are likely to lack significant speech semantics, making them atypical for speech understanding. However, if L4 included tasks involving doctor-patient conversations with consultative dialogues, it could be more suitable for speech understanding.
- The paper overlooks relevant existing speech and audio benchmarks, such as the SUPERB and HEAR benchmarks, which could strengthen the study’s scope and positioning.
- The benchmark section itself feels incomplete. As noted in the footnote, tasks for Levels 4 and 5 have not been fully identified or outlined.
- Finally, the analysis sections appear somewhat disconnected from the roadmap (especially Sections 4.3 and 4.4), with limited cohesion between the analysis content and the proposed roadmap.
Questions
- Page 2, Itemization of Levels: The three levels listed here conflict with the five levels in the roadmap, which may create confusion. Given the noted ambiguity between Levels 2 and 3, I recommend refining the levels and corresponding tasks for clearer distinction.
- Page 4, Level 2 Terminology: Consider using terminology more specific to physical or acoustic features rather than "paralinguistics," as this section pertains more to physical or acoustic information than traditional paralinguistic elements.
- Page 4, Level 3 - Citation Needed: The statement, “Interestingly, even some higher animals, like pet dogs, can perceive these types of non-semantic information,” would benefit from supporting citations to strengthen the claim.
- Page 5, Lines 218-219 - Citation for Gladwell Reference: Please add a citation for “The 10,000-Hour Rule” from Malcolm Gladwell’s Outliers.
- Table 2 - License Information: Including license details in Table 2 would be helpful, as this is essential for widely-used benchmarks.
- Table 3 - Indicate Superiority with Symbols: Adding arrows (↑, ↓) in Table 3 to indicate the direction of performance metrics (accuracy vs. error) would make it easier to distinguish superior and inferior results.
- Page 6, Section 3.2 - Clarify “Manually” in Testing GPT-4o: The term “manually” in this context is vague. Please clarify the process of how GPT-4o was used in the advanced speech mode.
- Experimental Findings - Limited Novelty: The experimental findings seem somewhat elementary. For instance, GAMA’s superior performance in the audio task and speech LLMs’ strengths in term recognition (Section 4.1) are unsurprising.
- Tables 3 and 4 - Consistency of Units: There are inconsistencies in the representation of units (e.g., omission of % in some tasks). If this is intentional, please clarify the rationale.