Summary
The paper introduces ivrit.ai, which is a comprehensive Hebrew speech dataset designed to address the lack of extensive and high-quality resources for Automated Speech Recognition (ASR) technology in Hebrew. The dataset consists of over 10,000 speech hours from a diverse set of speakers, covering various contexts. It is provided in three different forms: raw unprocessed audio, data post-Voice Activity Detection, and partially transcribed data. What sets ivrit.ai apart is its legal accessibility, allowing free usage, making it a valuable resource for researchers, developers, and commercial entities. The dataset has the potential to enhance AI capabilities in Hebrew.
Strengths
Language-specific datasets play a crucial role in advancing ASR technology for specific languages. Hebrew, as a complex and unique language, presents its own set of challenges and nuances in speech recognition. By addressing the distinct lack of extensive and high-quality resources for Hebrew ASR, the paper fills an important gap in the research landscape.
Weaknesses
The primary concern with the paper is its narrow focus on Hebrew ASR dataset. Given the diverse range of topics covered at ICLR, which includes machine learning, representation learning, and various other areas, a paper solely dedicated to a specific language may struggle to attract a broad audience. The conference typically prioritizes research with broader applicability and impact across multiple domains. Although the paper addresses the need for extensive and high-quality resources for Hebrew ASR, the impact beyond the Hebrew language itself seems limited. While Hebrew is a unique language with its own complexities, the specific challenges faced in Hebrew ASR may not resonate with researchers working on other languages or broader speech recognition topics. This lack of broader impact may further diminish the paper's appeal to the ICLR audience. I suggest the authors submit the paper to specific workshop or speech conference.
Another concern is the absence of baseline numbers with well-established open-source frameworks such as ESPnet, k2, and fairseq. Others could easily build their system and have a relatively fair comparision if these numbers are provided.
Questions
Why not use Whisper large to generate transtriptions?
Rating
3: reject, not good enough
Confidence
4: You are confident in your assessment, but not absolutely certain. It is unlikely, but not impossible, that you did not understand some parts of the submission or that you are unfamiliar with some pieces of related work.