WER-BERT: Automatic WER Estimation with BERT in a Balanced Ordinal Classification Paradigm

Automatic Speech Recognition (ASR) systems are evaluated using Word Error\nRate (WER), which is calculated by comparing the number of errors between the\nground truth and the transcription of the ASR system. This calculation,\nhowever, requires manual transcription of the speech signal to obtain the\nground truth. Since transcribing audio signals is a costly process, Automatic\nWER Evaluation (e-WER) methods have been developed to automatically predict the\nWER of a speech system by only relying on the transcription and the speech\nsignal features. While WER is a continuous variable, previous works have shown\nthat positing e-WER as a classification problem is more effective than\nregression. However, while converting to a classification setting, these\napproaches suffer from heavy class imbalance. In this paper, we propose a new\nbalanced paradigm for e-WER in a classification setting. Within this paradigm,\nwe also propose WER-BERT, a BERT based architecture with speech features for\ne-WER. Furthermore, we introduce a distance loss function to tackle the ordinal\nnature of e-WER classification. The proposed approach and paradigm are\nevaluated on the Librispeech dataset and a commercial (black box) ASR system,\nGoogle Cloud's Speech-to-Text API. The results and experiments demonstrate that\nWER-BERT establishes a new state-of-the-art in automatic WER estimation.\n

Paper

References (25)

Scroll for more · 13 remaining

Similar papers

© 2026 NYSGPT2525 LLC