Towards Measuring Fairness in Speech Recognition: Casual Conversations\n Dataset Transcriptions

It is well known that many machine learning systems demonstrate bias towards\nspecific groups of individuals. This problem has been studied extensively in\nthe Facial Recognition area, but much less so in Automatic Speech Recognition\n(ASR). This paper presents initial Speech Recognition results on "Casual\nConversations" -- a publicly released 846 hour corpus designed to help\nresearchers evaluate their computer vision and audio models for accuracy across\na diverse set of metadata, including age, gender, and skin tone. The entire\ncorpus has been manually transcribed, allowing for detailed ASR evaluations\nacross these metadata. Multiple ASR models are evaluated, including models\ntrained on LibriSpeech, 14,000 hour transcribed, and over 2 million hour\nuntranscribed social media videos. Significant differences in word error rate\nacross gender and skin tone are observed at times for all models. We are\nreleasing human transcripts from the Casual Conversations dataset to encourage\nthe community to develop a variety of techniques to reduce these statistical\nbiases.\n

Paper

References (38)

Scroll for more · 26 remaining

Similar papers

© 2026 NYSGPT2525 LLC