The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

The People's Speech is a free-to-download 30,000-hour and growing supervised\nconversational English speech recognition dataset licensed for academic and\ncommercial usage under CC-BY-SA (with a CC-BY subset). The data is collected\nvia searching the Internet for appropriately licensed audio data with existing\ntranscriptions. We describe our data collection methodology and release our\ndata collection system under the Apache 2.0 license. We show that a model\ntrained on this dataset achieves a 9.98% word error rate on Librispeech's\ntest-clean test set.Finally, we discuss the legal and ethical issues\nsurrounding the creation of a sizable machine learning corpora and plans for\ncontinued maintenance of the project under MLCommons's sponsorship.\n

Paper

Similar papers

© 2026 NYSGPT2525 LLC