2021/11/17 by Daniel Galvez, Daniel Gálvez, Galvez, Daniel +18 · 27 citations
Computer Science · Mathematics · #FOS: Computer and information sciences #Machine Learning (cs.LG) #Machine Learning (stat.ML) #Natural Language Processing Techniques #Speech Recognition and Synthesis #cs.LG #stat.ML
paper · pdf · doi:10.48550/arxiv.2111.09344
Part of 2021 Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks
arxiv created 2021/11/17 · openalex publication_date 2021/11/17 · arxiv updated 2021/11/19 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The People's Speech is a free-to-download 30,000-hour and growing supervised conversational English speech recognition dataset licensed for academic and commercial usage under CC-BY-SA (with a CC-BY subset). The data is collected via searching the Internet for appropriately licensed audio data with existing transcriptions. We describe our data collection methodology and release our data collection system under the Apache 2.0 license. We show that a model trained on this dataset achieves a 9.98% word error rate on Librispeech's test-clean test set.Finally, we discuss the legal and ethical issues surrounding the creation of a sizable machine learning corpora and plans for continued maintenance of the project under MLCommons's sponsorship.