vix.ing · top · new · best · stats

End-to-end Lyrics Alignment for Polyphonic Music Using an\n Audio-to-Character Recognition Model

2019/02/18 by Daniel Stoller, Stoller, Daniel, Simon Durand +3 · 2 citations
Computer Science · Engineering · #Acoustics #Artificial intelligence #Audio and Speech Processing (eess.AS) #Character (mathematics) #Computer science #FOS: Computer and information sciences #FOS: Electrical engineering #Lyrics #Machine Learning (cs.LG) #Music and Audio Processing #Natural language processing #Polyphony #Sound (cs.SD) #Speech Recognition and Synthesis #Speech and Audio Processing #Speech recognition #cs.LG #cs.SD #eess.AS #electronic engineering #information engineering

paper · pdf · doi:10.48550/arxiv.1902.06797

published in arXiv (Cornell University) (Cornell University) · 5 pages (1 for references), 2 figures, 2 tables. Camera-ready version, accepted at the International Conference on Acoustics, Speech, and Signal Processing 2019 (ICASSP)

arxiv created 2019/02/18 · openalex publication_date 2019/02/18 · arxiv updated 2019/02/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/08/05

Abstract

Time-aligned lyrics can enrich the music listening experience by enabling\nkaraoke, text-based song retrieval and intra-song navigation, and other\napplications. Compared to text-to-speech alignment, lyrics alignment remains\nhighly challenging, despite many attempts to combine numerous sub-modules\nincluding vocal separation and detection in an effort to break down the\nproblem. Furthermore, training required fine-grained annotations to be\navailable in some form. Here, we present a novel system based on a modified\nWave-U-Net architecture, which predicts character probabilities directly from\nraw audio using learnt multi-scale representations of the various signal\ncomponents. There are no sub-modules whose interdependencies need to be\noptimized. Our training procedure is designed to work with weak, line-level\nannotations available in the real world. With a mean alignment error of 0.35s\non a standard dataset our system outperforms the state-of-the-art by an order\nof magnitude.\n

Citations

Cited by

Related