vix.ing · top · new · best · stats · spec

A Lightweight Architecture for Multi-instrument Transcription with Practical Optimizations

2025/09/16 by Li, Ruigang, Zhu, Yongxu
#FOS: Computer and information sciences #Information Retrieval (cs.IR) #Sound (cs.SD)

paper · doi:10.48550/arxiv.2509.12712

Abstract

Existing multi-timbre transcription models struggle with generalization beyond pre-trained instruments, rigid source-count constraints, and high computational demands that hinder deployment on low-resource devices. We address these limitations with a lightweight model that extends a timbre-agnostic transcription backbone with a dedicated timbre encoder and performs deep clustering at the note level, enabling joint transcription and dynamic separation of arbitrary instruments. Practical optimizations including spectral normalization, dilated convolutions, and contrastive clustering further improve efficiency and robustness. Despite its small size and fast inference, the model achieves competitive performance with heavier baselines in transcription accuracy and separation quality, and shows promising generalization ability, making it highly suitable for real-world deployment in practical and resource-constrained settings.

Citations

Related