2023/10/20 by Donghuo Zeng, Zeng, Donghuo, Kazushi Ikeda +1
Computer Science · Mathematics · #Advanced Image and Video Retrieval Techniques #Algorithm #Artificial intelligence #Audio and Speech Processing (eess.AS) #Audio visual #Computer Vision and Pattern Recognition (cs.CV) #Computer science #Embedding #FOS: Computer and information sciences #FOS: Electrical engineering #Image (mathematics) #Information Retrieval (cs.IR) #Interpolation (computer graphics) #Machine learning #Mathematics #Modal #Multimedia #Multimedia (cs.MM) #Music and Audio Processing #Process (computing) #Programming language #Scheme (mathematics) #Set (abstract data type) #Sound (cs.SD) #Speech and Audio Processing #Task (project management) #electronic engineering #information engineering
paper · pdf · doi:10.48550/arxiv.2310.13451
openalex publication_date 2023/10/20 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28
The cross-modal retrieval model leverages the potential of triple loss optimization to learn robust embedding spaces. However, existing methods often train these models in a singular pass, overlooking the distinction between semi-hard and hard triples in the optimization process. The oversight of not distinguishing between semi-hard and hard triples leads to suboptimal model performance. In this paper, we introduce a novel approach rooted in curriculum learning to address this problem. We propose a two-stage training paradigm that guides the model's learning process from semi-hard to hard triplets. In the first stage, the model is trained with a set of semi-hard triplets, starting from a low-loss base. Subsequently, in the second stage, we augment the embeddings using an interpolation technique. This process identifies potential hard negatives, alleviating issues arising from high-loss functions due to a scarcity of hard triples. Our approach then applies hard triplet mining in the augmented embedding space to further optimize the model. Extensive experimental results conducted on two audio-visual datasets show a significant improvement of approximately 9.8% in terms of average Mean Average Precision (MAP) over the current state-of-the-art method, MSNSCA, for the Audio-Visual Cross-Modal Retrieval (AV-CMR) task on the AVE dataset, indicating the effectiveness of our proposed method.