2025/02/21 by Zhang, Wenyu, Luo, Jie, Zhang, Xinming +1
#Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Information Retrieval (cs.IR)
paper · doi:10.48550/arxiv.2502.15542
With the explosive growth of multimodal content online, pre-trained visual-language models have shown great potential for multimodal recommendation. However, while these models achieve decent performance when applied in a frozen manner, surprisingly, due to significant domain gaps (e.g., feature distribution discrepancy and task objective misalignment) between pre-training and personalized recommendation, adopting a joint training approach instead leads to performance worse than baseline. Existing approaches either rely on simple feature extraction or require computationally expensive full model fine-tuning, struggling to balance effectiveness and efficiency. To tackle these challenges, we propose Parameter-efficient Tuning for Multimodal Recommendation (PTMRec), a novel framework that bridges the domain gap between pre-trained models and recommendation systems through a knowledge-guided dual-stage parameter-efficient training strategy. This framework not only eliminates the need for costly additional pre-training but also flexibly accommodates various parameter-efficient tuning methods.