2024/09/18 by Kalakonda, Sai Shashank, Shubh Maheshwari, Maheshwari, Shubh +2 · 3 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Hand Gesture Recognition Systems #Human Pose and Action Recognition #Multimedia (cs.MM)
paper · pdf · doi:10.48550/arxiv.2409.12140
openalex publication_date 2024/09/18 · openalex created_date 2024/10/25 · openalex updated_date 2026/07/28
We introduce MoRAG, a novel multi-part fusion based retrieval-augmented generation strategy for text-based human motion generation. The method enhances motion diffusion models by leveraging additional knowledge obtained through an improved motion retrieval process. By effectively prompting large language models (LLMs), we address spelling errors and rephrasing issues in motion retrieval. Our approach utilizes a multi-part retrieval strategy to improve the generalizability of motion retrieval across the language space. We create diverse samples through the spatial composition of the retrieved motions. Furthermore, by utilizing low-level, part-specific motion information, we can construct motion samples for unseen text descriptions. Our experiments demonstrate that our framework can serve as a plug-and-play module, improving the performance of motion diffusion models. Code, pretrained models and sample videos are available at: https://motion-rag.github.io/