vix.ing · top · new · best · stats · spec

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

2026/08/02 by Tongyan Wang, Zhengyuan Li, Muhan Lin +5
Computer Science · #cs.CV

paper · pdf

arxiv created 2026/08/02 · arxiv updated 2026/08/04

Abstract

Text-conditioned human motion generation has made rapid progress with the emergence of large-scale motion--language datasets. However, even datasets with rich long-form descriptions typically provide supervision only at the clip level, without explicit temporal correspondence between motion frames and language. This limits fine-grained motion--text grounding and temporally precise generation. We propose FineMoLA, a weakly supervised framework that learns fine-grained frame--phrase correspondence directly from clip-level annotations. Our method first segments long-form descriptions into action-bearing phrases, and then formulates motion--language alignment as an optimal transport problem, which naturally models many-to-many relations between motion frames and text under global constraints. With entropic regularization and Sinkhorn iterations, FineMoLA efficiently infers pseudo frame-level alignments without human labeling. Experiments on SnapMoGen demonstrate that the learned alignments outperform baselines in motion--text grounding.

Citations