2025/11/21 by Cheng, Shihan, Kulkarni, Nilesh, Hyde, David +1
Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #Human Motion and Animation #Music Technology and Sound Studies
paper · doi:10.48550/arxiv.2511.17844
Fine-tuning large-scale text-to-video diffusion models to add new generative controls, such as those over physical camera parameters (e.g., shutter speed or aperture), typically requires vast, high-fidelity datasets that are difficult to acquire. In this work, we propose a data-efficient fine-tuning strategy that learns these controls from sparse, low-quality synthetic data. We show that not only does fine-tuning on such simple data enable the desired controls, it actually yields superior results to models fine-tuned on photorealistic "real" data. Beyond demonstrating these results, we provide a framework that justifies this phenomenon both intuitively and quantitatively.