vix.ing · top · new · best · stats · spec

Video Diffusion Transformers are In-Context Learners

2024/12/14 by Zhengcong Fei, Di Qiu, Fei, Zhengcong +7 · 4 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Neural Networks and Applications

paper · pdf · doi:10.48550/arxiv.2412.10783

openalex publication_date 2024/12/14 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper investigates a solution for enabling in-context capabilities of video diffusion transformers, with minimal tuning required for activation. Specifically, we propose a simple pipeline to leverage in-context generation: (i) concatenate videos along spacial or time dimension, (ii) jointly caption multi-scene video clips from one source, and (iii) apply task-specific fine-tuning using carefully curated small datasets. Through a series of diverse controllable tasks, we demonstrate qualitatively that existing advanced text-to-video models can effectively perform in-context generation. Notably, it allows for the creation of consistent multi-scene videos exceeding 30 seconds in duration, without additional computational overhead. Importantly, this method requires no modifications to the original models, results in high-fidelity video outputs that better align with prompt specifications and maintain role consistency. Our framework presents a valuable tool for the research community and offers critical insights for advancing product-level controllable video generation systems. The data, code, and model weights are publicly available at: https://github.com/feizc/Video-In-Context.

Cited by

Related