vix.ing · top · new · best · stats · spec

Goku: Flow Based Video Generative Foundation Models

2025/02/07 by Shoufa Chen, Chongjian Ge, Chen, Shoufa +40 · 22 citations
Computer Science · #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization

paper · pdf · doi:10.48550/arxiv.2502.04896

openalex publication_date 2025/02/07 · openalex created_date 2025/10/10 · openalex updated_date 2026/07/28

Abstract

This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models.

Cited by

Related