vix.ing · top · new · best · stats · spec

Fei, Zhengcong

  1. SkyReels-V2: Infinite-length Film Generative Model
    2025/04/17 by Guibin Chen, Dixuan Lin, Chen, Guibin +46 · 57 citations
    Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #Human Motion and Animation
  2. A-JEPA: Joint-Embedding Predictive Architecture Can Listen
    2023/11/27 by Fei, Zhengcong, Fan, Mingyuan, Huang, Junshi · 9 citations
    #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  3. Scaling Diffusion Transformers to 16 Billion Parameters
    2024/07/16 by Zhengcong Fei, Mingyuan Fan, Fei, Zhengcong +7 · 11 citations
    Physics and Astronomy · Earth and Planetary Sciences · #Nuclear physics research studies #Cold Fusion and Nuclear Reactions #Cold Atom Physics and Bose-Einstein Condensates
  4. SkyReels-A2: Compose Anything in Video Diffusion Transformers
    2025/04/03 by Zhengcong Fei, Fei, Zhengcong, Debang Li +18 · 19 citations
    Computer Science · Engineering · #Advanced Optical Imaging Technologies #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  5. SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers
    2025/02/15 by Di Qiu, Zhengcong Fei, Qiu, Di +13 · 11 citations
    Earth and Planetary Sciences · #3D Surveying and Cultural Heritage
  6. Diffusion-RWKV: Scaling RWKV-Like Architectures for Diffusion Models
    2024/04/06 by Fei, Zhengcong, Fan, Mingyuan, Yu, Changqian +2 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Towards Expressive Communication with Internet Memes: A New Multimodal Conversation Dataset and Benchmark
    2021/09/04 by Zhengcong Fei, Fei, Zhengcong, Zekang Li +7 · 2 citations
    Computer Science · Psychology · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Digital Communication and Language #FOS: Computer and information sciences #Humor Studies and Applications #Sentiment Analysis and Opinion Mining
  8. FLUX that Plays Music
    2024/09/01 by Zhengcong Fei, Mingyuan Fan, Fei, Zhengcong +5 · 4 citations
    Computer Science · #Artificial Intelligence in Games #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  9. Video Diffusion Transformers are In-Context Learners
    2024/12/14 by Zhengcong Fei, Di Qiu, Fei, Zhengcong +7 · 4 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Neural Networks and Applications
  10. Dimba: Transformer-Mamba Diffusion Models
    2024/06/03 by Zhengcong Fei, Mingyuan Fan, Fei, Zhengcong +9 · 3 citations
    Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Vibration and Dynamic Analysis
  11. Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning
    2023/09/10 by Guisheng Liu, Yi Li, Liu, Guisheng +9 · 2 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  12. SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
    2025/06/01 by Zhengcong Fei, Hao Jiang, Fei, Zhengcong +19 · 6 citations
    Arts and Humanities · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia Communication and Technology #Radio, Podcasts, and Digital Media #Subtitles and Audiovisual Media
  13. Uncertainty-Aware Image Captioning
    2022/11/30 by Zhengcong Fei, Mingyuan Fan, Fei, Zhengcong +9 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Video Analysis and Summarization
  14. Music Consistency Models
    2024/04/20 by Fei, Zhengcong, Fan, Mingyuan, Huang, Junshi · 1 citation
    #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  15. Scalable Diffusion Models with State Space Backbone
    2024/02/08 by Zhengcong Fei, Fei, Zhengcong, Mingyuan Fan +5 · 1 citation
    Engineering · #Advanced Control Systems Optimization #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM)
  16. Ingredients: Blending Custom Photos with Video Diffusion Transformers
    2025/01/03 by Zhengcong Fei, Fei, Zhengcong, Debang Li +7 · 1 citation
    Earth and Planetary Sciences · #3D Surveying and Cultural Heritage