vix.ing · top · new · best · stats · spec

Ruihang Chu

  1. Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
    2024/03/27 by Yanwei Li, Li, Yanwei, Yuechen Zhang +13 · 2 voices · 51 citations
    Computer Science · #Multimodal Machine Learning Applications
  2. Wan: Open and Advanced Large-Scale Video Generative Models
    2025/03/26 by Team Wan, WanTeam, Wan, Team +125 · 2 voices · 592 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #cs.CV
  3. DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation
    2023/07/04 by Shentong Mo, Mo, Shentong, Enze Xie +11 · 25 citations
    Engineering · Computer Science · #3D Shape Modeling and Analysis #Generative Adversarial Networks and Image Synthesis #Advanced Vision and Imaging
  4. DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
    2024/03/13 by Minbin Huang, Yanxin Long, Huang, Minbin +15 · 1 voice · 7 citations
    Computer Science · #Speech and dialogue systems #Multimodal Machine Learning Applications #Topic Modeling
  5. DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
    2024/03/25 by Tingkai Wang, Enze Xie, Wang, Tianqi +7 · 14 citations
    Computer Science · #Advanced Text Analysis Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO) #Semantic Web and Ontologies
  6. A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
    2025/04/11 by Jiankai Sun, Chuanyang Zheng, Enze Xie +34 · 22 citations
    Computer Science · #Logic, Reasoning, and Knowledge #Multi-Agent Systems and Negotiation #Semantic Web and Ontologies
  7. Mask-Attention-Free Transformer for 3D Instance Segmentation
    2023/09/04 by Xin Lai, Lai, Xin, Yuhui Yuan +9 · 6 citations
    Computer Science · Engineering · #Advanced Neural Network Applications #Robotics and Sensor-Based Localization #Medical Image Segmentation Techniques
  8. DiffComplete: Diffusion-based Generative 3D Shape Completion
    2023/06/28 by Ruihang Chu, Chu, Ruihang, Enze Xie +11 · 4 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
  9. InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
    2025/03/04 by Dingdong Wang, Jin Xu, Wang, Dingdong +14 · 7 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
  10. SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
    2025/09/29 by Junsong Chen, Yuyang Zhao, Chen, Junsong +35 · 18 citations
    Computer Science · #Advanced Data Compression Techniques #Video Coding and Compression Technologies #Digital Filter Design and Implementation
  11. Teaching Your Models to Understand Code via Focal Preference Alignment
    2025/03/04 by Jie Wu, Haoling Li, Wu, Jie +15 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Speech and dialogue systems
  12. TriVol: Point Cloud Rendering via Triple Volumes
    2023/03/29 by Tao Hu, Hu, Tao, Xiaogang Xu +5 · 1 citation
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  13. Generative Universal Verifier as Multimodal Meta-Reasoner
    2025/10/15 by Xinchen Zhang, Xiaoying Zhang, Zhang, Xinchen +13 · 4 citations
    Computer Science · #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
  14. LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
    2024/07/21 by Li Yu, Li, Yu, Yifan Chen +14 · 1 citation
    Computer Science · #Video Analysis and Summarization #Semantic Web and Ontologies #Image Retrieval and Classification Techniques
  15. DreamVE: Unified Instruction-based Image and Video Editing
    2025/08/08 by Bin Xia, Xia, Bin, Jiyang Liu +15 · 3 citations
    Computer Science · #Video Analysis and Summarization #Advanced Vision and Imaging #Generative Adversarial Networks and Image Synthesis
  16. AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
    2025/07/17 by Yiming Ren, Ren, Yiming, Zhiqiang Lin +18 · 3 citations
    Arts and Humanities · Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media #Video Analysis and Summarization
  17. VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
    2025/12/26 by Yang Ding, Ding, Yang, Yizhen Zhang +7 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
  18. O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
    2025/09/01 by Yuqing Chen, Junjie Wang, Chen, Yuqing +11 · 3 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #Visual Attention and Saliency Detection