Ruihang Chu
- Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
2024/03/27 by Yanwei Li, Li, Yanwei, Yuechen Zhang +13 · 2 voices · 51 citations
Computer Science · #Multimodal Machine Learning Applications
- Wan: Open and Advanced Large-Scale Video Generative Models
2025/03/26 by Team Wan, WanTeam, Wan, Team +125 · 2 voices · 592 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #cs.CV
- DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation
2023/07/04 by Shentong Mo, Mo, Shentong, Enze Xie +11 · 25 citations
Engineering · Computer Science · #3D Shape Modeling and Analysis #Generative Adversarial Networks and Image Synthesis #Advanced Vision and Imaging
- DialogGen: Multi-modal Interactive Dialogue System for Multi-turn Text-to-Image Generation
2024/03/13 by Minbin Huang, Yanxin Long, Huang, Minbin +15 · 1 voice · 7 citations
Computer Science · #Speech and dialogue systems #Multimodal Machine Learning Applications #Topic Modeling
- DriveCoT: Integrating Chain-of-Thought Reasoning with End-to-End Driving
2024/03/25 by Tingkai Wang, Enze Xie, Wang, Tianqi +7 · 14 citations
Computer Science · #Advanced Text Analysis Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotics (cs.RO) #Semantic Web and Ontologies
- A Survey of Reasoning with Foundation Models: Concepts, Methodologies, and Outlook
2025/04/11 by Jiankai Sun, Chuanyang Zheng, Enze Xie +34 · 22 citations
Computer Science · #Logic, Reasoning, and Knowledge #Multi-Agent Systems and Negotiation #Semantic Web and Ontologies
- Mask-Attention-Free Transformer for 3D Instance Segmentation
2023/09/04 by Xin Lai, Lai, Xin, Yuhui Yuan +9 · 6 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Robotics and Sensor-Based Localization #Medical Image Segmentation Techniques
- DiffComplete: Diffusion-based Generative 3D Shape Completion
2023/06/28 by Ruihang Chu, Chu, Ruihang, Enze Xie +11 · 4 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
- InSerter: Speech Instruction Following with Unsupervised Interleaved Pre-training
2025/03/04 by Dingdong Wang, Jin Xu, Wang, Dingdong +14 · 7 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #FOS: Computer and information sciences #FOS: Electrical engineering #Human-Computer Interaction (cs.HC) #Natural Language Processing Techniques #Sound (cs.SD) #Speech Recognition and Synthesis #Topic Modeling #electronic engineering #information engineering
- SANA-Video: Efficient Video Generation with Block Linear Diffusion Transformer
2025/09/29 by Junsong Chen, Yuyang Zhao, Chen, Junsong +35 · 18 citations
Computer Science · #Advanced Data Compression Techniques #Video Coding and Compression Technologies #Digital Filter Design and Implementation
- Teaching Your Models to Understand Code via Focal Preference Alignment
2025/03/04 by Jie Wu, Haoling Li, Wu, Jie +15 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Software Engineering Research #Speech and dialogue systems
- TriVol: Point Cloud Rendering via Triple Volumes
2023/03/29 by Tao Hu, Hu, Tao, Xiaogang Xu +5 · 1 citation
Computer Science · Engineering · #3D Shape Modeling and Analysis #Advanced Vision and Imaging #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Generative Universal Verifier as Multimodal Meta-Reasoner
2025/10/15 by Xinchen Zhang, Xiaoying Zhang, Zhang, Xinchen +13 · 4 citations
Computer Science · #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
- LayoutDiT: Exploring Content-Graphic Balance in Layout Generation with Diffusion Transformer
2024/07/21 by Li Yu, Li, Yu, Yifan Chen +14 · 1 citation
Computer Science · #Video Analysis and Summarization #Semantic Web and Ontologies #Image Retrieval and Classification Techniques
- DreamVE: Unified Instruction-based Image and Video Editing
2025/08/08 by Bin Xia, Xia, Bin, Jiyang Liu +15 · 3 citations
Computer Science · #Video Analysis and Summarization #Advanced Vision and Imaging #Generative Adversarial Networks and Image Synthesis
- AnyCap Project: A Unified Framework, Dataset, and Benchmark for Controllable Omni-modal Captioning
2025/07/17 by Yiming Ren, Ren, Yiming, Zhiqiang Lin +18 · 3 citations
Arts and Humanities · Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media #Video Analysis and Summarization
- VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
2025/12/26 by Yang Ding, Ding, Yang, Yizhen Zhang +7 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications
- O-DisCo-Edit: Object Distortion Control for Unified Realistic Video Editing
2025/09/01 by Yuqing Chen, Junjie Wang, Chen, Yuqing +11 · 3 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #Visual Attention and Saliency Detection