Bai, Shuai
- Qwen2.5-VL Technical Report
2025/02/19 by Shuai Bai, Keqin Chen, Bai, Shuai +48 · 1702 citations
Engineering · #Semiconductor Lasers and Optical Devices
- Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
2024/09/18 by Peng Wang, Wang, Peng, Shuai Bai +35 · 1056 citations
Psychology · #Categorization, perception, and language
- Qwen Technical Report
2023/09/28 by Jinze Bai, Bai, Jinze, Shuai Bai +91 · 666 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
2023/08/24 by Jinze Bai, Shuai Bai, Bai, Jinze +15 · 559 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Qwen2 Technical Report
2024/07/15 by Yang An, Baosong Yang, Yang, An +120 · 377 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
2024/03/11 by Liang Chen, Chen, Liang, Haozhe Zhao +11 · 137 citations
Computer Science · Medicine · #Multimodal Machine Learning Applications #COVID-19 diagnosis using AI
- Qwen2.5-Omni Technical Report
2025/03/26 by Jin Xu, Zihan Guo, Xu, Jin +22 · 249 citations
Engineering · #Embedded Systems and FPGA Design
- Qwen-Image Technical Report
2025/08/04 by Wu, Chenfei, Li, Jiahao, Zhou, Jingren +36 · 243 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
2022/02/07 by Peng Wang, Wang, Peng, Yang An +17 · 37 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- Qwen3-Omni Technical Report
2025/09/22 by Xu Jin, Jin Xu, Zhifang Guo +78 · 1 voice · 66 citations
Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Speech and Audio Processing #cs.AI #cs.CL #cs.CV #eess.AS
- Qwen3-VL Technical Report
2025/11/26 by Shuai Bai, Bai, Shuai, Yuxuan Cai +120 · 128 citations
Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Generative Adversarial Networks and Image Synthesis
- ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
2023/05/18 by Wang, Peng, Wang, Shijie, Lin, Junyang +5 · 12 citations
#Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
- CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
2024/12/03 by Yang, Zhibo, Tang, Jun, Li, Zhaohai +9 · 20 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Single Stage Virtual Try-on via Deformable Attention Flows
2022/07/19 by Shuai Bai, Bai, Shuai, Huiling Zhou +7 · 4 citations
Computer Science · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
- TouchStone: Evaluating Vision-Language Models by Language Models
2023/08/31 by Shuai Bai, Bai, Shuai, Shusheng Yang +15 · 5 citations
Arts and Humanities · Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media
- Class-wise Dynamic Graph Convolution for Semantic Segmentation
2020/07/19 by Hanzhe Hu, Deyi Ji, Hu, Hanzhe +9 · 2 citations
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
2024/12/16 by Liang Chen, Chen, Liang, Zekun Wang +49 · 6 citations
Computer Science · #Natural Language Processing Techniques #Topic Modeling
- OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
2022/12/08 by Bai, Jinze, Men, Rui, Yang, Hao +15 · 2 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- Dense Relation Distillation with Context-aware Aggregation for Few-Shot Object Detection
2021/03/30 by Hanzhe Hu, Shuai Bai, Hu, Hanzhe +7 · 1 citation
Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
2025/09/11 by Fang, Rongyao, Yu, Aldrich, Duan, Chengqi +7 · 9 citations
#Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
2025/02/27 by Liang Chen, Chen, Liang, Shuai Bai +12 · 3 citations
Computer Science · Arts and Humanities · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Digital Humanities and Scholarship
- Pretrained Diffusion Models for Unified Human Motion Synthesis
2022/12/06 by Jianxin Ma, Ma, Jianxin, Shuai Bai +3 · 1 citation
Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graphics (cs.GR) #Human Motion and Animation #Human Pose and Action Recognition #Machine Learning (cs.LG) #Robotics (cs.RO)
- Soft Adaptive Policy Optimization
2025/11/25 by Chang Gao, Chujie Zheng, Gao, Chang +16 · 4 citations
Computer Science · #Reinforcement Learning in Robotics #Topic Modeling #Domain Adaptation and Few-Shot Learning