vix.ing · top · new · best · stats · spec

Bai, Shuai

  1. Qwen2.5-VL Technical Report
    2025/02/19 by Shuai Bai, Keqin Chen, Bai, Shuai +48 · 1702 citations
    Engineering · #Semiconductor Lasers and Optical Devices
  2. Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
    2024/09/18 by Peng Wang, Wang, Peng, Shuai Bai +35 · 1056 citations
    Psychology · #Categorization, perception, and language
  3. Qwen Technical Report
    2023/09/28 by Jinze Bai, Bai, Jinze, Shuai Bai +91 · 666 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques
  4. Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
    2023/08/24 by Jinze Bai, Shuai Bai, Bai, Jinze +15 · 559 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  5. Qwen2 Technical Report
    2024/07/15 by Yang An, Baosong Yang, Yang, An +120 · 377 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  6. An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
    2024/03/11 by Liang Chen, Chen, Liang, Haozhe Zhao +11 · 137 citations
    Computer Science · Medicine · #Multimodal Machine Learning Applications #COVID-19 diagnosis using AI
  7. Qwen2.5-Omni Technical Report
    2025/03/26 by Jin Xu, Zihan Guo, Xu, Jin +22 · 249 citations
    Engineering · #Embedded Systems and FPGA Design
  8. Qwen-Image Technical Report
    2025/08/04 by Wu, Chenfei, Li, Jiahao, Zhou, Jingren +36 · 243 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  9. OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning Framework
    2022/02/07 by Peng Wang, Wang, Peng, Yang An +17 · 37 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  10. Qwen3-Omni Technical Report
    2025/09/22 by Xu Jin, Jin Xu, Zhifang Guo +78 · 1 voice · 66 citations
    Computer Science · Engineering · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Speech and Audio Processing #cs.AI #cs.CL #cs.CV #eess.AS
  11. Qwen3-VL Technical Report
    2025/11/26 by Shuai Bai, Bai, Shuai, Yuxuan Cai +120 · 128 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Generative Adversarial Networks and Image Synthesis
  12. ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
    2023/05/18 by Wang, Peng, Wang, Shijie, Lin, Junyang +5 · 12 citations
    #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Sound (cs.SD) #electronic engineering #information engineering
  13. CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy
    2024/12/03 by Yang, Zhibo, Tang, Jun, Li, Zhaohai +9 · 20 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  14. Single Stage Virtual Try-on via Deformable Attention Flows
    2022/07/19 by Shuai Bai, Bai, Shuai, Huiling Zhou +7 · 4 citations
    Computer Science · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Human Pose and Action Recognition
  15. TouchStone: Evaluating Vision-Language Models by Language Models
    2023/08/31 by Shuai Bai, Bai, Shuai, Shusheng Yang +15 · 5 citations
    Arts and Humanities · Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media
  16. Class-wise Dynamic Graph Convolution for Semantic Segmentation
    2020/07/19 by Hanzhe Hu, Deyi Ji, Hu, Hanzhe +9 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  17. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
    2024/12/16 by Liang Chen, Chen, Liang, Zekun Wang +49 · 6 citations
    Computer Science · #Natural Language Processing Techniques #Topic Modeling
  18. OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models
    2022/12/08 by Bai, Jinze, Men, Rui, Yang, Hao +15 · 2 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  19. Dense Relation Distillation with Context-aware Aggregation for Few-Shot Object Detection
    2021/03/30 by Hanzhe Hu, Shuai Bai, Hu, Hanzhe +7 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  20. FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
    2025/09/11 by Fang, Rongyao, Yu, Aldrich, Duan, Chengqi +7 · 9 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  21. Multimodal Representation Alignment for Image Generation: Text-Image Interleaved Control Is Easier Than You Think
    2025/02/27 by Liang Chen, Chen, Liang, Shuai Bai +12 · 3 citations
    Computer Science · Arts and Humanities · #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Digital Humanities and Scholarship
  22. Pretrained Diffusion Models for Unified Human Motion Synthesis
    2022/12/06 by Jianxin Ma, Ma, Jianxin, Shuai Bai +3 · 1 citation
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graphics (cs.GR) #Human Motion and Animation #Human Pose and Action Recognition #Machine Learning (cs.LG) #Robotics (cs.RO)
  23. Soft Adaptive Policy Optimization
    2025/11/25 by Chang Gao, Chujie Zheng, Gao, Chang +16 · 4 citations
    Computer Science · #Reinforcement Learning in Robotics #Topic Modeling #Domain Adaptation and Few-Shot Learning