vix.ing · top · new · best · stats · spec

Zhengyuan Yang

  1. The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
    2023/09/29 by Zhengyuan Yang, Yang, Zhengyuan, Linjie Li +11 · 11 voices · 70 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  2. GIT: A Generative Image-to-text Transformer for Vision and Language
    2022/05/27 by Jianfeng Wang, Wang, Jianfeng, Zhengyuan Yang +15 · 2 voices · 47 citations
    Computer Science · #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications #Natural Language Processing Techniques #cs.CV
  3. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
    2023/08/04 by Weihao Yu, Zhengyuan Yang, Yu, Weihao +13 · 184 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
  4. MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
    2023/03/20 by Zhengyuan Yang, Yang, Zhengyuan, Linjie Li +16 · 59 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  5. GenXD: Generating Any 3D and 4D Scenes
    2024/11/04 by Yuyang Zhao, Chung-Ching Lin, Zhao, Yuyang +15 · 1 voice · 17 citations
    #cs.CV #cs.AI
  6. Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
    2024/03/05 by Chenglei Si, Yanzhe Zhang, Si, Chenglei +8 · 39 citations
    Engineering · #BIM and Construction Integration #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Manufacturing Process and Optimization
  7. ShowUI: One Vision-Language-Action Model for GUI Visual Agent
    2024/11/26 by Kevin Qinghong Lin, Linjie Li, Lin, Kevin Qinghong +15 · 54 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Multimodal Machine Learning Applications
  8. RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
    2025/04/24 by Zihan Wang, Wang, Zihan, Kangrui Wang +33 · 77 citations
    Computer Science · #Reinforcement Learning in Robotics #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning
  9. Multimodal Foundation Models: From Specialists to General-Purpose Assistants
    2023/09/18 by Chunyuan Li, Zhe Gan, Li, Chunyuan +11 · 28 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  10. ReCo: Region-Controlled Text-to-Image Generation
    2022/11/23 by Zhengyuan Yang, Jianfeng Wang, Yang, Zhengyuan +19 · 19 citations
    Computer Science · #Multimodal Machine Learning Applications #Image Retrieval and Classification Techniques #Generative Adversarial Networks and Image Synthesis
  11. NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
    2023/03/22 by Shengming Yin, Chenfei Wu, Yin, Shengming +28 · 18 citations
    Computer Science · #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #Video Coding and Compression Technologies
  12. An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
    2021/09/10 by Zhengyuan Yang, Yang, Zhengyuan, Zhe Gan +11 · 14 citations
    Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
  13. Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
    2025/01/09 by Yi Hao, Hao, Yunzhuo, Jiawei Gu +11 · 37 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
  14. SAT: 2D Semantics Assisted Training for 3D Visual Grounding
    2021/05/24 by Zhengyuan Yang, Yang, Zhengyuan, Songyang Zhang +5 · 13 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Domain Adaptation and Few-Shot Learning
  15. GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
    2023/11/13 by An Yan, Yan, An, Zhengyuan Yang +21 · 18 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Social Robot Interaction and HRI
  16. Scaling Up Vision-Language Pre-training for Image Captioning
    2021/11/24 by Xiaowei Hu, Hu, Xiaowei, Zhe Gan +11 · 12 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  17. SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
    2025/04/10 by Xiyao Wang, Zhengyuan Yang, Wang, Xiyao +15 · 40 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning and Data Classification #Multimodal Machine Learning Applications
  18. ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
    2025/01/09 by Xingyu Fu, Fu, Xingyu, Min‐Qian Liu +15 · 27 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques
  19. MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
    2024/08/01 by Weihao Yu, Yu, Weihao, Zhengyuan Yang +17 · 17 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Speech and dialogue systems
  20. GRiT: A Generative Region-to-text Transformer for Object Understanding
    2022/12/01 by Jialian Wu, Jianfeng Wang, Wu, Jialian +11 · 8 citations
    Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  21. Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
    2024/06/11 by Yuanhao Zhai, Kevin Lin, Zhai, Yuanhao +15 · 11 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Video Quality Assessment
  22. PromptCap: Prompt-Guided Task-Aware Image Captioning
    2022/11/15 by Yushi Hu, Hang Hua, Hu, Yushi +9 · 6 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Image and Video Retrieval Techniques
  23. TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
    2020/12/08 by Zhengyuan Yang, Yang, Zhengyuan, Yijuan Lu +15 · 5 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  24. MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
    2024/06/12 by Xuehai He, Weixi Feng, He, Xuehai +25 · 9 citations
    Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
  25. MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning
    2023/11/29 by Chaoyi Zhang, Zhang, Chaoyi, Kevin Lin +13 · 7 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Video Analysis and Summarization
  26. List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
    2024/04/25 by Yan An, Yan, An, Zhengyuan Yang +19 · 7 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Lexicography and Language Studies #Natural Language Processing Techniques #Semantic Web and Ontologies
  27. End-to-end Multi-Modal Multi-Task Vehicle Control for Self-Driving Cars with Visual Perception
    2018/01/20 by Zhengyuan Yang, Yixuan Zhang, Yang, Zhengyuan +7 · 4 citations
    Computer Science · Engineering · #Advanced Neural Network Applications #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotic Path Planning Algorithms
  28. Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
    2020/07/03 by Liwei Wang, Wang, Liwei, Jing Huang +9 · 3 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  29. SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
    2024/10/30 by Yining Hong, Baiyu Liu, Hong, Yining +21 · 7 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Cell Image Analysis Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Robotics (cs.RO)
  30. EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
    2024/10/03 by Zheng, Kaizhi, Xiaotong Chen, Xuehai He +16 · 7 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Artificial Intelligence (cs.AI) #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graph Theory and Algorithms #Graphics (cs.GR) #Human-Computer Interaction (cs.HC)
  31. Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
    2023/10/12 by Zhengyuan Yang, Yang, Zhengyuan, Jianfeng Wang +11 · 4 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  32. Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample Optimization
    2024/10/04 by Zichen Miao, Miao, Zichen, Zhengyuan Yang +11 · 6 citations
    Computer Science · #Neural Networks and Applications
  33. MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
    2024/10/13 by Hang Hua, Yunlong Tang, Hua, Hang +13 · 7 citations
    Arts and Humanities · #Media, Religion, Digital Communication
  34. Audio-Aware Large Language Models as Judges for Speaking Styles
    2025/06/06 by Cheng-Han Chiang, Chiang, Cheng-Han, Xiaofei Wang +18 · 10 citations
    Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #electronic engineering #information engineering
  35. MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos
    2023/06/07 by Jielin Qiu, Qiu, Jielin, Jiacheng Zhu +21 · 3 citations
    Computer Science · #Video Analysis and Summarization #Natural Language Processing Techniques #Music and Audio Processing
  36. ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
    2025/06/11 by Xiyao Wang, Zhengyuan Yang, Wang, Xiyao +22 · 11 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  37. COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
    2024/01/01 by Alex Jinpeng Wang, Wang, Alex Jinpeng, Linjie Li +13 · 3 citations
    Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
  38. Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
    2025/05/26 by Minheng Ni, Ni, Minheng, Zhengyuan Yang +11 · 10 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
  39. StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
    2024/01/30 by Zecheng Tang, Chenfei Wu, Tang, Zecheng +19 · 3 citations
    Engineering · #Additive Manufacturing and 3D Printing Technologies #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Manufacturing Process and Optimization
  40. VideoGUI: A Benchmark for GUI Automation from Instructional Videos
    2024/06/14 by Kevin Qinghong Lin, Lin, Kevin Qinghong, Linjie Li +13 · 5 citations
    Computer Science · #Video Analysis and Summarization
  41. VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
    2025/10/19 by Kangrui Wang, Pingyue Zhang, Wang, Kangrui +27 · 10 citations
    Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
  42. Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
    2024/12/12 by Jitesh Jain, Zhengyuan Yang, Jain, Jitesh +7 · 3 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Semantic Web and Ontologies #Video Analysis and Summarization
  43. What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
    2025/06/08 by Ming Li, Zhengyuan Yang, Li, Ming +11 · 5 citations
    Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Constraint Satisfaction and Optimization
  44. GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
    2025/07/13 by Yiyang Zhou, Zhou, Yiyang, Linjie Li +21 · 6 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  45. Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors
    2026/07/17 by Yilin Wang, Xiangxi Zheng, Dongxing Mao +6
    #cs.CV