Zhengyuan Yang
- The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
2023/09/29 by Zhengyuan Yang, Yang, Zhengyuan, Linjie Li +11 · 11 voices · 70 citations
Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
- GIT: A Generative Image-to-text Transformer for Vision and Language
2022/05/27 by Jianfeng Wang, Wang, Jianfeng, Zhengyuan Yang +15 · 2 voices · 47 citations
Computer Science · #Handwritten Text Recognition Techniques #Multimodal Machine Learning Applications #Natural Language Processing Techniques #cs.CV
- MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
2023/08/04 by Weihao Yu, Zhengyuan Yang, Yu, Weihao +13 · 184 citations
Computer Science · #Topic Modeling #Natural Language Processing Techniques #Multimodal Machine Learning Applications
- MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action
2023/03/20 by Zhengyuan Yang, Yang, Zhengyuan, Linjie Li +16 · 59 citations
Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
- GenXD: Generating Any 3D and 4D Scenes
2024/11/04 by Yuyang Zhao, Chung-Ching Lin, Zhao, Yuyang +15 · 1 voice · 17 citations
#cs.CV #cs.AI
- Design2Code: Benchmarking Multimodal Code Generation for Automated Front-End Engineering
2024/03/05 by Chenglei Si, Yanzhe Zhang, Si, Chenglei +8 · 39 citations
Engineering · #BIM and Construction Integration #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Computers and Society (cs.CY) #FOS: Computer and information sciences #Manufacturing Process and Optimization
- ShowUI: One Vision-Language-Action Model for GUI Visual Agent
2024/11/26 by Kevin Qinghong Lin, Linjie Li, Lin, Kevin Qinghong +15 · 54 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human-Computer Interaction (cs.HC) #Multimodal Machine Learning Applications
- RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
2025/04/24 by Zihan Wang, Wang, Zihan, Kangrui Wang +33 · 77 citations
Computer Science · #Reinforcement Learning in Robotics #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning
- Multimodal Foundation Models: From Specialists to General-Purpose Assistants
2023/09/18 by Chunyuan Li, Zhe Gan, Li, Chunyuan +11 · 28 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- ReCo: Region-Controlled Text-to-Image Generation
2022/11/23 by Zhengyuan Yang, Jianfeng Wang, Yang, Zhengyuan +19 · 19 citations
Computer Science · #Multimodal Machine Learning Applications #Image Retrieval and Classification Techniques #Generative Adversarial Networks and Image Synthesis
- NUWA-XL: Diffusion over Diffusion for eXtremely Long Video Generation
2023/03/22 by Shengming Yin, Chenfei Wu, Yin, Shengming +28 · 18 citations
Computer Science · #Generative Adversarial Networks and Image Synthesis #Video Analysis and Summarization #Video Coding and Compression Technologies
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA
2021/09/10 by Zhengyuan Yang, Yang, Zhengyuan, Zhe Gan +11 · 14 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
- Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
2025/01/09 by Yi Hao, Hao, Yunzhuo, Jiawei Gu +11 · 37 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
- SAT: 2D Semantics Assisted Training for 3D Visual Grounding
2021/05/24 by Zhengyuan Yang, Yang, Zhengyuan, Songyang Zhang +5 · 13 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Domain Adaptation and Few-Shot Learning
- GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
2023/11/13 by An Yan, Yan, An, Zhengyuan Yang +21 · 18 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Social Robot Interaction and HRI
- Scaling Up Vision-Language Pre-training for Image Captioning
2021/11/24 by Xiaowei Hu, Hu, Xiaowei, Zhe Gan +11 · 12 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement
2025/04/10 by Xiyao Wang, Zhengyuan Yang, Wang, Xiyao +15 · 40 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Machine Learning and Data Classification #Multimodal Machine Learning Applications
- ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
2025/01/09 by Xingyu Fu, Fu, Xingyu, Min‐Qian Liu +15 · 27 citations
Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques
- MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
2024/08/01 by Weihao Yu, Yu, Weihao, Zhengyuan Yang +17 · 17 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Speech and dialogue systems
- GRiT: A Generative Region-to-text Transformer for Object Understanding
2022/12/01 by Jialian Wu, Jianfeng Wang, Wu, Jialian +11 · 8 citations
Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Motion Consistency Model: Accelerating Video Diffusion with Disentangled Motion-Appearance Distillation
2024/06/11 by Yuanhao Zhai, Kevin Lin, Zhai, Yuanhao +15 · 11 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image and Video Quality Assessment
- PromptCap: Prompt-Guided Task-Aware Image Captioning
2022/11/15 by Yushi Hu, Hang Hua, Hu, Yushi +9 · 6 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Image and Video Retrieval Techniques
- TAP: Text-Aware Pre-training for Text-VQA and Text-Caption
2020/12/08 by Zhengyuan Yang, Yang, Zhengyuan, Yijuan Lu +15 · 5 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
- MMWorld: Towards Multi-discipline Multi-faceted World Model Evaluation in Videos
2024/06/12 by Xuehai He, Weixi Feng, He, Xuehai +25 · 9 citations
Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Video Surveillance and Tracking Methods
- MM-Narrator: Narrating Long-form Videos with Multimodal In-Context Learning
2023/11/29 by Chaoyi Zhang, Zhang, Chaoyi, Kevin Lin +13 · 7 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Music and Audio Processing #Natural Language Processing Techniques #Video Analysis and Summarization
- List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
2024/04/25 by Yan An, Yan, An, Zhengyuan Yang +19 · 7 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Lexicography and Language Studies #Natural Language Processing Techniques #Semantic Web and Ontologies
- End-to-end Multi-Modal Multi-Task Vehicle Control for Self-Driving Cars with Visual Perception
2018/01/20 by Zhengyuan Yang, Yixuan Zhang, Yang, Zhengyuan +7 · 4 citations
Computer Science · Engineering · #Advanced Neural Network Applications #Autonomous Vehicle Technology and Safety #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Robotic Path Planning Algorithms
- Improving Weakly Supervised Visual Grounding by Contrastive Knowledge Distillation
2020/07/03 by Liwei Wang, Wang, Liwei, Jing Huang +9 · 3 citations
Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- SlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video Generation
2024/10/30 by Yining Hong, Baiyu Liu, Hong, Yining +21 · 7 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Advanced Vision and Imaging #Artificial Intelligence (cs.AI) #Cell Image Analysis Techniques #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Machine Learning (cs.LG) #Robotics (cs.RO)
- EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing
2024/10/03 by Zheng, Kaizhi, Xiaotong Chen, Xuehai He +16 · 7 citations
Computer Science · Engineering · #3D Shape Modeling and Analysis #Artificial Intelligence (cs.AI) #Computer Graphics and Visualization Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Graph Theory and Algorithms #Graphics (cs.GR) #Human-Computer Interaction (cs.HC)
- Idea2Img: Iterative Self-Refinement with GPT-4V(ision) for Automatic Image Design and Generation
2023/10/12 by Zhengyuan Yang, Yang, Zhengyuan, Jianfeng Wang +11 · 4 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
- Tuning Timestep-Distilled Diffusion Model Using Pairwise Sample Optimization
2024/10/04 by Zichen Miao, Miao, Zichen, Zhengyuan Yang +11 · 6 citations
Computer Science · #Neural Networks and Applications
- MMCOMPOSITION: Revisiting the Compositionality of Pre-trained Vision-Language Models
2024/10/13 by Hang Hua, Yunlong Tang, Hua, Hang +13 · 7 citations
Arts and Humanities · #Media, Religion, Digital Communication
- Audio-Aware Large Language Models as Judges for Speaking Styles
2025/06/06 by Cheng-Han Chiang, Chiang, Cheng-Han, Xiaofei Wang +18 · 10 citations
Computer Science · Psychology · #Artificial Intelligence (cs.AI) #Audio and Speech Processing (eess.AS) #Computation and Language (cs.CL) #Emotion and Mood Recognition #FOS: Computer and information sciences #FOS: Electrical engineering #Music and Audio Processing #Speech Recognition and Synthesis #electronic engineering #information engineering
- MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos
2023/06/07 by Jielin Qiu, Qiu, Jielin, Jiacheng Zhu +21 · 3 citations
Computer Science · #Video Analysis and Summarization #Natural Language Processing Techniques #Music and Audio Processing
- ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs
2025/06/11 by Xiyao Wang, Zhengyuan Yang, Wang, Xiyao +22 · 11 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
- COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
2024/01/01 by Alex Jinpeng Wang, Wang, Alex Jinpeng, Linjie Li +13 · 3 citations
Computer Science · #Multimodal Machine Learning Applications #Topic Modeling #Natural Language Processing Techniques
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
2025/05/26 by Minheng Ni, Ni, Minheng, Zhengyuan Yang +11 · 10 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Semantic Web and Ontologies #Speech and dialogue systems
- StrokeNUWA: Tokenizing Strokes for Vector Graphic Synthesis
2024/01/30 by Zecheng Tang, Chenfei Wu, Tang, Zecheng +19 · 3 citations
Engineering · #Additive Manufacturing and 3D Printing Technologies #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Manufacturing Process and Optimization
- VideoGUI: A Benchmark for GUI Automation from Instructional Videos
2024/06/14 by Kevin Qinghong Lin, Lin, Kevin Qinghong, Linjie Li +13 · 5 citations
Computer Science · #Video Analysis and Summarization
- VAGEN: Reinforcing World Model Reasoning for Multi-Turn VLM Agents
2025/10/19 by Kangrui Wang, Pingyue Zhang, Wang, Kangrui +27 · 10 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Topic Modeling
- Elevating Visual Perception in Multimodal LLMs with Visual Embedding Distillation
2024/12/12 by Jitesh Jain, Zhengyuan Yang, Jain, Jitesh +7 · 3 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Semantic Web and Ontologies #Video Analysis and Summarization
- What makes Reasoning Models Different? Follow the Reasoning Leader for Efficient Decoding
2025/06/08 by Ming Li, Zhengyuan Yang, Li, Ming +11 · 5 citations
Computer Science · #Topic Modeling #Multimodal Machine Learning Applications #Constraint Satisfaction and Optimization
- GLIMPSE: Do Large Vision-Language Models Truly Think With Videos or Just Glimpse at Them?
2025/07/13 by Yiyang Zhou, Zhou, Yiyang, Linjie Li +21 · 6 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors
2026/07/17 by Yilin Wang, Xiangxi Zheng, Dongxing Mao +6
#cs.CV