Zhaoyang Liu
- Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
2024/12/06 by Zhe Chen, Chen, Zhe, Weiyun Wang +86 · 3 voices · 439 citations
Computer Science · Decision Sciences · #Topic Modeling #Scientific Computing and Data Management #Machine Learning and Data Classification
- MotionBERT: A Unified Perspective on Learning Human Motion Representations
2022/10/12 by Wentao Zhu, Zhu, Wentao, Xiaoxuan Ma +9 · 38 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Motion and Animation #Human Pose and Action Recognition
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
2024/06/12 by Jiannan Wu, Wu, Jiannan, Muyan Zhong +23 · 42 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
- What is Wrong with Perplexity for Long-context Language Modeling?
2024/10/31 by Yifei Wang, Fang, Lizhe, Wang, Yifei +12 · 16 citations
Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
- MASTER: Market-Guided Stock Transformer for Stock Price Forecasting
2023/12/23 by Tong Li, Li, Tong, Zhaoyang Liu +9 · 10 citations
Computer Science · Decision Sciences · Economics, Econometrics and Finance · #Complex Systems and Time Series Analysis #Computational Engineering #FOS: Computer and information sciences #Finance #Stock Market Forecasting Methods #Time Series Analysis and Forecasting #and Science (cs.CE)
- TAM: Temporal Adaptive Module for Video Recognition
2020/05/14 by Zhaoyang Liu, Liu, Zhaoyang, Limin Wang +7 · 8 citations
Computer Science · #Human Pose and Action Recognition #Advanced Vision and Imaging #Video Analysis and Summarization
- Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary Detection
2021/12/09 by Jiaqi Tang, Tang, Jiaqi, Zhaoyang Liu +7 · 5 citations
Computer Science · #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #Generative Adversarial Networks and Image Synthesis
- InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
2023/05/09 by Zhaoyang Liu, Yinan He, Liu, Zhaoyang +36 · 7 citations
Computer Science · Neuroscience · #AI in Service Interactions #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Tactile and Sensory Interactions
- ControlLLM: Augment Language Models with Tools by Searching on Graphs
2023/10/26 by Zhaoyang Liu, Zeqiang Lai, Liu, Zhaoyang +18 · 7 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Topic Modeling
- VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
2024/06/06 by Zeyue Tian, Zhaoyang Liu, Tian, Zeyue +15 · 7 citations
Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Multimedia Communication and Technology #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD)
- ZeroGUI: Automating Online GUI Learning at Zero Human Cost
2025/05/29 by Chenyu Yang, Shiqian Su, Yang, Chenyu +21 · 14 citations
Computer Science · Engineering · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Social Robot Interaction and HRI
- AudioX: A Unified Framework for Anything-to-Audio Generation
2025/03/13 by Zhaoyang Liu, Tian, Zeyue, Jin, Yizhu +10 · 13 citations
Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- Joint-Modal Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
2022/04/25 by Haoyue Cheng, Cheng, Haoyue, Zhaoyang Liu +8 · 3 citations
Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Speech and Audio Processing #Video Analysis and Summarization
- LLMs Meet Multimodal Generation and Editing: A Survey
2024/05/29 by Yingqing He, He, Yingqing, Zhaoyang Liu +29 · 5 citations
Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Natural Language Processing Techniques #Semantic Web and Ontologies #Sound (cs.SD) #Translation Studies and Practices
- TEINet: Towards an Efficient Architecture for Video Recognition
2019/11/21 by Zhaoyang Liu, Donghao Luo, Liu, Zhaoyang +15 · 2 citations
Computer Science · Engineering · #Human Pose and Action Recognition #Gait Recognition and Analysis #Anomaly Detection Techniques and Applications
- Enhancing the Carrier Transport in Monolayer MoS2 through Interlayer Coupling with 2D Covalent Organic Frameworks
2023/09/10 by Can Wang, Luca Cusin, Chun Ma +15 · 1 voice · 2 citations
Materials Science · #2D Materials and Applications #Covalent Organic Framework Applications #Graphene research and applications
- VLG: General Video Recognition with Web Textual Knowledge
2022/12/03 by Jintao Lin, Zhaoyang Liu, Lin, Jintao +7 · 1 citation
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
- AgentEvolver: Towards Efficient Self-Evolving Agent System
2025/11/13 by Yunpeng Zhai, Shuchang Tao, Zhai, Yunpeng +23 · 4 citations
Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Topic Modeling
- MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
2024/07/30 by Xiaowei Chi, Yatian Wang, Chi, Xiaowei +35 · 1 citation
Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
- ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
2024/12/25 by Zhefan Rao, Lili Ji, Rao, Zhefan +15 · 1 citation
Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling