vix.ing · top · new · best · stats · spec

Zhaoyang Liu

  1. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
    2024/12/06 by Zhe Chen, Chen, Zhe, Weiyun Wang +86 · 3 voices · 439 citations
    Computer Science · Decision Sciences · #Topic Modeling #Scientific Computing and Data Management #Machine Learning and Data Classification
  2. MotionBERT: A Unified Perspective on Learning Human Motion Representations
    2022/10/12 by Wentao Zhu, Zhu, Wentao, Xiaoxuan Ma +9 · 38 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Motion and Animation #Human Pose and Action Recognition
  3. VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
    2024/06/12 by Jiannan Wu, Wu, Jiannan, Muyan Zhong +23 · 42 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  4. What is Wrong with Perplexity for Long-context Language Modeling?
    2024/10/31 by Yifei Wang, Fang, Lizhe, Wang, Yifei +12 · 16 citations
    Computer Science · #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Natural Language Processing Techniques #Topic Modeling
  5. MASTER: Market-Guided Stock Transformer for Stock Price Forecasting
    2023/12/23 by Tong Li, Li, Tong, Zhaoyang Liu +9 · 10 citations
    Computer Science · Decision Sciences · Economics, Econometrics and Finance · #Complex Systems and Time Series Analysis #Computational Engineering #FOS: Computer and information sciences #Finance #Stock Market Forecasting Methods #Time Series Analysis and Forecasting #and Science (cs.CE)
  6. TAM: Temporal Adaptive Module for Video Recognition
    2020/05/14 by Zhaoyang Liu, Liu, Zhaoyang, Limin Wang +7 · 8 citations
    Computer Science · #Human Pose and Action Recognition #Advanced Vision and Imaging #Video Analysis and Summarization
  7. Progressive Attention on Multi-Level Dense Difference Maps for Generic Event Boundary Detection
    2021/12/09 by Jiaqi Tang, Tang, Jiaqi, Zhaoyang Liu +7 · 5 citations
    Computer Science · #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #Generative Adversarial Networks and Image Synthesis
  8. InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language
    2023/05/09 by Zhaoyang Liu, Yinan He, Liu, Zhaoyang +36 · 7 citations
    Computer Science · Neuroscience · #AI in Service Interactions #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Tactile and Sensory Interactions
  9. ControlLLM: Augment Language Models with Tools by Searching on Graphs
    2023/10/26 by Zhaoyang Liu, Zeqiang Lai, Liu, Zhaoyang +18 · 7 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimedia (cs.MM) #Multimodal Machine Learning Applications #Topic Modeling
  10. VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling
    2024/06/06 by Zeyue Tian, Zhaoyang Liu, Tian, Zeyue +15 · 7 citations
    Computer Science · Social Sciences · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimedia (cs.MM) #Multimedia Communication and Technology #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD)
  11. ZeroGUI: Automating Online GUI Learning at Zero Human Cost
    2025/05/29 by Chenyu Yang, Shiqian Su, Yang, Chenyu +21 · 14 citations
    Computer Science · Engineering · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Social Robot Interaction and HRI
  12. AudioX: A Unified Framework for Anything-to-Audio Generation
    2025/03/13 by Zhaoyang Liu, Tian, Zeyue, Jin, Yizhu +10 · 13 citations
    Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #FOS: Electrical engineering #Generative Adversarial Networks and Image Synthesis #Machine Learning (cs.LG) #Multimedia (cs.MM) #Music Technology and Sound Studies #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  13. Joint-Modal Label Denoising for Weakly-Supervised Audio-Visual Video Parsing
    2022/04/25 by Haoyue Cheng, Cheng, Haoyue, Zhaoyang Liu +8 · 3 citations
    Biochemistry, Genetics and Molecular Biology · Computer Science · #Cancer-related molecular mechanisms research #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Speech and Audio Processing #Video Analysis and Summarization
  14. LLMs Meet Multimodal Generation and Editing: A Survey
    2024/05/29 by Yingqing He, He, Yingqing, Zhaoyang Liu +29 · 5 citations
    Arts and Humanities · Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimedia (cs.MM) #Natural Language Processing Techniques #Semantic Web and Ontologies #Sound (cs.SD) #Translation Studies and Practices
  15. TEINet: Towards an Efficient Architecture for Video Recognition
    2019/11/21 by Zhaoyang Liu, Donghao Luo, Liu, Zhaoyang +15 · 2 citations
    Computer Science · Engineering · #Human Pose and Action Recognition #Gait Recognition and Analysis #Anomaly Detection Techniques and Applications
  16. Enhancing the Carrier Transport in Monolayer MoS2 through Interlayer Coupling with 2D Covalent Organic Frameworks
    2023/09/10 by Can Wang, Luca Cusin, Chun Ma +15 · 1 voice · 2 citations
    Materials Science · #2D Materials and Applications #Covalent Organic Framework Applications #Graphene research and applications
  17. VLG: General Video Recognition with Web Textual Knowledge
    2022/12/03 by Jintao Lin, Zhaoyang Liu, Lin, Jintao +7 · 1 citation
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  18. AgentEvolver: Towards Efficient Self-Evolving Agent System
    2025/11/13 by Yunpeng Zhai, Shuchang Tao, Zhai, Yunpeng +23 · 4 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Reinforcement Learning in Robotics #Topic Modeling
  19. MMTrail: A Multimodal Trailer Video Dataset with Language and Music Descriptions
    2024/07/30 by Xiaowei Chi, Yatian Wang, Chi, Xiaowei +35 · 1 citation
    Arts and Humanities · Computer Science · #Audio and Speech Processing (eess.AS) #Computer Vision and Pattern Recognition (cs.CV) #Diverse Musicological Studies #FOS: Computer and information sciences #FOS: Electrical engineering #Multimedia (cs.MM) #Music and Audio Processing #Sound (cs.SD) #electronic engineering #information engineering
  20. ModelGrow: Continual Text-to-Video Pre-training with Model Expansion and Language Understanding Enhancement
    2024/12/25 by Zhefan Rao, Lili Ji, Rao, Zhefan +15 · 1 citation
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling