vix.ing · top · new · best · stats · spec

Xu, Jiazheng

  1. CogVLM: Visual Expert for Pretrained Language Models
    2023/11/06 by Weihan Wang, Qingsong Lv, Wang, Weihan +29 · 2 voices · 83 citations
    Computer Science · #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  2. CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
    2024/08/12 by Yang, Zhuoyi, Teng, Jiayan, Zheng, Wendi +15 · 500 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  3. ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation
    2023/04/12 by Jiazheng Xu, Xiao Liu, Xu, Jiazheng +13 · 215 citations
    Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Artificial Intelligence in Games
  4. CogAgent: A Visual Language Model for GUI Agents
    2023/12/14 by Wenyi Hong, Hong, Wenyi, Weihan Wang +19 · 127 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Natural Language Processing Techniques #Topic Modeling
  5. GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
    2025/07/01 by V Team, Hong, Wenyi, Yu, Wenmeng +85 · 169 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  6. LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks
    2024/12/19 by Yushi Bai, Shangqing Tu, Bai, Yushi +21 · 53 citations
    Computer Science · #Topic Modeling #Semantic Web and Ontologies
  7. VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
    2024/12/30 by Jiazheng Xu, Yu Huang, Xu, Jiazheng +39 · 41 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Video Surveillance and Tracking Methods
  8. VPO: Aligning Text-to-Video Generation Models with Prompt Optimization
    2025/03/26 by Cheng, Jiale, Lyu, Ruiliang, Gu, Xiaotao +9 · 11 citations
    #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. On the Out-Of-Distribution Generalization of Multimodal Large Language Models
    2024/02/09 by Xingxuan Zhang, Jiansheng Li, Zhang, Xingxuan +15 · 3 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques
  10. AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
    2024/06/13 by Yuhang Wu, Wenmeng Yu, Wu, Yuhang +13 · 2 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications