vix.ing · top · new · best · stats · spec

Jinguo Zhu

  1. Attention Residuals
    2026/03/16 by Kimi Team, Guangyu Chen, Yu Zhang +34 · 13 voices · 4 citations
    #cs.CL
  2. Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
    2024/12/06 by Zhe Chen, Chen, Zhe, Weiyun Wang +86 · 3 voices · 421 citations
    Computer Science · Decision Sciences · #Topic Modeling #Scientific Computing and Data Management #Machine Learning and Data Classification
  3. SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
    2024/04/22 by Yuying Ge, Ge, Yuying, Sijie Zhao +15 · 63 citations
    Computer Science · #Topic Modeling #Natural Language Processing Techniques #Speech and dialogue systems
  4. Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
    2024/11/15 by Wei‐Yun Wang, Zhe Chen, Wang, Weiyun +19 · 59 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling
  5. VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
    2025/03/13 by Wei‐Yun Wang, Zhangwei Gao, Wang, Weiyun +26 · 41 citations
    Computer Science · #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #Explainable Artificial Intelligence (XAI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
  6. Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
    2021/12/02 by Xizhou Zhu, Jinguo Zhu, Zhu, Xizhou +13 · 8 citations
    Computer Science · Biochemistry, Genetics and Molecular Biology · #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications #Cell Image Analysis Techniques
  7. VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pre-trained Models
    2023/10/07 by Ziyi Yin, Yin, Ziyi, Muchao Ye +15 · 9 citations
    Computer Science · #Adversarial Robustness in Machine Learning #Multimodal Machine Learning Applications #Topic Modeling
  8. Uni-Perceiver v2: A Generalist Model for Large-Scale Vision and Vision-Language Tasks
    2022/11/17 by Hao Li, Li, Hao, Jinguo Zhu +19 · 6 citations
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  9. SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
    2024/12/12 by Hao Li, Li, Hao, Changyao Tian +19 · 11 citations
    Computer Science · #Image Retrieval and Classification Techniques #Advanced Image and Video Retrieval Techniques
  10. ZeroGUI: Automating Online GUI Learning at Zero Human Cost
    2025/05/29 by Chenyu Yang, Yang, Chenyu, Shiqian Su +21 · 14 citations
    Computer Science · Engineering · Psychology · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Social Robot Interaction and HRI
  11. Layerwise Optimization by Gradient Decomposition for Continual Learning
    2021/05/17 by Shixiang Tang, Tang, Shixiang, Dapeng Chen +7 · 3 citations
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  12. Complementary Relation Contrastive Distillation
    2021/03/29 by Jinguo Zhu, Shixiang Tang, Zhu, Jinguo +13 · 3 citations
    Computer Science · #Domain Adaptation and Few-Shot Learning #Advanced Neural Network Applications #Topic Modeling
  13. Kimi K3: Open Frontier Intelligence
    2026/07/27 by Kimi Team, Tongtong Bai, Yifan Bai +398 · 1 voice
    #cs.CL #cs.LG
  14. PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
    2026/07/27 by Zichao Lin, Yifeng Xie, Bowen Qu +30
    #cs.CV