vix.ing · top · new · best · stats · spec

Chengyao Wang

  1. Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
    2024/03/27 by Yanwei Li, Yuechen Zhang, Li, Yanwei +13 · 2 voices · 81 citations
    Computer Science · #Multimodal Machine Learning Applications
  2. LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
    2023/11/28 by Yanwei Li, Chengyao Wang, Li, Yanwei +3 · 164 citations
    Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
  3. VisionZip: Longer is Better but Not Necessary in Vision Language Models
    2024/12/05 by Senqiao Yang, Yukang Chen, Yang, Senqiao +11 · 86 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Multimodal Machine Learning Applications
  4. GroupContrast: Semantic-aware Self-supervised Representation Learning for 3D Understanding
    2024/03/14 by Chengyao Wang, Wang, Chengyao, Li Jiang +11 · 8 citations
    Computer Science · Engineering · #3D Shape Modeling and Analysis #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Image Processing and 3D Reconstruction
  5. DreamOmni: Unified Image Generation and Editing
    2024/12/22 by Bin Xia, Yuechen Zhang, Xia, Bin +13 · 13 citations
    Computer Science · #Computer Graphics and Visualization Techniques
  6. DreamOmni2: Multimodal Instruction-based Editing and Generation
    2025/10/08 by Bin Xia, Xia, Bin, Bohao Peng +22 · 6 citations
    Computer Science · #Natural Language Processing Techniques #Speech and dialogue systems #Topic Modeling