vix.ing · top · new · best · stats · spec

Kehan Li

  1. VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
    2025/01/22 by Boqiang Zhang, Kehan Li, Zhang, Boqiang +26 · 141 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques
  2. Parallel Vertex Diffusion for Unified Visual Grounding
    2023/03/13 by Zesen Cheng, Kehan Li, Cheng, Zesen +11 · 5 citations
    Computer Science · #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  3. FreestyleRet: Retrieving Images from Style-Diversified Queries
    2023/12/05 by Hao Li, Li, Hao, Curise Jia +13 · 5 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Retrieval and Classification Techniques #Information Retrieval (cs.IR) #Multimodal Machine Learning Applications
  4. Locality Guidance for Improving Vision Transformers on Tiny Datasets
    2022/07/20 by Kehan Li, Runyi Yu, Li, Kehan +9 · 3 citations
    Computer Science · Engineering · #Advanced Neural Network Applications #CCD and CMOS Imaging Sensors #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Image Enhancement Techniques
  5. LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
    2025/09/03 by Xingxuan Zhang, Gang Ren, Zhang, Xingxuan +72 · 10 citations
    Computer Science · #Machine Learning and Data Classification #Big Data and Digital Economy #Explainable Artificial Intelligence (XAI)
  6. VoxCPM: Tokenizer-Free TTS for Context-Aware Speech Generation and True-to-Life Voice Cloning
    2025/09/29 by Yixuan Zhou, Zhou, Yixuan, Guoqing Zeng +21 · 3 citations
    Computer Science · #Speech Recognition and Synthesis #Speech and dialogue systems #Speech and Audio Processing
  7. Multi-granularity Interaction Simulation for Unsupervised Interactive Segmentation
    2023/03/23 by Kehan Li, Y. Zhao, Li, Kehan +15 · 1 citation
    Computer Science · #Advanced Neural Network Applications #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  8. RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
    2025/09/18 by Yuming Jiang, Jiang, Yuming, Siteng Huang +23 · 8 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO)
  9. Instance Brownian Bridge as Texts for Open-vocabulary Video Instance Segmentation
    2024/01/18 by Zesen Cheng, Kehan Li, Cheng, Zesen +13 · 1 citation
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #Domain Adaptation and Few-Shot Learning #FOS: Computer and information sciences #Multimodal Machine Learning Applications
  10. Forensic Histopathological Recognition via a Context-Aware MIL Network Powered by Self-Supervised Contrastive Learning
    2023/08/27 by Chen Shen, Jun Zhang, Shen, Chen +13 · 1 citation
    Computer Science · Medicine · Arts and Humanities · #AI in cancer detection #Autopsy Techniques and Outcomes #Forensic Anthropology and Bioarchaeology Studies
  11. RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
    2026/07/20 by Kehan Li, Bohan Hou, Minghao Zhu +27 · 1 voice
    #cs.RO