vix.ing · top · new · best · stats · spec

Yuan, Zhihao

  1. InstanceRefer: Cooperative Holistic Understanding for Visual Grounding on Point Clouds through Instance Multi-level Contextual Referring
    2021/03/01 by Zhihao Yuan, Yan Xu, Yuan, Zhihao +11 · 19 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Neural Network Applications
  2. Visual Programming for Zero-shot Open-Vocabulary 3D Visual Grounding
    2023/11/26 by Zhihao Yuan, Yuan, Zhihao, Jinke Ren +9 · 27 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  3. X-Trans2Cap: Cross-Modal Knowledge Transfer using Transformer for 3D Dense Captioning
    2022/03/02 by Zhihao Yuan, Yan Xu, Yuan, Zhihao +11 · 8 citations
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  4. Comprehensive Visual Question Answering on Point Clouds through Compositional Scene Manipulation
    2021/12/22 by Yan Xu, Zhihao Yuan, Yan, Xu +11 · 5 citations
    Computer Science · #Advanced Image and Video Retrieval Techniques #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications
  5. Sciunits: Reusable Research Objects
    2017/07/18 by That, Dai Hai Ton, Fils, Gabriel, Yuan, Zhihao +1 · 2 citations
    #Digital Libraries (cs.DL) #FOS: Computer and information sciences
  6. Toward Explainable and Fine-Grained 3D Grounding through Referring Textual Phrases
    2022/07/05 by Yuan, Zhihao, Yan, Xu, Li, Zhuo +4 · 1 citation
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
    2025/06/21 by Yuan, Zhihao, Jiang, Shuyi, Feng, Chun-Mei +4 · 5 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  8. Agent Data Protocol: Unifying Datasets for Diverse, Effective Fine-tuning of LLM Agents
    2025/10/28 by Song, Yueqi, Ramaneti, Ketan, Sheikh, Zaid +18 · 3 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #FOS: Computer and information sciences
  9. PiSA: A Self-Augmented Data Engine and Training Strategy for 3D Understanding with Large Models
    2025/03/13 by Guo, Zilu, Lin, Hongbin, Yuan, Zhihao +6 · 2 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  10. Empowering Large Language Models with 3D Situation Awareness
    2025/03/29 by Yuan, Zhihao, Jinke Ren, Peng, Yibo +13 · 2 citations
    Computer Science · #Advanced Neural Network Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Topic Modeling
  11. See the Forest and the Trees: A Synergistic Reasoning Framework for Knowledge-Based Visual Question Answering
    2025/07/23 by Wang, Junjie, Tang, Yunhan, Wang, Yijie +4 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences