vix.ing · top · new · best · stats · spec

Huang, Shijia

  1. Towards Learning a Generalist Model for Embodied Navigation
    2023/12/04 by Zheng, Duo, Huang, Shijia, Zhao, Lin +2 · 35 citations
    #Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  2. Multi-View Transformer for 3D Visual Grounding
    2022/04/05 by Shijia Huang, Yilun Chen, Huang, Shijia +5 · 20 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Neural Network Applications
  3. Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
    2024/11/30 by Duo Zheng, Zheng, Duo, Shijia Huang +3 · 38 citations
    Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Robotics and Sensor-Based Localization #Advanced Vision and Imaging
  4. LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
    2023/12/05 by Hao Zhang, Hongyang Li, Zhang, Hao +19 · 23 citations
    Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media #Speech and dialogue systems
  5. Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
    2025/05/30 by Duo Zheng, Shijia Huang, Zheng, Duo +5 · 23 citations
    Computer Science · #Open Education and E-Learning
  6. DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding
    2022/11/28 by Liu, Shilong, Liang, Yaoyuan, Li, Feng +5 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  7. MP-Former: Mask-Piloted Transformer for Image Segmentation
    2023/03/13 by Zhang, Hao, Li, Feng, Xu, Huaizhe +4 · 3 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  8. Enhancing Temporal Modeling of Video LLMs via Time Gating
    2024/10/08 by Hu, Zi-Yuan, Zhong, Yiwu, Huang, Shijia +2 · 5 citations
    #Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
  9. DSGN++: Exploiting Visual-Spatial Relation for Stereo-based 3D Detectors
    2022/04/06 by Chen, Yilun, Huang, Shijia, Liu, Shu +2 · 2 citations
    #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
  10. CLEVA: Chinese Language Models EVAluation Platform
    2023/08/09 by Li, Yanyang, Zhao, Jianqiao, Zheng, Duo +8 · 2 citations
    #Computation and Language (cs.CL) #FOS: Computer and information sciences