Huang, Shijia
- Towards Learning a Generalist Model for Embodied Navigation
2023/12/04 by Zheng, Duo, Huang, Shijia, Zhao, Lin +2 · 35 citations
#Artificial Intelligence (cs.AI) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Multi-View Transformer for 3D Visual Grounding
2022/04/05 by Shijia Huang, Yilun Chen, Huang, Shijia +5 · 20 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Neural Network Applications
- Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
2024/11/30 by Duo Zheng, Zheng, Duo, Shijia Huang +3 · 38 citations
Computer Science · Engineering · #Advanced Image and Video Retrieval Techniques #Robotics and Sensor-Based Localization #Advanced Vision and Imaging
- LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
2023/12/05 by Hao Zhang, Hongyang Li, Zhang, Hao +19 · 23 citations
Computer Science · Arts and Humanities · #Multimodal Machine Learning Applications #Subtitles and Audiovisual Media #Speech and dialogue systems
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
2025/05/30 by Duo Zheng, Shijia Huang, Zheng, Duo +5 · 23 citations
Computer Science · #Open Education and E-Learning
- DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and Grounding
2022/11/28 by Liu, Shilong, Liang, Yaoyuan, Li, Feng +5 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- MP-Former: Mask-Piloted Transformer for Image Segmentation
2023/03/13 by Zhang, Hao, Li, Feng, Xu, Huaizhe +4 · 3 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- Enhancing Temporal Modeling of Video LLMs via Time Gating
2024/10/08 by Hu, Zi-Yuan, Zhong, Yiwu, Huang, Shijia +2 · 5 citations
#Artificial Intelligence (cs.AI) #Computation and Language (cs.CL) #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG)
- DSGN++: Exploiting Visual-Spatial Relation for Stereo-based 3D Detectors
2022/04/06 by Chen, Yilun, Huang, Shijia, Liu, Shu +2 · 2 citations
#Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences
- CLEVA: Chinese Language Models EVAluation Platform
2023/08/09 by Li, Yanyang, Zhao, Jianqiao, Zheng, Duo +8 · 2 citations
#Computation and Language (cs.CL) #FOS: Computer and information sciences