Siteng Huang
- WorldVLA: Towards Autoregressive Action World Model
2025/06/26 by Jun Cen, Cen, Jun, Chaohui Yu +21 · 4 voices · 78 citations
Computer Science · #Data Visualization and Analytics
- Accelerating Diffusion Transformers with Token-wise Feature Caching
2024/10/05 by Chang Zou, Zou, Chang, Xuyang Liu +6 · 31 citations
Computer Science · #Caching and Content Delivery #Advanced Data Storage Technologies #Stochastic Gradient Optimization Techniques
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
2025/05/06 by Can Cui, Pengxiang Ding, Cui, Can +22 · 32 citations
Computer Science · Engineering · Psychology · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO) #Social Robot Interaction and HRI
- VoP: Text-Video Co-operative Prompt Tuning for Cross-Modal Retrieval
2022/11/23 by Siteng Huang, Biao Gong, Huang, Siteng +11 · 6 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Domain Adaptation and Few-Shot Learning
- Learning Disentangled Identifiers for Action-Customized Text-to-Image Generation
2023/11/27 by Siteng Huang, Biao Gong, Huang, Siteng +11 · 1 voice · 3 citations
#cs.CV
- Global Compression Commander: Plug-and-Play Inference Acceleration for High-Resolution Large Vision-Language Models
2025/01/09 by Xuyang Liu, Ziming Wang, Liu, Xuyang +16 · 15 citations
Computer Science · Physics and Astronomy · #Algorithms and Data Compression #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Magnetic confinement fusion research
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
2025/05/18 by Yang Liu, Liu, Yang, Ming Ma +13 · 18 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Constraint Satisfaction and Optimization
- VGDiffZero: Text-to-image Diffusion Models Can Be Zero-shot Visual Grounders
2023/09/03 by Xuyang Liu, Siteng Huang, Liu, Xuyang +7 · 2 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Natural Language Processing Techniques
- M2IST: Multi-Modal Interactive Side-Tuning for Efficient Referring Expression Comprehension
2024/07/01 by Xuyang Liu, Ting Liu, Liu, Xuyang +13 · 3 citations
Computer Science · #Speech and dialogue systems #Topic Modeling #Natural Language Processing Techniques
- ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
2024/09/30 by Can Cui, Siteng Huang, Cui, Can +9 · 2 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition #Multimedia (cs.MM) #Video Surveillance and Tracking Methods
- RynnVLA-001: Using Human Demonstrations to Improve Robot Manipulation
2025/09/18 by Yuming Jiang, Siteng Huang, Jiang, Yuming +23 · 8 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Generative Adversarial Networks and Image Synthesis #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO)
- Variation-aware Vision Token Dropping for Faster Large Vision-Language Models
2025/09/01 by Junjie Chen, Xuyang Liu, Chen, Junjie +9 · 7 citations
Computer Science · #Multimodal Machine Learning Applications #Domain Adaptation and Few-Shot Learning #Generative Adversarial Networks and Image Synthesis
- DARA: Domain- and Relation-aware Adapters Make Parameter-efficient Tuning for Visual Grounding
2024/05/10 by Ting Liu, Liu, Ting, Xuyang Liu +12 · 1 citation
Computer Science · #Advanced Vision and Imaging #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimedia (cs.MM) #Multimodal Machine Learning Applications
- RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model
2026/07/20 by Kehan Li, Bohan Hou, Minghao Zhu +27 · 1 voice
#cs.RO