Pengxiang Ding
- QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
2023/12/22 by Pengxiang Ding, Ding, Pengxiang, Han Zhao +7 · 14 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Robotics (cs.RO) #Robotics and Sensor-Based Localization
- OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
2025/05/06 by Can Cui, Cui, Can, Pengxiang Ding +22 · 32 citations
Computer Science · Engineering · Psychology · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO) #Social Robot Interaction and HRI
- Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
2025/02/20 by Pengxiang Ding, Jianfei Ma, Ding, Pengxiang +26 · 18 citations
Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Robotic Path Planning Algorithms #Robotics (cs.RO) #Video Surveillance and Tracking Methods
- VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
2025/02/19 by Wei Zhao, Pengxiang Ding, Zhao, Wei +10 · 17 citations
Computer Science · Engineering · #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO) #Robotics and Automated Systems
- SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
2025/05/18 by Yang Liu, Ming Ma, Liu, Yang +13 · 18 citations
Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Constraint Satisfaction and Optimization
- TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion Prediction
2020/09/03 by Xiaoli Liu, Jianqin Yin, Jin Liu +3 · 4 citations
Computer Science · Engineering · #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #Human Motion and Animation
- MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
2025/03/11 by Han Zhao, Zhao, Han, Song, Wenxuan +10 · 13 citations
Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotics (cs.RO)
- GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
2025/02/13 by Hongyin Zhang, Zhang, Hongyin, Pengxiang Ding +7 · 10 citations
Computer Science · #Advanced Vision and Imaging #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO) #Visual Attention and Saliency Detection
- GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
2024/03/20 by Wenxuan Song, Han Zhao, Song, Wenxuan +11 · 7 citations
Engineering · #Artificial Immune Systems Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Modular Robots and Swarm Intelligence #Robotics (cs.RO) #Robotics and Sensor-Based Localization
- Exploring the Evolution of Physics Cognition in Video Generation: A Survey
2025/03/27 by Minghui Lin, Lin, Minghui, Xiang Wang +17 · 9 citations
Computer Science · Neuroscience · #Cognitive Science and Education Research #Computer Vision and Pattern Recognition (cs.CV) #Data Visualization and Analytics #FOS: Computer and information sciences #Video Analysis and Summarization
- Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
2025/10/14 by Fuhao Li, Li, Fuhao, Song, Wenxuan +11 · 12 citations
Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Image and Video Retrieval Techniques
- Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
2025/05/16 by Wei Zhao, Gongsheng Li, Zhao, Wei +8 · 7 citations
Computer Science · Psychology · #Advanced Neural Network Applications #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotics (cs.RO) #Social Robot Interaction and HRI
- ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
2024/09/30 by Can Cui, Cui, Can, Siteng Huang +9 · 2 citations
Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition #Multimedia (cs.MM) #Video Surveillance and Tracking Methods
- VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
2025/10/01 by Hanyang Li, Pengxiang Ding, Li, Hengtao +19 · 6 citations
Computer Science · Engineering · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Robot Manipulation and Learning
- Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
2025/02/18 by Mingyang Sun, Sun, Mingyang, Pengxiang Ding +5 · 4 citations
Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Traffic control and management
- Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
2025/11/03 by Jiayi Chen, Wenxuan Song, Chen, Jiayi +13 · 5 citations
Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
- Enhancing Adversarial Transferability via Component-Wise Transformation
2025/01/21 by Hangyu Liu, Liu, Hangyu, Bo Peng +6 · 1 citation
Computer Science · Engineering · Physics and Astronomy · #Advanced Optical Sensing Technologies #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Fire Detection and Safety Systems