vix.ing · top · new · best · stats · spec

Pengxiang Ding

  1. QUAR-VLA: Vision-Language-Action Model for Quadruped Robots
    2023/12/22 by Pengxiang Ding, Ding, Pengxiang, Han Zhao +7 · 14 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Multimodal Machine Learning Applications #Robotics (cs.RO) #Robotics and Sensor-Based Localization
  2. OpenHelix: A Short Survey, Empirical Analysis, and Open-Source Dual-System VLA Model for Robotic Manipulation
    2025/05/06 by Can Cui, Cui, Can, Pengxiang Ding +22 · 32 citations
    Computer Science · Engineering · Psychology · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO) #Social Robot Interaction and HRI
  3. Humanoid-VLA: Towards Universal Humanoid Control with Visual Integration
    2025/02/20 by Pengxiang Ding, Jianfei Ma, Ding, Pengxiang +26 · 18 citations
    Computer Science · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Human Pose and Action Recognition #Robotic Path Planning Algorithms #Robotics (cs.RO) #Video Surveillance and Tracking Methods
  4. VLAS: Vision-Language-Action Model With Speech Instructions For Customized Robot Manipulation
    2025/02/19 by Wei Zhao, Pengxiang Ding, Zhao, Wei +10 · 17 citations
    Computer Science · Engineering · #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robot Manipulation and Learning #Robotics (cs.RO) #Robotics and Automated Systems
  5. SSR: Enhancing Depth Perception in Vision-Language Models via Rationale-Guided Spatial Reasoning
    2025/05/18 by Yang Liu, Ming Ma, Liu, Yang +13 · 18 citations
    Computer Science · #Multimodal Machine Learning Applications #Advanced Image and Video Retrieval Techniques #Constraint Satisfaction and Optimization
  6. TrajectoryCNN: A New Spatio-Temporal Feature Learning Network for Human Motion Prediction
    2020/09/03 by Xiaoli Liu, Jianqin Yin, Jin Liu +3 · 4 citations
    Computer Science · Engineering · #Human Pose and Action Recognition #Video Surveillance and Tracking Methods #Human Motion and Animation
  7. MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
    2025/03/11 by Han Zhao, Zhao, Han, Song, Wenxuan +10 · 13 citations
    Computer Science · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotics (cs.RO)
  8. GEVRM: Goal-Expressive Video Generation Model For Robust Visual Manipulation
    2025/02/13 by Hongyin Zhang, Zhang, Hongyin, Pengxiang Ding +7 · 10 citations
    Computer Science · #Advanced Vision and Imaging #FOS: Computer and information sciences #Machine Learning (cs.LG) #Reinforcement Learning in Robotics #Robotics (cs.RO) #Visual Attention and Saliency Detection
  9. GeRM: A Generalist Robotic Model with Mixture-of-experts for Quadruped Robot
    2024/03/20 by Wenxuan Song, Han Zhao, Song, Wenxuan +11 · 7 citations
    Engineering · #Artificial Immune Systems Applications #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Modular Robots and Swarm Intelligence #Robotics (cs.RO) #Robotics and Sensor-Based Localization
  10. Exploring the Evolution of Physics Cognition in Video Generation: A Survey
    2025/03/27 by Minghui Lin, Lin, Minghui, Xiang Wang +17 · 9 citations
    Computer Science · Neuroscience · #Cognitive Science and Education Research #Computer Vision and Pattern Recognition (cs.CV) #Data Visualization and Analytics #FOS: Computer and information sciences #Video Analysis and Summarization
  11. Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
    2025/10/14 by Fuhao Li, Li, Fuhao, Song, Wenxuan +11 · 12 citations
    Computer Science · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Advanced Image and Video Retrieval Techniques
  12. Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions
    2025/05/16 by Wei Zhao, Gongsheng Li, Zhao, Wei +8 · 7 citations
    Computer Science · Psychology · #Advanced Neural Network Applications #FOS: Computer and information sciences #Multimodal Machine Learning Applications #Robotics (cs.RO) #Social Robot Interaction and HRI
  13. ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
    2024/09/30 by Can Cui, Cui, Can, Siteng Huang +9 · 2 citations
    Computer Science · Engineering · #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Gait Recognition and Analysis #Human Pose and Action Recognition #Multimedia (cs.MM) #Video Surveillance and Tracking Methods
  14. VLA-RFT: Vision-Language-Action Reinforcement Fine-tuning with Verified Rewards in World Simulators
    2025/10/01 by Hanyang Li, Pengxiang Ding, Li, Hengtao +19 · 6 citations
    Computer Science · Engineering · #Multimodal Machine Learning Applications #Human Pose and Action Recognition #Robot Manipulation and Learning
  15. Score-Based Diffusion Policy Compatible with Reinforcement Learning via Optimal Transport
    2025/02/18 by Mingyang Sun, Sun, Mingyang, Pengxiang Ding +5 · 4 citations
    Engineering · #Artificial Intelligence (cs.AI) #FOS: Computer and information sciences #Machine Learning (cs.LG) #Traffic control and management
  16. Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denoising Diffusion Process
    2025/11/03 by Jiayi Chen, Wenxuan Song, Chen, Jiayi +13 · 5 citations
    Computer Science · #Multimodal Machine Learning Applications #Generative Adversarial Networks and Image Synthesis #Domain Adaptation and Few-Shot Learning
  17. Enhancing Adversarial Transferability via Component-Wise Transformation
    2025/01/21 by Hangyu Liu, Liu, Hangyu, Bo Peng +6 · 1 citation
    Computer Science · Engineering · Physics and Astronomy · #Advanced Optical Sensing Technologies #Adversarial Robustness in Machine Learning #Computer Vision and Pattern Recognition (cs.CV) #FOS: Computer and information sciences #Fire Detection and Safety Systems